Files
sentiment-engine/MALKHUT/README.md

1399 lines
67 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# MALKHUT — Adversarial Self-Play Order-Fulfilment Pipeline
**Lock-free. No GC. Fully async. GraalVM-compatible.**
MALKHUT is an order-fulfilment engine that optimises a **distribution over execution
actions** against a diverse counterparty ecology, not a single deterministic quote.
## Core Thesis
> Sequential/pure optimisation ("what is the best quote?") leads to
> deterministic quote → predictable fill pattern → adverse selection.
>
> Correct simultaneous-market framing: "what mixed order-placement policy
> survives a diverse counterparty ecology?"
## Architecture
### Corrected Stack (T19 ANNEX B-CORRECTION)
> **iceoryx2 = the wires; ASEx = the state cells; UV clock = the conductor.**
| Layer | Component | Role |
|-------|-----------|------|
| Transport | iceoryx2/zinc topics | Inter-component event fabric. Observable in `/dev/shm`. Single-writer per topic. |
| State | ASEx (ASExGuardedState + ASExWorker) | Validate-before-mutate state serialization. One writer thread per state. NOT a message bus. |
| Dispatch | UV Clock (asyncio event loop) | Event-driven reactor. Subscribes to event types. Nothing sleeps-and-polls. |
**Per T19:** ASExWatch is **DEFERRED** from the critical path until heavy-test hangs are fixed.
Use ASExWorker (FulfilmentWorker) for production hot-path mutations.
```
┌──────────────────────────────────────────────────────────────────────┐
│ Zinc SHM (POSIX /dev/shm/zinc_*) │
│ malkhut_book_state | malkhut_account_state | malkhut_fulfilment_out │
│ malkhut_risk_gate | malkhut_control (CONTROL_PLANE) │
└───────────────────────────────┬──────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────────────┐
│ ASEx VALIDATE-BEFORE-MUTATE │
│ GuardedFulfilmentState — book/account/intent/policy mutations │
│ GuardedRiskState — risk gate + kill switch │
│ FulfilmentWorker — single-threaded ASExWorker per state │
│ FulfilmentWatch — zero-overhead ring buffer for hot path │
│ ShardedWorker — per-symbol state partitioning │
└───────────────────────────────┬──────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────────────┐
│ CODE WORLD MODEL (CWM) │
│ deterministic exchange transition function │
│ f(state, joint_actions, latency_model, venue_rules) → next_state │
│ wraps hftbacktest for replay verification │
└───────────────────────────────┬──────────────────────────────────────┘
│
┌──────────┴──────────┐
▼ ▼
┌─────────────────────────┐ ┌──────────────────────────────────┐
│ LIVE FULFILMENT PLANNER │ │ OFFLINE CMA-ES TRAINER │
│ Decoupled UCB/UCT │ │ pycma + nevergrad │
│ ≤ 25 ms cadence │ │ growing self-play pool │
└───────────┬──────────────┘ └─────────────────┬────────────────┘
│ │
▼ ▼
┌─────────────────────────┐ ┌──────────────────────────────────┐
│ RISK / COMPLIANCE GATE │ │ POLICY REGISTRY / POOL │
│ leverage, notional, │ │ accepted candidates, weights │
│ self-trade, kill switch │ │ diversity-preserving eviction │
└───────────┬──────────────┘ └──────────────────────────────────┘
│
▼
┌─────────────────────────┐
│ VENUE ADAPTER (BingX) │
│ create/cancel/replace │
└───────────┬──────────────┘
│
▼
┌─────────────────────────┐
│ ClickHouse Persistence │
│ decisions, episodes, │
│ snapshots, discrepancies│
└─────────────────────────┘
```
## Package Structure
```
MALKHUT/
├── malkhut/
│ ├── __init__.py # package version
│ ├── state.py # frozen data model (42 dataclasses)
│ ├── actions.py # action model, PlannedPolicy, RiskDecision
│ ├── features.py # feature extraction (17 features)
│ ├── counterparties.py # adversarial agent ecology (4 agents)
│ ├── engine.py # FulfilmentEngine (hot-path orchestrator)
│ ├── bench_numba.py # numba speedup benchmark
│ ├── launch_smoke_test.py # 10-min smoke test launcher
│ ├── cwm/
│ │ ├── __init__.py
│ │ ├── core.py # CWM (deterministic exchange simulator)
│ │ ├── numba_core.py # numba-accelerated hot loops
│ │ └── replay_verify.py # replay verification (mandatory before trust)
│ ├── planner/
│ │ ├── __init__.py
│ │ ├── action_menu.py # compact action space construction
│ │ └── sm_mcts.py # Decoupled UCB/UCT simultaneous-move MCTS
│ ├── risk/
│ │ ├── __init__.py
│ │ └── gate.py # risk/compliance gate (hard invariants)
│ ├── venue/
│ │ ├── __init__.py
│ │ └── bingx/
│ │ ├── __init__.py
│ │ └── adapter.py # BingX VST/Live venue adapter (wraps DITAv2)
│ ├── ipc/
│ │ ├── __init__.py
│ │ ├── zinc_plane.py # Zinc shared memory IPC (POSIX SHM)
│ │ └── control_plane.py # CONTROL_PLANE shared memory region
│ ├── storage/
│ │ ├── __init__.py
│ │ ├── ch_store.py # ClickHouse persistence (5 tables)
│ │ └── asset_store.py # DuckDB file-backed asset universe store
│ ├── execution/
│ │ ├── __init__.py
│ │ └── asex_integration.py # ASEx validate-before-mutate kernel
│ ├── clock/
│ │ ├── __init__.py
│ │ ├── host.py # UVClock — event-driven reactor (T19)
│ │ ├── events.py # Typed events: Scan, Tick, Timer, BarFire, Stale
│ │ ├── staleness.py # Freshness watchdog (T19 Staleness Law)
│ │ └── deadnode.py # DeadNode reaper for iox2 orphans
│ ├── training/
│ │ ├── __init__.py
│ │ ├── cma_trainer.py # CMA-ES + self-play + parallel workers
│ │ ├── registry.py # Policy lifecycle (CANDIDATE → ACTIVE)
│ │ ├── pipeline.py # Bounded continuous learning loop + logger
│ │ ├── dsl.py # Strategy DSL v2 (40+ primitives, 40+ sensors)
│ │ ├── generator.py # Genetic programming strategy evolution
│ │ ├── order_types.py # Standardized order type enum (FIX/CCXT-aligned)
│ │ ├── selector.py # Regime → strategy mapping + performance matrix
│ │ ├── asset_classification.py # Multi-label taxonomy + exchange registry
│ │ ├── asset_behavior.py # 10-dimension behavior DSL, research-validated
│ │ ├── asset_compiler.py # Auto-fetch from Binance/BingX, compile profiles
│ │ ├── asset_bridge.py # Directory ↔ classification sync
│ │ ├── parallel_eval.py # ProcessPoolExecutor episode runner
│ │ ├── ray_eval.py # Ray-based eval (industrial alternative)
│ │ ├── vbt_analysis.py # Post-sim metrics: Sharpe, Sortino, VaR
│ │ ├── advantage_scorer.py # Advantage estimation for offline analysis
│ │ ├── cognition.py # Rate-limited market regime research
│ │ ├── regime_expansion.py # 200+ regimes from dimension combinations
│ │ ├── news_sources.py # 12 industry-standard news sources
│ │ └── monitor.py # Cognition metrics, health, alerts
│ ├── daat/ # Direction-Anchored Ambiguity Triage
│ │ ├── __init__.py
│ │ └── core.py # DaatQuery, DaatVerdict, daat_classify
│ ├── cognition_launcher.py # Standalone long-run cognition service
│ ├── continuous_pipeline.py # Continuous training pipeline
│ └── tests/ # 1178 tests across 50 test files
├── specs/
│ └── MALKHUT_ADVERSARIAL_SELFPLAY_SPEC.py # full spec
└── README.md # this file
```
## Shared Memory IPC (Zinc)
All IPC uses **real Zinc shared memory** (POSIX SHM via `/dev/shm/zinc_*`), NOT
the file-based transport. This is lock-free, zero-copy, GraalVM-compatible.
### Regions
| Region | Writer | Readers | Purpose |
|--------|--------|---------|---------|
| `malkhut_book_state` | data feed | planner, CWM | canonical order book snapshot |
| `malkhut_account_state` | data feed | planner, risk | account/position snapshot |
| `malkhut_fulfilment_out` | engine | other systems | planner output + distribution |
| `malkhut_risk_gate` | engine | other systems | risk gate decisions |
| `malkhut_control` | external | engine | commands, symbols, lifecycle |
### Framing
All regions use `UVZINC01` seqlock framing:
```
[8B magic "UVZINC01"] [8B seq] [8B json_size] [UTF-8 JSON body]
```
### CONTROL_PLANE
The `malkhut_control` region is the management interface:
- External systems write `ControlCommand` frames
- Engine reads and processes commands
- ACKs written back for `ack_required` commands
Commands: `START`, `STOP`, `PAUSE`, `RESUME`, `SET_SYMBOLS`, `CONNECT_VENUE`,
`DISCONNECT_VENUE`, `SET_MODE`, `EMERGENCY_STOP`, `HOT_RELOAD_POLICY`,
`STATUS_REQUEST`
## ClickHouse Persistence
Database: `dolphin_malkhut` on `localhost:8123`
| Table | Purpose |
|-------|---------|
| `replay_steps` | CWM replay traces |
| `fulfilment_decisions` | every live decision |
| `self_play_episodes` | training episode results |
| `policy_snapshots` | versioned policy registry |
| `live_discrepancies` | simulation vs actual |
## External Dependencies (Spec-Mandated)
| Library | Purpose | Used For |
|---------|---------|----------|
| **pycma** (`cma`) | CMA-ES optimisation | `training/cma_trainer.py` |
| **hftbacktest** | replay/queue/latency fill simulation | CWM replay verification |
| **nevergrad** | gradient-free optimiser zoo | trainer comparison |
| **msgspec** | fast typed serialization | state/event encoding |
| **numba** | hot-loop acceleration | CWM transitions, feature extraction |
| **polars** | offline analytics | large trade-path corpora |
| **lmdb** | existing SILOQY storage | ticks, candles, hd_vectors |
| **numpy** | vectorized feature extraction | parameter decoding |
| **scipy** | stats tests, bootstrap | CI, promotion gates |
## Quick Start
```python
from malkhut.engine import FulfilmentEngine
from malkhut.state import FulfilmentPolicyParams
def my_params():
return FulfilmentPolicyParams(
version="v0.1", ucb_c=1.414, max_sims=64, max_depth=3,
rollout_depth=3, root_temperature=0.5, min_root_entropy=0.25,
quote_offsets_ticks=(0, 1, 2), quote_size_fractions=(0.10, 0.25, 0.50),
passive_ttl_ms=200, aggressive_ttl_ms=50,
maker_edge_min_bps=0.5, cross_spread_edge_min_bps=5.0,
adverse_toxicity_cancel_threshold=0.5, queue_churn_cancel_threshold=0.5,
mae_tail_cut_bps=50.0, mfe_giveback_cut_fraction=0.5,
max_time_in_loss_s=300.0, failed_recovery_cut_count=3,
recovery_velocity_min_bps_per_s=0.0,
max_symbol_notional_fraction=0.20, max_single_order_notional_fraction=0.05,
reduce_when_global_up_fraction=0.30, session_profit_lock_fraction=0.02,
w_expected_pnl=1.0, w_fill_probability=0.5, w_adverse_selection=2.0,
w_queue_priority=0.5, w_inventory_risk=1.5, w_tail_loss=5.0,
w_fee_quality=0.5, w_time_decay=0.3, w_policy_entropy=0.5,
robust_tail_weight=2.0, toxic_counterparty_weight=3.0,
low_liquidity_weight=2.0, latency_stress_weight=1.0,
)
engine = FulfilmentEngine(params_provider=my_params)
# engine.on_state(latest_market_state) # hot path
```
## ASEx Integration (Validate-Before-Mutate)
MALKHUT uses **ASEx** (Advanced Serial Executor) for all mutable state mutations.
This is the same kernel used by VIOLET/DOLPHIN for race-proof state management.
### Core Pattern
```python
class GuardedFulfilmentState(ASExGuardedState):
def _validate(self, mutation) -> bool:
# Pure predicate — MUST NOT mutate self
return mutation.ts_ns > 0 and len(mutation.bids) > 0
def _apply(self, mutation):
# Only called after _validate() passes
# MAY mutate self
self._state = new_state
return result
```
### Components
| Component | Purpose | Thread Safety |
|-----------|---------|---------------|
| `GuardedFulfilmentState` | Engine state mutations (book, account, intent, policy) | One writer thread via ASExWorker |
| `GuardedRiskState` | Risk gate + kill switch | One writer thread via ASExWorker |
| `FulfilmentWorker` | Single-threaded ASExWorker wrapping GuardedFulfilmentState | Serialised via queue.Queue |
| `RiskWorker` | Single-threaded ASExWorker wrapping GuardedRiskState | Serialised via queue.Queue |
| `FulfilmentWatch` | Zero-overhead ring buffer for hot-path mutations | Producer/consumer lock-free |
| `ShardedWorker` | Per-symbol state partitioning (N independent workers) | No shared state between partitions |
### Why ASEx
- **No locks**: One thread per state object. Races impossible by construction.
- **No GC pressure**: Frozen dataclasses + queue.Queue (CPython C-level ops).
- **Validate before apply**: Invalid mutations rejected without touching state.
- **GraalVM compatible**: No CPython-only APIs in the hot path.
- **Battle-tested**: 110 ASEx tests + 1109 MALKHUT tests = 1219 total.
## Design Rules (from spec)
1. **Optimise an ecology-resistant distribution, not a quote.**
2. **The book is not passive scenery.** Ask "why did I get filled?"
3. **Replay correctness before search depth.** Wrong CWM + deep search = confident nonsense.
4. **CMA-ES does not promote from noisy wins.** Survive scenario suite + bootstrap CI.
5. **Preserve feature humility.** Allow CMA-ES to discover non-obvious interactions.
6. **Path-aware SL/TP is part of fulfilment.** Exiting = order-placement under adversarial liquidity.
7. **Never let the optimiser learn illegal behavior.** Hard-code constraints.
8. **Keep hot path bounded.** ≤ 100 ms. No live CMA, no large neural nets.
9. **Version everything.** Every decision explainable by policy + CWM + feature + state + action + outcome.
10. **If shadow live diverges, trust live.** Simulation is only a tool.
---
# ADDENDUM — EXEC-STACK PLUGGABILITY + SPEC HARDENING
*[Fable, integrator, 2026-07-06 — operator-ordered extension. Mimo: this section binds
MALKHUT into the UV/DITAv2 stack; nothing above is overridden except where marked FIX.]*
## A. Where MALKHUT plugs in (the seat at the table)
MALKHUT **is the execution layer under the UV Clock** (see
`prod/docs/uv_subspecs/UV_TASK_T19_UV_CLOCK_HOST.md` + annexes):
```
UV Clock (T19) → PromotionBridge → KernelIntent (DITAv2 contracts)
→ MALKHUT FulfilmentEngine (replaces naive kernel-submit)
→ RISK GATE → BingX venue adapter → fills
```
Your Design Rule 6 ("path-aware SL/TP is part of fulfilment") is CONVERGENT with the
operator's T19 ruling ("the whole point of UV = faster-than-scan-cadence TP/SL paths").
MALKHUT's planner is the *intelligent* form of the C11 tick-exit guardian. Sequencing:
simple doctrinal tick-exits (C11) ship first via T19 step 3; MALKHUT supersedes them
**when certified** — same seat, smarter occupant.
## B. Intake contract (BINDING)
1. **Input = DITAv2 `KernelIntent`** (`prod/clean_arch/dita_v2/contracts.py`): asset,
side (TradeSide), action (ENTER/EXIT), reference_price, target_size, leverage,
metadata (carries `promo_client_id`). Accept KernelIntent natively; do not invent a
parallel intent type.
2. **`u-` clientOrderId prefix is LAW** (T9 seam): every venue order MALKHUT places
carries the u- prefix from intent metadata. Non-negotiable — it is the attribution
boundary vs BLUE and vs legacy VIOLET orders.
3. **DUAL-LEVERAGE LAW**: `KernelIntent.leverage` is CONVICTION leverage [0.5, 9.0] —
it sizes QUANTITY only. Venue leverage = the mapping in `prod/bingx/leverage.py`
(linear → ROUND_HALF_EVEN → int clamp [1,3]). WIRE the existing typed L3 wrapper
(`prod/clean_arch/violet/exchange_leverage.py`); litmus: 0.5→1, 4.75→2, 8.0→3, 9.0→3.
The venue adapter MUST NOT pass conviction leverage raw to set-leverage.
4. **Arming/two-man/kill**: MALKHUT honors `UV_PROMOTED=1` + arm-file
(`/root/uv-wt/prime-live/UV_PROMOTED.arm`), re-evaluated PER EVENT (never cached —
the T12 kill-switch lesson). DARK default = ObserveOnly venue. Map arm-file-delete
→ your `EMERGENCY_STOP` semantics: inert next event, no restart.
## C. Fabric + state alignment (the tonight-rulings)
- **Stack law: iceoryx2 = wires, ASEx = state cells, UV clock = conductor.** Your ASEx
usage (GuardedState + Worker per state object) is CORRECT — it is exactly what ASEx
is for. Your zinc regions already speak UVZINC01 seqlock framing — good; when T19's
event topics land, consume ScanEvent/TickEvent (with PROVENANCE: scan_number,
scan_ts, ingest_ts, AGE) from the shared LiveSource instead of a private feed.
- **Account state**: read the DITAv2 ASEx account core (asex_account /
AccountProjectionV2) as a CONSUMER only. MALKHUT never becomes a second writer to
account state (two-cores-share-nothing law); its own guarded states are its own.
- **Staleness law (T19 Annex A) — extend your Rule 2**: a stale book is not scenery
either, it is an ADVERSARY. Planner receives AGE on every input and must
discount/refuse plans over stale state; freshness watchdog emits StaleInput.
- **ASEx pre-condition inherited**: DeadNode reaper / orphan-sweep on start (iox2
segments orphan on crash — known gap).
## D. Persistence + observability alignment
- `dolphin_malkhut` namespace: approved (own-namespace law). **NO TTL on any table**
(retention doctrine — audit surfaces never expire).
- **Cross-reference law**: every `fulfilment_decisions` row carries `intent_id` +
`trade_id` from the source KernelIntent → end-to-end trace joins
`dolphin_uv.exec_journal` ⇄ `dolphin_malkhut.*` ⇄ venue order history. This is what
makes MALKHUT certifiable instead of a black box.
- Publish planner pulse (decision rate, plan entropy, gate verdicts, AGE-at-decision)
into the zinc snapshot → extant TUIs render it (TUI-first doctrine).
## E. Ecology + reward grounding (FIX-grade improvements)
1. **Ground the counterparty ecology in OUR tape**: fit/validate adversary populations
against `dolphin.obf_universe` recorded sub-second tape and the WT2 stress
manifold. The CLASS-M/CLASS-X hour sets (WT2D) are your ready-made adversarial
scenario suites — corrupt/stress hours ARE the hostile ecology, measured.
CLASS-X law applies: extreme magnitudes are a FEATURE channel, never filtered.
2. **Latency model = our measured tails**, not defaults: scan_to_fill/step_bar
distributions from the WT1 latency audit (median 47ms/1.58ms, stop-tail 7.8s) —
calibrate the CWM's latency injections from these empirical curves.
3. **Reconcile the budget statement (FIX)**: planner box says ≤25ms, Rule 8 says
≤100ms — bind them: ≤25ms planning inside a ≤100ms end-to-end intent→venue-call
budget, both enforced by test.
4. **Promotion gate binds to the cert conveyor**: CMA-ES candidate promotion (Rule 4)
emits artifacts via the T4 cert reporter into the gate ledger; policy snapshots
versioned in `dolphin_malkhut.policy_snapshots` AND referenced in
`prod/docs/uv_cert/`. No policy trades live without a ledger row.
## F. Certification pathway (Gate-M — how MALKHUT reaches money)
1. **CWM replay verification** (your Rule 3) vs hftbacktest — already specced; add
determinism ×2 (byte-identical minus ts).
2. **Dual-leverage litmus** (B.3 values) as a named test.
3. **DARK shadow window**: MALKHUT plans on live intents while the simple path
executes; diff planned-vs-naive on fills/slippage/adverse-selection — the
improvement claim must be MEASURED on VST before MALKHUT takes the seat.
4. **Risk-gate mutation litmus**: drop a constraint → a test goes RED (looks-finished
doctrine; 222 green tests ≠ verification until mutations bite).
5. **Rule 10 stands supreme**: shadow-vs-live divergence → trust live, stand down,
file the discrepancy row.
*End addendum. — F.*
---
## DEVELOPMENT STATUS (2026-07-13)
**1204 test functions. 50 test files. All green. 0 failures. 0 regressions.**
### Fable Spec Items — All Complete
| # | Item | Module | Status |
|---|------|--------|--------|
| 1 | Mutation-litmus test | `test_mutation_litmus.py` | ✅ PASSES |
| 2 | Re-run benchmarks at corrected fees | (long run) | ⏸️ Deferred |
| 3 | Maker fee UNVERIFIED comment | `asset_classification.py` | ✅ Done |
| 4 | ScenarioLibrary sweep | `scenario_library.py` | ✅ 7,488 grid points |
| 5 | PerformanceMatrix manifold | `selector.py` | ✅ confidence+support |
| 6 | ActualsLoader | `actuals_loader.py` | ✅ CH reader + fallback |
| 7 | OOD verdict | `risk/gate.py` | ✅ daat_verdict parameter |
| 8 | Manifold query | `manifold_query.py` | ✅ three-phase recommend |
| 9 | DAAT package | `daat/core.py` | ✅ DaatVerdict{KNOWN,MARGINAL,OOD} |
| 10 | Book fidelity | `book_fidelity.py` | ✅ OBF → OrderBookState |
### Fee Correction (CRITICAL — all prior training was at 10x too-cheap fees)
Source of truth: `dolphin.trade_execution_quality` (8006 rows, avg taker=5.016 bps).
| Venue | Old taker (WRONG) | Correct taker | Old maker (WRONG) | Correct maker |
|-------|-------------------|---------------|-------------------|---------------|
| BingX | 0.5 bps | **5.0 bps** | -0.2 (rebate) | **+2.0 (POSITIVE)** |
| Binance | 0.4 bps | **4.5 bps** | -0.2 | -0.2 (UNVERIFIED) |
| Bybit | 0.06 bps | **5.5 bps** | -0.01 | -0.1 (UNVERIFIED) |
**Every policy trained before this fix is suspect** — cheap fees reward overtrading.
Re-measurement at correct fees is required for production deployment.
### Completed subsystems
| Subsystem | Module | Tests | Status |
|-----------|--------|-------|--------|
| **State Model** | `state.py` | 17 | 42 frozen dataclasses, immutable |
| **CWM** | `cwm/core.py` + `cwm/numba_core.py` | 103 | Exchange mechanics + numba JIT (5.3µs/transition) + vectorized reward + vectorized UCB + batch MCTS kernel |
| **Replay Verification** | `cwm/replay_verify.py` | 65 | Deep comparison, binary search, trajectory recording |
| **Planner** | `planner/sm_mcts.py` | 11 | Decoupled UCB/UCT, ≤25ms budget |
| **Action Menu** | `planner/action_menu.py` | (in planner) | Compact action space construction |
| **Counterparties** | `counterparties.py` | 19 | 4 adversarial agent policies |
| **Risk Gate** | `risk/gate.py` | 4 | Hard invariants: leverage, self-trade, post-only |
| **ASEx Integration** | `execution/asex_integration.py` | 33 | Validate-before-mutate, single-writer, no locks |
| **Zinc IPC** | `ipc/zinc_plane.py` | 8 | Real POSIX SHM, UVZINC01 framing |
| **Control Plane** | `ipc/control_plane.py` | 3 | HOT_RELOAD_POLICY, kill switch |
| **ClickHouse** | `storage/ch_store.py` | 9 | 5 tables, HTTP API |
| **DuckDB Asset Store** | `storage/asset_store.py` | 30 | File-backed + in-memory materialization, sub-µs reads |
| **Asset Bridge** | `training/asset_bridge.py` | 49 | Directory ↔ classification sync |
| **CMA-ES Training** | `training/cma_trainer.py` | 65 | Behavior-driven, auto-compile, parallel workers, 7x speedup |
| **Parallel Eval** | `training/parallel_eval.py` | 16 | ProcessPoolExecutor, 7x CMA-ES speedup |
| **Ray Eval** | `training/ray_eval.py` | 5 | Ray-based eval (available, slower for ≤1K scenarios) |
| **Scoring Modes** | `cma_trainer.py` + `advantage_scorer.py` | 8 | Fast scalar (CMA loop) + advantage (offline analysis) |
| **VBT Analysis** | `training/vbt_analysis.py` | 8 | Post-sim trade metrics: Sharpe, Sortino, VaR, cross-asset |
| **ScenarioLibrary** | `training/scenario_library.py` | 7 | Sweep 7,488 grid points (spread×depth×tox×regime) |
| **Manifold Query** | `training/manifold_query.py` | (new) | Three-phase: DAAT → manifold → recommend |
| **ActualsLoader** | `training/actuals_loader.py` | (new) | Reads CH tables for Mode 2 live query |
| **Book Fidelity** | `training/book_fidelity.py` | (new) | OBF 15B rows → OrderBookState synthesis |
| **DAAT** | `daat/core.py` | 9 | Direction-Anchored Ambiguity Triage |
| **OOD Verdict** | `risk/gate.py` | +1 | daat_verdict parameter, OOD → doctrinal fallback |
| **Policy Registry** | `training/registry.py` | 14 | CANDIDATE → ACTIVE lifecycle |
| **Training Pipeline** | `training/pipeline.py` | 21 | Bounded continuous learning loop |
| **Strategy DSL v2** | `training/dsl.py` | 69 | 40+ primitives, 40+ sensors, 16 builtins |
| **Order Types** | `training/order_types.py` | (new) | Three-dimensional: OrderType × TimeInForce × Instructions (FIX-aligned) |
| **Strategy Generator** | `training/generator.py` | 20 | Genetic programming: crossover, mutation, tournament |
| **Strategy Selector** | `training/selector.py` | 24 | Regime × strategy × venue matrix, cross-exchange comparison |
| **Asset Classification** | `training/asset_classification.py` | 190 | Multi-label taxonomy + exchange registry, 13 assets, system-wide store |
| **Asset Behavior DSL** | `training/asset_behavior.py` | (in classification) | 10 orthogonal dimensions, 3 templates, 13 behaviors, research-validated |
| **Asset Compiler** | `training/asset_compiler.py` | (new) | Auto-fetch from Binance/BingX API, compile profiles, rate-limited |
| **Cognition Pipeline** | `training/cognition.py` | 29 | Rate-limited, 8 sources, dedup, perm-run |
| **Regime Expansion** | `training/regime_expansion.py` | 29 | 200+ regimes from 4×4×4×4 dimensions |
| **News Sources** | `training/news_sources.py` | 12 | 12 industry-standard sources, ranking |
| **Cognition Monitor** | `training/monitor.py` | 9 | Metrics, health scoring, alerts, JSONL logging |
| **Cognition Launcher** | `cognition_launcher.py` | 21 | Standalone long-run cognition service |
| **UV Clock** | `clock/host.py` | 30 | T19 event-driven reactor |
| **Numba Acceleration** | `cwm/numba_core.py` | 13 | JIT on hot loops, 1.8x fill speedup |
| **BingX Adapter** | `venue/bingx/adapter.py` | 28 | Wraps DITAv2, order tracking |
| **Hypothesis** | `test_hypothesis_properties.py` | 6 | Property-based fuzzing |
| **Fuzz/Adversarial** | `test_fuzz.py`, `test_adversarial.py` | 17 | Random stress + adversarial scenarios |
| **Concurrency** | `test_concurrency.py` | 8 | Thread safety, race detection |
| **Sync/Async Seams** | `test_sync_async_seams.py` | 7 | Zinc latency, engine budget |
| **E2E Integration** | `test_e2e_integration.py` | 2 | Full pipeline: train→register→plan→risk→venue→zinc→ch→reload |
| **HftBacktestCWM** | `cwm/hft_cwm.py` | (new) | Queue-model fills (PowerProb), fill quality tracking, drop-in CWM |
### Fill Quality — The Core Optimization Target
MALKHUT is an **execution improvement engine**. Fill quality IS the primary aim — not PnL, not Sharpe, not win rate. The system learns to get **better fills**: faster, better-priced, with less adverse selection.
#### Fill Quality Metrics (per CWM transition)
| Metric | Source | What it measures |
|--------|--------|------------------|
| `slippage_bps` | `abs(fill_price - mid) / mid * 10000` | How far from mid did we fill? (aggressive) |
| `price_improvement_bps` | `(best_bid - fill_price) / best_bid * 10000` | How much better than touch? (passive) |
| `levels_consumed` | `fill_qty / avg_level_qty` | Queue depth of fill |
| `is_maker_fill` | `order_type==LIMIT or post_only` | Passive vs aggressive |
| `rolling_fill_rate` | EMA(0.8, 0.2) over recent fills | Recent fill success rate |
| `post_fill_adverse_bps` | `(new_mid - old_mid) / old_mid * 10000` | Price movement after fill |
| `fill_value_score` | `quality - abs(adverse) * 0.5` | **Composite optimization metric** |
#### Fill Quality in the Reward Function
```
reward = w_fill_probability * fill_value_score ← PRIMARY (fill quality)
+ w_expected_pnl * pnl ← secondary (PnL)
- w_adverse_selection * toxicity
- w_inventory_risk * inventory_risk
- w_tail_loss * tail_risk
- w_time_decay * time_in_loss
+ w_fee_quality * fee_savings ← maker saves (taker-maker) bps
- w_fee_quality * taker_fee ← taker pays full fee
- w_fee_quality * markout_cost * 0.3 ← markout = honest execution cost
```
**Fee model (BingX, no rebates):**
- Maker fee: 2.0 bps (you pay)
- Taker fee: 5.0 bps (you pay)
- Fee savings: 3.0 bps (maker saves 3bp vs taker)
**Markout = quality:** slippage_bps + post_fill_adverse_bps = honest execution cost.
The system learns: pay the friction (fee + slippage) when urgency × (fee + slippage) < threshold.
#### Fill Quality in the PerformanceMatrix
`RegimeStrategyScore` now tracks 4 fill quality metrics per (regime, strategy, venue):
- `avg_fill_rate`: rolling fill success rate
- `avg_slippage_bps`: average slippage for aggressive fills
- `avg_price_improvement_bps`: average improvement for passive fills
- `avg_fill_value_score`: composite fill quality metric
This enables: "Which strategy achieves the best fill quality in regime X on venue Y?"
#### Fill Quality in CMA-ES
The CMA-ES optimizer now receives fill quality metrics in each `EpisodeResult`:
- `avg_fill_value_score`: average fill value across the episode
- `avg_price_improvement_bps`: average price improvement
- `avg_post_fill_adverse_bps`: average adverse selection
The CMA-ES objective is: maximize fill quality (primary) while maintaining positive PnL.
#### Urgency-Driven Maker/Taker Decision
The system learns WHEN to use maker (passive) vs taker (aggressive) as a function of urgency:
| Urgency | Behavior | Penalty |
|---------|----------|---------|
| 0.0 – 0.3 | Passive only (maker). No taker actions generated. | — |
| 0.3 – 0.65 | IOC partial taker allowed. Small sizes, partial fills. | Low |
| 0.65 – 1.0 | Full taker allowed. Aggressive crossing at full size. | None |
Two CMA-ES optimizable parameters:
- `urgency_taker_threshold` (default 0.65): crosses from passive to aggressive
- `urgency_taker_penalty_bps` (default 2.0): penalty for taker at low urgency
The system learns: "For ADAUSDT in stress regime, threshold=0.4 (thin book → try maker,
but resort to taker when liquidity appears). For BTCUSDT in normal regime, threshold=0.8
(deep book → maker always works)."
#### Calibrated Slippage (Flight7 VST + Mainnet Anchors)
Slippage model per-asset, configurable, overridable per-run:
| Book Type | Model | Assets |
|-----------|-------|--------|
| **Deep book** | `slippage = alpha * levels + beta * depth_ratio` | BTC, ETH, BNB |
| **Thin book** | `slippage = intercept + adverse_selection` | DOGE, ADA, UNI, AAVE, ATOM |
Switch: `if book_depth_usd < thin_book_threshold_usd → thin mode`
Fable's Flight7 calibration anchors (testnet + mainnet):
- Majors: BTC ~0.1 bps, ETH ~0.4 bps (testnet ≈ mainnet)
- Liquid alts: 3-6 bps testnet, 10-30 bps mainnet (3-6× gap)
- Thin alts: 14-25 bps testnet, ~140 bps round-trip mainnet
#### Chase Mechanics
Cancel → wait → retry with configurable parameters (CMA-ES optimizable):
- `wait_to_retry_ms`: delay before re-quoting after cancel (0-2000ms)
- `chase_enabled`: enable chase-follow behavior
- `chase_offset_ticks`: ticks from target price to chase (0-10)
- `chase_max_retries`: max cancel-retry cycles (0-5)
CHASE in the DSL places a limit order at target offset with TTL=wait_to_retry_ms.
If not filled, the order expires and next step re-places at a new offset.
### Performance Benchmarks
| Metric | Value |
|--------|-------|
| CWM throughput | 189K calls/sec (numba JIT) |
| CWM per-call latency | 5.3 µs |
| CWM reward (numba vectorized) | ~0.3µs |
| UCB selection (numba vectorized) | 1.13µs (was 8.7µs, 7.7x speedup) |
| Numba fill speedup | 1.8x (batch 100) |
| DuckDB asset reads | 0.2µs (in-memory) |
| Scenario generation | 390 scenarios in 0.8s |
| CMA-ES parallel (8 workers) | 7× speedup, 87% efficiency |
| Best CMA-ES score (48 evals) | 8,628 (fast scalar, 100.4 bps PnL) |
| Score/min (8 workers) | 3,043 |
| Episode throughput (parallel) | 16ms/ep |
| Episode throughput (sequential) | 23ms/ep |
### Bugs found and fixed (22 total)
1. Empty book crash in `materialize_price_from_action` (SELL side)
2. Empty book crash in `total_notional` calculation
3. Empty book crash in mark-to-market
4. Empty book crash in `_inventory_risk`
5. Empty book crash in counterparty processing
6. Empty book crash in `_update_path_state`
7. Empty book crash in feature extractor (`b.mid`)
8. Counterparty fills used `max()` instead of consuming levels
9. Position overwritten instead of accumulated on add
10. Test helper `_state()` not forwarding `positions` kwarg
11. `_make_open_order` set empty symbol instead of venue symbol
12. `FulfilmentPolicyParams` frozen dataclass `__dict__` access
13. Banker's rounding in tick tests
14. `HOT_RELOAD_POLICY` compared float to string
15. Engine init order (workers before hot_reload)
16. `_tp()` helper missing `failed_recovery_count` parameter
17. CWM: SELL position flip didn't reset avg_entry
18. CWM: Counterparty fills overwrote our fill tracking
19. CWM: Double-counting unrealized PnL in equity
20. UnboundLocalError when max_steps=0
21. fill_count never incremented in training
22. Cancel rate limiter incorrectly gated order placements
### Key design rules enforced
1. ✅ Optimise distribution, not single quote
2. ✅ Book is not passive scenery
3. ✅ **Replay correctness before search depth** (mandatory gate)
4. ✅ CMA-ES does not promote from noisy wins (bootstrap CI)
5. ✅ Preserve feature humility (CMA discovers interactions)
6. ✅ Path-aware SL/TP is part of fulfilment
7. ✅ Never let optimiser learn illegal behavior
8. ✅ Hot path bounded (≤100ms)
9. ✅ Version everything
10. ✅ Shadow-vs-live divergence → trust live
### Strategy Selection (completed)
The system classifies strategies per market condition and floats the best to top:
```
market_fingerprint → Regime Classifier (8 regimes)
↓
Performance Matrix (regime × strategy → score)
↓
Strategy Selector → Engine uses best strategy
```
### Upstream integration path
- **ExoF** — funding rates, open interest, liquidation levels
- **MARAS** — 5-tier ensemble → 8-regime classification
- **DOLPHINNG7 vel_div** — velocity divergence signals
- **DOLPHINNG7 eigenscan** — eigenspace correlation analysis
---
## G. Strategy DSL v2 — Third-Party Strategy Development
MALKHUT includes a **Domain-Specific Language (DSL)** for strategy composition.
Third parties can write, evolve, and share strategies without modifying core code.
### DSL Format
```
STRATEGY "passive_maker" {
PRIORITY 1: IF spread_bps < 5.0 AND orderflow_toxicity < 0.3 THEN QUOTE(BUY, 0, 0.25)
PRIORITY 2: IF spread_bps < 5.0 AND orderflow_toxicity < 0.3 THEN QUOTE(SELL, 0, 0.25)
PRIORITY 3: IF orderflow_toxicity > 0.7 THEN CANCEL_ALL
PRIORITY 4: IF time_in_trade > 300 THEN EXIT
PRIORITY 5: IF unrealized_pnl < -50bps THEN STOP_LOSS
PRIORITY 6: NOOP
}
```
### Action Primitives (40+)
| Category | Primitives |
|----------|-----------|
| Passive | QUOTE, JOIN_QUEUE, STEP_BACK, LADDER, GRID, ICEBERG, TWAP |
| Aggressive | CROSS, SNIPER, PING |
| Cancellation | CANCEL, CANCEL_ALL, CANCEL_AND_HOLD, REQUOTE |
| Position | EXIT, HALF_EXIT, QUARTER_EXIT, TAKE_PARTIAL, STOP_LOSS, TAKE_PROFIT, TRAILING_STOP, EMERGENCY_EXIT, FLAT_ALL |
| Sizing | SCALE_IN, SCALE_OUT, INCREASE_SIZE, REDUCE_SIZE |
| Stop/Target | MOVE_STOP, MOVE_TAKE_PROFIT, BRACKET, OCO |
| Hedging | HEDGE, PAIR_TRADE |
| Waiting | HOLD, WAIT_FOR_FILL, WAIT_FOR_PRICE, WAIT_FOR_SPREAD |
### Market Sensors (40+)
| Category | Sensors |
|----------|---------|
| Book | spread_bps, imbalance_3/5/10, bid_depth_3/5/10, bid_ask_ratio |
| Flow | toxicity, queue_churn, cross_venue_lead, order_book_toxicity |
| Momentum | price_momentum_1s/5s/15s/1m |
| Position | position_qty, pnl_bps, leverage, position_age_s |
| Path | mae_bps, mfe_bps, time_in_loss, failed_recoveries, recovery_velocity |
| Regime | volatility, atr_14/50, rsi_14, funding, regime_score |
| Account | equity, risk_budget_used, session_pnl, drawdown |
| Performance | profit_factor, sharpe_ratio, consecutive_wins/losses |
| Time | current_hour, day_of_week, is_weekend, is_liquid_hours |
### Comparison Operators (12)
`>`, `<`, `>=`, `<=`, `==`, `!=`, `abs>`, `abs<`, `changing`, `stable`, `crossing_above`, `crossing_below`
### Builtin Strategies (16)
| Strategy | Description |
|----------|-------------|
| `passive_maker` | Passive limit orders, toxicity avoidance |
| `aggressive_taker` | Momentum-based aggressive crosses |
| `toxicity_avoider` | Cancels on high toxicity, quotes when safe |
| `path_risk_exit` | Path-aware SL/TP exits |
| `regime_adaptive` | Adapts to MARAS regime classification |
| `momentum_catcher` | Catches momentum moves, partial exits |
| `mean_reversion` | Wide-spread mean reversion |
| `scalper` | Ultra-tight spreads, small sizes |
| `inventory_manager` | Position size control |
| `session_guard` | Time-based position management |
| `liquidity_hunter` | Deep book exploitation |
| `volatility_breakout` | Volatility-based breakouts |
| `funding_arb` | Funding rate arbitrage |
| `grid_trader` | Grid-based accumulation |
| `risk_parity` | Risk budget management |
| `hybrid_adaptive` | Multi-condition adaptive |
### Strategy Generator (Genetic Programming)
The system evolves strategies via genetic operators:
```python
from malkhut.training.generator import StrategyGenerator, GeneratorConfig
config = GeneratorConfig(
population_size=20, generations=5, tournament_size=3,
mutation_rate=0.15, crossover_rate=0.7,
)
generator = StrategyGenerator(config=config)
population = generator.evolve(baseline_params, scenarios)
# Get successful strategies
successful = generator.get_successful_strategies(min_fitness=0.0)
for genome in successful:
generator.add_to_pool(genome)
```
### Strategy Selector (Market-Adaptive)
The system classifies strategies per market regime and selects the best:
```python
from malkhut.training.selector import StrategySelector, MarketRegime
selector = StrategySelector()
result = selector.select(current_state, available_strategies)
# result.strategy_id → best strategy for current regime
# result.confidence → how confident we are
# result.alternatives → other considered strategies
```
Regimes: `trending_up`, `trending_down`, `high_volatility`, `low_volatility`,
`mean_reverting`, `momentum`, `choppy`, `liquidity_hole`, `normal`
### Asset Classification (Invariant Taxonomy)
Multi-label asset classification using **only invariant characteristics** — properties
intrinsic to the token that predict price behaviour and order-book performance.
**Three-layer identifier architecture** for cross-system interoperability.
**Design principle:** Classify by WHAT an asset IS (fundamental), not how it's trading
right now. Volatility/liquidity appear as long-run statistical profiles, not as
classification axes that change minute-to-minute.
#### Three-Layer Identifier Architecture
| Layer | Fields | Purpose | Example |
|-------|--------|---------|---------|
| **1. Canonical** | `symbol`, `base_asset`, `name`, `unified_symbol`, `quote_currency` | Venue-independent identity | BTCUSDT / BTC / "Bitcoin" / BTC/USDT / USDT |
| **2. Cross-system** | `coingecko_id`, `cmc_id`, `blockchain`, `contract_address` | Data provider + on-chain mapping | "bitcoin" / 1 / "ethereum" / "" |
| **3. Exchange mapping** | `exchanges: tuple[str, ...]` | Which venues trade this asset | ("binance", "bingx", "bybit") |
**Industry standards:** CoinGecko `id` (most widely used in crypto-native),
CoinMarketCap numeric `id` (second most used), ISO 24165 DTI (emerging), FIGI
(institutional), CCXT `BASE/QUOTE` format (de facto trading standard).
#### Fundamental Dimensions (intrinsic, never change)
| Dimension | Values | Multi-label | Purpose |
|-----------|--------|-------------|---------|
| **Sector** | CURRENCY, LAYER1, LAYER2, DEFI, ORACLE, EXCHANGE, MEME, PRIVACY, STORAGE, GAMING_NFT | Yes | Primary use-case / vertical |
| **TokenRole** | GAS, STORE_OF_VALUE, GOVERNANCE, UTILITY, MEME, EXCHANGE_FEE | Yes | Functional role → demand elasticity |
| **SupplyModel** | FIXED_CAP, DISINFLATIONARY, INFLATIONARY, BURN_MECHANISM | No | Supply pressure dynamics |
| **ConsensusFamily** | POW, POS, DPOS | No | Miner/validator selling behaviour |
| **SmartContractCapability** | FULL, PARTIAL, NONE | No | DeFi composability |
#### Technical Dimensions (invariant market-structure properties)
| Dimension | Values | Purpose |
|-----------|--------|---------|
| **MarketCapTier** | MEGA (>500B), LARGE (50-500B), MID (5-50B), SMALL (500M-5B), MICRO (<500M) | Price impact per dollar |
| **VolatilityProfile** | LOW (<30%), MEDIUM (30-80%), HIGH (80-150%), EXTREME (>150%) | Long-run average vol |
| **LiquidityProfile** | DEEP (>100M), NORMAL (10-100M), THIN (1-10M), ILLIQUID (<1M) | Long-run avg depth |
| **DerivativeAccess** | PERPS_AND_OPTIONS, PERPS_ONLY, NONE | Shorting / funding availability |
| **tick_size / lot_size** | Exchange-set | Minimum price/order size |
| **maker/taker fee_bps** | Exchange-set | Trading cost |
| **typical_spread_bps / depth / volume** | Long-run averages | Order-book fingerprint |
#### Predictive Properties (derived from invariants)
| Property | Logic | Value |
|----------|-------|-------|
| `supply_pressure` | PoW → forced selling (miners cover electricity) | "forced" / "optional" |
| `demand_elasticity` | GAS or STORE_OF_VALUE → inelastic (must hold) | "inelastic" / "elastic" |
| `can_be_shorted` | DerivativeAccess != NONE | bool |
#### Multi-Label Examples
```
BTC: sectors=(CURRENCY), roles=(STORE_OF_VALUE)
ETH: sectors=(LAYER1, DEFI), roles=(GAS, STORE_OF_VALUE, GOVERNANCE)
BNB: sectors=(EXCHANGE, LAYER1), roles=(EXCHANGE_FEE, GAS)
DOGE: sectors=(MEME, CURRENCY), roles=(MEME, GAS)
DOT: sectors=(LAYER1), roles=(GAS, GOVERNANCE)
UNI: sectors=(DEFI), roles=(GOVERNANCE, UTILITY)
```
#### Querying
```python
from malkhut.training.asset_classification import *
# Multi-label queries (match if queried value is ANY of the asset's labels)
layer1s = get_assets_by_sector(Sector.LAYER1) # ETH, SOL, ADA, AVAX, BNB, DOT, ATOM
gas_tokens = get_assets_by_token_role(TokenRole.GAS) # ETH, SOL, ADA, AVAX, DOGE, BNB, MATIC, DOT, ATOM
# Predictive properties
forced = get_pov_assets() # BTC, DOGE (PoW miners with forced selling)
inelastic = [p for p in ASSET_PROFILES.values() if p.demand_elasticity == "inelastic"]
# Backward-compatible single-label access (first element = primary)
btc = get_asset_profile("BTCUSDT")
btc.sector # Sector.CURRENCY
btc.token_role # TokenRole.STORE_OF_VALUE
```
### Exchange Registry (Multi-Exchange Asset Universe)
MALKHUT serves as the **system-wide asset universe store** for BLUE, VIOLET, UV, and all
downstream systems. Each asset knows which exchanges trade it; each exchange has a
standardized profile.
#### ExchangeProfile
| Field | Type | Purpose |
|-------|------|---------|
| `exchange_id` | str | Canonical key: `"binance"`, `"bingx"`, `"bybit"` |
| `display_name` | str | Human-readable name |
| `has_spot` | bool | Spot trading available |
| `has_perps` | bool | Perpetual futures available |
| `has_options` | bool | Options available |
| `api_base_url` | str | REST API root URL |
| `ws_base_url` | str | WebSocket root URL |
| `default_taker_fee_bps` | float | Default taker fee |
| `default_maker_fee_bps` | float | Default maker fee |
| `typical_latency_ms` | float | Typical API latency |
#### Pre-defined Exchanges
| Exchange | Spot | Perps | Options | Taker Fee | Latency |
|----------|------|-------|---------|-----------|---------|
| **Binance** | ✓ | ✓ | ✓ | 0.4 bps | 40 ms |
| **BingX** | ✓ | ✓ | ✗ | 0.5 bps | 100 ms |
| **Bybit** | ✓ | ✓ | ✓ | 0.06 bps | 50 ms |
#### Exchange-Asset Mapping
Each `AssetProfile` carries an `exchanges: tuple[str, ...]` field listing which venues
trade the asset. Default: `("binance",)`. Other systems (BLUE/VIOLET/UV) import their
asset universes into this store.
```python
from malkhut.training.asset_classification import *
# Which exchanges trade BTC?
btc = get_asset_profile("BTCUSDT")
print(btc.exchanges) # ('binance',)
# All assets on Binance
binance_assets = get_assets_on_exchange("binance")
# Assets traded on BOTH Binance and BingX
common = get_common_assets("binance", "bingx")
# Which exchanges trade a given asset?
venues = get_exchange_for_asset("ETHUSDT")
# Exchange metadata
ex = get_exchange("binance")
print(ex.default_taker_fee_bps) # 0.4
print(ex.typical_latency_ms) # 40
```
### Asset Behavior DSL (10-Dimension Research-Validated Model)
The Asset Behavior DSL decomposes each asset's **trading behavior** into 10 orthogonal
dimensions, each with empirically-validated parameters from live Binance/BingX API data
and academic literature.
**Source:** Binance live REST API (depth 500 levels, 1h klines, ticker, funding rates),
BingX open API (depth, funding, premium index), academic literature (Bouchaud "Trades
Quotes and Prices" 2018, Cont/Stoikov/Talreja 2010, Cartea/Jaimungal/Penalva 2015),
CoinGlass OI data, industry MM disclosures.
#### 10 Dimensions
| Dimension | What it captures | Example: BTC vs DOGE |
|-----------|------------------|---------------------|
| **DepthProfile** | Book shape: amplitude, decay exponent α, fragility | BTC: $750K/0.70/0.10 vs DOGE: $22K/1.00/0.30 |
| **SpreadProfile** | Normal spread, stress multiplier | BTC: 0.01bps/50x vs DOGE: 1.35bps/15x |
| **FlowProfile** | Order rate, size distribution, cancel ratio | BTC: 300/s/$643/20x vs DOGE: 80/s/$96/8x |
| **VolatilityProfile** | Annualized vol, GARCH params, half-life | BTC: 35%/48h vs DOGE: 78%/24h |
| **IntradayProfile** | Peak/trough hours, ratio | BTC: 15:00 UTC, 7.4x ratio |
| **WeekendProfile** | Vol/volume/spread multipliers | 0.65x vol, 0.52x volume, 1.20x spread |
| **CorrelationProfile** | ETH beta, BTC corr (normal vs crash) | BTC: 1.0/1.0 vs UNI: 0.70/0.50 |
| **MarketMakerProfile** | Inventory limits, pull speed, margins | BTC: $10M/3ms/0.5bps vs DOGE: $500K/25ms/2bps |
| **LiquidationProfile** | OI/MCap, trigger %, cascade speed | BTC: 0.5%/6.5%/slow vs DOGE: 1.4%/4%/fast |
| **FundingProfile** | Mean/std funding, basis | BTC: 0.59bps/0.22bps vs DOGE: 0.39bps/0.34bps |
| **RetailProfile** | Retail ratio, inst gap | BTC: 35%/0.04 vs DOGE: 80%/0.31 |
| **BingxProfile** | BingX spread/depth multiplier, fees, latency | BTC: 12.6x/0.048x vs DOGE: 7x/1.27x |
#### 3 Composable Templates
| Template | Assets | Key traits |
|----------|--------|-----------|
| `institutional_blue_chip` | BTC, ETH, BNB | Deep book, low spread, slow cascade, high institutional |
| `mid_cap_l1` | SOL, ADA, AVAX, DOT, ATOM, UNI, LINK, MATIC, AAVE | Moderate depth, higher vol, medium cascade |
| `retail_meme` | DOGE | Thin book, high vol, fast cascade, retail-dominated |
#### 13 Per-Asset Behaviors (research-validated)
| Asset | Price | Depth@1bps | Spread | Vol | Cascade | Retail | BingX mult |
|-------|-------|-----------|--------|-----|---------|--------|------------|
| BTC | $64K | $750K | 0.01 bps | 35% | slow/fast | 35% | 12.6x |
| ETH | $1.8K | $600K | 0.01 bps | 66.5% | medium/medium | 35% | 9.0x |
| SOL | $80 | $400K | 1.26 bps | 72% | medium/medium | 72% | 1.7x |
| DOGE | $0.07 | $22K | 1.35 bps | 78% | fast/slow | 80% | 7.0x |
| ADA | $0.17 | $60K | 5.95 bps | 79% | medium/medium | 65% | 8.0x |
| UNI | $3.6 | $30K | 2.76 bps | 95% | medium/medium | 60% | 10x |
#### Usage
```python
from malkhut.training.asset_behavior import *
# Get behavior for any asset
btc = get_behavior("BTCUSDT")
print(btc.depth.amplitude_usd) # $750,000
print(btc.vol.annualized_normal) # 35.0%
print(btc.liquidation.speed) # "slow"
# Compute depth at distance
depth_10bps = btc.depth_at_bps(10) # $1.5M
# Estimate slippage for a $100K order
slippage = btc.expected_slippage_bps(100_000) # ~2 bps
# Query by template, sector, or role
institutional = get_behaviors_by_template("institutional_blue_chip")
gas_tokens = get_behaviors_by_role(TokenRole.GAS)
thin_book = get_thin_book_assets()
```
### Asset Compiler (Auto-Fetch from Live APIs)
The Asset Compiler automatically fetches market data from Binance and BingX public APIs
and produces system-ready `AssetProfile` + `AssetBehavior` for any tradeable symbol.
**Rate-limited** (1 req/sec, well under Binance 1200/min limit), **cached** (avoids
re-fetching within a session), **resumable**.
#### What it auto-fetches
| Endpoint | Data | Computes |
|----------|------|----------|
| `/api/v3/ticker/24hr` | Price, volume, high/low | Daily volume, price reference |
| `/api/v3/depth?limit=100` | Order book (100 levels) | Spread, depth amplitude, decay exponent α |
| `/api/v3/klines?interval=1h&limit=168` | 7 days hourly OHLCV | Annualized volatility, order flow stats |
| `/api/v3/exchangeInfo` | Symbol filters | Tick size, lot size, price decimals |
| `/fapi/v1/fundingRate` | Funding rates (20 intervals) | Mean/std/positive% funding |
| `/fapi/v1/openInterest` | Open interest | OI/MCap ratio for liquidation modeling |
#### Usage
```python
from malkhut.training.asset_compiler import AssetCompiler
compiler = AssetCompiler()
# Single asset — auto-fetches from Binance, computes all params
result = compiler.compile("XRPUSDT")
print(result.asset_profile.symbol) # "XRPUSDT"
print(result.asset_behavior.reference_price) # $1.1056 (live Binance price)
print(result.asset_behavior.spread.normal_bps) # 0.90 bps (computed from depth)
print(result.asset_behavior.vol.annualized_normal) # 46.6% (computed from klines)
# Register into global registries
compiler.register(result)
# Now XRPUSDT is available for ScenarioFactory, queries, etc.
# One-step compile and register
result = compiler.compile_and_register("HBARUSDT")
# Batch compile (rate-limited sequentially)
results = compiler.compile_batch(["APTUSDT", "SEIUSDT", "INJUSDT"])
for r in results:
compiler.register(r)
```
#### Known classifications
The compiler includes pre-built classifications for 28 assets:
BTC, ETH, SOL, DOGE, ADA, AVAX, UNI, LINK, BNB, MATIC, AAVE, DOT, ATOM,
XRP, TRX, LTC, NEAR, APT, OP, ARB, SUI, PEPE, WIF, SEI, INJ, FIL, HBAR, IMX
Unknown assets get heuristic defaults (LAYER1/GAS/INFLATIONARY/POS/FULL/PERPS_ONLY).
### Scenario Factory (Behavior-Driven, Label-Queryable)
The ScenarioFactory builds evaluation scenarios using **real per-asset prices,
depth profiles, and spread characteristics** — no more hardcoded BTC prices.
#### Behavior-driven state creation
Every scenario derives its parameters from the asset's `AssetBehavior`:
| Scenario type | Spread multiplier | Depth fraction | Description |
|---------------|------------------|----------------|-------------|
| `normal` | 1.0x | 100% | Calm, liquid market |
| `thin_book` | 1.5x | 10% | Low participation |
| `wide_spread` | 100x | 50% | High volatility |
| `toxic_stress` | 2.0x | 30% | Adverse selection |
| `flash_crash` | 3.0x | 5% | Book collapse |
| `liquidity_vacuum` | 5.0x | 1% | Near-empty book |
| `weekend_thin` | 5.0x | 5% | Weekend low-volume |
| `trending` | 1.0x | 80% | Momentum move |
| `mean_reverting` | 20x | 60% | Reversal setup |
| `funding_shock` | 1.5x | 30% | Deleveraging event |
| `liquidation_cascade` | 1.5x | 20% | Chain liquidation |
| `market_maker_withdrawal` | 3.0x | 15% | MM pull-out |
#### Realistic per-asset pricing
```
Scenario: "normal_BTCUSDT_42"
BTC bid: $63,999.97 (11.7 BTC = $750K)
BTC ask: $64,000.03
Spread: 0.01 bps
Scenario: "normal_DOGEUSDT_42"
DOGE bid: $0.07 (314K DOGE = $22K)
DOGE ask: $0.07
Spread: 1.35 bps
Scenario: "flash_crash_SOLUSDT_47"
SOL bid: $80.00 (depth = 5% of normal = $20K)
SOL ask: $80.02
Spread: 3.78 bps (3x normal)
```
#### Auto-compilation for unknown assets
When you request a symbol that isn't pre-defined, the factory **automatically compiles
it from live Binance API data** before building scenarios:
```python
factory = ScenarioFactory()
# XRPUSDT is NOT in our pre-defined 13 assets
# But this auto-compiles it from Binance in ~5 seconds:
suite = factory.build_suite_for_symbols(("XRPUSDT",), steps_per_scenario=5)
# XRP behavior is now registered for future use too
```
#### Label-based query interfaces
```python
factory = ScenarioFactory()
# By sector
suite = factory.build_suite_for_sector(Sector.LAYER1) # ETH, SOL, ADA, AVAX, BNB, DOT, ATOM
suite = factory.build_suite_for_sector(Sector.DEFI) # ETH, AVAX, UNI, AAVE
suite = factory.build_suite_for_sector(Sector.MEME) # DOGE
# By token role
suite = factory.build_suite_for_role(TokenRole.GAS) # 9 assets
suite = factory.build_suite_for_role(TokenRole.GOVERNANCE) # ETH, UNI, AAVE, DOT
# By behavior template
suite = factory.build_suite_for_template("institutional_blue_chip") # BTC, ETH, BNB
suite = factory.build_suite_for_template("retail_meme") # DOGE
# By volatility band
suite = factory.build_suite_for_volatility(min_ann=75, max_ann=200) # High-vol assets
# Multi-label union (any matching)
suite = factory.build_suite_for_labels(
sectors=[Sector.DEFI],
roles=[TokenRole.GOVERNANCE]
)
# Arbitrary symbols (auto-compiles unknowns)
suite = factory.build_suite_for_symbols(("BTCUSDT", "XRPUSDT", "HBARUSDT"))
```
#### Research-validated scenario diversity
Each asset's 30 scenario types are parameterized by its behavior profile:
| Scenario | BTC params | DOGE params | Why it matters |
|----------|-----------|-------------|----------------|
| Normal | $750K depth, 0.01bps spread | $22K depth, 1.35bps spread | Baseline for scoring |
| Thin book | $75K depth | $2.2K depth | Tests fill rate under low liquidity |
| Flash crash | $37.5K depth, 0.03bps spread | $1.1K depth, 4bps spread | Tests adverse selection survival |
| Weekend thin | $37.5K depth, 5bps spread | $1.1K depth, 6.75bps spread | Tests off-hours strategy fitness |
| Liquidation cascade | $150K depth | $4.4K depth | Tests position sizing under stress |
| Market maker withdrawal | $112.5K depth | $3.3K depth | Tests quote quality when MMs pull |
### Numba Performance
Hot-path functions JIT-compiled with numba:
| Function | Speedup |
|----------|---------|
| fill_from_levels (batch 100) | 1.8x |
| round_tick | JIT-compiled |
| round_lot | JIT-compiled |
| clip_lots | JIT-compiled |
| extract_features | Vectorized numpy |
CWM throughput: **157K calls/sec**, **6.4 µs/transition**, **0.64 ms/100-step episode**.
### Design Principles
1. **Hardcoded baseline is NEVER replaced** — always available as reference
2. **Generated strategies are ADDED** — system GROWS its repertoire
3. **Crossover** combines two strategies (uniform crossover on parameters)
4. **Mutation** randomly modifies parameters (Gaussian noise)
5. **Selection** uses tournament selection (fitter genomes win more often)
6. **Third parties** can write strategies in DSL text format
7. **Market adaptation**: selector chooses strategy based on ExoF/MARAS regime
8. **Behavior-driven scenarios**: every asset behaves like its real-world self
### Strategy Naming Convention
```
{strategy_type}_gen{generation}_{timestamp}
```
Examples: `UCB1_gen1_20260707_223754`, `THOMPSON_gen2_20260707_223754`
---
## H. Cambrian Expansion
### Tie-In Fix
PerformanceMatrix now wired to evaluator. Strategies scored by regime fitness:
- ALL strategies tested in ALL scenarios
- Score tracked per strategy × regime
- Selector queries matrix for best strategy per regime
- `PerformanceMatrix.record(strategy_id, regime, score)` called during evaluation
### Phase 0.1: Cognition Pipeline
Rate-limited market regime research:
- `RateLimiter`: token bucket, 30 RPM default
- `SourceCatalogue`: tracks sources, fetch counts, error rates
- `RegimeExtractor`: keyword matching + sentiment analysis
- `CognitionPipeline`: rate-limited, deduplication, long/perm-run
- 8 default sources: CoinDesk, Cointelegraph, The Block, CryptoQuant, Glassnode, Coinglass, Binance Research, Messari
### Phase 0.2: Exponential Regime Expansion
Orthogonal to cognition pipeline — generates 200+ regimes from dimensions:
- 4 liquidity (vacuum, thin, normal, deep)
- 4 volatility (tight, normal, wide, extreme)
- 4 flow (balanced, buy_pressure, sell_pressure, toxic)
- 4 structure (single_toxic, mm_toxic, multi_toxic, retail_mm)
Each combination = distinct, testable regime. Total: 256 theoretical, ~200 practical.
### Phase 0.3: Multi-Asset & Invariant Classification
13 assets across 10 sectors, classified by **invariant characteristics only**:
| Asset | Sectors | Roles | Supply | Consensus | Vol | Liq |
|-------|---------|-------|--------|-----------|-----|-----|
| BTC | CURRENCY | STORE_OF_VALUE | FIXED_CAP | POW | LOW | DEEP |
| ETH | LAYER1, DEFI | GAS, SOV, GOV | DISINFLATIONARY | POS | MEDIUM | DEEP |
| SOL | LAYER1 | GAS | INFLATIONARY | POS | HIGH | NORMAL |
| DOGE | MEME, CURRENCY | MEME, GAS | INFLATIONARY | POW | HIGH | NORMAL |
| ADA | LAYER1 | GAS | INFLATIONARY | DPOS | MEDIUM | NORMAL |
| AVAX | LAYER1, DEFI | GAS | INFLATIONARY | POS | HIGH | NORMAL |
| UNI | DEFI | GOV, UTILITY | INFLATIONARY | POS | HIGH | THIN |
| LINK | ORACLE | UTILITY | INFLATIONARY | POS | MEDIUM | THIN |
| BNB | EXCHANGE, LAYER1 | EXCHANGE_FEE, GAS | BURN | DPOS | MEDIUM | DEEP |
| MATIC | LAYER2 | GAS | INFLATIONARY | POS | HIGH | THIN |
| AAVE | DEFI | GOV | FIXED_CAP | POS | HIGH | THIN |
| DOT | LAYER1 | GAS, GOV | INFLATIONARY | DPOS | MEDIUM | NORMAL |
| ATOM | LAYER1 | GAS | INFLATIONARY | POS | HIGH | THIN |
Multi-asset = linear multiplication: 30 scenarios × 13 assets = 390+ testable scenarios.
### Phase 0.4: Asset Behavior DSL + Compiler + Behavior-Driven Scenarios
**Research-validated** from Binance live API, BingX API, CoinGlass, and academic literature:
| Component | Purpose | File |
|-----------|---------|------|
| **AssetBehavior** | 10 orthogonal behavior dimensions | `training/asset_behavior.py` |
| **BehaviorTemplate** | Reusable class-level behavior profiles | `training/asset_behavior.py` |
| **AssetCompiler** | Auto-fetch from Binance/BingX, compute profiles | `training/asset_compiler.py` |
| **ScenarioFactory** | Behavior-driven scenarios with label queries | `training/cma_trainer.py` |
**Key research findings wired into the system:**
- Depth decay: D(d) = A * d^(1-alpha). BTC alpha=0.7, DOGE alpha=1.0
- Spread ranges: BTC 0.01bps, SOL 1.26bps, DOGE 1.35bps, ADA 5.95bps
- Order sizes: BTC P50=$643, DOGE P50=$96 (7x smaller)
- Cascade dynamics: BTC 5-8% trigger/slow/fast recovery, DOGE 3-5%/fast/slow
- BingX specifics: 1.7-12.6x wider spreads, 0.05-9x variable depth
- Weekend effects: -35% vol, -48% volume, +20% spread
- Intraday patterns: 7.4x peak/trough ratio, US session dominant
- GARCH persistence: 0.95-0.99, vol half-life 2-5 days
- Correlation: BTC-ETH 0.90 normal, 0.97 crash; BTC-DOGE 0.45 normal, 0.80 crash
### Training Scaling Results
CMA-ES evaluation on behavior-driven multi-asset scenarios (BTC/ETH/SOL, 3 assets × 30
scenario types = 90 scenarios). Each eval = one CMA candidate × 90 multi-step episodes
through the CWM.
#### Parallel Training Benchmark (48 evals, pop=12)
| Workers | Time | Best Score | Score/min | Speedup | Efficiency |
|---------|------|-----------|-----------|---------|-----------|
| 1 (sequential) | 632s | 1,727 | 164 | 1.0× | — |
| 2 | 320s | 2,262 | 425 | 1.98× | 99% |
| 4 | 163s | 2,152 | 791 | 3.87× | 97% |
| **8** | **91s** | **2,594** | **1,718** | **6.98×** | **87%** |
**87% parallel efficiency at 8 workers.** Independent workers explore more diverse
strategy space — parallel runs find BETTER scores than sequential (2,594 vs 1,727).
#### Budget Scaling (ProcessPool, 3 assets × 30 scenarios)
| Budget | Gens | Best Score | Time | Score/min |
|--------|------|-----------|------|-----------|
| 48 evals | 4 | 2,594 | 91s | 1,718 |
| 96 evals | 8 | 7,202 | 1246s | 346 |
| 192 evals | 16 | ~14K (est) | ~45min | ~311 |
#### Vectorized Reward (numba)
The CWM `reward()` method now uses `compute_reward_vectorized` (numba JIT) when
available, bypassing the Python FeatureVector dict allocation. This eliminates
~2M dict allocations per CMA generation.
#### Training Performance
| Metric | Sequential | Parallel (8 workers) | Speedup |
|--------|-----------|---------------------|---------|
| CWM throughput | 189K calls/sec | (bottleneck is planner, not CWM) | — |
| CWM per-call latency | 5.3 µs | (numba + vectorized reward) | — |
| Scenario generation | 390 scenarios in 0.8s | — | — |
| CMA-ES per-generation (pop=12) | ~155s | **~23s** | **~7×** |
| CMA-ES 48 evals | 632s | **91s** | **7×** |
| Best score at 48 evals | 1,727 | **2,594** | +50% |
| Score/min | 164 | **1,718** | **10.5×** |
Parallel evaluation achieves **9x speedup on single evals** (embarrassingly parallel,
zero fidelity loss). CMA-ES training speedup is ~1.5x because multiprocessing overhead
is amortized across 90 scenarios per eval. Bottleneck is planner (MCTS), not CWM —
numba already accelerates the CWM hot path (`_HAS_NUMBA = True`).
### Order Type Standardization (Multi-Exchange, FIX/CCXT-Aligned)
Three **orthogonal dimensions** (not one flat enum), each mapped independently:
**Dimension 1 — Order Types (FIX Tag 40 OrdType):** what the order IS
`MARKET`, `LIMIT`, `STOP_MARKET`, `STOP_LIMIT`, `TRIGGER_MARKET`, `TRIGGER_LIMIT`,
`TRAILING_STOP`, `OCO`, `TP_SL`
**Dimension 2 — Time-in-Force (FIX Tag 59):** how long the order LIVES
`GTC`, `IOC`, `FOK`, `GTD`
**Dimension 3 — Instructions (FIX Tag 18 ExecInst):** behavioral modifiers
`POST_ONLY`, `REDUCE_ONLY`, `HIDDEN`, `ICEBERG`
> CRITICAL: `POST_ONLY` is an instruction on a `LIMIT` order, not a standalone type.
> `IOC`/`FOK` are TimeInForce values applied to a `LIMIT` order, not order types.
> This aligns with FIX Protocol: Tag 40 (OrdType), Tag 59 (TimeInForce), Tag 18 (ExecInst).
#### Cross-Exchange Order Type Mapping
| Normalized | BingX | Binance Spot | Binance Futures | Bybit |
|-----------|-------|-------------|-----------------|-------|
| `LIMIT` | LIMIT | LIMIT | LIMIT | LIMIT |
| `MARKET` | MARKET | MARKET | MARKET | MARKET |
| `STOP_MARKET` | TRIGGER_MARKET | STOP_MARKET | STOP_MARKET | STOP_MARKET |
| `STOP_LIMIT` | TRIGGER_LIMIT | STOP_LOSS_LIMIT | STOP | STOP_LIMIT |
| `TRIGGER_MARKET` | TRIGGER_MARKET | TAKE_PROFIT | TAKE_PROFIT | TAKE_PROFIT_MARKET |
| `TRIGGER_LIMIT` | TRIGGER_LIMIT | TAKE_PROFIT_LIMIT | TAKE_PROFIT | TAKE_PROFIT_LIMIT |
| `TRAILING_STOP` | TRAILING_STOP_MARKET | TRAILING_STOP_MARKET | TRAILING_STOP_MARKET | TRAILING_STOP |
#### TimeInForce Mapping
| Normalized | BingX | Binance | Bybit |
|-----------|-------|---------|-------|
| `GTC` | GTC | GTC | GTC |
| `IOC` | IOC | IOC | IOC |
| `FOK` | FOK | FOK | FOK |
#### Transferability Principle
Strategy **PARAMETERS** transfer across exchanges. Order type **NAMES** are venue-specific
but semantics are identical. A `LIMIT` on BingX = `LIMIT` on Binance = `LIMIT` on Bybit
(same fill behavior). Only the API string differs. The venue adapter translates
normalized → exchange-native at submission time.
#### Cross-Exchange Learning
ScenarioFactory is venue-aware: `ScenarioFactory(exchange_id='bingx')` tags every
scenario with its venue. Cross-exchange transfer re-tags for a different venue:
```python
factory = ScenarioFactory(exchange_id='bingx')
scenarios_bingx = factory.build_suite(symbols=['BTCUSDT', 'ETHUSDT'])
strategy = train(scenarios_bingx) # evolve on BingX
scenarios_binance = factory.cross_exchange_transfer(scenarios_bingx, 'binance')
score = evaluate(strategy, scenarios_binance) # test on Binance
```
PerformanceMatrix keys are `(regime, strategy_id, venue)`:
- `get_best(regime, venue='bingx')` — per-venue best
- `get_venue_comparison(regime, strategy_id)` — `{venue: score}`
- `get_best(regime)` — venue-agnostic (backward compatible)
#### Adversary Ecology
Counterparties operate at the ActionKind level (CROSS_SPREAD, PLACE, CANCEL),
not at order-type level. The CWM translates ActionKind to venue-native order types:
- `CROSS_SPREAD` → fills aggressively → equivalent to MARKET
- `PLACE` → passive quote → equivalent to LIMIT
Fee calculation uses VenueRules (per-exchange fees). The ecology is venue-independent.
#### Standards Referenced
- FIX 4.4: Tag 40 (OrdType), Tag 59 (TimeInForce), Tag 18 (ExecInst)
- CCXT: unified `limit`/`market` + parameter composition (triggerPrice, stopPrice)
- ISO 10383: MIC for venue identification
- MiFID II / MiCA: defers to FIX for order type classification
- FIA: references FIX Tag 40 in digital asset guidelines
### Prod Tooling
| Component | Purpose | File |
|-----------|---------|------|
| **CognitionLauncher** | Standalone long-run script | `cognition_launcher.py` |
| **CognitionMonitor** | Metrics, health scoring, alerts | `training/monitor.py` |
| **NewsSourceRepository** | 12 industry-standard sources, ranking | `training/news_sources.py` |
| **SourceCatalogue** | Source persistence, fetch tracking | `training/cognition.py` |
| **RateLimiter** | Token bucket, 30 RPM | `training/cognition.py` |
| **RegimeExpander** | 200 regimes from dimensions | `training/regime_expansion.py` |
### Running
```bash
# Training pipeline (continuous)
python -m malkhut.continuous_pipeline
# Cognition pipeline (continuous regime research)
python -m malkhut.cognition_launcher
```