# MALKHUT — Adversarial Self-Play Order-Fulfilment Pipeline **Lock-free. No GC. Fully async. GraalVM-compatible.** MALKHUT is an order-fulfilment engine that optimises a **distribution over execution actions** against a diverse counterparty ecology, not a single deterministic quote. ## Core Thesis > Sequential/pure optimisation ("what is the best quote?") leads to > deterministic quote → predictable fill pattern → adverse selection. > > Correct simultaneous-market framing: "what mixed order-placement policy > survives a diverse counterparty ecology?" ## Architecture ### Corrected Stack (T19 ANNEX B-CORRECTION) > **iceoryx2 = the wires; ASEx = the state cells; UV clock = the conductor.** | Layer | Component | Role | |-------|-----------|------| | Transport | iceoryx2/zinc topics | Inter-component event fabric. Observable in `/dev/shm`. Single-writer per topic. | | State | ASEx (ASExGuardedState + ASExWorker) | Validate-before-mutate state serialization. One writer thread per state. NOT a message bus. | | Dispatch | UV Clock (asyncio event loop) | Event-driven reactor. Subscribes to event types. Nothing sleeps-and-polls. | **Per T19:** ASExWatch is **DEFERRED** from the critical path until heavy-test hangs are fixed. Use ASExWorker (FulfilmentWorker) for production hot-path mutations. ``` ┌──────────────────────────────────────────────────────────────────────┐ │ Zinc SHM (POSIX /dev/shm/zinc_*) │ │ malkhut_book_state | malkhut_account_state | malkhut_fulfilment_out │ │ malkhut_risk_gate | malkhut_control (CONTROL_PLANE) │ └───────────────────────────────┬──────────────────────────────────────┘ │ ▼ ┌──────────────────────────────────────────────────────────────────────┐ │ ASEx VALIDATE-BEFORE-MUTATE │ │ GuardedFulfilmentState — book/account/intent/policy mutations │ │ GuardedRiskState — risk gate + kill switch │ │ FulfilmentWorker — single-threaded ASExWorker per state │ │ FulfilmentWatch — zero-overhead ring buffer for hot path │ │ ShardedWorker — per-symbol state partitioning │ └───────────────────────────────┬──────────────────────────────────────┘ │ ▼ ┌──────────────────────────────────────────────────────────────────────┐ │ CODE WORLD MODEL (CWM) │ │ deterministic exchange transition function │ │ f(state, joint_actions, latency_model, venue_rules) → next_state │ │ wraps hftbacktest for replay verification │ └───────────────────────────────┬──────────────────────────────────────┘ │ ┌──────────┴──────────┐ ▼ ▼ ┌─────────────────────────┐ ┌──────────────────────────────────┐ │ LIVE FULFILMENT PLANNER │ │ OFFLINE CMA-ES TRAINER │ │ Decoupled UCB/UCT │ │ pycma + nevergrad │ │ ≤ 25 ms cadence │ │ growing self-play pool │ └───────────┬──────────────┘ └─────────────────┬────────────────┘ │ │ ▼ ▼ ┌─────────────────────────┐ ┌──────────────────────────────────┐ │ RISK / COMPLIANCE GATE │ │ POLICY REGISTRY / POOL │ │ leverage, notional, │ │ accepted candidates, weights │ │ self-trade, kill switch │ │ diversity-preserving eviction │ └───────────┬──────────────┘ └──────────────────────────────────┘ │ ▼ ┌─────────────────────────┐ │ VENUE ADAPTER (BingX) │ │ create/cancel/replace │ └───────────┬──────────────┘ │ ▼ ┌─────────────────────────┐ │ ClickHouse Persistence │ │ decisions, episodes, │ │ snapshots, discrepancies│ └─────────────────────────┘ ``` ## Package Structure ``` MALKHUT/ ├── malkhut/ │ ├── __init__.py # package version │ ├── state.py # frozen data model (42 dataclasses) │ ├── actions.py # action model, PlannedPolicy, RiskDecision │ ├── features.py # feature extraction (17 features) │ ├── counterparties.py # adversarial agent ecology (4 agents) │ ├── engine.py # FulfilmentEngine (hot-path orchestrator) │ ├── bench_numba.py # numba speedup benchmark │ ├── launch_smoke_test.py # 10-min smoke test launcher │ ├── cwm/ │ │ ├── __init__.py │ │ ├── core.py # CWM (deterministic exchange simulator) │ │ ├── numba_core.py # numba-accelerated hot loops │ │ └── replay_verify.py # replay verification (mandatory before trust) │ ├── planner/ │ │ ├── __init__.py │ │ ├── action_menu.py # compact action space construction │ │ └── sm_mcts.py # Decoupled UCB/UCT simultaneous-move MCTS │ ├── risk/ │ │ ├── __init__.py │ │ └── gate.py # risk/compliance gate (hard invariants) │ ├── venue/ │ │ ├── __init__.py │ │ └── bingx/ │ │ ├── __init__.py │ │ └── adapter.py # BingX VST/Live venue adapter (wraps DITAv2) │ ├── ipc/ │ │ ├── __init__.py │ │ ├── zinc_plane.py # Zinc shared memory IPC (POSIX SHM) │ │ └── control_plane.py # CONTROL_PLANE shared memory region │ ├── storage/ │ │ ├── __init__.py │ │ ├── ch_store.py # ClickHouse persistence (5 tables) │ │ └── asset_store.py # DuckDB file-backed asset universe store │ ├── execution/ │ │ ├── __init__.py │ │ └── asex_integration.py # ASEx validate-before-mutate kernel │ ├── clock/ │ │ ├── __init__.py │ │ ├── host.py # UVClock — event-driven reactor (T19) │ │ ├── events.py # Typed events: Scan, Tick, Timer, BarFire, Stale │ │ ├── staleness.py # Freshness watchdog (T19 Staleness Law) │ │ └── deadnode.py # DeadNode reaper for iox2 orphans │ ├── training/ │ │ ├── __init__.py │ │ ├── cma_trainer.py # CMA-ES + self-play + behavior-driven scenarios │ │ ├── registry.py # Policy lifecycle (CANDIDATE → ACTIVE) │ │ ├── pipeline.py # Bounded continuous learning loop + logger │ │ ├── dsl.py # Strategy DSL v2 (40+ primitives, 40+ sensors) │ │ ├── generator.py # Genetic programming strategy evolution │ │ ├── selector.py # Regime → strategy mapping + performance matrix │ │ ├── asset_classification.py # Multi-label taxonomy + exchange registry (system-wide) │ │ ├── asset_behavior.py # 10-dimension behavior DSL, research-validated │ │ ├── asset_compiler.py # Auto-fetch from Binance/BingX, compile profiles │ │ ├── parallel_eval.py # ProcessPoolExecutor episode runner, 9x speedup │ │ ├── cognition.py # Rate-limited market regime research │ │ ├── regime_expansion.py # 200+ regimes from dimension combinations │ │ ├── news_sources.py # 12 industry-standard news sources │ │ └── monitor.py # Cognition metrics, health, alerts │ ├── cognition_launcher.py # Standalone long-run cognition service │ └── tests/ # 1109+ tests across 46+ test files ├── specs/ │ └── MALKHUT_ADVERSARIAL_SELFPLAY_SPEC.py # full spec └── README.md # this file ``` ## Shared Memory IPC (Zinc) All IPC uses **real Zinc shared memory** (POSIX SHM via `/dev/shm/zinc_*`), NOT the file-based transport. This is lock-free, zero-copy, GraalVM-compatible. ### Regions | Region | Writer | Readers | Purpose | |--------|--------|---------|---------| | `malkhut_book_state` | data feed | planner, CWM | canonical order book snapshot | | `malkhut_account_state` | data feed | planner, risk | account/position snapshot | | `malkhut_fulfilment_out` | engine | other systems | planner output + distribution | | `malkhut_risk_gate` | engine | other systems | risk gate decisions | | `malkhut_control` | external | engine | commands, symbols, lifecycle | ### Framing All regions use `UVZINC01` seqlock framing: ``` [8B magic "UVZINC01"] [8B seq] [8B json_size] [UTF-8 JSON body] ``` ### CONTROL_PLANE The `malkhut_control` region is the management interface: - External systems write `ControlCommand` frames - Engine reads and processes commands - ACKs written back for `ack_required` commands Commands: `START`, `STOP`, `PAUSE`, `RESUME`, `SET_SYMBOLS`, `CONNECT_VENUE`, `DISCONNECT_VENUE`, `SET_MODE`, `EMERGENCY_STOP`, `HOT_RELOAD_POLICY`, `STATUS_REQUEST` ## ClickHouse Persistence Database: `dolphin_malkhut` on `localhost:8123` | Table | Purpose | |-------|---------| | `replay_steps` | CWM replay traces | | `fulfilment_decisions` | every live decision | | `self_play_episodes` | training episode results | | `policy_snapshots` | versioned policy registry | | `live_discrepancies` | simulation vs actual | ## External Dependencies (Spec-Mandated) | Library | Purpose | Used For | |---------|---------|----------| | **pycma** (`cma`) | CMA-ES optimisation | `training/cma_trainer.py` | | **hftbacktest** | replay/queue/latency fill simulation | CWM replay verification | | **nevergrad** | gradient-free optimiser zoo | trainer comparison | | **msgspec** | fast typed serialization | state/event encoding | | **numba** | hot-loop acceleration | CWM transitions, feature extraction | | **polars** | offline analytics | large trade-path corpora | | **lmdb** | existing SILOQY storage | ticks, candles, hd_vectors | | **numpy** | vectorized feature extraction | parameter decoding | | **scipy** | stats tests, bootstrap | CI, promotion gates | ## Quick Start ```python from malkhut.engine import FulfilmentEngine from malkhut.state import FulfilmentPolicyParams def my_params(): return FulfilmentPolicyParams( version="v0.1", ucb_c=1.414, max_sims=64, max_depth=3, rollout_depth=3, root_temperature=0.5, min_root_entropy=0.25, quote_offsets_ticks=(0, 1, 2), quote_size_fractions=(0.10, 0.25, 0.50), passive_ttl_ms=200, aggressive_ttl_ms=50, maker_edge_min_bps=0.5, cross_spread_edge_min_bps=5.0, adverse_toxicity_cancel_threshold=0.5, queue_churn_cancel_threshold=0.5, mae_tail_cut_bps=50.0, mfe_giveback_cut_fraction=0.5, max_time_in_loss_s=300.0, failed_recovery_cut_count=3, recovery_velocity_min_bps_per_s=0.0, max_symbol_notional_fraction=0.20, max_single_order_notional_fraction=0.05, reduce_when_global_up_fraction=0.30, session_profit_lock_fraction=0.02, w_expected_pnl=1.0, w_fill_probability=0.5, w_adverse_selection=2.0, w_queue_priority=0.5, w_inventory_risk=1.5, w_tail_loss=5.0, w_fee_quality=0.5, w_time_decay=0.3, w_policy_entropy=0.5, robust_tail_weight=2.0, toxic_counterparty_weight=3.0, low_liquidity_weight=2.0, latency_stress_weight=1.0, ) engine = FulfilmentEngine(params_provider=my_params) # engine.on_state(latest_market_state) # hot path ``` ## ASEx Integration (Validate-Before-Mutate) MALKHUT uses **ASEx** (Advanced Serial Executor) for all mutable state mutations. This is the same kernel used by VIOLET/DOLPHIN for race-proof state management. ### Core Pattern ```python class GuardedFulfilmentState(ASExGuardedState): def _validate(self, mutation) -> bool: # Pure predicate — MUST NOT mutate self return mutation.ts_ns > 0 and len(mutation.bids) > 0 def _apply(self, mutation): # Only called after _validate() passes # MAY mutate self self._state = new_state return result ``` ### Components | Component | Purpose | Thread Safety | |-----------|---------|---------------| | `GuardedFulfilmentState` | Engine state mutations (book, account, intent, policy) | One writer thread via ASExWorker | | `GuardedRiskState` | Risk gate + kill switch | One writer thread via ASExWorker | | `FulfilmentWorker` | Single-threaded ASExWorker wrapping GuardedFulfilmentState | Serialised via queue.Queue | | `RiskWorker` | Single-threaded ASExWorker wrapping GuardedRiskState | Serialised via queue.Queue | | `FulfilmentWatch` | Zero-overhead ring buffer for hot-path mutations | Producer/consumer lock-free | | `ShardedWorker` | Per-symbol state partitioning (N independent workers) | No shared state between partitions | ### Why ASEx - **No locks**: One thread per state object. Races impossible by construction. - **No GC pressure**: Frozen dataclasses + queue.Queue (CPython C-level ops). - **Validate before apply**: Invalid mutations rejected without touching state. - **GraalVM compatible**: No CPython-only APIs in the hot path. - **Battle-tested**: 110 ASEx tests + 1109 MALKHUT tests = 1219 total. ## Design Rules (from spec) 1. **Optimise an ecology-resistant distribution, not a quote.** 2. **The book is not passive scenery.** Ask "why did I get filled?" 3. **Replay correctness before search depth.** Wrong CWM + deep search = confident nonsense. 4. **CMA-ES does not promote from noisy wins.** Survive scenario suite + bootstrap CI. 5. **Preserve feature humility.** Allow CMA-ES to discover non-obvious interactions. 6. **Path-aware SL/TP is part of fulfilment.** Exiting = order-placement under adversarial liquidity. 7. **Never let the optimiser learn illegal behavior.** Hard-code constraints. 8. **Keep hot path bounded.** ≤ 100 ms. No live CMA, no large neural nets. 9. **Version everything.** Every decision explainable by policy + CWM + feature + state + action + outcome. 10. **If shadow live diverges, trust live.** Simulation is only a tool. --- # ADDENDUM — EXEC-STACK PLUGGABILITY + SPEC HARDENING *[Fable, integrator, 2026-07-06 — operator-ordered extension. Mimo: this section binds MALKHUT into the UV/DITAv2 stack; nothing above is overridden except where marked FIX.]* ## A. Where MALKHUT plugs in (the seat at the table) MALKHUT **is the execution layer under the UV Clock** (see `prod/docs/uv_subspecs/UV_TASK_T19_UV_CLOCK_HOST.md` + annexes): ``` UV Clock (T19) → PromotionBridge → KernelIntent (DITAv2 contracts) → MALKHUT FulfilmentEngine (replaces naive kernel-submit) → RISK GATE → BingX venue adapter → fills ``` Your Design Rule 6 ("path-aware SL/TP is part of fulfilment") is CONVERGENT with the operator's T19 ruling ("the whole point of UV = faster-than-scan-cadence TP/SL paths"). MALKHUT's planner is the *intelligent* form of the C11 tick-exit guardian. Sequencing: simple doctrinal tick-exits (C11) ship first via T19 step 3; MALKHUT supersedes them **when certified** — same seat, smarter occupant. ## B. Intake contract (BINDING) 1. **Input = DITAv2 `KernelIntent`** (`prod/clean_arch/dita_v2/contracts.py`): asset, side (TradeSide), action (ENTER/EXIT), reference_price, target_size, leverage, metadata (carries `promo_client_id`). Accept KernelIntent natively; do not invent a parallel intent type. 2. **`u-` clientOrderId prefix is LAW** (T9 seam): every venue order MALKHUT places carries the u- prefix from intent metadata. Non-negotiable — it is the attribution boundary vs BLUE and vs legacy VIOLET orders. 3. **DUAL-LEVERAGE LAW**: `KernelIntent.leverage` is CONVICTION leverage [0.5, 9.0] — it sizes QUANTITY only. Venue leverage = the mapping in `prod/bingx/leverage.py` (linear → ROUND_HALF_EVEN → int clamp [1,3]). WIRE the existing typed L3 wrapper (`prod/clean_arch/violet/exchange_leverage.py`); litmus: 0.5→1, 4.75→2, 8.0→3, 9.0→3. The venue adapter MUST NOT pass conviction leverage raw to set-leverage. 4. **Arming/two-man/kill**: MALKHUT honors `UV_PROMOTED=1` + arm-file (`/root/uv-wt/prime-live/UV_PROMOTED.arm`), re-evaluated PER EVENT (never cached — the T12 kill-switch lesson). DARK default = ObserveOnly venue. Map arm-file-delete → your `EMERGENCY_STOP` semantics: inert next event, no restart. ## C. Fabric + state alignment (the tonight-rulings) - **Stack law: iceoryx2 = wires, ASEx = state cells, UV clock = conductor.** Your ASEx usage (GuardedState + Worker per state object) is CORRECT — it is exactly what ASEx is for. Your zinc regions already speak UVZINC01 seqlock framing — good; when T19's event topics land, consume ScanEvent/TickEvent (with PROVENANCE: scan_number, scan_ts, ingest_ts, AGE) from the shared LiveSource instead of a private feed. - **Account state**: read the DITAv2 ASEx account core (asex_account / AccountProjectionV2) as a CONSUMER only. MALKHUT never becomes a second writer to account state (two-cores-share-nothing law); its own guarded states are its own. - **Staleness law (T19 Annex A) — extend your Rule 2**: a stale book is not scenery either, it is an ADVERSARY. Planner receives AGE on every input and must discount/refuse plans over stale state; freshness watchdog emits StaleInput. - **ASEx pre-condition inherited**: DeadNode reaper / orphan-sweep on start (iox2 segments orphan on crash — known gap). ## D. Persistence + observability alignment - `dolphin_malkhut` namespace: approved (own-namespace law). **NO TTL on any table** (retention doctrine — audit surfaces never expire). - **Cross-reference law**: every `fulfilment_decisions` row carries `intent_id` + `trade_id` from the source KernelIntent → end-to-end trace joins `dolphin_uv.exec_journal` ⇄ `dolphin_malkhut.*` ⇄ venue order history. This is what makes MALKHUT certifiable instead of a black box. - Publish planner pulse (decision rate, plan entropy, gate verdicts, AGE-at-decision) into the zinc snapshot → extant TUIs render it (TUI-first doctrine). ## E. Ecology + reward grounding (FIX-grade improvements) 1. **Ground the counterparty ecology in OUR tape**: fit/validate adversary populations against `dolphin.obf_universe` recorded sub-second tape and the WT2 stress manifold. The CLASS-M/CLASS-X hour sets (WT2D) are your ready-made adversarial scenario suites — corrupt/stress hours ARE the hostile ecology, measured. CLASS-X law applies: extreme magnitudes are a FEATURE channel, never filtered. 2. **Latency model = our measured tails**, not defaults: scan_to_fill/step_bar distributions from the WT1 latency audit (median 47ms/1.58ms, stop-tail 7.8s) — calibrate the CWM's latency injections from these empirical curves. 3. **Reconcile the budget statement (FIX)**: planner box says ≤25ms, Rule 8 says ≤100ms — bind them: ≤25ms planning inside a ≤100ms end-to-end intent→venue-call budget, both enforced by test. 4. **Promotion gate binds to the cert conveyor**: CMA-ES candidate promotion (Rule 4) emits artifacts via the T4 cert reporter into the gate ledger; policy snapshots versioned in `dolphin_malkhut.policy_snapshots` AND referenced in `prod/docs/uv_cert/`. No policy trades live without a ledger row. ## F. Certification pathway (Gate-M — how MALKHUT reaches money) 1. **CWM replay verification** (your Rule 3) vs hftbacktest — already specced; add determinism ×2 (byte-identical minus ts). 2. **Dual-leverage litmus** (B.3 values) as a named test. 3. **DARK shadow window**: MALKHUT plans on live intents while the simple path executes; diff planned-vs-naive on fills/slippage/adverse-selection — the improvement claim must be MEASURED on VST before MALKHUT takes the seat. 4. **Risk-gate mutation litmus**: drop a constraint → a test goes RED (looks-finished doctrine; 222 green tests ≠ verification until mutations bite). 5. **Rule 10 stands supreme**: shadow-vs-live divergence → trust live, stand down, file the discrepancy row. *End addendum. — F.* --- ## DEVELOPMENT STATUS (2026-07-11) **1276 test functions. 50 test files. All green. 0 failures. 0 regressions.** ### Completed subsystems | Subsystem | Module | Tests | Status | |-----------|--------|-------|--------| | **State Model** | `state.py` | 17 | 42 frozen dataclasses, immutable | | **CWM** | `cwm/core.py` + `cwm/numba_core.py` | 103 | Full exchange mechanics + numba JIT (5.3 µs/transition) | | **Replay Verification** | `cwm/replay_verify.py` | 65 | Deep comparison, binary search, trajectory recording | | **Planner** | `planner/sm_mcts.py` | 11 | Decoupled UCB/UCT, ≤25ms budget | | **Action Menu** | `planner/action_menu.py` | (in planner) | Compact action space construction | | **Counterparties** | `counterparties.py` | 19 | 4 adversarial agent policies | | **Risk Gate** | `risk/gate.py` | 4 | Hard invariants: leverage, self-trade, post-only | | **ASEx Integration** | `execution/asex_integration.py` | 33 | Validate-before-mutate, single-writer, no locks | | **Zinc IPC** | `ipc/zinc_plane.py` | 8 | Real POSIX SHM, UVZINC01 framing | | **Control Plane** | `ipc/control_plane.py` | 3 | HOT_RELOAD_POLICY, kill switch | | **ClickHouse** | `storage/ch_store.py` | 9 | 5 tables, HTTP API | | **DuckDB Asset Store** | `storage/asset_store.py` | 30 | File-backed + in-memory materialization, sub-µs reads | | **Asset Bridge** | `training/asset_bridge.py` | 49 | Directory ↔ classification sync | | **CMA-ES Training** | `training/cma_trainer.py` | 65 | Behavior-driven scenarios, auto-compile, label queries | | **Parallel Eval** | `training/parallel_eval.py` | 16 | 9x speedup, zero fidelity loss, ProcessPoolExecutor | | **Policy Registry** | `training/registry.py` | 14 | CANDIDATE → ACTIVE lifecycle | | **Training Pipeline** | `training/pipeline.py` | 21 | Bounded continuous learning loop | | **Strategy DSL v2** | `training/dsl.py` | 69 | 40+ primitives, 40+ sensors, 16 builtins | | **Strategy Generator** | `training/generator.py` | 20 | Genetic programming: crossover, mutation, tournament | | **Strategy Selector** | `training/selector.py` | 24 | Regime → strategy mapping, performance matrix | | **Asset Classification** | `training/asset_classification.py` | 190 | Multi-label taxonomy + exchange registry, 13 assets, system-wide store | | **Asset Behavior DSL** | `training/asset_behavior.py` | (in classification) | 10 orthogonal dimensions, 3 templates, 13 behaviors, research-validated | | **Asset Compiler** | `training/asset_compiler.py` | (new) | Auto-fetch from Binance/BingX API, compile profiles, rate-limited | | **Cognition Pipeline** | `training/cognition.py` | 29 | Rate-limited, 8 sources, dedup, perm-run | | **Regime Expansion** | `training/regime_expansion.py` | 29 | 200+ regimes from 4×4×4×4 dimensions | | **News Sources** | `training/news_sources.py` | 12 | 12 industry-standard sources, ranking | | **Cognition Monitor** | `training/monitor.py` | 9 | Metrics, health scoring, alerts, JSONL logging | | **Cognition Launcher** | `cognition_launcher.py` | 21 | Standalone long-run cognition service | | **UV Clock** | `clock/host.py` | 30 | T19 event-driven reactor | | **Numba Acceleration** | `cwm/numba_core.py` | 13 | JIT on hot loops, 1.8x fill speedup | | **BingX Adapter** | `venue/bingx/adapter.py` | 28 | Wraps DITAv2, order tracking | | **Hypothesis** | `test_hypothesis_properties.py` | 6 | Property-based fuzzing | | **Fuzz/Adversarial** | `test_fuzz.py`, `test_adversarial.py` | 17 | Random stress + adversarial scenarios | | **Concurrency** | `test_concurrency.py` | 8 | Thread safety, race detection | | **Sync/Async Seams** | `test_sync_async_seams.py` | 7 | Zinc latency, engine budget | | **E2E Integration** | `test_e2e_integration.py` | 2 | Full pipeline: train→register→plan→risk→venue→zinc→ch→reload | ### Performance Benchmarks | Metric | Value | |--------|-------| | CWM transition | 6.4 µs/call | | CWM throughput | 157K calls/sec | | CWM 100-step episode | 0.64 ms | | Numba fill speedup | 1.8x (batch 100) | | Peak RAM | 146 MB | | Smoke test (10min) | 174s, 5 gens, 50 evals, 11 strategies | ### Bugs found and fixed (22 total) 1. Empty book crash in `materialize_price_from_action` (SELL side) 2. Empty book crash in `total_notional` calculation 3. Empty book crash in mark-to-market 4. Empty book crash in `_inventory_risk` 5. Empty book crash in counterparty processing 6. Empty book crash in `_update_path_state` 7. Empty book crash in feature extractor (`b.mid`) 8. Counterparty fills used `max()` instead of consuming levels 9. Position overwritten instead of accumulated on add 10. Test helper `_state()` not forwarding `positions` kwarg 11. `_make_open_order` set empty symbol instead of venue symbol 12. `FulfilmentPolicyParams` frozen dataclass `__dict__` access 13. Banker's rounding in tick tests 14. `HOT_RELOAD_POLICY` compared float to string 15. Engine init order (workers before hot_reload) 16. `_tp()` helper missing `failed_recovery_count` parameter 17. CWM: SELL position flip didn't reset avg_entry 18. CWM: Counterparty fills overwrote our fill tracking 19. CWM: Double-counting unrealized PnL in equity 20. UnboundLocalError when max_steps=0 21. fill_count never incremented in training 22. Cancel rate limiter incorrectly gated order placements ### Key design rules enforced 1. ✅ Optimise distribution, not single quote 2. ✅ Book is not passive scenery 3. ✅ **Replay correctness before search depth** (mandatory gate) 4. ✅ CMA-ES does not promote from noisy wins (bootstrap CI) 5. ✅ Preserve feature humility (CMA discovers interactions) 6. ✅ Path-aware SL/TP is part of fulfilment 7. ✅ Never let optimiser learn illegal behavior 8. ✅ Hot path bounded (≤100ms) 9. ✅ Version everything 10. ✅ Shadow-vs-live divergence → trust live ### Strategy Selection (completed) The system classifies strategies per market condition and floats the best to top: ``` market_fingerprint → Regime Classifier (8 regimes) ↓ Performance Matrix (regime × strategy → score) ↓ Strategy Selector → Engine uses best strategy ``` ### Upstream integration path - **ExoF** — funding rates, open interest, liquidation levels - **MARAS** — 5-tier ensemble → 8-regime classification - **DOLPHINNG7 vel_div** — velocity divergence signals - **DOLPHINNG7 eigenscan** — eigenspace correlation analysis --- ## G. Strategy DSL v2 — Third-Party Strategy Development MALKHUT includes a **Domain-Specific Language (DSL)** for strategy composition. Third parties can write, evolve, and share strategies without modifying core code. ### DSL Format ``` STRATEGY "passive_maker" { PRIORITY 1: IF spread_bps < 5.0 AND orderflow_toxicity < 0.3 THEN QUOTE(BUY, 0, 0.25) PRIORITY 2: IF spread_bps < 5.0 AND orderflow_toxicity < 0.3 THEN QUOTE(SELL, 0, 0.25) PRIORITY 3: IF orderflow_toxicity > 0.7 THEN CANCEL_ALL PRIORITY 4: IF time_in_trade > 300 THEN EXIT PRIORITY 5: IF unrealized_pnl < -50bps THEN STOP_LOSS PRIORITY 6: NOOP } ``` ### Action Primitives (40+) | Category | Primitives | |----------|-----------| | Passive | QUOTE, JOIN_QUEUE, STEP_BACK, LADDER, GRID, ICEBERG, TWAP | | Aggressive | CROSS, SNIPER, PING | | Cancellation | CANCEL, CANCEL_ALL, CANCEL_AND_HOLD, REQUOTE | | Position | EXIT, HALF_EXIT, QUARTER_EXIT, TAKE_PARTIAL, STOP_LOSS, TAKE_PROFIT, TRAILING_STOP, EMERGENCY_EXIT, FLAT_ALL | | Sizing | SCALE_IN, SCALE_OUT, INCREASE_SIZE, REDUCE_SIZE | | Stop/Target | MOVE_STOP, MOVE_TAKE_PROFIT, BRACKET, OCO | | Hedging | HEDGE, PAIR_TRADE | | Waiting | HOLD, WAIT_FOR_FILL, WAIT_FOR_PRICE, WAIT_FOR_SPREAD | ### Market Sensors (40+) | Category | Sensors | |----------|---------| | Book | spread_bps, imbalance_3/5/10, bid_depth_3/5/10, bid_ask_ratio | | Flow | toxicity, queue_churn, cross_venue_lead, order_book_toxicity | | Momentum | price_momentum_1s/5s/15s/1m | | Position | position_qty, pnl_bps, leverage, position_age_s | | Path | mae_bps, mfe_bps, time_in_loss, failed_recoveries, recovery_velocity | | Regime | volatility, atr_14/50, rsi_14, funding, regime_score | | Account | equity, risk_budget_used, session_pnl, drawdown | | Performance | profit_factor, sharpe_ratio, consecutive_wins/losses | | Time | current_hour, day_of_week, is_weekend, is_liquid_hours | ### Comparison Operators (12) `>`, `<`, `>=`, `<=`, `==`, `!=`, `abs>`, `abs<`, `changing`, `stable`, `crossing_above`, `crossing_below` ### Builtin Strategies (16) | Strategy | Description | |----------|-------------| | `passive_maker` | Passive limit orders, toxicity avoidance | | `aggressive_taker` | Momentum-based aggressive crosses | | `toxicity_avoider` | Cancels on high toxicity, quotes when safe | | `path_risk_exit` | Path-aware SL/TP exits | | `regime_adaptive` | Adapts to MARAS regime classification | | `momentum_catcher` | Catches momentum moves, partial exits | | `mean_reversion` | Wide-spread mean reversion | | `scalper` | Ultra-tight spreads, small sizes | | `inventory_manager` | Position size control | | `session_guard` | Time-based position management | | `liquidity_hunter` | Deep book exploitation | | `volatility_breakout` | Volatility-based breakouts | | `funding_arb` | Funding rate arbitrage | | `grid_trader` | Grid-based accumulation | | `risk_parity` | Risk budget management | | `hybrid_adaptive` | Multi-condition adaptive | ### Strategy Generator (Genetic Programming) The system evolves strategies via genetic operators: ```python from malkhut.training.generator import StrategyGenerator, GeneratorConfig config = GeneratorConfig( population_size=20, generations=5, tournament_size=3, mutation_rate=0.15, crossover_rate=0.7, ) generator = StrategyGenerator(config=config) population = generator.evolve(baseline_params, scenarios) # Get successful strategies successful = generator.get_successful_strategies(min_fitness=0.0) for genome in successful: generator.add_to_pool(genome) ``` ### Strategy Selector (Market-Adaptive) The system classifies strategies per market regime and selects the best: ```python from malkhut.training.selector import StrategySelector, MarketRegime selector = StrategySelector() result = selector.select(current_state, available_strategies) # result.strategy_id → best strategy for current regime # result.confidence → how confident we are # result.alternatives → other considered strategies ``` Regimes: `trending_up`, `trending_down`, `high_volatility`, `low_volatility`, `mean_reverting`, `momentum`, `choppy`, `liquidity_hole`, `normal` ### Asset Classification (Invariant Taxonomy) Multi-label asset classification using **only invariant characteristics** — properties intrinsic to the token that predict price behaviour and order-book performance. **Three-layer identifier architecture** for cross-system interoperability. **Design principle:** Classify by WHAT an asset IS (fundamental), not how it's trading right now. Volatility/liquidity appear as long-run statistical profiles, not as classification axes that change minute-to-minute. #### Three-Layer Identifier Architecture | Layer | Fields | Purpose | Example | |-------|--------|---------|---------| | **1. Canonical** | `symbol`, `base_asset`, `name`, `unified_symbol`, `quote_currency` | Venue-independent identity | BTCUSDT / BTC / "Bitcoin" / BTC/USDT / USDT | | **2. Cross-system** | `coingecko_id`, `cmc_id`, `blockchain`, `contract_address` | Data provider + on-chain mapping | "bitcoin" / 1 / "ethereum" / "" | | **3. Exchange mapping** | `exchanges: tuple[str, ...]` | Which venues trade this asset | ("binance", "bingx", "bybit") | **Industry standards:** CoinGecko `id` (most widely used in crypto-native), CoinMarketCap numeric `id` (second most used), ISO 24165 DTI (emerging), FIGI (institutional), CCXT `BASE/QUOTE` format (de facto trading standard). #### Fundamental Dimensions (intrinsic, never change) | Dimension | Values | Multi-label | Purpose | |-----------|--------|-------------|---------| | **Sector** | CURRENCY, LAYER1, LAYER2, DEFI, ORACLE, EXCHANGE, MEME, PRIVACY, STORAGE, GAMING_NFT | Yes | Primary use-case / vertical | | **TokenRole** | GAS, STORE_OF_VALUE, GOVERNANCE, UTILITY, MEME, EXCHANGE_FEE | Yes | Functional role → demand elasticity | | **SupplyModel** | FIXED_CAP, DISINFLATIONARY, INFLATIONARY, BURN_MECHANISM | No | Supply pressure dynamics | | **ConsensusFamily** | POW, POS, DPOS | No | Miner/validator selling behaviour | | **SmartContractCapability** | FULL, PARTIAL, NONE | No | DeFi composability | #### Technical Dimensions (invariant market-structure properties) | Dimension | Values | Purpose | |-----------|--------|---------| | **MarketCapTier** | MEGA (>500B), LARGE (50-500B), MID (5-50B), SMALL (500M-5B), MICRO (<500M) | Price impact per dollar | | **VolatilityProfile** | LOW (<30%), MEDIUM (30-80%), HIGH (80-150%), EXTREME (>150%) | Long-run average vol | | **LiquidityProfile** | DEEP (>100M), NORMAL (10-100M), THIN (1-10M), ILLIQUID (<1M) | Long-run avg depth | | **DerivativeAccess** | PERPS_AND_OPTIONS, PERPS_ONLY, NONE | Shorting / funding availability | | **tick_size / lot_size** | Exchange-set | Minimum price/order size | | **maker/taker fee_bps** | Exchange-set | Trading cost | | **typical_spread_bps / depth / volume** | Long-run averages | Order-book fingerprint | #### Predictive Properties (derived from invariants) | Property | Logic | Value | |----------|-------|-------| | `supply_pressure` | PoW → forced selling (miners cover electricity) | "forced" / "optional" | | `demand_elasticity` | GAS or STORE_OF_VALUE → inelastic (must hold) | "inelastic" / "elastic" | | `can_be_shorted` | DerivativeAccess != NONE | bool | #### Multi-Label Examples ``` BTC: sectors=(CURRENCY), roles=(STORE_OF_VALUE) ETH: sectors=(LAYER1, DEFI), roles=(GAS, STORE_OF_VALUE, GOVERNANCE) BNB: sectors=(EXCHANGE, LAYER1), roles=(EXCHANGE_FEE, GAS) DOGE: sectors=(MEME, CURRENCY), roles=(MEME, GAS) DOT: sectors=(LAYER1), roles=(GAS, GOVERNANCE) UNI: sectors=(DEFI), roles=(GOVERNANCE, UTILITY) ``` #### Querying ```python from malkhut.training.asset_classification import * # Multi-label queries (match if queried value is ANY of the asset's labels) layer1s = get_assets_by_sector(Sector.LAYER1) # ETH, SOL, ADA, AVAX, BNB, DOT, ATOM gas_tokens = get_assets_by_token_role(TokenRole.GAS) # ETH, SOL, ADA, AVAX, DOGE, BNB, MATIC, DOT, ATOM # Predictive properties forced = get_pov_assets() # BTC, DOGE (PoW miners with forced selling) inelastic = [p for p in ASSET_PROFILES.values() if p.demand_elasticity == "inelastic"] # Backward-compatible single-label access (first element = primary) btc = get_asset_profile("BTCUSDT") btc.sector # Sector.CURRENCY btc.token_role # TokenRole.STORE_OF_VALUE ``` ### Exchange Registry (Multi-Exchange Asset Universe) MALKHUT serves as the **system-wide asset universe store** for BLUE, VIOLET, UV, and all downstream systems. Each asset knows which exchanges trade it; each exchange has a standardized profile. #### ExchangeProfile | Field | Type | Purpose | |-------|------|---------| | `exchange_id` | str | Canonical key: `"binance"`, `"bingx"`, `"bybit"` | | `display_name` | str | Human-readable name | | `has_spot` | bool | Spot trading available | | `has_perps` | bool | Perpetual futures available | | `has_options` | bool | Options available | | `api_base_url` | str | REST API root URL | | `ws_base_url` | str | WebSocket root URL | | `default_taker_fee_bps` | float | Default taker fee | | `default_maker_fee_bps` | float | Default maker fee | | `typical_latency_ms` | float | Typical API latency | #### Pre-defined Exchanges | Exchange | Spot | Perps | Options | Taker Fee | Latency | |----------|------|-------|---------|-----------|---------| | **Binance** | ✓ | ✓ | ✓ | 0.4 bps | 40 ms | | **BingX** | ✓ | ✓ | ✗ | 0.5 bps | 100 ms | | **Bybit** | ✓ | ✓ | ✓ | 0.06 bps | 50 ms | #### Exchange-Asset Mapping Each `AssetProfile` carries an `exchanges: tuple[str, ...]` field listing which venues trade the asset. Default: `("binance",)`. Other systems (BLUE/VIOLET/UV) import their asset universes into this store. ```python from malkhut.training.asset_classification import * # Which exchanges trade BTC? btc = get_asset_profile("BTCUSDT") print(btc.exchanges) # ('binance',) # All assets on Binance binance_assets = get_assets_on_exchange("binance") # Assets traded on BOTH Binance and BingX common = get_common_assets("binance", "bingx") # Which exchanges trade a given asset? venues = get_exchange_for_asset("ETHUSDT") # Exchange metadata ex = get_exchange("binance") print(ex.default_taker_fee_bps) # 0.4 print(ex.typical_latency_ms) # 40 ``` ### Asset Behavior DSL (10-Dimension Research-Validated Model) The Asset Behavior DSL decomposes each asset's **trading behavior** into 10 orthogonal dimensions, each with empirically-validated parameters from live Binance/BingX API data and academic literature. **Source:** Binance live REST API (depth 500 levels, 1h klines, ticker, funding rates), BingX open API (depth, funding, premium index), academic literature (Bouchaud "Trades Quotes and Prices" 2018, Cont/Stoikov/Talreja 2010, Cartea/Jaimungal/Penalva 2015), CoinGlass OI data, industry MM disclosures. #### 10 Dimensions | Dimension | What it captures | Example: BTC vs DOGE | |-----------|------------------|---------------------| | **DepthProfile** | Book shape: amplitude, decay exponent α, fragility | BTC: $750K/0.70/0.10 vs DOGE: $22K/1.00/0.30 | | **SpreadProfile** | Normal spread, stress multiplier | BTC: 0.01bps/50x vs DOGE: 1.35bps/15x | | **FlowProfile** | Order rate, size distribution, cancel ratio | BTC: 300/s/$643/20x vs DOGE: 80/s/$96/8x | | **VolatilityProfile** | Annualized vol, GARCH params, half-life | BTC: 35%/48h vs DOGE: 78%/24h | | **IntradayProfile** | Peak/trough hours, ratio | BTC: 15:00 UTC, 7.4x ratio | | **WeekendProfile** | Vol/volume/spread multipliers | 0.65x vol, 0.52x volume, 1.20x spread | | **CorrelationProfile** | ETH beta, BTC corr (normal vs crash) | BTC: 1.0/1.0 vs UNI: 0.70/0.50 | | **MarketMakerProfile** | Inventory limits, pull speed, margins | BTC: $10M/3ms/0.5bps vs DOGE: $500K/25ms/2bps | | **LiquidationProfile** | OI/MCap, trigger %, cascade speed | BTC: 0.5%/6.5%/slow vs DOGE: 1.4%/4%/fast | | **FundingProfile** | Mean/std funding, basis | BTC: 0.59bps/0.22bps vs DOGE: 0.39bps/0.34bps | | **RetailProfile** | Retail ratio, inst gap | BTC: 35%/0.04 vs DOGE: 80%/0.31 | | **BingxProfile** | BingX spread/depth multiplier, fees, latency | BTC: 12.6x/0.048x vs DOGE: 7x/1.27x | #### 3 Composable Templates | Template | Assets | Key traits | |----------|--------|-----------| | `institutional_blue_chip` | BTC, ETH, BNB | Deep book, low spread, slow cascade, high institutional | | `mid_cap_l1` | SOL, ADA, AVAX, DOT, ATOM, UNI, LINK, MATIC, AAVE | Moderate depth, higher vol, medium cascade | | `retail_meme` | DOGE | Thin book, high vol, fast cascade, retail-dominated | #### 13 Per-Asset Behaviors (research-validated) | Asset | Price | Depth@1bps | Spread | Vol | Cascade | Retail | BingX mult | |-------|-------|-----------|--------|-----|---------|--------|------------| | BTC | $64K | $750K | 0.01 bps | 35% | slow/fast | 35% | 12.6x | | ETH | $1.8K | $600K | 0.01 bps | 66.5% | medium/medium | 35% | 9.0x | | SOL | $80 | $400K | 1.26 bps | 72% | medium/medium | 72% | 1.7x | | DOGE | $0.07 | $22K | 1.35 bps | 78% | fast/slow | 80% | 7.0x | | ADA | $0.17 | $60K | 5.95 bps | 79% | medium/medium | 65% | 8.0x | | UNI | $3.6 | $30K | 2.76 bps | 95% | medium/medium | 60% | 10x | #### Usage ```python from malkhut.training.asset_behavior import * # Get behavior for any asset btc = get_behavior("BTCUSDT") print(btc.depth.amplitude_usd) # $750,000 print(btc.vol.annualized_normal) # 35.0% print(btc.liquidation.speed) # "slow" # Compute depth at distance depth_10bps = btc.depth_at_bps(10) # $1.5M # Estimate slippage for a $100K order slippage = btc.expected_slippage_bps(100_000) # ~2 bps # Query by template, sector, or role institutional = get_behaviors_by_template("institutional_blue_chip") gas_tokens = get_behaviors_by_role(TokenRole.GAS) thin_book = get_thin_book_assets() ``` ### Asset Compiler (Auto-Fetch from Live APIs) The Asset Compiler automatically fetches market data from Binance and BingX public APIs and produces system-ready `AssetProfile` + `AssetBehavior` for any tradeable symbol. **Rate-limited** (1 req/sec, well under Binance 1200/min limit), **cached** (avoids re-fetching within a session), **resumable**. #### What it auto-fetches | Endpoint | Data | Computes | |----------|------|----------| | `/api/v3/ticker/24hr` | Price, volume, high/low | Daily volume, price reference | | `/api/v3/depth?limit=100` | Order book (100 levels) | Spread, depth amplitude, decay exponent α | | `/api/v3/klines?interval=1h&limit=168` | 7 days hourly OHLCV | Annualized volatility, order flow stats | | `/api/v3/exchangeInfo` | Symbol filters | Tick size, lot size, price decimals | | `/fapi/v1/fundingRate` | Funding rates (20 intervals) | Mean/std/positive% funding | | `/fapi/v1/openInterest` | Open interest | OI/MCap ratio for liquidation modeling | #### Usage ```python from malkhut.training.asset_compiler import AssetCompiler compiler = AssetCompiler() # Single asset — auto-fetches from Binance, computes all params result = compiler.compile("XRPUSDT") print(result.asset_profile.symbol) # "XRPUSDT" print(result.asset_behavior.reference_price) # $1.1056 (live Binance price) print(result.asset_behavior.spread.normal_bps) # 0.90 bps (computed from depth) print(result.asset_behavior.vol.annualized_normal) # 46.6% (computed from klines) # Register into global registries compiler.register(result) # Now XRPUSDT is available for ScenarioFactory, queries, etc. # One-step compile and register result = compiler.compile_and_register("HBARUSDT") # Batch compile (rate-limited sequentially) results = compiler.compile_batch(["APTUSDT", "SEIUSDT", "INJUSDT"]) for r in results: compiler.register(r) ``` #### Known classifications The compiler includes pre-built classifications for 28 assets: BTC, ETH, SOL, DOGE, ADA, AVAX, UNI, LINK, BNB, MATIC, AAVE, DOT, ATOM, XRP, TRX, LTC, NEAR, APT, OP, ARB, SUI, PEPE, WIF, SEI, INJ, FIL, HBAR, IMX Unknown assets get heuristic defaults (LAYER1/GAS/INFLATIONARY/POS/FULL/PERPS_ONLY). ### Scenario Factory (Behavior-Driven, Label-Queryable) The ScenarioFactory builds evaluation scenarios using **real per-asset prices, depth profiles, and spread characteristics** — no more hardcoded BTC prices. #### Behavior-driven state creation Every scenario derives its parameters from the asset's `AssetBehavior`: | Scenario type | Spread multiplier | Depth fraction | Description | |---------------|------------------|----------------|-------------| | `normal` | 1.0x | 100% | Calm, liquid market | | `thin_book` | 1.5x | 10% | Low participation | | `wide_spread` | 100x | 50% | High volatility | | `toxic_stress` | 2.0x | 30% | Adverse selection | | `flash_crash` | 3.0x | 5% | Book collapse | | `liquidity_vacuum` | 5.0x | 1% | Near-empty book | | `weekend_thin` | 5.0x | 5% | Weekend low-volume | | `trending` | 1.0x | 80% | Momentum move | | `mean_reverting` | 20x | 60% | Reversal setup | | `funding_shock` | 1.5x | 30% | Deleveraging event | | `liquidation_cascade` | 1.5x | 20% | Chain liquidation | | `market_maker_withdrawal` | 3.0x | 15% | MM pull-out | #### Realistic per-asset pricing ``` Scenario: "normal_BTCUSDT_42" BTC bid: $63,999.97 (11.7 BTC = $750K) BTC ask: $64,000.03 Spread: 0.01 bps Scenario: "normal_DOGEUSDT_42" DOGE bid: $0.07 (314K DOGE = $22K) DOGE ask: $0.07 Spread: 1.35 bps Scenario: "flash_crash_SOLUSDT_47" SOL bid: $80.00 (depth = 5% of normal = $20K) SOL ask: $80.02 Spread: 3.78 bps (3x normal) ``` #### Auto-compilation for unknown assets When you request a symbol that isn't pre-defined, the factory **automatically compiles it from live Binance API data** before building scenarios: ```python factory = ScenarioFactory() # XRPUSDT is NOT in our pre-defined 13 assets # But this auto-compiles it from Binance in ~5 seconds: suite = factory.build_suite_for_symbols(("XRPUSDT",), steps_per_scenario=5) # XRP behavior is now registered for future use too ``` #### Label-based query interfaces ```python factory = ScenarioFactory() # By sector suite = factory.build_suite_for_sector(Sector.LAYER1) # ETH, SOL, ADA, AVAX, BNB, DOT, ATOM suite = factory.build_suite_for_sector(Sector.DEFI) # ETH, AVAX, UNI, AAVE suite = factory.build_suite_for_sector(Sector.MEME) # DOGE # By token role suite = factory.build_suite_for_role(TokenRole.GAS) # 9 assets suite = factory.build_suite_for_role(TokenRole.GOVERNANCE) # ETH, UNI, AAVE, DOT # By behavior template suite = factory.build_suite_for_template("institutional_blue_chip") # BTC, ETH, BNB suite = factory.build_suite_for_template("retail_meme") # DOGE # By volatility band suite = factory.build_suite_for_volatility(min_ann=75, max_ann=200) # High-vol assets # Multi-label union (any matching) suite = factory.build_suite_for_labels( sectors=[Sector.DEFI], roles=[TokenRole.GOVERNANCE] ) # Arbitrary symbols (auto-compiles unknowns) suite = factory.build_suite_for_symbols(("BTCUSDT", "XRPUSDT", "HBARUSDT")) ``` #### Research-validated scenario diversity Each asset's 30 scenario types are parameterized by its behavior profile: | Scenario | BTC params | DOGE params | Why it matters | |----------|-----------|-------------|----------------| | Normal | $750K depth, 0.01bps spread | $22K depth, 1.35bps spread | Baseline for scoring | | Thin book | $75K depth | $2.2K depth | Tests fill rate under low liquidity | | Flash crash | $37.5K depth, 0.03bps spread | $1.1K depth, 4bps spread | Tests adverse selection survival | | Weekend thin | $37.5K depth, 5bps spread | $1.1K depth, 6.75bps spread | Tests off-hours strategy fitness | | Liquidation cascade | $150K depth | $4.4K depth | Tests position sizing under stress | | Market maker withdrawal | $112.5K depth | $3.3K depth | Tests quote quality when MMs pull | ### Numba Performance Hot-path functions JIT-compiled with numba: | Function | Speedup | |----------|---------| | fill_from_levels (batch 100) | 1.8x | | round_tick | JIT-compiled | | round_lot | JIT-compiled | | clip_lots | JIT-compiled | | extract_features | Vectorized numpy | CWM throughput: **157K calls/sec**, **6.4 µs/transition**, **0.64 ms/100-step episode**. ### Design Principles 1. **Hardcoded baseline is NEVER replaced** — always available as reference 2. **Generated strategies are ADDED** — system GROWS its repertoire 3. **Crossover** combines two strategies (uniform crossover on parameters) 4. **Mutation** randomly modifies parameters (Gaussian noise) 5. **Selection** uses tournament selection (fitter genomes win more often) 6. **Third parties** can write strategies in DSL text format 7. **Market adaptation**: selector chooses strategy based on ExoF/MARAS regime 8. **Behavior-driven scenarios**: every asset behaves like its real-world self ### Strategy Naming Convention ``` {strategy_type}_gen{generation}_{timestamp} ``` Examples: `UCB1_gen1_20260707_223754`, `THOMPSON_gen2_20260707_223754` --- ## H. Cambrian Expansion ### Tie-In Fix PerformanceMatrix now wired to evaluator. Strategies scored by regime fitness: - ALL strategies tested in ALL scenarios - Score tracked per strategy × regime - Selector queries matrix for best strategy per regime - `PerformanceMatrix.record(strategy_id, regime, score)` called during evaluation ### Phase 0.1: Cognition Pipeline Rate-limited market regime research: - `RateLimiter`: token bucket, 30 RPM default - `SourceCatalogue`: tracks sources, fetch counts, error rates - `RegimeExtractor`: keyword matching + sentiment analysis - `CognitionPipeline`: rate-limited, deduplication, long/perm-run - 8 default sources: CoinDesk, Cointelegraph, The Block, CryptoQuant, Glassnode, Coinglass, Binance Research, Messari ### Phase 0.2: Exponential Regime Expansion Orthogonal to cognition pipeline — generates 200+ regimes from dimensions: - 4 liquidity (vacuum, thin, normal, deep) - 4 volatility (tight, normal, wide, extreme) - 4 flow (balanced, buy_pressure, sell_pressure, toxic) - 4 structure (single_toxic, mm_toxic, multi_toxic, retail_mm) Each combination = distinct, testable regime. Total: 256 theoretical, ~200 practical. ### Phase 0.3: Multi-Asset & Invariant Classification 13 assets across 10 sectors, classified by **invariant characteristics only**: | Asset | Sectors | Roles | Supply | Consensus | Vol | Liq | |-------|---------|-------|--------|-----------|-----|-----| | BTC | CURRENCY | STORE_OF_VALUE | FIXED_CAP | POW | LOW | DEEP | | ETH | LAYER1, DEFI | GAS, SOV, GOV | DISINFLATIONARY | POS | MEDIUM | DEEP | | SOL | LAYER1 | GAS | INFLATIONARY | POS | HIGH | NORMAL | | DOGE | MEME, CURRENCY | MEME, GAS | INFLATIONARY | POW | HIGH | NORMAL | | ADA | LAYER1 | GAS | INFLATIONARY | DPOS | MEDIUM | NORMAL | | AVAX | LAYER1, DEFI | GAS | INFLATIONARY | POS | HIGH | NORMAL | | UNI | DEFI | GOV, UTILITY | INFLATIONARY | POS | HIGH | THIN | | LINK | ORACLE | UTILITY | INFLATIONARY | POS | MEDIUM | THIN | | BNB | EXCHANGE, LAYER1 | EXCHANGE_FEE, GAS | BURN | DPOS | MEDIUM | DEEP | | MATIC | LAYER2 | GAS | INFLATIONARY | POS | HIGH | THIN | | AAVE | DEFI | GOV | FIXED_CAP | POS | HIGH | THIN | | DOT | LAYER1 | GAS, GOV | INFLATIONARY | DPOS | MEDIUM | NORMAL | | ATOM | LAYER1 | GAS | INFLATIONARY | POS | HIGH | THIN | Multi-asset = linear multiplication: 30 scenarios × 13 assets = 390+ testable scenarios. ### Phase 0.4: Asset Behavior DSL + Compiler + Behavior-Driven Scenarios **Research-validated** from Binance live API, BingX API, CoinGlass, and academic literature: | Component | Purpose | File | |-----------|---------|------| | **AssetBehavior** | 10 orthogonal behavior dimensions | `training/asset_behavior.py` | | **BehaviorTemplate** | Reusable class-level behavior profiles | `training/asset_behavior.py` | | **AssetCompiler** | Auto-fetch from Binance/BingX, compute profiles | `training/asset_compiler.py` | | **ScenarioFactory** | Behavior-driven scenarios with label queries | `training/cma_trainer.py` | **Key research findings wired into the system:** - Depth decay: D(d) = A * d^(1-alpha). BTC alpha=0.7, DOGE alpha=1.0 - Spread ranges: BTC 0.01bps, SOL 1.26bps, DOGE 1.35bps, ADA 5.95bps - Order sizes: BTC P50=$643, DOGE P50=$96 (7x smaller) - Cascade dynamics: BTC 5-8% trigger/slow/fast recovery, DOGE 3-5%/fast/slow - BingX specifics: 1.7-12.6x wider spreads, 0.05-9x variable depth - Weekend effects: -35% vol, -48% volume, +20% spread - Intraday patterns: 7.4x peak/trough ratio, US session dominant - GARCH persistence: 0.95-0.99, vol half-life 2-5 days - Correlation: BTC-ETH 0.90 normal, 0.97 crash; BTC-DOGE 0.45 normal, 0.80 crash ### Training Scaling Results CMA-ES evaluation on behavior-driven multi-asset scenarios (BTC/ETH/SOL, 3 assets × 30 scenario types = 90 scenarios). Each eval = one CMA candidate × 90 multi-step episodes through the CWM. #### Population × Budget Grid | Config | Pop | Gens | Evals | Time | Best Score | Mean PnL | Eval Rate | |--------|-----|------|-------|------|-----------|----------|-----------| | pop12_b48 | 12 | 4 | 48 | 615s | 3,253 | 31.1 bps | 0.08 e/s | | pop12_b96 | 12 | 8 | 96 | 1246s | 7,202 | 58.8 bps | 0.08 e/s | | pop20_b48 | 20 | 2 | 48 | 639s | 4,145 | 35.8 bps | 0.08 e/s | #### Scaling Laws - **Budget (evals) is the primary driver.** 48→96 evals → 2.2× score (near-linear). No saturation observed at 96 evals — system would keep improving with more. - **Population helps at the margin.** pop 12→20 at same budget → 1.3× score. CMA explores more diverse candidates per generation. - **Eval rate is constant at 0.08 e/s** regardless of pop — bottleneck is CWM episode execution (13s/eval for 90 scenarios), not CMA overhead. - **PnL tracks score closely.** 31 bps → 59 bps (2× budget → 2× PnL). - **Score is not saturated at 96 evals.** Extrapolation: ~74 score/eval at pop12. 500 evals ≈ $37K score, ~110 min. 5000 evals ≈ ~18 hours. #### Training Performance | Metric | Sequential | Parallel (8 workers) | Speedup | |--------|-----------|---------------------|---------| | CWM throughput | 189K calls/sec | (bottleneck is planner, not CWM) | — | | CWM per-call latency | 5.3 µs | (already numba-optimized) | — | | Scenario generation | 390 scenarios in 0.8s | — | — | | Single eval (90 scenarios) | 3.3s | **0.4s** | **9.1x** | | CMA-ES per-generation (pop=12) | ~155s | **~21s** | **~7x** | | Best score achieved | 7,202 (96 evals, pop=12) | same (fidelity preserved) | — | | Best mean PnL | 58.8 bps | same | — | | CMA-ES 48 evals (4 gens) | ~125s (est) | **85s** | **1.5x** | Parallel evaluation achieves **9x speedup on single evals** (embarrassingly parallel, zero fidelity loss). CMA-ES training speedup is ~1.5x because multiprocessing overhead is amortized across 90 scenarios per eval. Bottleneck is planner (MCTS), not CWM — numba already accelerates the CWM hot path (`_HAS_NUMBA = True`). ### Prod Tooling | Component | Purpose | File | |-----------|---------|------| | **CognitionLauncher** | Standalone long-run script | `cognition_launcher.py` | | **CognitionMonitor** | Metrics, health scoring, alerts | `training/monitor.py` | | **NewsSourceRepository** | 12 industry-standard sources, ranking | `training/news_sources.py` | | **SourceCatalogue** | Source persistence, fetch tracking | `training/cognition.py` | | **RateLimiter** | Token bucket, 30 RPM | `training/cognition.py` | | **RegimeExpander** | 200 regimes from dimensions | `training/regime_expansion.py` | ### Running ```bash # Training pipeline (continuous) python -m malkhut.continuous_pipeline # Cognition pipeline (continuous regime research) python -m malkhut.cognition_launcher ```