HftBacktestCWM (cwm/hft_cwm.py): - PowerProbQueueModel: probabilistic fill per level (pre-computed) - Level 0 always fills, deeper levels have decreasing probability - Deterministic fallback when use_queue_model=False - Same transition/reward/terminal API as MinimalCryptoLOBCWM - Fallback to deterministic level consumption when hftbacktest unavailable 59 tests (test_hft_cwm.py) covering 15 test classes: 1. Queue model correctness (fill probs, monotonic, bounds, determinism) 2. Determinism & reproducibility 3. CWM interface compatibility (cross, place, cancel, post_only, reduce) 4. Reward function (profit, noop, maker bonus) 5. Edge cases (empty book, zero qty, extreme price, many levels) 6. Position tracking (buy, sell, flip) 7. Fee application (taker fee reduces equity) 8. Counterparty ecology (toxic taker hits book, noop preserves) 9. CWM comparison (Hft vs Minimal agree on noop) 10. Venue propagation (scenario tagging, cross-exchange transfer) 11. PerformanceMatrix venue keying (record, per-venue best, comparison) 12. Risk gate integration (approve, leverage, OOD, kill switch, self-trade) 13. Stress tests (rapid transitions, 20 open orders, cancel all) 14. Full episode integration (single episode runs, policy evaluator) 15. hftbacktest availability check
MALKHUT — Adversarial Self-Play Order-Fulfilment Pipeline
Lock-free. No GC. Fully async. GraalVM-compatible.
MALKHUT is an order-fulfilment engine that optimises a distribution over execution actions against a diverse counterparty ecology, not a single deterministic quote.
Core Thesis
Sequential/pure optimisation ("what is the best quote?") leads to deterministic quote → predictable fill pattern → adverse selection.
Correct simultaneous-market framing: "what mixed order-placement policy survives a diverse counterparty ecology?"
Architecture
Corrected Stack (T19 ANNEX B-CORRECTION)
iceoryx2 = the wires; ASEx = the state cells; UV clock = the conductor.
| Layer | Component | Role |
|---|---|---|
| Transport | iceoryx2/zinc topics | Inter-component event fabric. Observable in /dev/shm. Single-writer per topic. |
| State | ASEx (ASExGuardedState + ASExWorker) | Validate-before-mutate state serialization. One writer thread per state. NOT a message bus. |
| Dispatch | UV Clock (asyncio event loop) | Event-driven reactor. Subscribes to event types. Nothing sleeps-and-polls. |
Per T19: ASExWatch is DEFERRED from the critical path until heavy-test hangs are fixed. Use ASExWorker (FulfilmentWorker) for production hot-path mutations.
┌──────────────────────────────────────────────────────────────────────┐
│ Zinc SHM (POSIX /dev/shm/zinc_*) │
│ malkhut_book_state | malkhut_account_state | malkhut_fulfilment_out │
│ malkhut_risk_gate | malkhut_control (CONTROL_PLANE) │
└───────────────────────────────┬──────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────────────┐
│ ASEx VALIDATE-BEFORE-MUTATE │
│ GuardedFulfilmentState — book/account/intent/policy mutations │
│ GuardedRiskState — risk gate + kill switch │
│ FulfilmentWorker — single-threaded ASExWorker per state │
│ FulfilmentWatch — zero-overhead ring buffer for hot path │
│ ShardedWorker — per-symbol state partitioning │
└───────────────────────────────┬──────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────────────┐
│ CODE WORLD MODEL (CWM) │
│ deterministic exchange transition function │
│ f(state, joint_actions, latency_model, venue_rules) → next_state │
│ wraps hftbacktest for replay verification │
└───────────────────────────────┬──────────────────────────────────────┘
│
┌──────────┴──────────┐
▼ ▼
┌─────────────────────────┐ ┌──────────────────────────────────┐
│ LIVE FULFILMENT PLANNER │ │ OFFLINE CMA-ES TRAINER │
│ Decoupled UCB/UCT │ │ pycma + nevergrad │
│ ≤ 25 ms cadence │ │ growing self-play pool │
└───────────┬──────────────┘ └─────────────────┬────────────────┘
│ │
▼ ▼
┌─────────────────────────┐ ┌──────────────────────────────────┐
│ RISK / COMPLIANCE GATE │ │ POLICY REGISTRY / POOL │
│ leverage, notional, │ │ accepted candidates, weights │
│ self-trade, kill switch │ │ diversity-preserving eviction │
└───────────┬──────────────┘ └──────────────────────────────────┘
│
▼
┌─────────────────────────┐
│ VENUE ADAPTER (BingX) │
│ create/cancel/replace │
└───────────┬──────────────┘
│
▼
┌─────────────────────────┐
│ ClickHouse Persistence │
│ decisions, episodes, │
│ snapshots, discrepancies│
└─────────────────────────┘
Package Structure
MALKHUT/
├── malkhut/
│ ├── __init__.py # package version
│ ├── state.py # frozen data model (42 dataclasses)
│ ├── actions.py # action model, PlannedPolicy, RiskDecision
│ ├── features.py # feature extraction (17 features)
│ ├── counterparties.py # adversarial agent ecology (4 agents)
│ ├── engine.py # FulfilmentEngine (hot-path orchestrator)
│ ├── bench_numba.py # numba speedup benchmark
│ ├── launch_smoke_test.py # 10-min smoke test launcher
│ ├── cwm/
│ │ ├── __init__.py
│ │ ├── core.py # CWM (deterministic exchange simulator)
│ │ ├── numba_core.py # numba-accelerated hot loops
│ │ └── replay_verify.py # replay verification (mandatory before trust)
│ ├── planner/
│ │ ├── __init__.py
│ │ ├── action_menu.py # compact action space construction
│ │ └── sm_mcts.py # Decoupled UCB/UCT simultaneous-move MCTS
│ ├── risk/
│ │ ├── __init__.py
│ │ └── gate.py # risk/compliance gate (hard invariants)
│ ├── venue/
│ │ ├── __init__.py
│ │ └── bingx/
│ │ ├── __init__.py
│ │ └── adapter.py # BingX VST/Live venue adapter (wraps DITAv2)
│ ├── ipc/
│ │ ├── __init__.py
│ │ ├── zinc_plane.py # Zinc shared memory IPC (POSIX SHM)
│ │ └── control_plane.py # CONTROL_PLANE shared memory region
│ ├── storage/
│ │ ├── __init__.py
│ │ ├── ch_store.py # ClickHouse persistence (5 tables)
│ │ └── asset_store.py # DuckDB file-backed asset universe store
│ ├── execution/
│ │ ├── __init__.py
│ │ └── asex_integration.py # ASEx validate-before-mutate kernel
│ ├── clock/
│ │ ├── __init__.py
│ │ ├── host.py # UVClock — event-driven reactor (T19)
│ │ ├── events.py # Typed events: Scan, Tick, Timer, BarFire, Stale
│ │ ├── staleness.py # Freshness watchdog (T19 Staleness Law)
│ │ └── deadnode.py # DeadNode reaper for iox2 orphans
│ ├── training/
│ │ ├── __init__.py
│ │ ├── cma_trainer.py # CMA-ES + self-play + parallel workers
│ │ ├── registry.py # Policy lifecycle (CANDIDATE → ACTIVE)
│ │ ├── pipeline.py # Bounded continuous learning loop + logger
│ │ ├── dsl.py # Strategy DSL v2 (40+ primitives, 40+ sensors)
│ │ ├── generator.py # Genetic programming strategy evolution
│ │ ├── order_types.py # Standardized order type enum (FIX/CCXT-aligned)
│ │ ├── selector.py # Regime → strategy mapping + performance matrix
│ │ ├── asset_classification.py # Multi-label taxonomy + exchange registry
│ │ ├── asset_behavior.py # 10-dimension behavior DSL, research-validated
│ │ ├── asset_compiler.py # Auto-fetch from Binance/BingX, compile profiles
│ │ ├── asset_bridge.py # Directory ↔ classification sync
│ │ ├── parallel_eval.py # ProcessPoolExecutor episode runner
│ │ ├── ray_eval.py # Ray-based eval (industrial alternative)
│ │ ├── vbt_analysis.py # Post-sim metrics: Sharpe, Sortino, VaR
│ │ ├── advantage_scorer.py # Advantage estimation for offline analysis
│ │ ├── cognition.py # Rate-limited market regime research
│ │ ├── regime_expansion.py # 200+ regimes from dimension combinations
│ │ ├── news_sources.py # 12 industry-standard news sources
│ │ └── monitor.py # Cognition metrics, health, alerts
│ ├── daat/ # Direction-Anchored Ambiguity Triage
│ │ ├── __init__.py
│ │ └── core.py # DaatQuery, DaatVerdict, daat_classify
│ ├── cognition_launcher.py # Standalone long-run cognition service
│ ├── continuous_pipeline.py # Continuous training pipeline
│ └── tests/ # 1178 tests across 50 test files
├── specs/
│ └── MALKHUT_ADVERSARIAL_SELFPLAY_SPEC.py # full spec
└── README.md # this file
Shared Memory IPC (Zinc)
All IPC uses real Zinc shared memory (POSIX SHM via /dev/shm/zinc_*), NOT
the file-based transport. This is lock-free, zero-copy, GraalVM-compatible.
Regions
| Region | Writer | Readers | Purpose |
|---|---|---|---|
malkhut_book_state |
data feed | planner, CWM | canonical order book snapshot |
malkhut_account_state |
data feed | planner, risk | account/position snapshot |
malkhut_fulfilment_out |
engine | other systems | planner output + distribution |
malkhut_risk_gate |
engine | other systems | risk gate decisions |
malkhut_control |
external | engine | commands, symbols, lifecycle |
Framing
All regions use UVZINC01 seqlock framing:
[8B magic "UVZINC01"] [8B seq] [8B json_size] [UTF-8 JSON body]
CONTROL_PLANE
The malkhut_control region is the management interface:
- External systems write
ControlCommandframes - Engine reads and processes commands
- ACKs written back for
ack_requiredcommands
Commands: START, STOP, PAUSE, RESUME, SET_SYMBOLS, CONNECT_VENUE,
DISCONNECT_VENUE, SET_MODE, EMERGENCY_STOP, HOT_RELOAD_POLICY,
STATUS_REQUEST
ClickHouse Persistence
Database: dolphin_malkhut on localhost:8123
| Table | Purpose |
|---|---|
replay_steps |
CWM replay traces |
fulfilment_decisions |
every live decision |
self_play_episodes |
training episode results |
policy_snapshots |
versioned policy registry |
live_discrepancies |
simulation vs actual |
External Dependencies (Spec-Mandated)
| Library | Purpose | Used For |
|---|---|---|
pycma (cma) |
CMA-ES optimisation | training/cma_trainer.py |
| hftbacktest | replay/queue/latency fill simulation | CWM replay verification |
| nevergrad | gradient-free optimiser zoo | trainer comparison |
| msgspec | fast typed serialization | state/event encoding |
| numba | hot-loop acceleration | CWM transitions, feature extraction |
| polars | offline analytics | large trade-path corpora |
| lmdb | existing SILOQY storage | ticks, candles, hd_vectors |
| numpy | vectorized feature extraction | parameter decoding |
| scipy | stats tests, bootstrap | CI, promotion gates |
Quick Start
from malkhut.engine import FulfilmentEngine
from malkhut.state import FulfilmentPolicyParams
def my_params():
return FulfilmentPolicyParams(
version="v0.1", ucb_c=1.414, max_sims=64, max_depth=3,
rollout_depth=3, root_temperature=0.5, min_root_entropy=0.25,
quote_offsets_ticks=(0, 1, 2), quote_size_fractions=(0.10, 0.25, 0.50),
passive_ttl_ms=200, aggressive_ttl_ms=50,
maker_edge_min_bps=0.5, cross_spread_edge_min_bps=5.0,
adverse_toxicity_cancel_threshold=0.5, queue_churn_cancel_threshold=0.5,
mae_tail_cut_bps=50.0, mfe_giveback_cut_fraction=0.5,
max_time_in_loss_s=300.0, failed_recovery_cut_count=3,
recovery_velocity_min_bps_per_s=0.0,
max_symbol_notional_fraction=0.20, max_single_order_notional_fraction=0.05,
reduce_when_global_up_fraction=0.30, session_profit_lock_fraction=0.02,
w_expected_pnl=1.0, w_fill_probability=0.5, w_adverse_selection=2.0,
w_queue_priority=0.5, w_inventory_risk=1.5, w_tail_loss=5.0,
w_fee_quality=0.5, w_time_decay=0.3, w_policy_entropy=0.5,
robust_tail_weight=2.0, toxic_counterparty_weight=3.0,
low_liquidity_weight=2.0, latency_stress_weight=1.0,
)
engine = FulfilmentEngine(params_provider=my_params)
# engine.on_state(latest_market_state) # hot path
ASEx Integration (Validate-Before-Mutate)
MALKHUT uses ASEx (Advanced Serial Executor) for all mutable state mutations. This is the same kernel used by VIOLET/DOLPHIN for race-proof state management.
Core Pattern
class GuardedFulfilmentState(ASExGuardedState):
def _validate(self, mutation) -> bool:
# Pure predicate — MUST NOT mutate self
return mutation.ts_ns > 0 and len(mutation.bids) > 0
def _apply(self, mutation):
# Only called after _validate() passes
# MAY mutate self
self._state = new_state
return result
Components
| Component | Purpose | Thread Safety |
|---|---|---|
GuardedFulfilmentState |
Engine state mutations (book, account, intent, policy) | One writer thread via ASExWorker |
GuardedRiskState |
Risk gate + kill switch | One writer thread via ASExWorker |
FulfilmentWorker |
Single-threaded ASExWorker wrapping GuardedFulfilmentState | Serialised via queue.Queue |
RiskWorker |
Single-threaded ASExWorker wrapping GuardedRiskState | Serialised via queue.Queue |
FulfilmentWatch |
Zero-overhead ring buffer for hot-path mutations | Producer/consumer lock-free |
ShardedWorker |
Per-symbol state partitioning (N independent workers) | No shared state between partitions |
Why ASEx
- No locks: One thread per state object. Races impossible by construction.
- No GC pressure: Frozen dataclasses + queue.Queue (CPython C-level ops).
- Validate before apply: Invalid mutations rejected without touching state.
- GraalVM compatible: No CPython-only APIs in the hot path.
- Battle-tested: 110 ASEx tests + 1109 MALKHUT tests = 1219 total.
Design Rules (from spec)
- Optimise an ecology-resistant distribution, not a quote.
- The book is not passive scenery. Ask "why did I get filled?"
- Replay correctness before search depth. Wrong CWM + deep search = confident nonsense.
- CMA-ES does not promote from noisy wins. Survive scenario suite + bootstrap CI.
- Preserve feature humility. Allow CMA-ES to discover non-obvious interactions.
- Path-aware SL/TP is part of fulfilment. Exiting = order-placement under adversarial liquidity.
- Never let the optimiser learn illegal behavior. Hard-code constraints.
- Keep hot path bounded. ≤ 100 ms. No live CMA, no large neural nets.
- Version everything. Every decision explainable by policy + CWM + feature + state + action + outcome.
- If shadow live diverges, trust live. Simulation is only a tool.
ADDENDUM — EXEC-STACK PLUGGABILITY + SPEC HARDENING
[Fable, integrator, 2026-07-06 — operator-ordered extension. Mimo: this section binds MALKHUT into the UV/DITAv2 stack; nothing above is overridden except where marked FIX.]
A. Where MALKHUT plugs in (the seat at the table)
MALKHUT is the execution layer under the UV Clock (see
prod/docs/uv_subspecs/UV_TASK_T19_UV_CLOCK_HOST.md + annexes):
UV Clock (T19) → PromotionBridge → KernelIntent (DITAv2 contracts)
→ MALKHUT FulfilmentEngine (replaces naive kernel-submit)
→ RISK GATE → BingX venue adapter → fills
Your Design Rule 6 ("path-aware SL/TP is part of fulfilment") is CONVERGENT with the operator's T19 ruling ("the whole point of UV = faster-than-scan-cadence TP/SL paths"). MALKHUT's planner is the intelligent form of the C11 tick-exit guardian. Sequencing: simple doctrinal tick-exits (C11) ship first via T19 step 3; MALKHUT supersedes them when certified — same seat, smarter occupant.
B. Intake contract (BINDING)
- Input = DITAv2
KernelIntent(prod/clean_arch/dita_v2/contracts.py): asset, side (TradeSide), action (ENTER/EXIT), reference_price, target_size, leverage, metadata (carriespromo_client_id). Accept KernelIntent natively; do not invent a parallel intent type. u-clientOrderId prefix is LAW (T9 seam): every venue order MALKHUT places carries the u- prefix from intent metadata. Non-negotiable — it is the attribution boundary vs BLUE and vs legacy VIOLET orders.- DUAL-LEVERAGE LAW:
KernelIntent.leverageis CONVICTION leverage [0.5, 9.0] — it sizes QUANTITY only. Venue leverage = the mapping inprod/bingx/leverage.py(linear → ROUND_HALF_EVEN → int clamp [1,3]). WIRE the existing typed L3 wrapper (prod/clean_arch/violet/exchange_leverage.py); litmus: 0.5→1, 4.75→2, 8.0→3, 9.0→3. The venue adapter MUST NOT pass conviction leverage raw to set-leverage. - Arming/two-man/kill: MALKHUT honors
UV_PROMOTED=1+ arm-file (/root/uv-wt/prime-live/UV_PROMOTED.arm), re-evaluated PER EVENT (never cached — the T12 kill-switch lesson). DARK default = ObserveOnly venue. Map arm-file-delete → yourEMERGENCY_STOPsemantics: inert next event, no restart.
C. Fabric + state alignment (the tonight-rulings)
- Stack law: iceoryx2 = wires, ASEx = state cells, UV clock = conductor. Your ASEx usage (GuardedState + Worker per state object) is CORRECT — it is exactly what ASEx is for. Your zinc regions already speak UVZINC01 seqlock framing — good; when T19's event topics land, consume ScanEvent/TickEvent (with PROVENANCE: scan_number, scan_ts, ingest_ts, AGE) from the shared LiveSource instead of a private feed.
- Account state: read the DITAv2 ASEx account core (asex_account / AccountProjectionV2) as a CONSUMER only. MALKHUT never becomes a second writer to account state (two-cores-share-nothing law); its own guarded states are its own.
- Staleness law (T19 Annex A) — extend your Rule 2: a stale book is not scenery either, it is an ADVERSARY. Planner receives AGE on every input and must discount/refuse plans over stale state; freshness watchdog emits StaleInput.
- ASEx pre-condition inherited: DeadNode reaper / orphan-sweep on start (iox2 segments orphan on crash — known gap).
D. Persistence + observability alignment
dolphin_malkhutnamespace: approved (own-namespace law). NO TTL on any table (retention doctrine — audit surfaces never expire).- Cross-reference law: every
fulfilment_decisionsrow carriesintent_id+trade_idfrom the source KernelIntent → end-to-end trace joinsdolphin_uv.exec_journal⇄dolphin_malkhut.*⇄ venue order history. This is what makes MALKHUT certifiable instead of a black box. - Publish planner pulse (decision rate, plan entropy, gate verdicts, AGE-at-decision) into the zinc snapshot → extant TUIs render it (TUI-first doctrine).
E. Ecology + reward grounding (FIX-grade improvements)
- Ground the counterparty ecology in OUR tape: fit/validate adversary populations
against
dolphin.obf_universerecorded sub-second tape and the WT2 stress manifold. The CLASS-M/CLASS-X hour sets (WT2D) are your ready-made adversarial scenario suites — corrupt/stress hours ARE the hostile ecology, measured. CLASS-X law applies: extreme magnitudes are a FEATURE channel, never filtered. - Latency model = our measured tails, not defaults: scan_to_fill/step_bar distributions from the WT1 latency audit (median 47ms/1.58ms, stop-tail 7.8s) — calibrate the CWM's latency injections from these empirical curves.
- Reconcile the budget statement (FIX): planner box says ≤25ms, Rule 8 says ≤100ms — bind them: ≤25ms planning inside a ≤100ms end-to-end intent→venue-call budget, both enforced by test.
- Promotion gate binds to the cert conveyor: CMA-ES candidate promotion (Rule 4)
emits artifacts via the T4 cert reporter into the gate ledger; policy snapshots
versioned in
dolphin_malkhut.policy_snapshotsAND referenced inprod/docs/uv_cert/. No policy trades live without a ledger row.
F. Certification pathway (Gate-M — how MALKHUT reaches money)
- CWM replay verification (your Rule 3) vs hftbacktest — already specced; add determinism ×2 (byte-identical minus ts).
- Dual-leverage litmus (B.3 values) as a named test.
- DARK shadow window: MALKHUT plans on live intents while the simple path executes; diff planned-vs-naive on fills/slippage/adverse-selection — the improvement claim must be MEASURED on VST before MALKHUT takes the seat.
- Risk-gate mutation litmus: drop a constraint → a test goes RED (looks-finished doctrine; 222 green tests ≠ verification until mutations bite).
- Rule 10 stands supreme: shadow-vs-live divergence → trust live, stand down, file the discrepancy row.
End addendum. — F.
DEVELOPMENT STATUS (2026-07-13)
1204 test functions. 50 test files. All green. 0 failures. 0 regressions.
Fable Spec Items — All Complete
| # | Item | Module | Status |
|---|---|---|---|
| 1 | Mutation-litmus test | test_mutation_litmus.py |
✅ PASSES |
| 2 | Re-run benchmarks at corrected fees | (long run) | ⏸️ Deferred |
| 3 | Maker fee UNVERIFIED comment | asset_classification.py |
✅ Done |
| 4 | ScenarioLibrary sweep | scenario_library.py |
✅ 7,488 grid points |
| 5 | PerformanceMatrix manifold | selector.py |
✅ confidence+support |
| 6 | ActualsLoader | actuals_loader.py |
✅ CH reader + fallback |
| 7 | OOD verdict | risk/gate.py |
✅ daat_verdict parameter |
| 8 | Manifold query | manifold_query.py |
✅ three-phase recommend |
| 9 | DAAT package | daat/core.py |
✅ DaatVerdict{KNOWN,MARGINAL,OOD} |
| 10 | Book fidelity | book_fidelity.py |
✅ OBF → OrderBookState |
Fee Correction (CRITICAL — all prior training was at 10x too-cheap fees)
Source of truth: dolphin.trade_execution_quality (8006 rows, avg taker=5.016 bps).
| Venue | Old taker (WRONG) | Correct taker | Old maker (WRONG) | Correct maker |
|---|---|---|---|---|
| BingX | 0.5 bps | 5.0 bps | -0.2 (rebate) | +2.0 (POSITIVE) |
| Binance | 0.4 bps | 4.5 bps | -0.2 | -0.2 (UNVERIFIED) |
| Bybit | 0.06 bps | 5.5 bps | -0.01 | -0.1 (UNVERIFIED) |
Every policy trained before this fix is suspect — cheap fees reward overtrading. Re-measurement at correct fees is required for production deployment.
Completed subsystems
| Subsystem | Module | Tests | Status |
|---|---|---|---|
| State Model | state.py |
17 | 42 frozen dataclasses, immutable |
| CWM | cwm/core.py + cwm/numba_core.py |
103 | Exchange mechanics + numba JIT (5.3µs/transition) + vectorized reward + vectorized UCB + batch MCTS kernel |
| Replay Verification | cwm/replay_verify.py |
65 | Deep comparison, binary search, trajectory recording |
| Planner | planner/sm_mcts.py |
11 | Decoupled UCB/UCT, ≤25ms budget |
| Action Menu | planner/action_menu.py |
(in planner) | Compact action space construction |
| Counterparties | counterparties.py |
19 | 4 adversarial agent policies |
| Risk Gate | risk/gate.py |
4 | Hard invariants: leverage, self-trade, post-only |
| ASEx Integration | execution/asex_integration.py |
33 | Validate-before-mutate, single-writer, no locks |
| Zinc IPC | ipc/zinc_plane.py |
8 | Real POSIX SHM, UVZINC01 framing |
| Control Plane | ipc/control_plane.py |
3 | HOT_RELOAD_POLICY, kill switch |
| ClickHouse | storage/ch_store.py |
9 | 5 tables, HTTP API |
| DuckDB Asset Store | storage/asset_store.py |
30 | File-backed + in-memory materialization, sub-µs reads |
| Asset Bridge | training/asset_bridge.py |
49 | Directory ↔ classification sync |
| CMA-ES Training | training/cma_trainer.py |
65 | Behavior-driven, auto-compile, parallel workers, 7x speedup |
| Parallel Eval | training/parallel_eval.py |
16 | ProcessPoolExecutor, 7x CMA-ES speedup |
| Ray Eval | training/ray_eval.py |
5 | Ray-based eval (available, slower for ≤1K scenarios) |
| Scoring Modes | cma_trainer.py + advantage_scorer.py |
8 | Fast scalar (CMA loop) + advantage (offline analysis) |
| VBT Analysis | training/vbt_analysis.py |
8 | Post-sim trade metrics: Sharpe, Sortino, VaR, cross-asset |
| ScenarioLibrary | training/scenario_library.py |
7 | Sweep 7,488 grid points (spread×depth×tox×regime) |
| Manifold Query | training/manifold_query.py |
(new) | Three-phase: DAAT → manifold → recommend |
| ActualsLoader | training/actuals_loader.py |
(new) | Reads CH tables for Mode 2 live query |
| Book Fidelity | training/book_fidelity.py |
(new) | OBF 15B rows → OrderBookState synthesis |
| DAAT | daat/core.py |
9 | Direction-Anchored Ambiguity Triage |
| OOD Verdict | risk/gate.py |
+1 | daat_verdict parameter, OOD → doctrinal fallback |
| Policy Registry | training/registry.py |
14 | CANDIDATE → ACTIVE lifecycle |
| Training Pipeline | training/pipeline.py |
21 | Bounded continuous learning loop |
| Strategy DSL v2 | training/dsl.py |
69 | 40+ primitives, 40+ sensors, 16 builtins |
| Order Types | training/order_types.py |
(new) | Three-dimensional: OrderType × TimeInForce × Instructions (FIX-aligned) |
| Strategy Generator | training/generator.py |
20 | Genetic programming: crossover, mutation, tournament |
| Strategy Selector | training/selector.py |
24 | Regime × strategy × venue matrix, cross-exchange comparison |
| Asset Classification | training/asset_classification.py |
190 | Multi-label taxonomy + exchange registry, 13 assets, system-wide store |
| Asset Behavior DSL | training/asset_behavior.py |
(in classification) | 10 orthogonal dimensions, 3 templates, 13 behaviors, research-validated |
| Asset Compiler | training/asset_compiler.py |
(new) | Auto-fetch from Binance/BingX API, compile profiles, rate-limited |
| Cognition Pipeline | training/cognition.py |
29 | Rate-limited, 8 sources, dedup, perm-run |
| Regime Expansion | training/regime_expansion.py |
29 | 200+ regimes from 4×4×4×4 dimensions |
| News Sources | training/news_sources.py |
12 | 12 industry-standard sources, ranking |
| Cognition Monitor | training/monitor.py |
9 | Metrics, health scoring, alerts, JSONL logging |
| Cognition Launcher | cognition_launcher.py |
21 | Standalone long-run cognition service |
| UV Clock | clock/host.py |
30 | T19 event-driven reactor |
| Numba Acceleration | cwm/numba_core.py |
13 | JIT on hot loops, 1.8x fill speedup |
| BingX Adapter | venue/bingx/adapter.py |
28 | Wraps DITAv2, order tracking |
| Hypothesis | test_hypothesis_properties.py |
6 | Property-based fuzzing |
| Fuzz/Adversarial | test_fuzz.py, test_adversarial.py |
17 | Random stress + adversarial scenarios |
| Concurrency | test_concurrency.py |
8 | Thread safety, race detection |
| Sync/Async Seams | test_sync_async_seams.py |
7 | Zinc latency, engine budget |
| E2E Integration | test_e2e_integration.py |
2 | Full pipeline: train→register→plan→risk→venue→zinc→ch→reload |
Performance Benchmarks
| Metric | Value |
|---|---|
| CWM throughput | 189K calls/sec (numba JIT) |
| CWM per-call latency | 5.3 µs |
| CWM reward (numba vectorized) | ~0.3µs |
| UCB selection (numba vectorized) | 1.13µs (was 8.7µs, 7.7x speedup) |
| Numba fill speedup | 1.8x (batch 100) |
| DuckDB asset reads | 0.2µs (in-memory) |
| Scenario generation | 390 scenarios in 0.8s |
| CMA-ES parallel (8 workers) | 7× speedup, 87% efficiency |
| Best CMA-ES score (48 evals) | 8,628 (fast scalar, 100.4 bps PnL) |
| Score/min (8 workers) | 3,043 |
| Episode throughput (parallel) | 16ms/ep |
| Episode throughput (sequential) | 23ms/ep |
Bugs found and fixed (22 total)
- Empty book crash in
materialize_price_from_action(SELL side) - Empty book crash in
total_notionalcalculation - Empty book crash in mark-to-market
- Empty book crash in
_inventory_risk - Empty book crash in counterparty processing
- Empty book crash in
_update_path_state - Empty book crash in feature extractor (
b.mid) - Counterparty fills used
max()instead of consuming levels - Position overwritten instead of accumulated on add
- Test helper
_state()not forwardingpositionskwarg _make_open_orderset empty symbol instead of venue symbolFulfilmentPolicyParamsfrozen dataclass__dict__access- Banker's rounding in tick tests
HOT_RELOAD_POLICYcompared float to string- Engine init order (workers before hot_reload)
_tp()helper missingfailed_recovery_countparameter- CWM: SELL position flip didn't reset avg_entry
- CWM: Counterparty fills overwrote our fill tracking
- CWM: Double-counting unrealized PnL in equity
- UnboundLocalError when max_steps=0
- fill_count never incremented in training
- Cancel rate limiter incorrectly gated order placements
Key design rules enforced
- ✅ Optimise distribution, not single quote
- ✅ Book is not passive scenery
- ✅ Replay correctness before search depth (mandatory gate)
- ✅ CMA-ES does not promote from noisy wins (bootstrap CI)
- ✅ Preserve feature humility (CMA discovers interactions)
- ✅ Path-aware SL/TP is part of fulfilment
- ✅ Never let optimiser learn illegal behavior
- ✅ Hot path bounded (≤100ms)
- ✅ Version everything
- ✅ Shadow-vs-live divergence → trust live
Strategy Selection (completed)
The system classifies strategies per market condition and floats the best to top:
market_fingerprint → Regime Classifier (8 regimes)
↓
Performance Matrix (regime × strategy → score)
↓
Strategy Selector → Engine uses best strategy
Upstream integration path
- ExoF — funding rates, open interest, liquidation levels
- MARAS — 5-tier ensemble → 8-regime classification
- DOLPHINNG7 vel_div — velocity divergence signals
- DOLPHINNG7 eigenscan — eigenspace correlation analysis
G. Strategy DSL v2 — Third-Party Strategy Development
MALKHUT includes a Domain-Specific Language (DSL) for strategy composition. Third parties can write, evolve, and share strategies without modifying core code.
DSL Format
STRATEGY "passive_maker" {
PRIORITY 1: IF spread_bps < 5.0 AND orderflow_toxicity < 0.3 THEN QUOTE(BUY, 0, 0.25)
PRIORITY 2: IF spread_bps < 5.0 AND orderflow_toxicity < 0.3 THEN QUOTE(SELL, 0, 0.25)
PRIORITY 3: IF orderflow_toxicity > 0.7 THEN CANCEL_ALL
PRIORITY 4: IF time_in_trade > 300 THEN EXIT
PRIORITY 5: IF unrealized_pnl < -50bps THEN STOP_LOSS
PRIORITY 6: NOOP
}
Action Primitives (40+)
| Category | Primitives |
|---|---|
| Passive | QUOTE, JOIN_QUEUE, STEP_BACK, LADDER, GRID, ICEBERG, TWAP |
| Aggressive | CROSS, SNIPER, PING |
| Cancellation | CANCEL, CANCEL_ALL, CANCEL_AND_HOLD, REQUOTE |
| Position | EXIT, HALF_EXIT, QUARTER_EXIT, TAKE_PARTIAL, STOP_LOSS, TAKE_PROFIT, TRAILING_STOP, EMERGENCY_EXIT, FLAT_ALL |
| Sizing | SCALE_IN, SCALE_OUT, INCREASE_SIZE, REDUCE_SIZE |
| Stop/Target | MOVE_STOP, MOVE_TAKE_PROFIT, BRACKET, OCO |
| Hedging | HEDGE, PAIR_TRADE |
| Waiting | HOLD, WAIT_FOR_FILL, WAIT_FOR_PRICE, WAIT_FOR_SPREAD |
Market Sensors (40+)
| Category | Sensors |
|---|---|
| Book | spread_bps, imbalance_3/5/10, bid_depth_3/5/10, bid_ask_ratio |
| Flow | toxicity, queue_churn, cross_venue_lead, order_book_toxicity |
| Momentum | price_momentum_1s/5s/15s/1m |
| Position | position_qty, pnl_bps, leverage, position_age_s |
| Path | mae_bps, mfe_bps, time_in_loss, failed_recoveries, recovery_velocity |
| Regime | volatility, atr_14/50, rsi_14, funding, regime_score |
| Account | equity, risk_budget_used, session_pnl, drawdown |
| Performance | profit_factor, sharpe_ratio, consecutive_wins/losses |
| Time | current_hour, day_of_week, is_weekend, is_liquid_hours |
Comparison Operators (12)
>, <, >=, <=, ==, !=, abs>, abs<, changing, stable, crossing_above, crossing_below
Builtin Strategies (16)
| Strategy | Description |
|---|---|
passive_maker |
Passive limit orders, toxicity avoidance |
aggressive_taker |
Momentum-based aggressive crosses |
toxicity_avoider |
Cancels on high toxicity, quotes when safe |
path_risk_exit |
Path-aware SL/TP exits |
regime_adaptive |
Adapts to MARAS regime classification |
momentum_catcher |
Catches momentum moves, partial exits |
mean_reversion |
Wide-spread mean reversion |
scalper |
Ultra-tight spreads, small sizes |
inventory_manager |
Position size control |
session_guard |
Time-based position management |
liquidity_hunter |
Deep book exploitation |
volatility_breakout |
Volatility-based breakouts |
funding_arb |
Funding rate arbitrage |
grid_trader |
Grid-based accumulation |
risk_parity |
Risk budget management |
hybrid_adaptive |
Multi-condition adaptive |
Strategy Generator (Genetic Programming)
The system evolves strategies via genetic operators:
from malkhut.training.generator import StrategyGenerator, GeneratorConfig
config = GeneratorConfig(
population_size=20, generations=5, tournament_size=3,
mutation_rate=0.15, crossover_rate=0.7,
)
generator = StrategyGenerator(config=config)
population = generator.evolve(baseline_params, scenarios)
# Get successful strategies
successful = generator.get_successful_strategies(min_fitness=0.0)
for genome in successful:
generator.add_to_pool(genome)
Strategy Selector (Market-Adaptive)
The system classifies strategies per market regime and selects the best:
from malkhut.training.selector import StrategySelector, MarketRegime
selector = StrategySelector()
result = selector.select(current_state, available_strategies)
# result.strategy_id → best strategy for current regime
# result.confidence → how confident we are
# result.alternatives → other considered strategies
Regimes: trending_up, trending_down, high_volatility, low_volatility,
mean_reverting, momentum, choppy, liquidity_hole, normal
Asset Classification (Invariant Taxonomy)
Multi-label asset classification using only invariant characteristics — properties intrinsic to the token that predict price behaviour and order-book performance. Three-layer identifier architecture for cross-system interoperability.
Design principle: Classify by WHAT an asset IS (fundamental), not how it's trading right now. Volatility/liquidity appear as long-run statistical profiles, not as classification axes that change minute-to-minute.
Three-Layer Identifier Architecture
| Layer | Fields | Purpose | Example |
|---|---|---|---|
| 1. Canonical | symbol, base_asset, name, unified_symbol, quote_currency |
Venue-independent identity | BTCUSDT / BTC / "Bitcoin" / BTC/USDT / USDT |
| 2. Cross-system | coingecko_id, cmc_id, blockchain, contract_address |
Data provider + on-chain mapping | "bitcoin" / 1 / "ethereum" / "" |
| 3. Exchange mapping | exchanges: tuple[str, ...] |
Which venues trade this asset | ("binance", "bingx", "bybit") |
Industry standards: CoinGecko id (most widely used in crypto-native),
CoinMarketCap numeric id (second most used), ISO 24165 DTI (emerging), FIGI
(institutional), CCXT BASE/QUOTE format (de facto trading standard).
Fundamental Dimensions (intrinsic, never change)
| Dimension | Values | Multi-label | Purpose |
|---|---|---|---|
| Sector | CURRENCY, LAYER1, LAYER2, DEFI, ORACLE, EXCHANGE, MEME, PRIVACY, STORAGE, GAMING_NFT | Yes | Primary use-case / vertical |
| TokenRole | GAS, STORE_OF_VALUE, GOVERNANCE, UTILITY, MEME, EXCHANGE_FEE | Yes | Functional role → demand elasticity |
| SupplyModel | FIXED_CAP, DISINFLATIONARY, INFLATIONARY, BURN_MECHANISM | No | Supply pressure dynamics |
| ConsensusFamily | POW, POS, DPOS | No | Miner/validator selling behaviour |
| SmartContractCapability | FULL, PARTIAL, NONE | No | DeFi composability |
Technical Dimensions (invariant market-structure properties)
| Dimension | Values | Purpose |
|---|---|---|
| MarketCapTier | MEGA (>500B), LARGE (50-500B), MID (5-50B), SMALL (500M-5B), MICRO (<500M) | Price impact per dollar |
| VolatilityProfile | LOW (<30%), MEDIUM (30-80%), HIGH (80-150%), EXTREME (>150%) | Long-run average vol |
| LiquidityProfile | DEEP (>100M), NORMAL (10-100M), THIN (1-10M), ILLIQUID (<1M) | Long-run avg depth |
| DerivativeAccess | PERPS_AND_OPTIONS, PERPS_ONLY, NONE | Shorting / funding availability |
| tick_size / lot_size | Exchange-set | Minimum price/order size |
| maker/taker fee_bps | Exchange-set | Trading cost |
| typical_spread_bps / depth / volume | Long-run averages | Order-book fingerprint |
Predictive Properties (derived from invariants)
| Property | Logic | Value |
|---|---|---|
supply_pressure |
PoW → forced selling (miners cover electricity) | "forced" / "optional" |
demand_elasticity |
GAS or STORE_OF_VALUE → inelastic (must hold) | "inelastic" / "elastic" |
can_be_shorted |
DerivativeAccess != NONE | bool |
Multi-Label Examples
BTC: sectors=(CURRENCY), roles=(STORE_OF_VALUE)
ETH: sectors=(LAYER1, DEFI), roles=(GAS, STORE_OF_VALUE, GOVERNANCE)
BNB: sectors=(EXCHANGE, LAYER1), roles=(EXCHANGE_FEE, GAS)
DOGE: sectors=(MEME, CURRENCY), roles=(MEME, GAS)
DOT: sectors=(LAYER1), roles=(GAS, GOVERNANCE)
UNI: sectors=(DEFI), roles=(GOVERNANCE, UTILITY)
Querying
from malkhut.training.asset_classification import *
# Multi-label queries (match if queried value is ANY of the asset's labels)
layer1s = get_assets_by_sector(Sector.LAYER1) # ETH, SOL, ADA, AVAX, BNB, DOT, ATOM
gas_tokens = get_assets_by_token_role(TokenRole.GAS) # ETH, SOL, ADA, AVAX, DOGE, BNB, MATIC, DOT, ATOM
# Predictive properties
forced = get_pov_assets() # BTC, DOGE (PoW miners with forced selling)
inelastic = [p for p in ASSET_PROFILES.values() if p.demand_elasticity == "inelastic"]
# Backward-compatible single-label access (first element = primary)
btc = get_asset_profile("BTCUSDT")
btc.sector # Sector.CURRENCY
btc.token_role # TokenRole.STORE_OF_VALUE
Exchange Registry (Multi-Exchange Asset Universe)
MALKHUT serves as the system-wide asset universe store for BLUE, VIOLET, UV, and all downstream systems. Each asset knows which exchanges trade it; each exchange has a standardized profile.
ExchangeProfile
| Field | Type | Purpose |
|---|---|---|
exchange_id |
str | Canonical key: "binance", "bingx", "bybit" |
display_name |
str | Human-readable name |
has_spot |
bool | Spot trading available |
has_perps |
bool | Perpetual futures available |
has_options |
bool | Options available |
api_base_url |
str | REST API root URL |
ws_base_url |
str | WebSocket root URL |
default_taker_fee_bps |
float | Default taker fee |
default_maker_fee_bps |
float | Default maker fee |
typical_latency_ms |
float | Typical API latency |
Pre-defined Exchanges
| Exchange | Spot | Perps | Options | Taker Fee | Latency |
|---|---|---|---|---|---|
| Binance | ✓ | ✓ | ✓ | 0.4 bps | 40 ms |
| BingX | ✓ | ✓ | ✗ | 0.5 bps | 100 ms |
| Bybit | ✓ | ✓ | ✓ | 0.06 bps | 50 ms |
Exchange-Asset Mapping
Each AssetProfile carries an exchanges: tuple[str, ...] field listing which venues
trade the asset. Default: ("binance",). Other systems (BLUE/VIOLET/UV) import their
asset universes into this store.
from malkhut.training.asset_classification import *
# Which exchanges trade BTC?
btc = get_asset_profile("BTCUSDT")
print(btc.exchanges) # ('binance',)
# All assets on Binance
binance_assets = get_assets_on_exchange("binance")
# Assets traded on BOTH Binance and BingX
common = get_common_assets("binance", "bingx")
# Which exchanges trade a given asset?
venues = get_exchange_for_asset("ETHUSDT")
# Exchange metadata
ex = get_exchange("binance")
print(ex.default_taker_fee_bps) # 0.4
print(ex.typical_latency_ms) # 40
Asset Behavior DSL (10-Dimension Research-Validated Model)
The Asset Behavior DSL decomposes each asset's trading behavior into 10 orthogonal dimensions, each with empirically-validated parameters from live Binance/BingX API data and academic literature.
Source: Binance live REST API (depth 500 levels, 1h klines, ticker, funding rates), BingX open API (depth, funding, premium index), academic literature (Bouchaud "Trades Quotes and Prices" 2018, Cont/Stoikov/Talreja 2010, Cartea/Jaimungal/Penalva 2015), CoinGlass OI data, industry MM disclosures.
10 Dimensions
| Dimension | What it captures | Example: BTC vs DOGE |
|---|---|---|
| DepthProfile | Book shape: amplitude, decay exponent α, fragility | BTC: $750K/0.70/0.10 vs DOGE: $22K/1.00/0.30 |
| SpreadProfile | Normal spread, stress multiplier | BTC: 0.01bps/50x vs DOGE: 1.35bps/15x |
| FlowProfile | Order rate, size distribution, cancel ratio | BTC: 300/s/$643/20x vs DOGE: 80/s/$96/8x |
| VolatilityProfile | Annualized vol, GARCH params, half-life | BTC: 35%/48h vs DOGE: 78%/24h |
| IntradayProfile | Peak/trough hours, ratio | BTC: 15:00 UTC, 7.4x ratio |
| WeekendProfile | Vol/volume/spread multipliers | 0.65x vol, 0.52x volume, 1.20x spread |
| CorrelationProfile | ETH beta, BTC corr (normal vs crash) | BTC: 1.0/1.0 vs UNI: 0.70/0.50 |
| MarketMakerProfile | Inventory limits, pull speed, margins | BTC: $10M/3ms/0.5bps vs DOGE: $500K/25ms/2bps |
| LiquidationProfile | OI/MCap, trigger %, cascade speed | BTC: 0.5%/6.5%/slow vs DOGE: 1.4%/4%/fast |
| FundingProfile | Mean/std funding, basis | BTC: 0.59bps/0.22bps vs DOGE: 0.39bps/0.34bps |
| RetailProfile | Retail ratio, inst gap | BTC: 35%/0.04 vs DOGE: 80%/0.31 |
| BingxProfile | BingX spread/depth multiplier, fees, latency | BTC: 12.6x/0.048x vs DOGE: 7x/1.27x |
3 Composable Templates
| Template | Assets | Key traits |
|---|---|---|
institutional_blue_chip |
BTC, ETH, BNB | Deep book, low spread, slow cascade, high institutional |
mid_cap_l1 |
SOL, ADA, AVAX, DOT, ATOM, UNI, LINK, MATIC, AAVE | Moderate depth, higher vol, medium cascade |
retail_meme |
DOGE | Thin book, high vol, fast cascade, retail-dominated |
13 Per-Asset Behaviors (research-validated)
| Asset | Price | Depth@1bps | Spread | Vol | Cascade | Retail | BingX mult |
|---|---|---|---|---|---|---|---|
| BTC | $64K | $750K | 0.01 bps | 35% | slow/fast | 35% | 12.6x |
| ETH | $1.8K | $600K | 0.01 bps | 66.5% | medium/medium | 35% | 9.0x |
| SOL | $80 | $400K | 1.26 bps | 72% | medium/medium | 72% | 1.7x |
| DOGE | $0.07 | $22K | 1.35 bps | 78% | fast/slow | 80% | 7.0x |
| ADA | $0.17 | $60K | 5.95 bps | 79% | medium/medium | 65% | 8.0x |
| UNI | $3.6 | $30K | 2.76 bps | 95% | medium/medium | 60% | 10x |
Usage
from malkhut.training.asset_behavior import *
# Get behavior for any asset
btc = get_behavior("BTCUSDT")
print(btc.depth.amplitude_usd) # $750,000
print(btc.vol.annualized_normal) # 35.0%
print(btc.liquidation.speed) # "slow"
# Compute depth at distance
depth_10bps = btc.depth_at_bps(10) # $1.5M
# Estimate slippage for a $100K order
slippage = btc.expected_slippage_bps(100_000) # ~2 bps
# Query by template, sector, or role
institutional = get_behaviors_by_template("institutional_blue_chip")
gas_tokens = get_behaviors_by_role(TokenRole.GAS)
thin_book = get_thin_book_assets()
Asset Compiler (Auto-Fetch from Live APIs)
The Asset Compiler automatically fetches market data from Binance and BingX public APIs
and produces system-ready AssetProfile + AssetBehavior for any tradeable symbol.
Rate-limited (1 req/sec, well under Binance 1200/min limit), cached (avoids re-fetching within a session), resumable.
What it auto-fetches
| Endpoint | Data | Computes |
|---|---|---|
/api/v3/ticker/24hr |
Price, volume, high/low | Daily volume, price reference |
/api/v3/depth?limit=100 |
Order book (100 levels) | Spread, depth amplitude, decay exponent α |
/api/v3/klines?interval=1h&limit=168 |
7 days hourly OHLCV | Annualized volatility, order flow stats |
/api/v3/exchangeInfo |
Symbol filters | Tick size, lot size, price decimals |
/fapi/v1/fundingRate |
Funding rates (20 intervals) | Mean/std/positive% funding |
/fapi/v1/openInterest |
Open interest | OI/MCap ratio for liquidation modeling |
Usage
from malkhut.training.asset_compiler import AssetCompiler
compiler = AssetCompiler()
# Single asset — auto-fetches from Binance, computes all params
result = compiler.compile("XRPUSDT")
print(result.asset_profile.symbol) # "XRPUSDT"
print(result.asset_behavior.reference_price) # $1.1056 (live Binance price)
print(result.asset_behavior.spread.normal_bps) # 0.90 bps (computed from depth)
print(result.asset_behavior.vol.annualized_normal) # 46.6% (computed from klines)
# Register into global registries
compiler.register(result)
# Now XRPUSDT is available for ScenarioFactory, queries, etc.
# One-step compile and register
result = compiler.compile_and_register("HBARUSDT")
# Batch compile (rate-limited sequentially)
results = compiler.compile_batch(["APTUSDT", "SEIUSDT", "INJUSDT"])
for r in results:
compiler.register(r)
Known classifications
The compiler includes pre-built classifications for 28 assets: BTC, ETH, SOL, DOGE, ADA, AVAX, UNI, LINK, BNB, MATIC, AAVE, DOT, ATOM, XRP, TRX, LTC, NEAR, APT, OP, ARB, SUI, PEPE, WIF, SEI, INJ, FIL, HBAR, IMX
Unknown assets get heuristic defaults (LAYER1/GAS/INFLATIONARY/POS/FULL/PERPS_ONLY).
Scenario Factory (Behavior-Driven, Label-Queryable)
The ScenarioFactory builds evaluation scenarios using real per-asset prices, depth profiles, and spread characteristics — no more hardcoded BTC prices.
Behavior-driven state creation
Every scenario derives its parameters from the asset's AssetBehavior:
| Scenario type | Spread multiplier | Depth fraction | Description |
|---|---|---|---|
normal |
1.0x | 100% | Calm, liquid market |
thin_book |
1.5x | 10% | Low participation |
wide_spread |
100x | 50% | High volatility |
toxic_stress |
2.0x | 30% | Adverse selection |
flash_crash |
3.0x | 5% | Book collapse |
liquidity_vacuum |
5.0x | 1% | Near-empty book |
weekend_thin |
5.0x | 5% | Weekend low-volume |
trending |
1.0x | 80% | Momentum move |
mean_reverting |
20x | 60% | Reversal setup |
funding_shock |
1.5x | 30% | Deleveraging event |
liquidation_cascade |
1.5x | 20% | Chain liquidation |
market_maker_withdrawal |
3.0x | 15% | MM pull-out |
Realistic per-asset pricing
Scenario: "normal_BTCUSDT_42"
BTC bid: $63,999.97 (11.7 BTC = $750K)
BTC ask: $64,000.03
Spread: 0.01 bps
Scenario: "normal_DOGEUSDT_42"
DOGE bid: $0.07 (314K DOGE = $22K)
DOGE ask: $0.07
Spread: 1.35 bps
Scenario: "flash_crash_SOLUSDT_47"
SOL bid: $80.00 (depth = 5% of normal = $20K)
SOL ask: $80.02
Spread: 3.78 bps (3x normal)
Auto-compilation for unknown assets
When you request a symbol that isn't pre-defined, the factory automatically compiles it from live Binance API data before building scenarios:
factory = ScenarioFactory()
# XRPUSDT is NOT in our pre-defined 13 assets
# But this auto-compiles it from Binance in ~5 seconds:
suite = factory.build_suite_for_symbols(("XRPUSDT",), steps_per_scenario=5)
# XRP behavior is now registered for future use too
Label-based query interfaces
factory = ScenarioFactory()
# By sector
suite = factory.build_suite_for_sector(Sector.LAYER1) # ETH, SOL, ADA, AVAX, BNB, DOT, ATOM
suite = factory.build_suite_for_sector(Sector.DEFI) # ETH, AVAX, UNI, AAVE
suite = factory.build_suite_for_sector(Sector.MEME) # DOGE
# By token role
suite = factory.build_suite_for_role(TokenRole.GAS) # 9 assets
suite = factory.build_suite_for_role(TokenRole.GOVERNANCE) # ETH, UNI, AAVE, DOT
# By behavior template
suite = factory.build_suite_for_template("institutional_blue_chip") # BTC, ETH, BNB
suite = factory.build_suite_for_template("retail_meme") # DOGE
# By volatility band
suite = factory.build_suite_for_volatility(min_ann=75, max_ann=200) # High-vol assets
# Multi-label union (any matching)
suite = factory.build_suite_for_labels(
sectors=[Sector.DEFI],
roles=[TokenRole.GOVERNANCE]
)
# Arbitrary symbols (auto-compiles unknowns)
suite = factory.build_suite_for_symbols(("BTCUSDT", "XRPUSDT", "HBARUSDT"))
Research-validated scenario diversity
Each asset's 30 scenario types are parameterized by its behavior profile:
| Scenario | BTC params | DOGE params | Why it matters |
|---|---|---|---|
| Normal | $750K depth, 0.01bps spread | $22K depth, 1.35bps spread | Baseline for scoring |
| Thin book | $75K depth | $2.2K depth | Tests fill rate under low liquidity |
| Flash crash | $37.5K depth, 0.03bps spread | $1.1K depth, 4bps spread | Tests adverse selection survival |
| Weekend thin | $37.5K depth, 5bps spread | $1.1K depth, 6.75bps spread | Tests off-hours strategy fitness |
| Liquidation cascade | $150K depth | $4.4K depth | Tests position sizing under stress |
| Market maker withdrawal | $112.5K depth | $3.3K depth | Tests quote quality when MMs pull |
Numba Performance
Hot-path functions JIT-compiled with numba:
| Function | Speedup |
|---|---|
| fill_from_levels (batch 100) | 1.8x |
| round_tick | JIT-compiled |
| round_lot | JIT-compiled |
| clip_lots | JIT-compiled |
| extract_features | Vectorized numpy |
CWM throughput: 157K calls/sec, 6.4 µs/transition, 0.64 ms/100-step episode.
Design Principles
- Hardcoded baseline is NEVER replaced — always available as reference
- Generated strategies are ADDED — system GROWS its repertoire
- Crossover combines two strategies (uniform crossover on parameters)
- Mutation randomly modifies parameters (Gaussian noise)
- Selection uses tournament selection (fitter genomes win more often)
- Third parties can write strategies in DSL text format
- Market adaptation: selector chooses strategy based on ExoF/MARAS regime
- Behavior-driven scenarios: every asset behaves like its real-world self
Strategy Naming Convention
{strategy_type}_gen{generation}_{timestamp}
Examples: UCB1_gen1_20260707_223754, THOMPSON_gen2_20260707_223754
H. Cambrian Expansion
Tie-In Fix
PerformanceMatrix now wired to evaluator. Strategies scored by regime fitness:
- ALL strategies tested in ALL scenarios
- Score tracked per strategy × regime
- Selector queries matrix for best strategy per regime
PerformanceMatrix.record(strategy_id, regime, score)called during evaluation
Phase 0.1: Cognition Pipeline
Rate-limited market regime research:
RateLimiter: token bucket, 30 RPM defaultSourceCatalogue: tracks sources, fetch counts, error ratesRegimeExtractor: keyword matching + sentiment analysisCognitionPipeline: rate-limited, deduplication, long/perm-run- 8 default sources: CoinDesk, Cointelegraph, The Block, CryptoQuant, Glassnode, Coinglass, Binance Research, Messari
Phase 0.2: Exponential Regime Expansion
Orthogonal to cognition pipeline — generates 200+ regimes from dimensions:
- 4 liquidity (vacuum, thin, normal, deep)
- 4 volatility (tight, normal, wide, extreme)
- 4 flow (balanced, buy_pressure, sell_pressure, toxic)
- 4 structure (single_toxic, mm_toxic, multi_toxic, retail_mm)
Each combination = distinct, testable regime. Total: 256 theoretical, ~200 practical.
Phase 0.3: Multi-Asset & Invariant Classification
13 assets across 10 sectors, classified by invariant characteristics only:
| Asset | Sectors | Roles | Supply | Consensus | Vol | Liq |
|---|---|---|---|---|---|---|
| BTC | CURRENCY | STORE_OF_VALUE | FIXED_CAP | POW | LOW | DEEP |
| ETH | LAYER1, DEFI | GAS, SOV, GOV | DISINFLATIONARY | POS | MEDIUM | DEEP |
| SOL | LAYER1 | GAS | INFLATIONARY | POS | HIGH | NORMAL |
| DOGE | MEME, CURRENCY | MEME, GAS | INFLATIONARY | POW | HIGH | NORMAL |
| ADA | LAYER1 | GAS | INFLATIONARY | DPOS | MEDIUM | NORMAL |
| AVAX | LAYER1, DEFI | GAS | INFLATIONARY | POS | HIGH | NORMAL |
| UNI | DEFI | GOV, UTILITY | INFLATIONARY | POS | HIGH | THIN |
| LINK | ORACLE | UTILITY | INFLATIONARY | POS | MEDIUM | THIN |
| BNB | EXCHANGE, LAYER1 | EXCHANGE_FEE, GAS | BURN | DPOS | MEDIUM | DEEP |
| MATIC | LAYER2 | GAS | INFLATIONARY | POS | HIGH | THIN |
| AAVE | DEFI | GOV | FIXED_CAP | POS | HIGH | THIN |
| DOT | LAYER1 | GAS, GOV | INFLATIONARY | DPOS | MEDIUM | NORMAL |
| ATOM | LAYER1 | GAS | INFLATIONARY | POS | HIGH | THIN |
Multi-asset = linear multiplication: 30 scenarios × 13 assets = 390+ testable scenarios.
Phase 0.4: Asset Behavior DSL + Compiler + Behavior-Driven Scenarios
Research-validated from Binance live API, BingX API, CoinGlass, and academic literature:
| Component | Purpose | File |
|---|---|---|
| AssetBehavior | 10 orthogonal behavior dimensions | training/asset_behavior.py |
| BehaviorTemplate | Reusable class-level behavior profiles | training/asset_behavior.py |
| AssetCompiler | Auto-fetch from Binance/BingX, compute profiles | training/asset_compiler.py |
| ScenarioFactory | Behavior-driven scenarios with label queries | training/cma_trainer.py |
Key research findings wired into the system:
- Depth decay: D(d) = A * d^(1-alpha). BTC alpha=0.7, DOGE alpha=1.0
- Spread ranges: BTC 0.01bps, SOL 1.26bps, DOGE 1.35bps, ADA 5.95bps
- Order sizes: BTC P50=$643, DOGE P50=$96 (7x smaller)
- Cascade dynamics: BTC 5-8% trigger/slow/fast recovery, DOGE 3-5%/fast/slow
- BingX specifics: 1.7-12.6x wider spreads, 0.05-9x variable depth
- Weekend effects: -35% vol, -48% volume, +20% spread
- Intraday patterns: 7.4x peak/trough ratio, US session dominant
- GARCH persistence: 0.95-0.99, vol half-life 2-5 days
- Correlation: BTC-ETH 0.90 normal, 0.97 crash; BTC-DOGE 0.45 normal, 0.80 crash
Training Scaling Results
CMA-ES evaluation on behavior-driven multi-asset scenarios (BTC/ETH/SOL, 3 assets × 30 scenario types = 90 scenarios). Each eval = one CMA candidate × 90 multi-step episodes through the CWM.
Parallel Training Benchmark (48 evals, pop=12)
| Workers | Time | Best Score | Score/min | Speedup | Efficiency |
|---|---|---|---|---|---|
| 1 (sequential) | 632s | 1,727 | 164 | 1.0× | — |
| 2 | 320s | 2,262 | 425 | 1.98× | 99% |
| 4 | 163s | 2,152 | 791 | 3.87× | 97% |
| 8 | 91s | 2,594 | 1,718 | 6.98× | 87% |
87% parallel efficiency at 8 workers. Independent workers explore more diverse strategy space — parallel runs find BETTER scores than sequential (2,594 vs 1,727).
Budget Scaling (ProcessPool, 3 assets × 30 scenarios)
| Budget | Gens | Best Score | Time | Score/min |
|---|---|---|---|---|
| 48 evals | 4 | 2,594 | 91s | 1,718 |
| 96 evals | 8 | 7,202 | 1246s | 346 |
| 192 evals | 16 | ~14K (est) | ~45min | ~311 |
Vectorized Reward (numba)
The CWM reward() method now uses compute_reward_vectorized (numba JIT) when
available, bypassing the Python FeatureVector dict allocation. This eliminates
~2M dict allocations per CMA generation.
Training Performance
| Metric | Sequential | Parallel (8 workers) | Speedup |
|---|---|---|---|
| CWM throughput | 189K calls/sec | (bottleneck is planner, not CWM) | — |
| CWM per-call latency | 5.3 µs | (numba + vectorized reward) | — |
| Scenario generation | 390 scenarios in 0.8s | — | — |
| CMA-ES per-generation (pop=12) | ~155s | ~23s | ~7× |
| CMA-ES 48 evals | 632s | 91s | 7× |
| Best score at 48 evals | 1,727 | 2,594 | +50% |
| Score/min | 164 | 1,718 | 10.5× |
Parallel evaluation achieves 9x speedup on single evals (embarrassingly parallel,
zero fidelity loss). CMA-ES training speedup is ~1.5x because multiprocessing overhead
is amortized across 90 scenarios per eval. Bottleneck is planner (MCTS), not CWM —
numba already accelerates the CWM hot path (_HAS_NUMBA = True).
Order Type Standardization (Multi-Exchange, FIX/CCXT-Aligned)
Three orthogonal dimensions (not one flat enum), each mapped independently:
Dimension 1 — Order Types (FIX Tag 40 OrdType): what the order IS
MARKET, LIMIT, STOP_MARKET, STOP_LIMIT, TRIGGER_MARKET, TRIGGER_LIMIT,
TRAILING_STOP, OCO, TP_SL
Dimension 2 — Time-in-Force (FIX Tag 59): how long the order LIVES
GTC, IOC, FOK, GTD
Dimension 3 — Instructions (FIX Tag 18 ExecInst): behavioral modifiers
POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
CRITICAL:
POST_ONLYis an instruction on aLIMITorder, not a standalone type.IOC/FOKare TimeInForce values applied to aLIMITorder, not order types. This aligns with FIX Protocol: Tag 40 (OrdType), Tag 59 (TimeInForce), Tag 18 (ExecInst).
Cross-Exchange Order Type Mapping
| Normalized | BingX | Binance Spot | Binance Futures | Bybit |
|---|---|---|---|---|
LIMIT |
LIMIT | LIMIT | LIMIT | LIMIT |
MARKET |
MARKET | MARKET | MARKET | MARKET |
STOP_MARKET |
TRIGGER_MARKET | STOP_MARKET | STOP_MARKET | STOP_MARKET |
STOP_LIMIT |
TRIGGER_LIMIT | STOP_LOSS_LIMIT | STOP | STOP_LIMIT |
TRIGGER_MARKET |
TRIGGER_MARKET | TAKE_PROFIT | TAKE_PROFIT | TAKE_PROFIT_MARKET |
TRIGGER_LIMIT |
TRIGGER_LIMIT | TAKE_PROFIT_LIMIT | TAKE_PROFIT | TAKE_PROFIT_LIMIT |
TRAILING_STOP |
TRAILING_STOP_MARKET | TRAILING_STOP_MARKET | TRAILING_STOP_MARKET | TRAILING_STOP |
TimeInForce Mapping
| Normalized | BingX | Binance | Bybit |
|---|---|---|---|
GTC |
GTC | GTC | GTC |
IOC |
IOC | IOC | IOC |
FOK |
FOK | FOK | FOK |
Transferability Principle
Strategy PARAMETERS transfer across exchanges. Order type NAMES are venue-specific
but semantics are identical. A LIMIT on BingX = LIMIT on Binance = LIMIT on Bybit
(same fill behavior). Only the API string differs. The venue adapter translates
normalized → exchange-native at submission time.
Cross-Exchange Learning
ScenarioFactory is venue-aware: ScenarioFactory(exchange_id='bingx') tags every
scenario with its venue. Cross-exchange transfer re-tags for a different venue:
factory = ScenarioFactory(exchange_id='bingx')
scenarios_bingx = factory.build_suite(symbols=['BTCUSDT', 'ETHUSDT'])
strategy = train(scenarios_bingx) # evolve on BingX
scenarios_binance = factory.cross_exchange_transfer(scenarios_bingx, 'binance')
score = evaluate(strategy, scenarios_binance) # test on Binance
PerformanceMatrix keys are (regime, strategy_id, venue):
get_best(regime, venue='bingx')— per-venue bestget_venue_comparison(regime, strategy_id)—{venue: score}get_best(regime)— venue-agnostic (backward compatible)
Adversary Ecology
Counterparties operate at the ActionKind level (CROSS_SPREAD, PLACE, CANCEL), not at order-type level. The CWM translates ActionKind to venue-native order types:
CROSS_SPREAD→ fills aggressively → equivalent to MARKETPLACE→ passive quote → equivalent to LIMIT Fee calculation uses VenueRules (per-exchange fees). The ecology is venue-independent.
Standards Referenced
- FIX 4.4: Tag 40 (OrdType), Tag 59 (TimeInForce), Tag 18 (ExecInst)
- CCXT: unified
limit/market+ parameter composition (triggerPrice, stopPrice) - ISO 10383: MIC for venue identification
- MiFID II / MiCA: defers to FIX for order type classification
- FIA: references FIX Tag 40 in digital asset guidelines
Prod Tooling
| Component | Purpose | File |
|---|---|---|
| CognitionLauncher | Standalone long-run script | cognition_launcher.py |
| CognitionMonitor | Metrics, health scoring, alerts | training/monitor.py |
| NewsSourceRepository | 12 industry-standard sources, ranking | training/news_sources.py |
| SourceCatalogue | Source persistence, fetch tracking | training/cognition.py |
| RateLimiter | Token bucket, 30 RPM | training/cognition.py |
| RegimeExpander | 200 regimes from dimensions | training/regime_expansion.py |
Running
# Training pipeline (continuous)
python -m malkhut.continuous_pipeline
# Cognition pipeline (continuous regime research)
python -m malkhut.cognition_launcher