README updated with: - Cross-exchange learning: ScenarioFactory exchange_id + cross_exchange_transfer - PerformanceMatrix keyed by (regime, strategy_id, venue) - Adversary ecology: ActionKind-level abstraction, venue-independent - Transferability principle: parameters transfer, names are venue-specific
MALKHUT — Adversarial Self-Play Order-Fulfilment Pipeline
Lock-free. No GC. Fully async. GraalVM-compatible.
MALKHUT is an order-fulfilment engine that optimises a distribution over execution actions against a diverse counterparty ecology, not a single deterministic quote.
Core Thesis
Sequential/pure optimisation ("what is the best quote?") leads to deterministic quote → predictable fill pattern → adverse selection.
Correct simultaneous-market framing: "what mixed order-placement policy survives a diverse counterparty ecology?"
Architecture
Corrected Stack (T19 ANNEX B-CORRECTION)
iceoryx2 = the wires; ASEx = the state cells; UV clock = the conductor.
| Layer | Component | Role |
|---|---|---|
| Transport | iceoryx2/zinc topics | Inter-component event fabric. Observable in /dev/shm. Single-writer per topic. |
| State | ASEx (ASExGuardedState + ASExWorker) | Validate-before-mutate state serialization. One writer thread per state. NOT a message bus. |
| Dispatch | UV Clock (asyncio event loop) | Event-driven reactor. Subscribes to event types. Nothing sleeps-and-polls. |
Per T19: ASExWatch is DEFERRED from the critical path until heavy-test hangs are fixed. Use ASExWorker (FulfilmentWorker) for production hot-path mutations.
┌──────────────────────────────────────────────────────────────────────┐
│ Zinc SHM (POSIX /dev/shm/zinc_*) │
│ malkhut_book_state | malkhut_account_state | malkhut_fulfilment_out │
│ malkhut_risk_gate | malkhut_control (CONTROL_PLANE) │
└───────────────────────────────┬──────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────────────┐
│ ASEx VALIDATE-BEFORE-MUTATE │
│ GuardedFulfilmentState — book/account/intent/policy mutations │
│ GuardedRiskState — risk gate + kill switch │
│ FulfilmentWorker — single-threaded ASExWorker per state │
│ FulfilmentWatch — zero-overhead ring buffer for hot path │
│ ShardedWorker — per-symbol state partitioning │
└───────────────────────────────┬──────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────────────┐
│ CODE WORLD MODEL (CWM) │
│ deterministic exchange transition function │
│ f(state, joint_actions, latency_model, venue_rules) → next_state │
│ wraps hftbacktest for replay verification │
└───────────────────────────────┬──────────────────────────────────────┘
│
┌──────────┴──────────┐
▼ ▼
┌─────────────────────────┐ ┌──────────────────────────────────┐
│ LIVE FULFILMENT PLANNER │ │ OFFLINE CMA-ES TRAINER │
│ Decoupled UCB/UCT │ │ pycma + nevergrad │
│ ≤ 25 ms cadence │ │ growing self-play pool │
└───────────┬──────────────┘ └─────────────────┬────────────────┘
│ │
▼ ▼
┌─────────────────────────┐ ┌──────────────────────────────────┐
│ RISK / COMPLIANCE GATE │ │ POLICY REGISTRY / POOL │
│ leverage, notional, │ │ accepted candidates, weights │
│ self-trade, kill switch │ │ diversity-preserving eviction │
└───────────┬──────────────┘ └──────────────────────────────────┘
│
▼
┌─────────────────────────┐
│ VENUE ADAPTER (BingX) │
│ create/cancel/replace │
└───────────┬──────────────┘
│
▼
┌─────────────────────────┐
│ ClickHouse Persistence │
│ decisions, episodes, │
│ snapshots, discrepancies│
└─────────────────────────┘
Package Structure
MALKHUT/
├── malkhut/
│ ├── __init__.py # package version
│ ├── state.py # frozen data model (42 dataclasses)
│ ├── actions.py # action model, PlannedPolicy, RiskDecision
│ ├── features.py # feature extraction (17 features)
│ ├── counterparties.py # adversarial agent ecology (4 agents)
│ ├── engine.py # FulfilmentEngine (hot-path orchestrator)
│ ├── bench_numba.py # numba speedup benchmark
│ ├── launch_smoke_test.py # 10-min smoke test launcher
│ ├── cwm/
│ │ ├── __init__.py
│ │ ├── core.py # CWM (deterministic exchange simulator)
│ │ ├── numba_core.py # numba-accelerated hot loops
│ │ └── replay_verify.py # replay verification (mandatory before trust)
│ ├── planner/
│ │ ├── __init__.py
│ │ ├── action_menu.py # compact action space construction
│ │ └── sm_mcts.py # Decoupled UCB/UCT simultaneous-move MCTS
│ ├── risk/
│ │ ├── __init__.py
│ │ └── gate.py # risk/compliance gate (hard invariants)
│ ├── venue/
│ │ ├── __init__.py
│ │ └── bingx/
│ │ ├── __init__.py
│ │ └── adapter.py # BingX VST/Live venue adapter (wraps DITAv2)
│ ├── ipc/
│ │ ├── __init__.py
│ │ ├── zinc_plane.py # Zinc shared memory IPC (POSIX SHM)
│ │ └── control_plane.py # CONTROL_PLANE shared memory region
│ ├── storage/
│ │ ├── __init__.py
│ │ ├── ch_store.py # ClickHouse persistence (5 tables)
│ │ └── asset_store.py # DuckDB file-backed asset universe store
│ ├── execution/
│ │ ├── __init__.py
│ │ └── asex_integration.py # ASEx validate-before-mutate kernel
│ ├── clock/
│ │ ├── __init__.py
│ │ ├── host.py # UVClock — event-driven reactor (T19)
│ │ ├── events.py # Typed events: Scan, Tick, Timer, BarFire, Stale
│ │ ├── staleness.py # Freshness watchdog (T19 Staleness Law)
│ │ └── deadnode.py # DeadNode reaper for iox2 orphans
│ ├── training/
│ │ ├── __init__.py
│ │ ├── cma_trainer.py # CMA-ES + self-play + parallel workers
│ │ ├── registry.py # Policy lifecycle (CANDIDATE → ACTIVE)
│ │ ├── pipeline.py # Bounded continuous learning loop + logger
│ │ ├── dsl.py # Strategy DSL v2 (40+ primitives, 40+ sensors)
│ │ ├── generator.py # Genetic programming strategy evolution
│ │ ├── order_types.py # Standardized order type enum (FIX/CCXT-aligned)
│ │ ├── selector.py # Regime → strategy mapping + performance matrix
│ │ ├── asset_classification.py # Multi-label taxonomy + exchange registry
│ │ ├── asset_behavior.py # 10-dimension behavior DSL, research-validated
│ │ ├── asset_compiler.py # Auto-fetch from Binance/BingX, compile profiles
│ │ ├── asset_bridge.py # Directory ↔ classification sync
│ │ ├── parallel_eval.py # ProcessPoolExecutor episode runner
│ │ ├── ray_eval.py # Ray-based eval (industrial alternative)
│ │ ├── vbt_analysis.py # Post-sim metrics: Sharpe, Sortino, VaR
│ │ ├── advantage_scorer.py # Advantage estimation for offline analysis
│ │ ├── cognition.py # Rate-limited market regime research
│ │ ├── regime_expansion.py # 200+ regimes from dimension combinations
│ │ ├── news_sources.py # 12 industry-standard news sources
│ │ └── monitor.py # Cognition metrics, health, alerts
│ ├── daat/ # Direction-Anchored Ambiguity Triage
│ │ ├── __init__.py
│ │ └── core.py # DaatQuery, DaatVerdict, daat_classify
│ ├── cognition_launcher.py # Standalone long-run cognition service
│ ├── continuous_pipeline.py # Continuous training pipeline
│ └── tests/ # 1178 tests across 50 test files
├── specs/
│ └── MALKHUT_ADVERSARIAL_SELFPLAY_SPEC.py # full spec
└── README.md # this file
Shared Memory IPC (Zinc)
All IPC uses real Zinc shared memory (POSIX SHM via /dev/shm/zinc_*), NOT
the file-based transport. This is lock-free, zero-copy, GraalVM-compatible.
Regions
| Region | Writer | Readers | Purpose |
|---|---|---|---|
malkhut_book_state |
data feed | planner, CWM | canonical order book snapshot |
malkhut_account_state |
data feed | planner, risk | account/position snapshot |
malkhut_fulfilment_out |
engine | other systems | planner output + distribution |
malkhut_risk_gate |
engine | other systems | risk gate decisions |
malkhut_control |
external | engine | commands, symbols, lifecycle |
Framing
All regions use UVZINC01 seqlock framing:
[8B magic "UVZINC01"] [8B seq] [8B json_size] [UTF-8 JSON body]
CONTROL_PLANE
The malkhut_control region is the management interface:
- External systems write
ControlCommandframes - Engine reads and processes commands
- ACKs written back for
ack_requiredcommands
Commands: START, STOP, PAUSE, RESUME, SET_SYMBOLS, CONNECT_VENUE,
DISCONNECT_VENUE, SET_MODE, EMERGENCY_STOP, HOT_RELOAD_POLICY,
STATUS_REQUEST
ClickHouse Persistence
Database: dolphin_malkhut on localhost:8123
| Table | Purpose |
|---|---|
replay_steps |
CWM replay traces |
fulfilment_decisions |
every live decision |
self_play_episodes |
training episode results |
policy_snapshots |
versioned policy registry |
live_discrepancies |
simulation vs actual |
External Dependencies (Spec-Mandated)
| Library | Purpose | Used For |
|---|---|---|
pycma (cma) |
CMA-ES optimisation | training/cma_trainer.py |
| hftbacktest | replay/queue/latency fill simulation | CWM replay verification |
| nevergrad | gradient-free optimiser zoo | trainer comparison |
| msgspec | fast typed serialization | state/event encoding |
| numba | hot-loop acceleration | CWM transitions, feature extraction |
| polars | offline analytics | large trade-path corpora |
| lmdb | existing SILOQY storage | ticks, candles, hd_vectors |
| numpy | vectorized feature extraction | parameter decoding |
| scipy | stats tests, bootstrap | CI, promotion gates |
Quick Start
from malkhut.engine import FulfilmentEngine
from malkhut.state import FulfilmentPolicyParams
def my_params():
return FulfilmentPolicyParams(
version="v0.1", ucb_c=1.414, max_sims=64, max_depth=3,
rollout_depth=3, root_temperature=0.5, min_root_entropy=0.25,
quote_offsets_ticks=(0, 1, 2), quote_size_fractions=(0.10, 0.25, 0.50),
passive_ttl_ms=200, aggressive_ttl_ms=50,
maker_edge_min_bps=0.5, cross_spread_edge_min_bps=5.0,
adverse_toxicity_cancel_threshold=0.5, queue_churn_cancel_threshold=0.5,
mae_tail_cut_bps=50.0, mfe_giveback_cut_fraction=0.5,
max_time_in_loss_s=300.0, failed_recovery_cut_count=3,
recovery_velocity_min_bps_per_s=0.0,
max_symbol_notional_fraction=0.20, max_single_order_notional_fraction=0.05,
reduce_when_global_up_fraction=0.30, session_profit_lock_fraction=0.02,
w_expected_pnl=1.0, w_fill_probability=0.5, w_adverse_selection=2.0,
w_queue_priority=0.5, w_inventory_risk=1.5, w_tail_loss=5.0,
w_fee_quality=0.5, w_time_decay=0.3, w_policy_entropy=0.5,
robust_tail_weight=2.0, toxic_counterparty_weight=3.0,
low_liquidity_weight=2.0, latency_stress_weight=1.0,
)
engine = FulfilmentEngine(params_provider=my_params)
# engine.on_state(latest_market_state) # hot path
ASEx Integration (Validate-Before-Mutate)
MALKHUT uses ASEx (Advanced Serial Executor) for all mutable state mutations. This is the same kernel used by VIOLET/DOLPHIN for race-proof state management.
Core Pattern
class GuardedFulfilmentState(ASExGuardedState):
def _validate(self, mutation) -> bool:
# Pure predicate — MUST NOT mutate self
return mutation.ts_ns > 0 and len(mutation.bids) > 0
def _apply(self, mutation):
# Only called after _validate() passes
# MAY mutate self
self._state = new_state
return result
Components
| Component | Purpose | Thread Safety |
|---|---|---|
GuardedFulfilmentState |
Engine state mutations (book, account, intent, policy) | One writer thread via ASExWorker |
GuardedRiskState |
Risk gate + kill switch | One writer thread via ASExWorker |
FulfilmentWorker |
Single-threaded ASExWorker wrapping GuardedFulfilmentState | Serialised via queue.Queue |
RiskWorker |
Single-threaded ASExWorker wrapping GuardedRiskState | Serialised via queue.Queue |
FulfilmentWatch |
Zero-overhead ring buffer for hot-path mutations | Producer/consumer lock-free |
ShardedWorker |
Per-symbol state partitioning (N independent workers) | No shared state between partitions |
Why ASEx
- No locks: One thread per state object. Races impossible by construction.
- No GC pressure: Frozen dataclasses + queue.Queue (CPython C-level ops).
- Validate before apply: Invalid mutations rejected without touching state.
- GraalVM compatible: No CPython-only APIs in the hot path.
- Battle-tested: 110 ASEx tests + 1109 MALKHUT tests = 1219 total.
Design Rules (from spec)
- Optimise an ecology-resistant distribution, not a quote.
- The book is not passive scenery. Ask "why did I get filled?"
- Replay correctness before search depth. Wrong CWM + deep search = confident nonsense.
- CMA-ES does not promote from noisy wins. Survive scenario suite + bootstrap CI.
- Preserve feature humility. Allow CMA-ES to discover non-obvious interactions.
- Path-aware SL/TP is part of fulfilment. Exiting = order-placement under adversarial liquidity.
- Never let the optimiser learn illegal behavior. Hard-code constraints.
- Keep hot path bounded. ≤ 100 ms. No live CMA, no large neural nets.
- Version everything. Every decision explainable by policy + CWM + feature + state + action + outcome.
- If shadow live diverges, trust live. Simulation is only a tool.
ADDENDUM — EXEC-STACK PLUGGABILITY + SPEC HARDENING
[Fable, integrator, 2026-07-06 — operator-ordered extension. Mimo: this section binds MALKHUT into the UV/DITAv2 stack; nothing above is overridden except where marked FIX.]
A. Where MALKHUT plugs in (the seat at the table)
MALKHUT is the execution layer under the UV Clock (see
prod/docs/uv_subspecs/UV_TASK_T19_UV_CLOCK_HOST.md + annexes):
UV Clock (T19) → PromotionBridge → KernelIntent (DITAv2 contracts)
→ MALKHUT FulfilmentEngine (replaces naive kernel-submit)
→ RISK GATE → BingX venue adapter → fills
Your Design Rule 6 ("path-aware SL/TP is part of fulfilment") is CONVERGENT with the operator's T19 ruling ("the whole point of UV = faster-than-scan-cadence TP/SL paths"). MALKHUT's planner is the intelligent form of the C11 tick-exit guardian. Sequencing: simple doctrinal tick-exits (C11) ship first via T19 step 3; MALKHUT supersedes them when certified — same seat, smarter occupant.
B. Intake contract (BINDING)
- Input = DITAv2
KernelIntent(prod/clean_arch/dita_v2/contracts.py): asset, side (TradeSide), action (ENTER/EXIT), reference_price, target_size, leverage, metadata (carriespromo_client_id). Accept KernelIntent natively; do not invent a parallel intent type. u-clientOrderId prefix is LAW (T9 seam): every venue order MALKHUT places carries the u- prefix from intent metadata. Non-negotiable — it is the attribution boundary vs BLUE and vs legacy VIOLET orders.- DUAL-LEVERAGE LAW:
KernelIntent.leverageis CONVICTION leverage [0.5, 9.0] — it sizes QUANTITY only. Venue leverage = the mapping inprod/bingx/leverage.py(linear → ROUND_HALF_EVEN → int clamp [1,3]). WIRE the existing typed L3 wrapper (prod/clean_arch/violet/exchange_leverage.py); litmus: 0.5→1, 4.75→2, 8.0→3, 9.0→3. The venue adapter MUST NOT pass conviction leverage raw to set-leverage. - Arming/two-man/kill: MALKHUT honors
UV_PROMOTED=1+ arm-file (/root/uv-wt/prime-live/UV_PROMOTED.arm), re-evaluated PER EVENT (never cached — the T12 kill-switch lesson). DARK default = ObserveOnly venue. Map arm-file-delete → yourEMERGENCY_STOPsemantics: inert next event, no restart.
C. Fabric + state alignment (the tonight-rulings)
- Stack law: iceoryx2 = wires, ASEx = state cells, UV clock = conductor. Your ASEx usage (GuardedState + Worker per state object) is CORRECT — it is exactly what ASEx is for. Your zinc regions already speak UVZINC01 seqlock framing — good; when T19's event topics land, consume ScanEvent/TickEvent (with PROVENANCE: scan_number, scan_ts, ingest_ts, AGE) from the shared LiveSource instead of a private feed.
- Account state: read the DITAv2 ASEx account core (asex_account / AccountProjectionV2) as a CONSUMER only. MALKHUT never becomes a second writer to account state (two-cores-share-nothing law); its own guarded states are its own.
- Staleness law (T19 Annex A) — extend your Rule 2: a stale book is not scenery either, it is an ADVERSARY. Planner receives AGE on every input and must discount/refuse plans over stale state; freshness watchdog emits StaleInput.
- ASEx pre-condition inherited: DeadNode reaper / orphan-sweep on start (iox2 segments orphan on crash — known gap).
D. Persistence + observability alignment
dolphin_malkhutnamespace: approved (own-namespace law). NO TTL on any table (retention doctrine — audit surfaces never expire).- Cross-reference law: every
fulfilment_decisionsrow carriesintent_id+trade_idfrom the source KernelIntent → end-to-end trace joinsdolphin_uv.exec_journal⇄dolphin_malkhut.*⇄ venue order history. This is what makes MALKHUT certifiable instead of a black box. - Publish planner pulse (decision rate, plan entropy, gate verdicts, AGE-at-decision) into the zinc snapshot → extant TUIs render it (TUI-first doctrine).
E. Ecology + reward grounding (FIX-grade improvements)
- Ground the counterparty ecology in OUR tape: fit/validate adversary populations
against
dolphin.obf_universerecorded sub-second tape and the WT2 stress manifold. The CLASS-M/CLASS-X hour sets (WT2D) are your ready-made adversarial scenario suites — corrupt/stress hours ARE the hostile ecology, measured. CLASS-X law applies: extreme magnitudes are a FEATURE channel, never filtered. - Latency model = our measured tails, not defaults: scan_to_fill/step_bar distributions from the WT1 latency audit (median 47ms/1.58ms, stop-tail 7.8s) — calibrate the CWM's latency injections from these empirical curves.
- Reconcile the budget statement (FIX): planner box says ≤25ms, Rule 8 says ≤100ms — bind them: ≤25ms planning inside a ≤100ms end-to-end intent→venue-call budget, both enforced by test.
- Promotion gate binds to the cert conveyor: CMA-ES candidate promotion (Rule 4)
emits artifacts via the T4 cert reporter into the gate ledger; policy snapshots
versioned in
dolphin_malkhut.policy_snapshotsAND referenced inprod/docs/uv_cert/. No policy trades live without a ledger row.
F. Certification pathway (Gate-M — how MALKHUT reaches money)
- CWM replay verification (your Rule 3) vs hftbacktest — already specced; add determinism ×2 (byte-identical minus ts).
- Dual-leverage litmus (B.3 values) as a named test.
- DARK shadow window: MALKHUT plans on live intents while the simple path executes; diff planned-vs-naive on fills/slippage/adverse-selection — the improvement claim must be MEASURED on VST before MALKHUT takes the seat.
- Risk-gate mutation litmus: drop a constraint → a test goes RED (looks-finished doctrine; 222 green tests ≠ verification until mutations bite).
- Rule 10 stands supreme: shadow-vs-live divergence → trust live, stand down, file the discrepancy row.
End addendum. — F.
DEVELOPMENT STATUS (2026-07-13)
1204 test functions. 50 test files. All green. 0 failures. 0 regressions.
Fable Spec Items — All Complete
| # | Item | Module | Status |
|---|---|---|---|
| 1 | Mutation-litmus test | test_mutation_litmus.py |
✅ PASSES |
| 2 | Re-run benchmarks at corrected fees | (long run) | ⏸️ Deferred |
| 3 | Maker fee UNVERIFIED comment | asset_classification.py |
✅ Done |
| 4 | ScenarioLibrary sweep | scenario_library.py |
✅ 7,488 grid points |
| 5 | PerformanceMatrix manifold | selector.py |
✅ confidence+support |
| 6 | ActualsLoader | actuals_loader.py |
✅ CH reader + fallback |
| 7 | OOD verdict | risk/gate.py |
✅ daat_verdict parameter |
| 8 | Manifold query | manifold_query.py |
✅ three-phase recommend |
| 9 | DAAT package | daat/core.py |
✅ DaatVerdict{KNOWN,MARGINAL,OOD} |
| 10 | Book fidelity | book_fidelity.py |
✅ OBF → OrderBookState |
Fee Correction (CRITICAL — all prior training was at 10x too-cheap fees)
Source of truth: dolphin.trade_execution_quality (8006 rows, avg taker=5.016 bps).
| Venue | Old taker (WRONG) | Correct taker | Old maker (WRONG) | Correct maker |
|---|---|---|---|---|
| BingX | 0.5 bps | 5.0 bps | -0.2 (rebate) | +2.0 (POSITIVE) |
| Binance | 0.4 bps | 4.5 bps | -0.2 | -0.2 (UNVERIFIED) |
| Bybit | 0.06 bps | 5.5 bps | -0.01 | -0.1 (UNVERIFIED) |
Every policy trained before this fix is suspect — cheap fees reward overtrading. Re-measurement at correct fees is required for production deployment.
Completed subsystems
| Subsystem | Module | Tests | Status |
|---|---|---|---|
| State Model | state.py |
17 | 42 frozen dataclasses, immutable |
| CWM | cwm/core.py + cwm/numba_core.py |
103 | Exchange mechanics + numba JIT (5.3µs/transition) + vectorized reward + vectorized UCB + batch MCTS kernel |
| Replay Verification | cwm/replay_verify.py |
65 | Deep comparison, binary search, trajectory recording |
| Planner | planner/sm_mcts.py |
11 | Decoupled UCB/UCT, ≤25ms budget |
| Action Menu | planner/action_menu.py |
(in planner) | Compact action space construction |
| Counterparties | counterparties.py |
19 | 4 adversarial agent policies |
| Risk Gate | risk/gate.py |
4 | Hard invariants: leverage, self-trade, post-only |
| ASEx Integration | execution/asex_integration.py |
33 | Validate-before-mutate, single-writer, no locks |
| Zinc IPC | ipc/zinc_plane.py |
8 | Real POSIX SHM, UVZINC01 framing |
| Control Plane | ipc/control_plane.py |
3 | HOT_RELOAD_POLICY, kill switch |
| ClickHouse | storage/ch_store.py |
9 | 5 tables, HTTP API |
| DuckDB Asset Store | storage/asset_store.py |
30 | File-backed + in-memory materialization, sub-µs reads |
| Asset Bridge | training/asset_bridge.py |
49 | Directory ↔ classification sync |
| CMA-ES Training | training/cma_trainer.py |
65 | Behavior-driven, auto-compile, parallel workers, 7x speedup |
| Parallel Eval | training/parallel_eval.py |
16 | ProcessPoolExecutor, 7x CMA-ES speedup |
| Ray Eval | training/ray_eval.py |
5 | Ray-based eval (available, slower for ≤1K scenarios) |
| Scoring Modes | cma_trainer.py + advantage_scorer.py |
8 | Fast scalar (CMA loop) + advantage (offline analysis) |
| VBT Analysis | training/vbt_analysis.py |
8 | Post-sim trade metrics: Sharpe, Sortino, VaR, cross-asset |
| ScenarioLibrary | training/scenario_library.py |
7 | Sweep 7,488 grid points (spread×depth×tox×regime) |
| Manifold Query | training/manifold_query.py |
(new) | Three-phase: DAAT → manifold → recommend |
| ActualsLoader | training/actuals_loader.py |
(new) | Reads CH tables for Mode 2 live query |
| Book Fidelity | training/book_fidelity.py |
(new) | OBF 15B rows → OrderBookState synthesis |
| DAAT | daat/core.py |
9 | Direction-Anchored Ambiguity Triage |
| OOD Verdict | risk/gate.py |
+1 | daat_verdict parameter, OOD → doctrinal fallback |
| Policy Registry | training/registry.py |
14 | CANDIDATE → ACTIVE lifecycle |
| Training Pipeline | training/pipeline.py |
21 | Bounded continuous learning loop |
| Strategy DSL v2 | training/dsl.py |
69 | 40+ primitives, 40+ sensors, 16 builtins |
| Order Types | training/order_types.py |
(new) | Three-dimensional: OrderType × TimeInForce × Instructions (FIX-aligned) |
| Strategy Generator | training/generator.py |
20 | Genetic programming: crossover, mutation, tournament |
| Strategy Selector | training/selector.py |
24 | Regime × strategy × venue matrix, cross-exchange comparison |
| Asset Classification | training/asset_classification.py |
190 | Multi-label taxonomy + exchange registry, 13 assets, system-wide store |
| Asset Behavior DSL | training/asset_behavior.py |
(in classification) | 10 orthogonal dimensions, 3 templates, 13 behaviors, research-validated |
| Asset Compiler | training/asset_compiler.py |
(new) | Auto-fetch from Binance/BingX API, compile profiles, rate-limited |
| Cognition Pipeline | training/cognition.py |
29 | Rate-limited, 8 sources, dedup, perm-run |
| Regime Expansion | training/regime_expansion.py |
29 | 200+ regimes from 4×4×4×4 dimensions |
| News Sources | training/news_sources.py |
12 | 12 industry-standard sources, ranking |
| Cognition Monitor | training/monitor.py |
9 | Metrics, health scoring, alerts, JSONL logging |
| Cognition Launcher | cognition_launcher.py |
21 | Standalone long-run cognition service |
| UV Clock | clock/host.py |
30 | T19 event-driven reactor |
| Numba Acceleration | cwm/numba_core.py |
13 | JIT on hot loops, 1.8x fill speedup |
| BingX Adapter | venue/bingx/adapter.py |
28 | Wraps DITAv2, order tracking |
| Hypothesis | test_hypothesis_properties.py |
6 | Property-based fuzzing |
| Fuzz/Adversarial | test_fuzz.py, test_adversarial.py |
17 | Random stress + adversarial scenarios |
| Concurrency | test_concurrency.py |
8 | Thread safety, race detection |
| Sync/Async Seams | test_sync_async_seams.py |
7 | Zinc latency, engine budget |
| E2E Integration | test_e2e_integration.py |
2 | Full pipeline: train→register→plan→risk→venue→zinc→ch→reload |
Performance Benchmarks
| Metric | Value |
|---|---|
| CWM throughput | 189K calls/sec (numba JIT) |
| CWM per-call latency | 5.3 µs |
| CWM reward (numba vectorized) | ~0.3µs |
| UCB selection (numba vectorized) | 1.13µs (was 8.7µs, 7.7x speedup) |
| Numba fill speedup | 1.8x (batch 100) |
| DuckDB asset reads | 0.2µs (in-memory) |
| Scenario generation | 390 scenarios in 0.8s |
| CMA-ES parallel (8 workers) | 7× speedup, 87% efficiency |
| Best CMA-ES score (48 evals) | 8,628 (fast scalar, 100.4 bps PnL) |
| Score/min (8 workers) | 3,043 |
| Episode throughput (parallel) | 16ms/ep |
| Episode throughput (sequential) | 23ms/ep |
Bugs found and fixed (22 total)
- Empty book crash in
materialize_price_from_action(SELL side) - Empty book crash in
total_notionalcalculation - Empty book crash in mark-to-market
- Empty book crash in
_inventory_risk - Empty book crash in counterparty processing
- Empty book crash in
_update_path_state - Empty book crash in feature extractor (
b.mid) - Counterparty fills used
max()instead of consuming levels - Position overwritten instead of accumulated on add
- Test helper
_state()not forwardingpositionskwarg _make_open_orderset empty symbol instead of venue symbolFulfilmentPolicyParamsfrozen dataclass__dict__access- Banker's rounding in tick tests
HOT_RELOAD_POLICYcompared float to string- Engine init order (workers before hot_reload)
_tp()helper missingfailed_recovery_countparameter- CWM: SELL position flip didn't reset avg_entry
- CWM: Counterparty fills overwrote our fill tracking
- CWM: Double-counting unrealized PnL in equity
- UnboundLocalError when max_steps=0
- fill_count never incremented in training
- Cancel rate limiter incorrectly gated order placements
Key design rules enforced
- ✅ Optimise distribution, not single quote
- ✅ Book is not passive scenery
- ✅ Replay correctness before search depth (mandatory gate)
- ✅ CMA-ES does not promote from noisy wins (bootstrap CI)
- ✅ Preserve feature humility (CMA discovers interactions)
- ✅ Path-aware SL/TP is part of fulfilment
- ✅ Never let optimiser learn illegal behavior
- ✅ Hot path bounded (≤100ms)
- ✅ Version everything
- ✅ Shadow-vs-live divergence → trust live
Strategy Selection (completed)
The system classifies strategies per market condition and floats the best to top:
market_fingerprint → Regime Classifier (8 regimes)
↓
Performance Matrix (regime × strategy → score)
↓
Strategy Selector → Engine uses best strategy
Upstream integration path
- ExoF — funding rates, open interest, liquidation levels
- MARAS — 5-tier ensemble → 8-regime classification
- DOLPHINNG7 vel_div — velocity divergence signals
- DOLPHINNG7 eigenscan — eigenspace correlation analysis
G. Strategy DSL v2 — Third-Party Strategy Development
MALKHUT includes a Domain-Specific Language (DSL) for strategy composition. Third parties can write, evolve, and share strategies without modifying core code.
DSL Format
STRATEGY "passive_maker" {
PRIORITY 1: IF spread_bps < 5.0 AND orderflow_toxicity < 0.3 THEN QUOTE(BUY, 0, 0.25)
PRIORITY 2: IF spread_bps < 5.0 AND orderflow_toxicity < 0.3 THEN QUOTE(SELL, 0, 0.25)
PRIORITY 3: IF orderflow_toxicity > 0.7 THEN CANCEL_ALL
PRIORITY 4: IF time_in_trade > 300 THEN EXIT
PRIORITY 5: IF unrealized_pnl < -50bps THEN STOP_LOSS
PRIORITY 6: NOOP
}
Action Primitives (40+)
| Category | Primitives |
|---|---|
| Passive | QUOTE, JOIN_QUEUE, STEP_BACK, LADDER, GRID, ICEBERG, TWAP |
| Aggressive | CROSS, SNIPER, PING |
| Cancellation | CANCEL, CANCEL_ALL, CANCEL_AND_HOLD, REQUOTE |
| Position | EXIT, HALF_EXIT, QUARTER_EXIT, TAKE_PARTIAL, STOP_LOSS, TAKE_PROFIT, TRAILING_STOP, EMERGENCY_EXIT, FLAT_ALL |
| Sizing | SCALE_IN, SCALE_OUT, INCREASE_SIZE, REDUCE_SIZE |
| Stop/Target | MOVE_STOP, MOVE_TAKE_PROFIT, BRACKET, OCO |
| Hedging | HEDGE, PAIR_TRADE |
| Waiting | HOLD, WAIT_FOR_FILL, WAIT_FOR_PRICE, WAIT_FOR_SPREAD |
Market Sensors (40+)
| Category | Sensors |
|---|---|
| Book | spread_bps, imbalance_3/5/10, bid_depth_3/5/10, bid_ask_ratio |
| Flow | toxicity, queue_churn, cross_venue_lead, order_book_toxicity |
| Momentum | price_momentum_1s/5s/15s/1m |
| Position | position_qty, pnl_bps, leverage, position_age_s |
| Path | mae_bps, mfe_bps, time_in_loss, failed_recoveries, recovery_velocity |
| Regime | volatility, atr_14/50, rsi_14, funding, regime_score |
| Account | equity, risk_budget_used, session_pnl, drawdown |
| Performance | profit_factor, sharpe_ratio, consecutive_wins/losses |
| Time | current_hour, day_of_week, is_weekend, is_liquid_hours |
Comparison Operators (12)
>, <, >=, <=, ==, !=, abs>, abs<, changing, stable, crossing_above, crossing_below
Builtin Strategies (16)
| Strategy | Description |
|---|---|
passive_maker |
Passive limit orders, toxicity avoidance |
aggressive_taker |
Momentum-based aggressive crosses |
toxicity_avoider |
Cancels on high toxicity, quotes when safe |
path_risk_exit |
Path-aware SL/TP exits |
regime_adaptive |
Adapts to MARAS regime classification |
momentum_catcher |
Catches momentum moves, partial exits |
mean_reversion |
Wide-spread mean reversion |
scalper |
Ultra-tight spreads, small sizes |
inventory_manager |
Position size control |
session_guard |
Time-based position management |
liquidity_hunter |
Deep book exploitation |
volatility_breakout |
Volatility-based breakouts |
funding_arb |
Funding rate arbitrage |
grid_trader |
Grid-based accumulation |
risk_parity |
Risk budget management |
hybrid_adaptive |
Multi-condition adaptive |
Strategy Generator (Genetic Programming)
The system evolves strategies via genetic operators:
from malkhut.training.generator import StrategyGenerator, GeneratorConfig
config = GeneratorConfig(
population_size=20, generations=5, tournament_size=3,
mutation_rate=0.15, crossover_rate=0.7,
)
generator = StrategyGenerator(config=config)
population = generator.evolve(baseline_params, scenarios)
# Get successful strategies
successful = generator.get_successful_strategies(min_fitness=0.0)
for genome in successful:
generator.add_to_pool(genome)
Strategy Selector (Market-Adaptive)
The system classifies strategies per market regime and selects the best:
from malkhut.training.selector import StrategySelector, MarketRegime
selector = StrategySelector()
result = selector.select(current_state, available_strategies)
# result.strategy_id → best strategy for current regime
# result.confidence → how confident we are
# result.alternatives → other considered strategies
Regimes: trending_up, trending_down, high_volatility, low_volatility,
mean_reverting, momentum, choppy, liquidity_hole, normal
Asset Classification (Invariant Taxonomy)
Multi-label asset classification using only invariant characteristics — properties intrinsic to the token that predict price behaviour and order-book performance. Three-layer identifier architecture for cross-system interoperability.
Design principle: Classify by WHAT an asset IS (fundamental), not how it's trading right now. Volatility/liquidity appear as long-run statistical profiles, not as classification axes that change minute-to-minute.
Three-Layer Identifier Architecture
| Layer | Fields | Purpose | Example |
|---|---|---|---|
| 1. Canonical | symbol, base_asset, name, unified_symbol, quote_currency |
Venue-independent identity | BTCUSDT / BTC / "Bitcoin" / BTC/USDT / USDT |
| 2. Cross-system | coingecko_id, cmc_id, blockchain, contract_address |
Data provider + on-chain mapping | "bitcoin" / 1 / "ethereum" / "" |
| 3. Exchange mapping | exchanges: tuple[str, ...] |
Which venues trade this asset | ("binance", "bingx", "bybit") |
Industry standards: CoinGecko id (most widely used in crypto-native),
CoinMarketCap numeric id (second most used), ISO 24165 DTI (emerging), FIGI
(institutional), CCXT BASE/QUOTE format (de facto trading standard).
Fundamental Dimensions (intrinsic, never change)
| Dimension | Values | Multi-label | Purpose |
|---|---|---|---|
| Sector | CURRENCY, LAYER1, LAYER2, DEFI, ORACLE, EXCHANGE, MEME, PRIVACY, STORAGE, GAMING_NFT | Yes | Primary use-case / vertical |
| TokenRole | GAS, STORE_OF_VALUE, GOVERNANCE, UTILITY, MEME, EXCHANGE_FEE | Yes | Functional role → demand elasticity |
| SupplyModel | FIXED_CAP, DISINFLATIONARY, INFLATIONARY, BURN_MECHANISM | No | Supply pressure dynamics |
| ConsensusFamily | POW, POS, DPOS | No | Miner/validator selling behaviour |
| SmartContractCapability | FULL, PARTIAL, NONE | No | DeFi composability |
Technical Dimensions (invariant market-structure properties)
| Dimension | Values | Purpose |
|---|---|---|
| MarketCapTier | MEGA (>500B), LARGE (50-500B), MID (5-50B), SMALL (500M-5B), MICRO (<500M) | Price impact per dollar |
| VolatilityProfile | LOW (<30%), MEDIUM (30-80%), HIGH (80-150%), EXTREME (>150%) | Long-run average vol |
| LiquidityProfile | DEEP (>100M), NORMAL (10-100M), THIN (1-10M), ILLIQUID (<1M) | Long-run avg depth |
| DerivativeAccess | PERPS_AND_OPTIONS, PERPS_ONLY, NONE | Shorting / funding availability |
| tick_size / lot_size | Exchange-set | Minimum price/order size |
| maker/taker fee_bps | Exchange-set | Trading cost |
| typical_spread_bps / depth / volume | Long-run averages | Order-book fingerprint |
Predictive Properties (derived from invariants)
| Property | Logic | Value |
|---|---|---|
supply_pressure |
PoW → forced selling (miners cover electricity) | "forced" / "optional" |
demand_elasticity |
GAS or STORE_OF_VALUE → inelastic (must hold) | "inelastic" / "elastic" |
can_be_shorted |
DerivativeAccess != NONE | bool |
Multi-Label Examples
BTC: sectors=(CURRENCY), roles=(STORE_OF_VALUE)
ETH: sectors=(LAYER1, DEFI), roles=(GAS, STORE_OF_VALUE, GOVERNANCE)
BNB: sectors=(EXCHANGE, LAYER1), roles=(EXCHANGE_FEE, GAS)
DOGE: sectors=(MEME, CURRENCY), roles=(MEME, GAS)
DOT: sectors=(LAYER1), roles=(GAS, GOVERNANCE)
UNI: sectors=(DEFI), roles=(GOVERNANCE, UTILITY)
Querying
from malkhut.training.asset_classification import *
# Multi-label queries (match if queried value is ANY of the asset's labels)
layer1s = get_assets_by_sector(Sector.LAYER1) # ETH, SOL, ADA, AVAX, BNB, DOT, ATOM
gas_tokens = get_assets_by_token_role(TokenRole.GAS) # ETH, SOL, ADA, AVAX, DOGE, BNB, MATIC, DOT, ATOM
# Predictive properties
forced = get_pov_assets() # BTC, DOGE (PoW miners with forced selling)
inelastic = [p for p in ASSET_PROFILES.values() if p.demand_elasticity == "inelastic"]
# Backward-compatible single-label access (first element = primary)
btc = get_asset_profile("BTCUSDT")
btc.sector # Sector.CURRENCY
btc.token_role # TokenRole.STORE_OF_VALUE
Exchange Registry (Multi-Exchange Asset Universe)
MALKHUT serves as the system-wide asset universe store for BLUE, VIOLET, UV, and all downstream systems. Each asset knows which exchanges trade it; each exchange has a standardized profile.
ExchangeProfile
| Field | Type | Purpose |
|---|---|---|
exchange_id |
str | Canonical key: "binance", "bingx", "bybit" |
display_name |
str | Human-readable name |
has_spot |
bool | Spot trading available |
has_perps |
bool | Perpetual futures available |
has_options |
bool | Options available |
api_base_url |
str | REST API root URL |
ws_base_url |
str | WebSocket root URL |
default_taker_fee_bps |
float | Default taker fee |
default_maker_fee_bps |
float | Default maker fee |
typical_latency_ms |
float | Typical API latency |
Pre-defined Exchanges
| Exchange | Spot | Perps | Options | Taker Fee | Latency |
|---|---|---|---|---|---|
| Binance | ✓ | ✓ | ✓ | 0.4 bps | 40 ms |
| BingX | ✓ | ✓ | ✗ | 0.5 bps | 100 ms |
| Bybit | ✓ | ✓ | ✓ | 0.06 bps | 50 ms |
Exchange-Asset Mapping
Each AssetProfile carries an exchanges: tuple[str, ...] field listing which venues
trade the asset. Default: ("binance",). Other systems (BLUE/VIOLET/UV) import their
asset universes into this store.
from malkhut.training.asset_classification import *
# Which exchanges trade BTC?
btc = get_asset_profile("BTCUSDT")
print(btc.exchanges) # ('binance',)
# All assets on Binance
binance_assets = get_assets_on_exchange("binance")
# Assets traded on BOTH Binance and BingX
common = get_common_assets("binance", "bingx")
# Which exchanges trade a given asset?
venues = get_exchange_for_asset("ETHUSDT")
# Exchange metadata
ex = get_exchange("binance")
print(ex.default_taker_fee_bps) # 0.4
print(ex.typical_latency_ms) # 40
Asset Behavior DSL (10-Dimension Research-Validated Model)
The Asset Behavior DSL decomposes each asset's trading behavior into 10 orthogonal dimensions, each with empirically-validated parameters from live Binance/BingX API data and academic literature.
Source: Binance live REST API (depth 500 levels, 1h klines, ticker, funding rates), BingX open API (depth, funding, premium index), academic literature (Bouchaud "Trades Quotes and Prices" 2018, Cont/Stoikov/Talreja 2010, Cartea/Jaimungal/Penalva 2015), CoinGlass OI data, industry MM disclosures.
10 Dimensions
| Dimension | What it captures | Example: BTC vs DOGE |
|---|---|---|
| DepthProfile | Book shape: amplitude, decay exponent α, fragility | BTC: $750K/0.70/0.10 vs DOGE: $22K/1.00/0.30 |
| SpreadProfile | Normal spread, stress multiplier | BTC: 0.01bps/50x vs DOGE: 1.35bps/15x |
| FlowProfile | Order rate, size distribution, cancel ratio | BTC: 300/s/$643/20x vs DOGE: 80/s/$96/8x |
| VolatilityProfile | Annualized vol, GARCH params, half-life | BTC: 35%/48h vs DOGE: 78%/24h |
| IntradayProfile | Peak/trough hours, ratio | BTC: 15:00 UTC, 7.4x ratio |
| WeekendProfile | Vol/volume/spread multipliers | 0.65x vol, 0.52x volume, 1.20x spread |
| CorrelationProfile | ETH beta, BTC corr (normal vs crash) | BTC: 1.0/1.0 vs UNI: 0.70/0.50 |
| MarketMakerProfile | Inventory limits, pull speed, margins | BTC: $10M/3ms/0.5bps vs DOGE: $500K/25ms/2bps |
| LiquidationProfile | OI/MCap, trigger %, cascade speed | BTC: 0.5%/6.5%/slow vs DOGE: 1.4%/4%/fast |
| FundingProfile | Mean/std funding, basis | BTC: 0.59bps/0.22bps vs DOGE: 0.39bps/0.34bps |
| RetailProfile | Retail ratio, inst gap | BTC: 35%/0.04 vs DOGE: 80%/0.31 |
| BingxProfile | BingX spread/depth multiplier, fees, latency | BTC: 12.6x/0.048x vs DOGE: 7x/1.27x |
3 Composable Templates
| Template | Assets | Key traits |
|---|---|---|
institutional_blue_chip |
BTC, ETH, BNB | Deep book, low spread, slow cascade, high institutional |
mid_cap_l1 |
SOL, ADA, AVAX, DOT, ATOM, UNI, LINK, MATIC, AAVE | Moderate depth, higher vol, medium cascade |
retail_meme |
DOGE | Thin book, high vol, fast cascade, retail-dominated |
13 Per-Asset Behaviors (research-validated)
| Asset | Price | Depth@1bps | Spread | Vol | Cascade | Retail | BingX mult |
|---|---|---|---|---|---|---|---|
| BTC | $64K | $750K | 0.01 bps | 35% | slow/fast | 35% | 12.6x |
| ETH | $1.8K | $600K | 0.01 bps | 66.5% | medium/medium | 35% | 9.0x |
| SOL | $80 | $400K | 1.26 bps | 72% | medium/medium | 72% | 1.7x |
| DOGE | $0.07 | $22K | 1.35 bps | 78% | fast/slow | 80% | 7.0x |
| ADA | $0.17 | $60K | 5.95 bps | 79% | medium/medium | 65% | 8.0x |
| UNI | $3.6 | $30K | 2.76 bps | 95% | medium/medium | 60% | 10x |
Usage
from malkhut.training.asset_behavior import *
# Get behavior for any asset
btc = get_behavior("BTCUSDT")
print(btc.depth.amplitude_usd) # $750,000
print(btc.vol.annualized_normal) # 35.0%
print(btc.liquidation.speed) # "slow"
# Compute depth at distance
depth_10bps = btc.depth_at_bps(10) # $1.5M
# Estimate slippage for a $100K order
slippage = btc.expected_slippage_bps(100_000) # ~2 bps
# Query by template, sector, or role
institutional = get_behaviors_by_template("institutional_blue_chip")
gas_tokens = get_behaviors_by_role(TokenRole.GAS)
thin_book = get_thin_book_assets()
Asset Compiler (Auto-Fetch from Live APIs)
The Asset Compiler automatically fetches market data from Binance and BingX public APIs
and produces system-ready AssetProfile + AssetBehavior for any tradeable symbol.
Rate-limited (1 req/sec, well under Binance 1200/min limit), cached (avoids re-fetching within a session), resumable.
What it auto-fetches
| Endpoint | Data | Computes |
|---|---|---|
/api/v3/ticker/24hr |
Price, volume, high/low | Daily volume, price reference |
/api/v3/depth?limit=100 |
Order book (100 levels) | Spread, depth amplitude, decay exponent α |
/api/v3/klines?interval=1h&limit=168 |
7 days hourly OHLCV | Annualized volatility, order flow stats |
/api/v3/exchangeInfo |
Symbol filters | Tick size, lot size, price decimals |
/fapi/v1/fundingRate |
Funding rates (20 intervals) | Mean/std/positive% funding |
/fapi/v1/openInterest |
Open interest | OI/MCap ratio for liquidation modeling |
Usage
from malkhut.training.asset_compiler import AssetCompiler
compiler = AssetCompiler()
# Single asset — auto-fetches from Binance, computes all params
result = compiler.compile("XRPUSDT")
print(result.asset_profile.symbol) # "XRPUSDT"
print(result.asset_behavior.reference_price) # $1.1056 (live Binance price)
print(result.asset_behavior.spread.normal_bps) # 0.90 bps (computed from depth)
print(result.asset_behavior.vol.annualized_normal) # 46.6% (computed from klines)
# Register into global registries
compiler.register(result)
# Now XRPUSDT is available for ScenarioFactory, queries, etc.
# One-step compile and register
result = compiler.compile_and_register("HBARUSDT")
# Batch compile (rate-limited sequentially)
results = compiler.compile_batch(["APTUSDT", "SEIUSDT", "INJUSDT"])
for r in results:
compiler.register(r)
Known classifications
The compiler includes pre-built classifications for 28 assets: BTC, ETH, SOL, DOGE, ADA, AVAX, UNI, LINK, BNB, MATIC, AAVE, DOT, ATOM, XRP, TRX, LTC, NEAR, APT, OP, ARB, SUI, PEPE, WIF, SEI, INJ, FIL, HBAR, IMX
Unknown assets get heuristic defaults (LAYER1/GAS/INFLATIONARY/POS/FULL/PERPS_ONLY).
Scenario Factory (Behavior-Driven, Label-Queryable)
The ScenarioFactory builds evaluation scenarios using real per-asset prices, depth profiles, and spread characteristics — no more hardcoded BTC prices.
Behavior-driven state creation
Every scenario derives its parameters from the asset's AssetBehavior:
| Scenario type | Spread multiplier | Depth fraction | Description |
|---|---|---|---|
normal |
1.0x | 100% | Calm, liquid market |
thin_book |
1.5x | 10% | Low participation |
wide_spread |
100x | 50% | High volatility |
toxic_stress |
2.0x | 30% | Adverse selection |
flash_crash |
3.0x | 5% | Book collapse |
liquidity_vacuum |
5.0x | 1% | Near-empty book |
weekend_thin |
5.0x | 5% | Weekend low-volume |
trending |
1.0x | 80% | Momentum move |
mean_reverting |
20x | 60% | Reversal setup |
funding_shock |
1.5x | 30% | Deleveraging event |
liquidation_cascade |
1.5x | 20% | Chain liquidation |
market_maker_withdrawal |
3.0x | 15% | MM pull-out |
Realistic per-asset pricing
Scenario: "normal_BTCUSDT_42"
BTC bid: $63,999.97 (11.7 BTC = $750K)
BTC ask: $64,000.03
Spread: 0.01 bps
Scenario: "normal_DOGEUSDT_42"
DOGE bid: $0.07 (314K DOGE = $22K)
DOGE ask: $0.07
Spread: 1.35 bps
Scenario: "flash_crash_SOLUSDT_47"
SOL bid: $80.00 (depth = 5% of normal = $20K)
SOL ask: $80.02
Spread: 3.78 bps (3x normal)
Auto-compilation for unknown assets
When you request a symbol that isn't pre-defined, the factory automatically compiles it from live Binance API data before building scenarios:
factory = ScenarioFactory()
# XRPUSDT is NOT in our pre-defined 13 assets
# But this auto-compiles it from Binance in ~5 seconds:
suite = factory.build_suite_for_symbols(("XRPUSDT",), steps_per_scenario=5)
# XRP behavior is now registered for future use too
Label-based query interfaces
factory = ScenarioFactory()
# By sector
suite = factory.build_suite_for_sector(Sector.LAYER1) # ETH, SOL, ADA, AVAX, BNB, DOT, ATOM
suite = factory.build_suite_for_sector(Sector.DEFI) # ETH, AVAX, UNI, AAVE
suite = factory.build_suite_for_sector(Sector.MEME) # DOGE
# By token role
suite = factory.build_suite_for_role(TokenRole.GAS) # 9 assets
suite = factory.build_suite_for_role(TokenRole.GOVERNANCE) # ETH, UNI, AAVE, DOT
# By behavior template
suite = factory.build_suite_for_template("institutional_blue_chip") # BTC, ETH, BNB
suite = factory.build_suite_for_template("retail_meme") # DOGE
# By volatility band
suite = factory.build_suite_for_volatility(min_ann=75, max_ann=200) # High-vol assets
# Multi-label union (any matching)
suite = factory.build_suite_for_labels(
sectors=[Sector.DEFI],
roles=[TokenRole.GOVERNANCE]
)
# Arbitrary symbols (auto-compiles unknowns)
suite = factory.build_suite_for_symbols(("BTCUSDT", "XRPUSDT", "HBARUSDT"))
Research-validated scenario diversity
Each asset's 30 scenario types are parameterized by its behavior profile:
| Scenario | BTC params | DOGE params | Why it matters |
|---|---|---|---|
| Normal | $750K depth, 0.01bps spread | $22K depth, 1.35bps spread | Baseline for scoring |
| Thin book | $75K depth | $2.2K depth | Tests fill rate under low liquidity |
| Flash crash | $37.5K depth, 0.03bps spread | $1.1K depth, 4bps spread | Tests adverse selection survival |
| Weekend thin | $37.5K depth, 5bps spread | $1.1K depth, 6.75bps spread | Tests off-hours strategy fitness |
| Liquidation cascade | $150K depth | $4.4K depth | Tests position sizing under stress |
| Market maker withdrawal | $112.5K depth | $3.3K depth | Tests quote quality when MMs pull |
Numba Performance
Hot-path functions JIT-compiled with numba:
| Function | Speedup |
|---|---|
| fill_from_levels (batch 100) | 1.8x |
| round_tick | JIT-compiled |
| round_lot | JIT-compiled |
| clip_lots | JIT-compiled |
| extract_features | Vectorized numpy |
CWM throughput: 157K calls/sec, 6.4 µs/transition, 0.64 ms/100-step episode.
Design Principles
- Hardcoded baseline is NEVER replaced — always available as reference
- Generated strategies are ADDED — system GROWS its repertoire
- Crossover combines two strategies (uniform crossover on parameters)
- Mutation randomly modifies parameters (Gaussian noise)
- Selection uses tournament selection (fitter genomes win more often)
- Third parties can write strategies in DSL text format
- Market adaptation: selector chooses strategy based on ExoF/MARAS regime
- Behavior-driven scenarios: every asset behaves like its real-world self
Strategy Naming Convention
{strategy_type}_gen{generation}_{timestamp}
Examples: UCB1_gen1_20260707_223754, THOMPSON_gen2_20260707_223754
H. Cambrian Expansion
Tie-In Fix
PerformanceMatrix now wired to evaluator. Strategies scored by regime fitness:
- ALL strategies tested in ALL scenarios
- Score tracked per strategy × regime
- Selector queries matrix for best strategy per regime
PerformanceMatrix.record(strategy_id, regime, score)called during evaluation
Phase 0.1: Cognition Pipeline
Rate-limited market regime research:
RateLimiter: token bucket, 30 RPM defaultSourceCatalogue: tracks sources, fetch counts, error ratesRegimeExtractor: keyword matching + sentiment analysisCognitionPipeline: rate-limited, deduplication, long/perm-run- 8 default sources: CoinDesk, Cointelegraph, The Block, CryptoQuant, Glassnode, Coinglass, Binance Research, Messari
Phase 0.2: Exponential Regime Expansion
Orthogonal to cognition pipeline — generates 200+ regimes from dimensions:
- 4 liquidity (vacuum, thin, normal, deep)
- 4 volatility (tight, normal, wide, extreme)
- 4 flow (balanced, buy_pressure, sell_pressure, toxic)
- 4 structure (single_toxic, mm_toxic, multi_toxic, retail_mm)
Each combination = distinct, testable regime. Total: 256 theoretical, ~200 practical.
Phase 0.3: Multi-Asset & Invariant Classification
13 assets across 10 sectors, classified by invariant characteristics only:
| Asset | Sectors | Roles | Supply | Consensus | Vol | Liq |
|---|---|---|---|---|---|---|
| BTC | CURRENCY | STORE_OF_VALUE | FIXED_CAP | POW | LOW | DEEP |
| ETH | LAYER1, DEFI | GAS, SOV, GOV | DISINFLATIONARY | POS | MEDIUM | DEEP |
| SOL | LAYER1 | GAS | INFLATIONARY | POS | HIGH | NORMAL |
| DOGE | MEME, CURRENCY | MEME, GAS | INFLATIONARY | POW | HIGH | NORMAL |
| ADA | LAYER1 | GAS | INFLATIONARY | DPOS | MEDIUM | NORMAL |
| AVAX | LAYER1, DEFI | GAS | INFLATIONARY | POS | HIGH | NORMAL |
| UNI | DEFI | GOV, UTILITY | INFLATIONARY | POS | HIGH | THIN |
| LINK | ORACLE | UTILITY | INFLATIONARY | POS | MEDIUM | THIN |
| BNB | EXCHANGE, LAYER1 | EXCHANGE_FEE, GAS | BURN | DPOS | MEDIUM | DEEP |
| MATIC | LAYER2 | GAS | INFLATIONARY | POS | HIGH | THIN |
| AAVE | DEFI | GOV | FIXED_CAP | POS | HIGH | THIN |
| DOT | LAYER1 | GAS, GOV | INFLATIONARY | DPOS | MEDIUM | NORMAL |
| ATOM | LAYER1 | GAS | INFLATIONARY | POS | HIGH | THIN |
Multi-asset = linear multiplication: 30 scenarios × 13 assets = 390+ testable scenarios.
Phase 0.4: Asset Behavior DSL + Compiler + Behavior-Driven Scenarios
Research-validated from Binance live API, BingX API, CoinGlass, and academic literature:
| Component | Purpose | File |
|---|---|---|
| AssetBehavior | 10 orthogonal behavior dimensions | training/asset_behavior.py |
| BehaviorTemplate | Reusable class-level behavior profiles | training/asset_behavior.py |
| AssetCompiler | Auto-fetch from Binance/BingX, compute profiles | training/asset_compiler.py |
| ScenarioFactory | Behavior-driven scenarios with label queries | training/cma_trainer.py |
Key research findings wired into the system:
- Depth decay: D(d) = A * d^(1-alpha). BTC alpha=0.7, DOGE alpha=1.0
- Spread ranges: BTC 0.01bps, SOL 1.26bps, DOGE 1.35bps, ADA 5.95bps
- Order sizes: BTC P50=$643, DOGE P50=$96 (7x smaller)
- Cascade dynamics: BTC 5-8% trigger/slow/fast recovery, DOGE 3-5%/fast/slow
- BingX specifics: 1.7-12.6x wider spreads, 0.05-9x variable depth
- Weekend effects: -35% vol, -48% volume, +20% spread
- Intraday patterns: 7.4x peak/trough ratio, US session dominant
- GARCH persistence: 0.95-0.99, vol half-life 2-5 days
- Correlation: BTC-ETH 0.90 normal, 0.97 crash; BTC-DOGE 0.45 normal, 0.80 crash
Training Scaling Results
CMA-ES evaluation on behavior-driven multi-asset scenarios (BTC/ETH/SOL, 3 assets × 30 scenario types = 90 scenarios). Each eval = one CMA candidate × 90 multi-step episodes through the CWM.
Parallel Training Benchmark (48 evals, pop=12)
| Workers | Time | Best Score | Score/min | Speedup | Efficiency |
|---|---|---|---|---|---|
| 1 (sequential) | 632s | 1,727 | 164 | 1.0× | — |
| 2 | 320s | 2,262 | 425 | 1.98× | 99% |
| 4 | 163s | 2,152 | 791 | 3.87× | 97% |
| 8 | 91s | 2,594 | 1,718 | 6.98× | 87% |
87% parallel efficiency at 8 workers. Independent workers explore more diverse strategy space — parallel runs find BETTER scores than sequential (2,594 vs 1,727).
Budget Scaling (ProcessPool, 3 assets × 30 scenarios)
| Budget | Gens | Best Score | Time | Score/min |
|---|---|---|---|---|
| 48 evals | 4 | 2,594 | 91s | 1,718 |
| 96 evals | 8 | 7,202 | 1246s | 346 |
| 192 evals | 16 | ~14K (est) | ~45min | ~311 |
Vectorized Reward (numba)
The CWM reward() method now uses compute_reward_vectorized (numba JIT) when
available, bypassing the Python FeatureVector dict allocation. This eliminates
~2M dict allocations per CMA generation.
Training Performance
| Metric | Sequential | Parallel (8 workers) | Speedup |
|---|---|---|---|
| CWM throughput | 189K calls/sec | (bottleneck is planner, not CWM) | — |
| CWM per-call latency | 5.3 µs | (numba + vectorized reward) | — |
| Scenario generation | 390 scenarios in 0.8s | — | — |
| CMA-ES per-generation (pop=12) | ~155s | ~23s | ~7× |
| CMA-ES 48 evals | 632s | 91s | 7× |
| Best score at 48 evals | 1,727 | 2,594 | +50% |
| Score/min | 164 | 1,718 | 10.5× |
Parallel evaluation achieves 9x speedup on single evals (embarrassingly parallel,
zero fidelity loss). CMA-ES training speedup is ~1.5x because multiprocessing overhead
is amortized across 90 scenarios per eval. Bottleneck is planner (MCTS), not CWM —
numba already accelerates the CWM hot path (_HAS_NUMBA = True).
Order Type Standardization (Multi-Exchange, FIX/CCXT-Aligned)
Three orthogonal dimensions (not one flat enum), each mapped independently:
Dimension 1 — Order Types (FIX Tag 40 OrdType): what the order IS
MARKET, LIMIT, STOP_MARKET, STOP_LIMIT, TRIGGER_MARKET, TRIGGER_LIMIT,
TRAILING_STOP, OCO, TP_SL
Dimension 2 — Time-in-Force (FIX Tag 59): how long the order LIVES
GTC, IOC, FOK, GTD
Dimension 3 — Instructions (FIX Tag 18 ExecInst): behavioral modifiers
POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
CRITICAL:
POST_ONLYis an instruction on aLIMITorder, not a standalone type.IOC/FOKare TimeInForce values applied to aLIMITorder, not order types. This aligns with FIX Protocol: Tag 40 (OrdType), Tag 59 (TimeInForce), Tag 18 (ExecInst).
Cross-Exchange Order Type Mapping
| Normalized | BingX | Binance Spot | Binance Futures | Bybit |
|---|---|---|---|---|
LIMIT |
LIMIT | LIMIT | LIMIT | LIMIT |
MARKET |
MARKET | MARKET | MARKET | MARKET |
STOP_MARKET |
TRIGGER_MARKET | STOP_MARKET | STOP_MARKET | STOP_MARKET |
STOP_LIMIT |
TRIGGER_LIMIT | STOP_LOSS_LIMIT | STOP | STOP_LIMIT |
TRIGGER_MARKET |
TRIGGER_MARKET | TAKE_PROFIT | TAKE_PROFIT | TAKE_PROFIT_MARKET |
TRIGGER_LIMIT |
TRIGGER_LIMIT | TAKE_PROFIT_LIMIT | TAKE_PROFIT | TAKE_PROFIT_LIMIT |
TRAILING_STOP |
TRAILING_STOP_MARKET | TRAILING_STOP_MARKET | TRAILING_STOP_MARKET | TRAILING_STOP |
TimeInForce Mapping
| Normalized | BingX | Binance | Bybit |
|---|---|---|---|
GTC |
GTC | GTC | GTC |
IOC |
IOC | IOC | IOC |
FOK |
FOK | FOK | FOK |
Transferability Principle
Strategy PARAMETERS transfer across exchanges. Order type NAMES are venue-specific
but semantics are identical. A LIMIT on BingX = LIMIT on Binance = LIMIT on Bybit
(same fill behavior). Only the API string differs. The venue adapter translates
normalized → exchange-native at submission time.
Cross-Exchange Learning
ScenarioFactory is venue-aware: ScenarioFactory(exchange_id='bingx') tags every
scenario with its venue. Cross-exchange transfer re-tags for a different venue:
factory = ScenarioFactory(exchange_id='bingx')
scenarios_bingx = factory.build_suite(symbols=['BTCUSDT', 'ETHUSDT'])
strategy = train(scenarios_bingx) # evolve on BingX
scenarios_binance = factory.cross_exchange_transfer(scenarios_bingx, 'binance')
score = evaluate(strategy, scenarios_binance) # test on Binance
PerformanceMatrix keys are (regime, strategy_id, venue):
get_best(regime, venue='bingx')— per-venue bestget_venue_comparison(regime, strategy_id)—{venue: score}get_best(regime)— venue-agnostic (backward compatible)
Adversary Ecology
Counterparties operate at the ActionKind level (CROSS_SPREAD, PLACE, CANCEL), not at order-type level. The CWM translates ActionKind to venue-native order types:
CROSS_SPREAD→ fills aggressively → equivalent to MARKETPLACE→ passive quote → equivalent to LIMIT Fee calculation uses VenueRules (per-exchange fees). The ecology is venue-independent.
Standards Referenced
- FIX 4.4: Tag 40 (OrdType), Tag 59 (TimeInForce), Tag 18 (ExecInst)
- CCXT: unified
limit/market+ parameter composition (triggerPrice, stopPrice) - ISO 10383: MIC for venue identification
- MiFID II / MiCA: defers to FIX for order type classification
- FIA: references FIX Tag 40 in digital asset guidelines
Prod Tooling
| Component | Purpose | File |
|---|---|---|
| CognitionLauncher | Standalone long-run script | cognition_launcher.py |
| CognitionMonitor | Metrics, health scoring, alerts | training/monitor.py |
| NewsSourceRepository | 12 industry-standard sources, ranking | training/news_sources.py |
| SourceCatalogue | Source persistence, fetch tracking | training/cognition.py |
| RateLimiter | Token bucket, 30 RPM | training/cognition.py |
| RegimeExpander | 200 regimes from dimensions | training/regime_expansion.py |
Running
# Training pipeline (continuous)
python -m malkhut.continuous_pipeline
# Cognition pipeline (continuous regime research)
python -m malkhut.cognition_launcher