malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
# MALKHUT — Adversarial Self-Play Order-Fulfilment Pipeline
**Lock-free. No GC. Fully async. GraalVM-compatible.**
MALKHUT is an order-fulfilment engine that optimises a **distribution over execution
actions** against a diverse counterparty ecology, not a single deterministic quote.
## Core Thesis
> Sequential/pure optimisation ("what is the best quote?") leads to
> deterministic quote → predictable fill pattern → adverse selection.
>
> Correct simultaneous-market framing: "what mixed order-placement policy
> survives a diverse counterparty ecology?"
## Architecture
### Corrected Stack (T19 ANNEX B-CORRECTION)
> **iceoryx2 = the wires; ASEx = the state cells; UV clock = the conductor.**
| Layer | Component | Role |
|-------|-----------|------|
| Transport | iceoryx2/zinc topics | Inter-component event fabric. Observable in `/dev/shm` . Single-writer per topic. |
| State | ASEx (ASExGuardedState + ASExWorker) | Validate-before-mutate state serialization. One writer thread per state. NOT a message bus. |
| Dispatch | UV Clock (asyncio event loop) | Event-driven reactor. Subscribes to event types. Nothing sleeps-and-polls. |
**Per T19:** ASExWatch is **DEFERRED** from the critical path until heavy-test hangs are fixed.
Use ASExWorker (FulfilmentWorker) for production hot-path mutations.
```
┌──────────────────────────────────────────────────────────────────────┐
│ Zinc SHM (POSIX /dev/shm/zinc_*) │
│ malkhut_book_state | malkhut_account_state | malkhut_fulfilment_out │
│ malkhut_risk_gate | malkhut_control (CONTROL_PLANE) │
└───────────────────────────────┬──────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────────────┐
│ ASEx VALIDATE-BEFORE-MUTATE │
│ GuardedFulfilmentState — book/account/intent/policy mutations │
│ GuardedRiskState — risk gate + kill switch │
│ FulfilmentWorker — single-threaded ASExWorker per state │
│ FulfilmentWatch — zero-overhead ring buffer for hot path │
│ ShardedWorker — per-symbol state partitioning │
└───────────────────────────────┬──────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────────────┐
│ CODE WORLD MODEL (CWM) │
│ deterministic exchange transition function │
│ f(state, joint_actions, latency_model, venue_rules) → next_state │
│ wraps hftbacktest for replay verification │
└───────────────────────────────┬──────────────────────────────────────┘
│
┌──────────┴──────────┐
▼ ▼
┌─────────────────────────┐ ┌──────────────────────────────────┐
│ LIVE FULFILMENT PLANNER │ │ OFFLINE CMA-ES TRAINER │
│ Decoupled UCB/UCT │ │ pycma + nevergrad │
│ ≤ 25 ms cadence │ │ growing self-play pool │
└───────────┬──────────────┘ └─────────────────┬────────────────┘
│ │
▼ ▼
┌─────────────────────────┐ ┌──────────────────────────────────┐
│ RISK / COMPLIANCE GATE │ │ POLICY REGISTRY / POOL │
│ leverage, notional, │ │ accepted candidates, weights │
│ self-trade, kill switch │ │ diversity-preserving eviction │
└───────────┬──────────────┘ └──────────────────────────────────┘
│
▼
┌─────────────────────────┐
│ VENUE ADAPTER (BingX) │
│ create/cancel/replace │
└───────────┬──────────────┘
│
▼
┌─────────────────────────┐
│ ClickHouse Persistence │
│ decisions, episodes, │
│ snapshots, discrepancies│
└─────────────────────────┘
```
## Package Structure
```
MALKHUT/
├── malkhut/
│ ├── __init__ .py # package version
│ ├── state.py # frozen data model (42 dataclasses)
│ ├── actions.py # action model, PlannedPolicy, RiskDecision
│ ├── features.py # feature extraction (17 features)
│ ├── counterparties.py # adversarial agent ecology (4 agents)
│ ├── engine.py # FulfilmentEngine (hot-path orchestrator)
│ ├── bench_numba.py # numba speedup benchmark
│ ├── launch_smoke_test.py # 10-min smoke test launcher
│ ├── cwm/
│ │ ├── __init__ .py
│ │ ├── core.py # CWM (deterministic exchange simulator)
│ │ ├── numba_core.py # numba-accelerated hot loops
│ │ └── replay_verify.py # replay verification (mandatory before trust)
│ ├── planner/
│ │ ├── __init__ .py
│ │ ├── action_menu.py # compact action space construction
│ │ └── sm_mcts.py # Decoupled UCB/UCT simultaneous-move MCTS
│ ├── risk/
│ │ ├── __init__ .py
│ │ └── gate.py # risk/compliance gate (hard invariants)
│ ├── venue/
│ │ ├── __init__ .py
│ │ └── bingx/
│ │ ├── __init__ .py
│ │ └── adapter.py # BingX VST/Live venue adapter (wraps DITAv2)
│ ├── ipc/
│ │ ├── __init__ .py
│ │ ├── zinc_plane.py # Zinc shared memory IPC (POSIX SHM)
│ │ └── control_plane.py # CONTROL_PLANE shared memory region
│ ├── storage/
│ │ ├── __init__ .py
malkhut: DuckDB file-backed asset store — schema, sync, queries, persistence
DuckDB store (asset_store.py):
- 4 tables: assets (28 cols), exchanges (10 cols), asset_exchanges (junction),
behavior_profiles (28 cols)
- Schema with indexes on blockchain, coingecko_id, cmc_id, exchange_id
- sync_from_profiles(): populate from in-memory dicts in one call
- Query API: query_assets(**filters), get_asset(), symbols_for_exchange(),
assets_on_blockchain(), asset_count(), exchange_count()
- Case-insensitive exchange lookup
- File-backed persistence: data survives connection close/reopen
- Performance: <1s sync, 100 queries in <1s, 1000 gets in <1s
30 tests covering:
- Schema creation and table structure (3 tests)
- Sync from profiles (6 tests, idempotent)
- Asset CRUD (7 tests, query by blockchain/coingecko/sector)
- Exchange mapping (5 tests, case-insensitive)
- Behavior profiles (4 tests)
- Persistence across connections (2 tests)
- Performance baseline (3 tests)
Total: 1276 tests across 50 files, all green, zero regressions.
2026-07-12 12:21:48 +02:00
│ │ ├── ch_store.py # ClickHouse persistence (5 tables)
│ │ └── asset_store.py # DuckDB file-backed asset universe store
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
│ ├── execution/
│ │ ├── __init__ .py
│ │ └── asex_integration.py # ASEx validate-before-mutate kernel
│ ├── clock/
│ │ ├── __init__ .py
│ │ ├── host.py # UVClock — event-driven reactor (T19)
│ │ ├── events.py # Typed events: Scan, Tick, Timer, BarFire, Stale
│ │ ├── staleness.py # Freshness watchdog (T19 Staleness Law)
│ │ └── deadnode.py # DeadNode reaper for iox2 orphans
│ ├── training/
│ │ ├── __init__ .py
2026-07-13 09:27:21 +02:00
│ │ ├── cma_trainer.py # CMA-ES + self-play + parallel workers
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
│ │ ├── registry.py # Policy lifecycle (CANDIDATE → ACTIVE)
│ │ ├── pipeline.py # Bounded continuous learning loop + logger
│ │ ├── dsl.py # Strategy DSL v2 (40+ primitives, 40+ sensors)
│ │ ├── generator.py # Genetic programming strategy evolution
malkhut: standardized order types — FIX/CCXT-aligned, multi-exchange mapping
order_types.py: Five-layer taxonomy normalized to industry standards:
Layer 1: Base types (FIX Tag 40) — MARKET, LIMIT
Layer 2: Time-in-force (FIX Tag 59) — GTC, IOC, FOK, GTD
Layer 3: Conditional/Trigger (FIX Tag 3/4+MIT) — STOP_MARKET, STOP_LIMIT,
TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
Layer 4: Instructions (FIX Tag 18) — POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
Layer 5: Compound (exchange-specific) — OCO, TP_SL
Cross-exchange mapping: BingX ↔ Binance ↔ Bybit (from CCXT source code).
Standards: FIX 4.4 Tag 40/59/18, CCXT unified API, ISO 10383 (MIC).
Transferability: strategy PARAMETERS transfer. ORDER TYPE NAMES are
venue-specific but semantics identical (LIMIT = LIMIT everywhere).
14 tests. README updated with full mapping table and standards references.
2026-07-14 12:02:01 +02:00
│ │ ├── order_types.py # Standardized order type enum (FIX/CCXT-aligned)
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
│ │ ├── selector.py # Regime → strategy mapping + performance matrix
2026-07-13 09:27:21 +02:00
│ │ ├── asset_classification.py # Multi-label taxonomy + exchange registry
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
│ │ ├── asset_behavior.py # 10-dimension behavior DSL, research-validated
│ │ ├── asset_compiler.py # Auto-fetch from Binance/BingX, compile profiles
2026-07-13 09:27:21 +02:00
│ │ ├── asset_bridge.py # Directory ↔ classification sync
│ │ ├── parallel_eval.py # ProcessPoolExecutor episode runner
│ │ ├── ray_eval.py # Ray-based eval (industrial alternative)
│ │ ├── vbt_analysis.py # Post-sim metrics: Sharpe, Sortino, VaR
malkhut(docs + bench): comprehensive update + smoke test script
README updated with:
- Vectorized UCB selection (7.7x speedup, 1.13µs/selection)
- Batch MCTS kernel (numba-accelerated)
- Fast scalar + advantage scoring modes
- Updated performance benchmarks (1186 tests, 390 scenarios, 3043 score/min)
- Advantage scorer module in package structure
smoke_1h.py: standalone training script for extended runs.
Total session: 19 commits, 1186 tests, all green.
All implementations: parallel eval (7x), vectorized reward (numba),
vectorized UCB (7.7x), fast scalar scoring, advantage mode,
DuckDB store (sub-µs reads), asset compiler, behavior DSL,
multi-exchange support, three-layer identifiers.
2026-07-13 17:04:40 +02:00
│ │ ├── advantage_scorer.py # Advantage estimation for offline analysis
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
│ │ ├── cognition.py # Rate-limited market regime research
│ │ ├── regime_expansion.py # 200+ regimes from dimension combinations
│ │ ├── news_sources.py # 12 industry-standard news sources
│ │ └── monitor.py # Cognition metrics, health, alerts
2026-07-14 09:14:31 +02:00
│ ├── daat/ # Direction-Anchored Ambiguity Triage
│ │ ├── __init__ .py
│ │ └── core.py # DaatQuery, DaatVerdict, daat_classify
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
│ ├── cognition_launcher.py # Standalone long-run cognition service
2026-07-14 09:14:31 +02:00
│ ├── continuous_pipeline.py # Continuous training pipeline
2026-07-13 09:27:21 +02:00
│ └── tests/ # 1178 tests across 50 test files
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
├── specs/
│ └── MALKHUT_ADVERSARIAL_SELFPLAY_SPEC.py # full spec
└── README.md # this file
```
## Shared Memory IPC (Zinc)
All IPC uses **real Zinc shared memory** (POSIX SHM via `/dev/shm/zinc_*` ), NOT
the file-based transport. This is lock-free, zero-copy, GraalVM-compatible.
### Regions
| Region | Writer | Readers | Purpose |
|--------|--------|---------|---------|
| `malkhut_book_state` | data feed | planner, CWM | canonical order book snapshot |
| `malkhut_account_state` | data feed | planner, risk | account/position snapshot |
| `malkhut_fulfilment_out` | engine | other systems | planner output + distribution |
| `malkhut_risk_gate` | engine | other systems | risk gate decisions |
| `malkhut_control` | external | engine | commands, symbols, lifecycle |
### Framing
All regions use `UVZINC01` seqlock framing:
```
[8B magic "UVZINC01"] [8B seq] [8B json_size] [UTF-8 JSON body]
```
### CONTROL_PLANE
The `malkhut_control` region is the management interface:
- External systems write `ControlCommand` frames
- Engine reads and processes commands
- ACKs written back for `ack_required` commands
Commands: `START` , `STOP` , `PAUSE` , `RESUME` , `SET_SYMBOLS` , `CONNECT_VENUE` ,
`DISCONNECT_VENUE` , `SET_MODE` , `EMERGENCY_STOP` , `HOT_RELOAD_POLICY` ,
`STATUS_REQUEST`
## ClickHouse Persistence
Database: `dolphin_malkhut` on `localhost:8123`
| Table | Purpose |
|-------|---------|
| `replay_steps` | CWM replay traces |
| `fulfilment_decisions` | every live decision |
| `self_play_episodes` | training episode results |
| `policy_snapshots` | versioned policy registry |
| `live_discrepancies` | simulation vs actual |
## External Dependencies (Spec-Mandated)
| Library | Purpose | Used For |
|---------|---------|----------|
| **pycma** (`cma` ) | CMA-ES optimisation | `training/cma_trainer.py` |
| **hftbacktest** | replay/queue/latency fill simulation | CWM replay verification |
| **nevergrad** | gradient-free optimiser zoo | trainer comparison |
| **msgspec** | fast typed serialization | state/event encoding |
| **numba** | hot-loop acceleration | CWM transitions, feature extraction |
| **polars** | offline analytics | large trade-path corpora |
| **lmdb** | existing SILOQY storage | ticks, candles, hd_vectors |
| **numpy** | vectorized feature extraction | parameter decoding |
| **scipy** | stats tests, bootstrap | CI, promotion gates |
## Quick Start
```python
from malkhut.engine import FulfilmentEngine
from malkhut.state import FulfilmentPolicyParams
def my_params():
return FulfilmentPolicyParams(
version="v0.1", ucb_c=1.414, max_sims=64, max_depth=3,
rollout_depth=3, root_temperature=0.5, min_root_entropy=0.25,
quote_offsets_ticks=(0, 1, 2), quote_size_fractions=(0.10, 0.25, 0.50),
passive_ttl_ms=200, aggressive_ttl_ms=50,
maker_edge_min_bps=0.5, cross_spread_edge_min_bps=5.0,
adverse_toxicity_cancel_threshold=0.5, queue_churn_cancel_threshold=0.5,
mae_tail_cut_bps=50.0, mfe_giveback_cut_fraction=0.5,
max_time_in_loss_s=300.0, failed_recovery_cut_count=3,
recovery_velocity_min_bps_per_s=0.0,
max_symbol_notional_fraction=0.20, max_single_order_notional_fraction=0.05,
reduce_when_global_up_fraction=0.30, session_profit_lock_fraction=0.02,
w_expected_pnl=1.0, w_fill_probability=0.5, w_adverse_selection=2.0,
w_queue_priority=0.5, w_inventory_risk=1.5, w_tail_loss=5.0,
w_fee_quality=0.5, w_time_decay=0.3, w_policy_entropy=0.5,
robust_tail_weight=2.0, toxic_counterparty_weight=3.0,
low_liquidity_weight=2.0, latency_stress_weight=1.0,
)
engine = FulfilmentEngine(params_provider=my_params)
# engine.on_state(latest_market_state) # hot path
```
## ASEx Integration (Validate-Before-Mutate)
MALKHUT uses **ASEx** (Advanced Serial Executor) for all mutable state mutations.
This is the same kernel used by VIOLET/DOLPHIN for race-proof state management.
### Core Pattern
```python
class GuardedFulfilmentState(ASExGuardedState):
def _validate(self, mutation) -> bool:
# Pure predicate — MUST NOT mutate self
return mutation.ts_ns > 0 and len(mutation.bids) > 0
def _apply(self, mutation):
# Only called after _validate() passes
# MAY mutate self
self._state = new_state
return result
```
### Components
| Component | Purpose | Thread Safety |
|-----------|---------|---------------|
| `GuardedFulfilmentState` | Engine state mutations (book, account, intent, policy) | One writer thread via ASExWorker |
| `GuardedRiskState` | Risk gate + kill switch | One writer thread via ASExWorker |
| `FulfilmentWorker` | Single-threaded ASExWorker wrapping GuardedFulfilmentState | Serialised via queue.Queue |
| `RiskWorker` | Single-threaded ASExWorker wrapping GuardedRiskState | Serialised via queue.Queue |
| `FulfilmentWatch` | Zero-overhead ring buffer for hot-path mutations | Producer/consumer lock-free |
| `ShardedWorker` | Per-symbol state partitioning (N independent workers) | No shared state between partitions |
### Why ASEx
- **No locks**: One thread per state object. Races impossible by construction.
- **No GC pressure**: Frozen dataclasses + queue.Queue (CPython C-level ops).
- **Validate before apply**: Invalid mutations rejected without touching state.
- **GraalVM compatible**: No CPython-only APIs in the hot path.
- **Battle-tested**: 110 ASEx tests + 1109 MALKHUT tests = 1219 total.
## Design Rules (from spec)
1. **Optimise an ecology-resistant distribution, not a quote.**
2. **The book is not passive scenery.** Ask "why did I get filled?"
3. **Replay correctness before search depth.** Wrong CWM + deep search = confident nonsense.
4. **CMA-ES does not promote from noisy wins.** Survive scenario suite + bootstrap CI.
5. **Preserve feature humility.** Allow CMA-ES to discover non-obvious interactions.
6. **Path-aware SL/TP is part of fulfilment.** Exiting = order-placement under adversarial liquidity.
7. **Never let the optimiser learn illegal behavior.** Hard-code constraints.
8. **Keep hot path bounded.** ≤ 100 ms. No live CMA, no large neural nets.
9. **Version everything.** Every decision explainable by policy + CWM + feature + state + action + outcome.
10. **If shadow live diverges, trust live.** Simulation is only a tool.
---
# ADDENDUM — EXEC-STACK PLUGGABILITY + SPEC HARDENING
*[Fable, integrator, 2026-07-06 — operator-ordered extension. Mimo: this section binds
MALKHUT into the UV/DITAv2 stack; nothing above is overridden except where marked FIX.]*
## A. Where MALKHUT plugs in (the seat at the table)
MALKHUT **is the execution layer under the UV Clock** (see
`prod/docs/uv_subspecs/UV_TASK_T19_UV_CLOCK_HOST.md` + annexes):
```
UV Clock (T19) → PromotionBridge → KernelIntent (DITAv2 contracts)
→ MALKHUT FulfilmentEngine (replaces naive kernel-submit)
→ RISK GATE → BingX venue adapter → fills
```
Your Design Rule 6 ("path-aware SL/TP is part of fulfilment") is CONVERGENT with the
operator's T19 ruling ("the whole point of UV = faster-than-scan-cadence TP/SL paths").
MALKHUT's planner is the *intelligent* form of the C11 tick-exit guardian. Sequencing:
simple doctrinal tick-exits (C11) ship first via T19 step 3; MALKHUT supersedes them
**when certified** — same seat, smarter occupant.
## B. Intake contract (BINDING)
1. **Input = DITAv2 `KernelIntent`** (`prod/clean_arch/dita_v2/contracts.py` ): asset,
side (TradeSide), action (ENTER/EXIT), reference_price, target_size, leverage,
metadata (carries `promo_client_id` ). Accept KernelIntent natively; do not invent a
parallel intent type.
2. ** `u-` clientOrderId prefix is LAW** (T9 seam): every venue order MALKHUT places
carries the u- prefix from intent metadata. Non-negotiable — it is the attribution
boundary vs BLUE and vs legacy VIOLET orders.
3. **DUAL-LEVERAGE LAW** : `KernelIntent.leverage` is CONVICTION leverage [0.5, 9.0] —
it sizes QUANTITY only. Venue leverage = the mapping in `prod/bingx/leverage.py`
(linear → ROUND_HALF_EVEN → int clamp [1,3]). WIRE the existing typed L3 wrapper
(`prod/clean_arch/violet/exchange_leverage.py` ); litmus: 0.5→1, 4.75→2, 8.0→3, 9.0→3.
The venue adapter MUST NOT pass conviction leverage raw to set-leverage.
4. **Arming/two-man/kill** : MALKHUT honors `UV_PROMOTED=1` + arm-file
(`/root/uv-wt/prime-live/UV_PROMOTED.arm` ), re-evaluated PER EVENT (never cached —
the T12 kill-switch lesson). DARK default = ObserveOnly venue. Map arm-file-delete
→ your `EMERGENCY_STOP` semantics: inert next event, no restart.
## C. Fabric + state alignment (the tonight-rulings)
- **Stack law: iceoryx2 = wires, ASEx = state cells, UV clock = conductor.** Your ASEx
usage (GuardedState + Worker per state object) is CORRECT — it is exactly what ASEx
is for. Your zinc regions already speak UVZINC01 seqlock framing — good; when T19's
event topics land, consume ScanEvent/TickEvent (with PROVENANCE: scan_number,
scan_ts, ingest_ts, AGE) from the shared LiveSource instead of a private feed.
- **Account state**: read the DITAv2 ASEx account core (asex_account /
AccountProjectionV2) as a CONSUMER only. MALKHUT never becomes a second writer to
account state (two-cores-share-nothing law); its own guarded states are its own.
- **Staleness law (T19 Annex A) — extend your Rule 2**: a stale book is not scenery
either, it is an ADVERSARY. Planner receives AGE on every input and must
discount/refuse plans over stale state; freshness watchdog emits StaleInput.
- **ASEx pre-condition inherited**: DeadNode reaper / orphan-sweep on start (iox2
segments orphan on crash — known gap).
## D. Persistence + observability alignment
- `dolphin_malkhut` namespace: approved (own-namespace law). **NO TTL on any table**
(retention doctrine — audit surfaces never expire).
- **Cross-reference law**: every `fulfilment_decisions` row carries `intent_id` +
`trade_id` from the source KernelIntent → end-to-end trace joins
`dolphin_uv.exec_journal` ⇄ `dolphin_malkhut.*` ⇄ venue order history. This is what
makes MALKHUT certifiable instead of a black box.
- Publish planner pulse (decision rate, plan entropy, gate verdicts, AGE-at-decision)
into the zinc snapshot → extant TUIs render it (TUI-first doctrine).
## E. Ecology + reward grounding (FIX-grade improvements)
1. **Ground the counterparty ecology in OUR tape** : fit/validate adversary populations
against `dolphin.obf_universe` recorded sub-second tape and the WT2 stress
manifold. The CLASS-M/CLASS-X hour sets (WT2D) are your ready-made adversarial
scenario suites — corrupt/stress hours ARE the hostile ecology, measured.
CLASS-X law applies: extreme magnitudes are a FEATURE channel, never filtered.
2. **Latency model = our measured tails** , not defaults: scan_to_fill/step_bar
distributions from the WT1 latency audit (median 47ms/1.58ms, stop-tail 7.8s) —
calibrate the CWM's latency injections from these empirical curves.
3. **Reconcile the budget statement (FIX)** : planner box says ≤25ms, Rule 8 says
≤100ms — bind them: ≤25ms planning inside a ≤100ms end-to-end intent→venue-call
budget, both enforced by test.
4. **Promotion gate binds to the cert conveyor** : CMA-ES candidate promotion (Rule 4)
emits artifacts via the T4 cert reporter into the gate ledger; policy snapshots
versioned in `dolphin_malkhut.policy_snapshots` AND referenced in
`prod/docs/uv_cert/` . No policy trades live without a ledger row.
## F. Certification pathway (Gate-M — how MALKHUT reaches money)
1. **CWM replay verification** (your Rule 3) vs hftbacktest — already specced; add
determinism × 2 (byte-identical minus ts).
2. **Dual-leverage litmus** (B.3 values) as a named test.
3. **DARK shadow window** : MALKHUT plans on live intents while the simple path
executes; diff planned-vs-naive on fills/slippage/adverse-selection — the
improvement claim must be MEASURED on VST before MALKHUT takes the seat.
4. **Risk-gate mutation litmus** : drop a constraint → a test goes RED (looks-finished
doctrine; 222 green tests ≠ verification until mutations bite).
5. **Rule 10 stands supreme** : shadow-vs-live divergence → trust live, stand down,
file the discrepancy row.
*End addendum. — F.*
---
2026-07-13 09:27:21 +02:00
## DEVELOPMENT STATUS (2026-07-13)
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
2026-07-14 09:14:31 +02:00
**1204 test functions. 50 test files. All green. 0 failures. 0 regressions.**
### Fable Spec Items — All Complete
| # | Item | Module | Status |
|---|------|--------|--------|
| 1 | Mutation-litmus test | `test_mutation_litmus.py` | ✅ PASSES |
| 2 | Re-run benchmarks at corrected fees | (long run) | ⏸️ Deferred |
| 3 | Maker fee UNVERIFIED comment | `asset_classification.py` | ✅ Done |
| 4 | ScenarioLibrary sweep | `scenario_library.py` | ✅ 7,488 grid points |
| 5 | PerformanceMatrix manifold | `selector.py` | ✅ confidence+support |
| 6 | ActualsLoader | `actuals_loader.py` | ✅ CH reader + fallback |
| 7 | OOD verdict | `risk/gate.py` | ✅ daat_verdict parameter |
| 8 | Manifold query | `manifold_query.py` | ✅ three-phase recommend |
| 9 | DAAT package | `daat/core.py` | ✅ DaatVerdict{KNOWN,MARGINAL,OOD} |
| 10 | Book fidelity | `book_fidelity.py` | ✅ OBF → OrderBookState |
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
2026-07-13 21:50:57 +02:00
### Fee Correction (CRITICAL — all prior training was at 10x too-cheap fees)
Source of truth: `dolphin.trade_execution_quality` (8006 rows, avg taker=5.016 bps).
| Venue | Old taker (WRONG) | Correct taker | Old maker (WRONG) | Correct maker |
|-------|-------------------|---------------|-------------------|---------------|
| BingX | 0.5 bps | **5.0 bps** | -0.2 (rebate) | ** +2.0 (POSITIVE)** |
| Binance | 0.4 bps | **4.5 bps** | -0.2 | -0.2 (UNVERIFIED) |
| Bybit | 0.06 bps | **5.5 bps** | -0.01 | -0.1 (UNVERIFIED) |
**Every policy trained before this fix is suspect** — cheap fees reward overtrading.
Re-measurement at correct fees is required for production deployment.
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
### Completed subsystems
| Subsystem | Module | Tests | Status |
|-----------|--------|-------|--------|
| **State Model** | `state.py` | 17 | 42 frozen dataclasses, immutable |
malkhut(docs + bench): comprehensive update + smoke test script
README updated with:
- Vectorized UCB selection (7.7x speedup, 1.13µs/selection)
- Batch MCTS kernel (numba-accelerated)
- Fast scalar + advantage scoring modes
- Updated performance benchmarks (1186 tests, 390 scenarios, 3043 score/min)
- Advantage scorer module in package structure
smoke_1h.py: standalone training script for extended runs.
Total session: 19 commits, 1186 tests, all green.
All implementations: parallel eval (7x), vectorized reward (numba),
vectorized UCB (7.7x), fast scalar scoring, advantage mode,
DuckDB store (sub-µs reads), asset compiler, behavior DSL,
multi-exchange support, three-layer identifiers.
2026-07-13 17:04:40 +02:00
| **CWM** | `cwm/core.py` + `cwm/numba_core.py` | 103 | Exchange mechanics + numba JIT (5.3µs/transition) + vectorized reward + vectorized UCB + batch MCTS kernel |
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
| **Replay Verification** | `cwm/replay_verify.py` | 65 | Deep comparison, binary search, trajectory recording |
| **Planner** | `planner/sm_mcts.py` | 11 | Decoupled UCB/UCT, ≤25ms budget |
| **Action Menu** | `planner/action_menu.py` | (in planner) | Compact action space construction |
| **Counterparties** | `counterparties.py` | 19 | 4 adversarial agent policies |
| **Risk Gate** | `risk/gate.py` | 4 | Hard invariants: leverage, self-trade, post-only |
| **ASEx Integration** | `execution/asex_integration.py` | 33 | Validate-before-mutate, single-writer, no locks |
| **Zinc IPC** | `ipc/zinc_plane.py` | 8 | Real POSIX SHM, UVZINC01 framing |
| **Control Plane** | `ipc/control_plane.py` | 3 | HOT_RELOAD_POLICY, kill switch |
| **ClickHouse** | `storage/ch_store.py` | 9 | 5 tables, HTTP API |
2026-07-12 19:38:24 +02:00
| **DuckDB Asset Store** | `storage/asset_store.py` | 30 | File-backed + in-memory materialization, sub-µs reads |
malkhut: DuckDB file-backed asset store — schema, sync, queries, persistence
DuckDB store (asset_store.py):
- 4 tables: assets (28 cols), exchanges (10 cols), asset_exchanges (junction),
behavior_profiles (28 cols)
- Schema with indexes on blockchain, coingecko_id, cmc_id, exchange_id
- sync_from_profiles(): populate from in-memory dicts in one call
- Query API: query_assets(**filters), get_asset(), symbols_for_exchange(),
assets_on_blockchain(), asset_count(), exchange_count()
- Case-insensitive exchange lookup
- File-backed persistence: data survives connection close/reopen
- Performance: <1s sync, 100 queries in <1s, 1000 gets in <1s
30 tests covering:
- Schema creation and table structure (3 tests)
- Sync from profiles (6 tests, idempotent)
- Asset CRUD (7 tests, query by blockchain/coingecko/sector)
- Exchange mapping (5 tests, case-insensitive)
- Behavior profiles (4 tests)
- Persistence across connections (2 tests)
- Performance baseline (3 tests)
Total: 1276 tests across 50 files, all green, zero regressions.
2026-07-12 12:21:48 +02:00
| **Asset Bridge** | `training/asset_bridge.py` | 49 | Directory ↔ classification sync |
2026-07-13 09:27:21 +02:00
| **CMA-ES Training** | `training/cma_trainer.py` | 65 | Behavior-driven, auto-compile, parallel workers, 7x speedup |
| **Parallel Eval** | `training/parallel_eval.py` | 16 | ProcessPoolExecutor, 7x CMA-ES speedup |
| **Ray Eval** | `training/ray_eval.py` | 5 | Ray-based eval (available, slower for ≤1K scenarios) |
malkhut(docs + bench): comprehensive update + smoke test script
README updated with:
- Vectorized UCB selection (7.7x speedup, 1.13µs/selection)
- Batch MCTS kernel (numba-accelerated)
- Fast scalar + advantage scoring modes
- Updated performance benchmarks (1186 tests, 390 scenarios, 3043 score/min)
- Advantage scorer module in package structure
smoke_1h.py: standalone training script for extended runs.
Total session: 19 commits, 1186 tests, all green.
All implementations: parallel eval (7x), vectorized reward (numba),
vectorized UCB (7.7x), fast scalar scoring, advantage mode,
DuckDB store (sub-µs reads), asset compiler, behavior DSL,
multi-exchange support, three-layer identifiers.
2026-07-13 17:04:40 +02:00
| **Scoring Modes** | `cma_trainer.py` + `advantage_scorer.py` | 8 | Fast scalar (CMA loop) + advantage (offline analysis) |
2026-07-13 09:27:21 +02:00
| **VBT Analysis** | `training/vbt_analysis.py` | 8 | Post-sim trade metrics: Sharpe, Sortino, VaR, cross-asset |
2026-07-14 09:14:31 +02:00
| **ScenarioLibrary** | `training/scenario_library.py` | 7 | Sweep 7,488 grid points (spread× depth× tox× regime) |
| **Manifold Query** | `training/manifold_query.py` | (new) | Three-phase: DAAT → manifold → recommend |
| **ActualsLoader** | `training/actuals_loader.py` | (new) | Reads CH tables for Mode 2 live query |
| **Book Fidelity** | `training/book_fidelity.py` | (new) | OBF 15B rows → OrderBookState synthesis |
| **DAAT** | `daat/core.py` | 9 | Direction-Anchored Ambiguity Triage |
| **OOD Verdict** | `risk/gate.py` | +1 | daat_verdict parameter, OOD → doctrinal fallback |
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
| **Policy Registry** | `training/registry.py` | 14 | CANDIDATE → ACTIVE lifecycle |
| **Training Pipeline** | `training/pipeline.py` | 21 | Bounded continuous learning loop |
| **Strategy DSL v2** | `training/dsl.py` | 69 | 40+ primitives, 40+ sensors, 16 builtins |
2026-07-14 14:52:50 +02:00
| **Order Types** | `training/order_types.py` | (new) | Three-dimensional: OrderType × TimeInForce × Instructions (FIX-aligned) |
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
| **Strategy Generator** | `training/generator.py` | 20 | Genetic programming: crossover, mutation, tournament |
2026-07-14 15:44:33 +02:00
| **Strategy Selector** | `training/selector.py` | 24 | Regime × strategy × venue matrix, cross-exchange comparison |
malkhut: multi-exchange asset universe + schema docs
ExchangeProfile: standardized exchange metadata (fees, latency, capabilities).
3 pre-defined exchanges: Binance, BingX, Bybit.
AssetProfile.exchanges: tuple[str] — which venues trade each asset.
New query functions: get_assets_on_exchange, get_common_assets,
get_exchange_for_asset, get_exchange, list_exchanges.
_DATA_STORAGE_SCHEMA_FORMATS.md: comprehensive reference for agent
consumption — data model, storage format, query interfaces, data flow
diagram, enum reference, import patterns for BLUE/VIOLET/UV integration.
README updated: exchange registry section, package structure, subsystems
table.
2026-07-11 22:37:15 +02:00
| **Asset Classification** | `training/asset_classification.py` | 190 | Multi-label taxonomy + exchange registry, 13 assets, system-wide store |
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
| **Asset Behavior DSL** | `training/asset_behavior.py` | (in classification) | 10 orthogonal dimensions, 3 templates, 13 behaviors, research-validated |
| **Asset Compiler** | `training/asset_compiler.py` | (new) | Auto-fetch from Binance/BingX API, compile profiles, rate-limited |
| **Cognition Pipeline** | `training/cognition.py` | 29 | Rate-limited, 8 sources, dedup, perm-run |
| **Regime Expansion** | `training/regime_expansion.py` | 29 | 200+ regimes from 4× 4× 4× 4 dimensions |
| **News Sources** | `training/news_sources.py` | 12 | 12 industry-standard sources, ranking |
| **Cognition Monitor** | `training/monitor.py` | 9 | Metrics, health scoring, alerts, JSONL logging |
| **Cognition Launcher** | `cognition_launcher.py` | 21 | Standalone long-run cognition service |
| **UV Clock** | `clock/host.py` | 30 | T19 event-driven reactor |
| **Numba Acceleration** | `cwm/numba_core.py` | 13 | JIT on hot loops, 1.8x fill speedup |
| **BingX Adapter** | `venue/bingx/adapter.py` | 28 | Wraps DITAv2, order tracking |
| **Hypothesis** | `test_hypothesis_properties.py` | 6 | Property-based fuzzing |
| **Fuzz/Adversarial** | `test_fuzz.py` , `test_adversarial.py` | 17 | Random stress + adversarial scenarios |
| **Concurrency** | `test_concurrency.py` | 8 | Thread safety, race detection |
| **Sync/Async Seams** | `test_sync_async_seams.py` | 7 | Zinc latency, engine budget |
| **E2E Integration** | `test_e2e_integration.py` | 2 | Full pipeline: train→register→plan→risk→venue→zinc→ch→reload |
malkhut(docs): fill quality documentation — README, integration, OB study
README: Fill Quality section (core optimization target, metrics, reward
function, PerformanceMatrix, CMA-ES integration)
HftBacktestCWM integration doc: FillQuality dataclass, fill_value_score
computation, reward function weighting, PerformanceMatrix tracking
OB microstructure study: Section 13 — Fill Quality Optimization,
per-asset expectations, optimization strategy, connection to OB dynamics
Fill quality is MALKHUT's core aim: the system learns to get better fills
(faster, better-priced, less adverse selection) across regimes and venues.
2026-07-15 15:36:22 +02:00
| **HftBacktestCWM** | `cwm/hft_cwm.py` | (new) | Queue-model fills (PowerProb), fill quality tracking, drop-in CWM |
### Fill Quality — The Core Optimization Target
MALKHUT is an **execution improvement engine** . Fill quality IS the primary aim — not PnL, not Sharpe, not win rate. The system learns to get **better fills** : faster, better-priced, with less adverse selection.
#### Fill Quality Metrics (per CWM transition)
| Metric | Source | What it measures |
|--------|--------|------------------|
| `slippage_bps` | `abs(fill_price - mid) / mid * 10000` | How far from mid did we fill? (aggressive) |
| `price_improvement_bps` | `(best_bid - fill_price) / best_bid * 10000` | How much better than touch? (passive) |
| `levels_consumed` | `fill_qty / avg_level_qty` | Queue depth of fill |
| `is_maker_fill` | `order_type==LIMIT or post_only` | Passive vs aggressive |
| `rolling_fill_rate` | EMA(0.8, 0.2) over recent fills | Recent fill success rate |
| `post_fill_adverse_bps` | `(new_mid - old_mid) / old_mid * 10000` | Price movement after fill |
| `fill_value_score` | `quality - abs(adverse) * 0.5` | **Composite optimization metric** |
#### Fill Quality in the Reward Function
```
reward = w_fill_probability * fill_value_score ← PRIMARY (fill quality)
+ w_expected_pnl * pnl ← secondary (PnL)
- w_adverse_selection * toxicity
- w_inventory_risk * inventory_risk
- w_tail_loss * tail_risk
- w_time_decay * time_in_loss
2026-07-18 19:14:34 +02:00
+ w_fee_quality * fee_savings ← maker saves (taker-maker) bps
- w_fee_quality * taker_fee ← taker pays full fee
- w_fee_quality * markout_cost * 0.3 ← markout = honest execution cost
malkhut(docs): fill quality documentation — README, integration, OB study
README: Fill Quality section (core optimization target, metrics, reward
function, PerformanceMatrix, CMA-ES integration)
HftBacktestCWM integration doc: FillQuality dataclass, fill_value_score
computation, reward function weighting, PerformanceMatrix tracking
OB microstructure study: Section 13 — Fill Quality Optimization,
per-asset expectations, optimization strategy, connection to OB dynamics
Fill quality is MALKHUT's core aim: the system learns to get better fills
(faster, better-priced, less adverse selection) across regimes and venues.
2026-07-15 15:36:22 +02:00
```
2026-07-18 19:14:34 +02:00
**Fee model (BingX, no rebates):**
- Maker fee: 2.0 bps (you pay)
- Taker fee: 5.0 bps (you pay)
- Fee savings: 3.0 bps (maker saves 3bp vs taker)
**Markout = quality:** slippage_bps + post_fill_adverse_bps = honest execution cost.
The system learns: pay the friction (fee + slippage) when urgency × (fee + slippage) < threshold.
malkhut(docs): fill quality documentation — README, integration, OB study
README: Fill Quality section (core optimization target, metrics, reward
function, PerformanceMatrix, CMA-ES integration)
HftBacktestCWM integration doc: FillQuality dataclass, fill_value_score
computation, reward function weighting, PerformanceMatrix tracking
OB microstructure study: Section 13 — Fill Quality Optimization,
per-asset expectations, optimization strategy, connection to OB dynamics
Fill quality is MALKHUT's core aim: the system learns to get better fills
(faster, better-priced, less adverse selection) across regimes and venues.
2026-07-15 15:36:22 +02:00
#### Fill Quality in the PerformanceMatrix
`RegimeStrategyScore` now tracks 4 fill quality metrics per (regime, strategy, venue):
- `avg_fill_rate` : rolling fill success rate
- `avg_slippage_bps` : average slippage for aggressive fills
- `avg_price_improvement_bps` : average improvement for passive fills
- `avg_fill_value_score` : composite fill quality metric
This enables: "Which strategy achieves the best fill quality in regime X on venue Y?"
#### Fill Quality in CMA-ES
The CMA-ES optimizer now receives fill quality metrics in each `EpisodeResult` :
- `avg_fill_value_score` : average fill value across the episode
- `avg_price_improvement_bps` : average price improvement
- `avg_post_fill_adverse_bps` : average adverse selection
The CMA-ES objective is: maximize fill quality (primary) while maintaining positive PnL.
2026-07-17 10:16:55 +02:00
#### Urgency-Driven Maker/Taker Decision
The system learns WHEN to use maker (passive) vs taker (aggressive) as a function of urgency:
| Urgency | Behavior | Penalty |
|---------|----------|---------|
| 0.0 – 0.3 | Passive only (maker). No taker actions generated. | — |
| 0.3 – 0.65 | IOC partial taker allowed. Small sizes, partial fills. | Low |
| 0.65 – 1.0 | Full taker allowed. Aggressive crossing at full size. | None |
Two CMA-ES optimizable parameters:
- `urgency_taker_threshold` (default 0.65): crosses from passive to aggressive
- `urgency_taker_penalty_bps` (default 2.0): penalty for taker at low urgency
The system learns: "For ADAUSDT in stress regime, threshold=0.4 (thin book → try maker,
but resort to taker when liquidity appears). For BTCUSDT in normal regime, threshold=0.8
(deep book → maker always works)."
#### Calibrated Slippage (Flight7 VST + Mainnet Anchors)
Slippage model per-asset, configurable, overridable per-run:
| Book Type | Model | Assets |
|-----------|-------|--------|
| **Deep book** | `slippage = alpha * levels + beta * depth_ratio` | BTC, ETH, BNB |
| **Thin book** | `slippage = intercept + adverse_selection` | DOGE, ADA, UNI, AAVE, ATOM |
Switch: `if book_depth_usd < thin_book_threshold_usd → thin mode`
Fable's Flight7 calibration anchors (testnet + mainnet):
- Majors: BTC ~0.1 bps, ETH ~0.4 bps (testnet ≈ mainnet)
- Liquid alts: 3-6 bps testnet, 10-30 bps mainnet (3-6× gap)
- Thin alts: 14-25 bps testnet, ~140 bps round-trip mainnet
#### Chase Mechanics
Cancel → wait → retry with configurable parameters (CMA-ES optimizable):
- `wait_to_retry_ms` : delay before re-quoting after cancel (0-2000ms)
- `chase_enabled` : enable chase-follow behavior
- `chase_offset_ticks` : ticks from target price to chase (0-10)
- `chase_max_retries` : max cancel-retry cycles (0-5)
CHASE in the DSL places a limit order at target offset with TTL=wait_to_retry_ms.
If not filled, the order expires and next step re-places at a new offset.
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
### Performance Benchmarks
| Metric | Value |
|--------|-------|
malkhut(docs + bench): comprehensive update + smoke test script
README updated with:
- Vectorized UCB selection (7.7x speedup, 1.13µs/selection)
- Batch MCTS kernel (numba-accelerated)
- Fast scalar + advantage scoring modes
- Updated performance benchmarks (1186 tests, 390 scenarios, 3043 score/min)
- Advantage scorer module in package structure
smoke_1h.py: standalone training script for extended runs.
Total session: 19 commits, 1186 tests, all green.
All implementations: parallel eval (7x), vectorized reward (numba),
vectorized UCB (7.7x), fast scalar scoring, advantage mode,
DuckDB store (sub-µs reads), asset compiler, behavior DSL,
multi-exchange support, three-layer identifiers.
2026-07-13 17:04:40 +02:00
| CWM throughput | 189K calls/sec (numba JIT) |
| CWM per-call latency | 5.3 µs |
| CWM reward (numba vectorized) | ~0.3µs |
| UCB selection (numba vectorized) | 1.13µs (was 8.7µs, 7.7x speedup) |
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
| Numba fill speedup | 1.8x (batch 100) |
2026-07-13 09:27:21 +02:00
| DuckDB asset reads | 0.2µs (in-memory) |
| Scenario generation | 390 scenarios in 0.8s |
| CMA-ES parallel (8 workers) | 7× speedup, 87% efficiency |
malkhut(docs + bench): comprehensive update + smoke test script
README updated with:
- Vectorized UCB selection (7.7x speedup, 1.13µs/selection)
- Batch MCTS kernel (numba-accelerated)
- Fast scalar + advantage scoring modes
- Updated performance benchmarks (1186 tests, 390 scenarios, 3043 score/min)
- Advantage scorer module in package structure
smoke_1h.py: standalone training script for extended runs.
Total session: 19 commits, 1186 tests, all green.
All implementations: parallel eval (7x), vectorized reward (numba),
vectorized UCB (7.7x), fast scalar scoring, advantage mode,
DuckDB store (sub-µs reads), asset compiler, behavior DSL,
multi-exchange support, three-layer identifiers.
2026-07-13 17:04:40 +02:00
| Best CMA-ES score (48 evals) | 8,628 (fast scalar, 100.4 bps PnL) |
| Score/min (8 workers) | 3,043 |
| Episode throughput (parallel) | 16ms/ep |
| Episode throughput (sequential) | 23ms/ep |
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
### Bugs found and fixed (22 total)
1. Empty book crash in `materialize_price_from_action` (SELL side)
2. Empty book crash in `total_notional` calculation
3. Empty book crash in mark-to-market
4. Empty book crash in `_inventory_risk`
5. Empty book crash in counterparty processing
6. Empty book crash in `_update_path_state`
7. Empty book crash in feature extractor (`b.mid` )
8. Counterparty fills used `max()` instead of consuming levels
9. Position overwritten instead of accumulated on add
10. Test helper `_state()` not forwarding `positions` kwarg
11. `_make_open_order` set empty symbol instead of venue symbol
12. `FulfilmentPolicyParams` frozen dataclass `__dict__` access
13. Banker's rounding in tick tests
14. `HOT_RELOAD_POLICY` compared float to string
15. Engine init order (workers before hot_reload)
16. `_tp()` helper missing `failed_recovery_count` parameter
17. CWM: SELL position flip didn't reset avg_entry
18. CWM: Counterparty fills overwrote our fill tracking
19. CWM: Double-counting unrealized PnL in equity
20. UnboundLocalError when max_steps=0
21. fill_count never incremented in training
22. Cancel rate limiter incorrectly gated order placements
### Key design rules enforced
1. ✅ Optimise distribution, not single quote
2. ✅ Book is not passive scenery
3. ✅ **Replay correctness before search depth** (mandatory gate)
4. ✅ CMA-ES does not promote from noisy wins (bootstrap CI)
5. ✅ Preserve feature humility (CMA discovers interactions)
6. ✅ Path-aware SL/TP is part of fulfilment
7. ✅ Never let optimiser learn illegal behavior
8. ✅ Hot path bounded (≤100ms)
9. ✅ Version everything
10. ✅ Shadow-vs-live divergence → trust live
### Strategy Selection (completed)
The system classifies strategies per market condition and floats the best to top:
```
market_fingerprint → Regime Classifier (8 regimes)
↓
Performance Matrix (regime × strategy → score)
↓
Strategy Selector → Engine uses best strategy
```
### Upstream integration path
- **ExoF** — funding rates, open interest, liquidation levels
- **MARAS** — 5-tier ensemble → 8-regime classification
- **DOLPHINNG7 vel_div** — velocity divergence signals
- **DOLPHINNG7 eigenscan** — eigenspace correlation analysis
---
## G. Strategy DSL v2 — Third-Party Strategy Development
MALKHUT includes a **Domain-Specific Language (DSL)** for strategy composition.
Third parties can write, evolve, and share strategies without modifying core code.
### DSL Format
```
STRATEGY "passive_maker" {
PRIORITY 1: IF spread_bps < 5.0 AND orderflow_toxicity < 0 . 3 THEN QUOTE ( BUY , 0 , 0 . 25 )
PRIORITY 2: IF spread_bps < 5.0 AND orderflow_toxicity < 0 . 3 THEN QUOTE ( SELL , 0 , 0 . 25 )
PRIORITY 3: IF orderflow_toxicity > 0.7 THEN CANCEL_ALL
PRIORITY 4: IF time_in_trade > 300 THEN EXIT
PRIORITY 5: IF unrealized_pnl < -50bps THEN STOP_LOSS
PRIORITY 6: NOOP
}
```
### Action Primitives (40+)
| Category | Primitives |
|----------|-----------|
| Passive | QUOTE, JOIN_QUEUE, STEP_BACK, LADDER, GRID, ICEBERG, TWAP |
| Aggressive | CROSS, SNIPER, PING |
| Cancellation | CANCEL, CANCEL_ALL, CANCEL_AND_HOLD, REQUOTE |
| Position | EXIT, HALF_EXIT, QUARTER_EXIT, TAKE_PARTIAL, STOP_LOSS, TAKE_PROFIT, TRAILING_STOP, EMERGENCY_EXIT, FLAT_ALL |
| Sizing | SCALE_IN, SCALE_OUT, INCREASE_SIZE, REDUCE_SIZE |
| Stop/Target | MOVE_STOP, MOVE_TAKE_PROFIT, BRACKET, OCO |
| Hedging | HEDGE, PAIR_TRADE |
| Waiting | HOLD, WAIT_FOR_FILL, WAIT_FOR_PRICE, WAIT_FOR_SPREAD |
### Market Sensors (40+)
| Category | Sensors |
|----------|---------|
| Book | spread_bps, imbalance_3/5/10, bid_depth_3/5/10, bid_ask_ratio |
| Flow | toxicity, queue_churn, cross_venue_lead, order_book_toxicity |
| Momentum | price_momentum_1s/5s/15s/1m |
| Position | position_qty, pnl_bps, leverage, position_age_s |
| Path | mae_bps, mfe_bps, time_in_loss, failed_recoveries, recovery_velocity |
| Regime | volatility, atr_14/50, rsi_14, funding, regime_score |
| Account | equity, risk_budget_used, session_pnl, drawdown |
| Performance | profit_factor, sharpe_ratio, consecutive_wins/losses |
| Time | current_hour, day_of_week, is_weekend, is_liquid_hours |
### Comparison Operators (12)
`>` , `<` , `>=` , `<=` , `==` , `!=` , `abs>` , `abs<` , `changing` , `stable` , `crossing_above` , `crossing_below`
### Builtin Strategies (16)
| Strategy | Description |
|----------|-------------|
| `passive_maker` | Passive limit orders, toxicity avoidance |
| `aggressive_taker` | Momentum-based aggressive crosses |
| `toxicity_avoider` | Cancels on high toxicity, quotes when safe |
| `path_risk_exit` | Path-aware SL/TP exits |
| `regime_adaptive` | Adapts to MARAS regime classification |
| `momentum_catcher` | Catches momentum moves, partial exits |
| `mean_reversion` | Wide-spread mean reversion |
| `scalper` | Ultra-tight spreads, small sizes |
| `inventory_manager` | Position size control |
| `session_guard` | Time-based position management |
| `liquidity_hunter` | Deep book exploitation |
| `volatility_breakout` | Volatility-based breakouts |
| `funding_arb` | Funding rate arbitrage |
| `grid_trader` | Grid-based accumulation |
| `risk_parity` | Risk budget management |
| `hybrid_adaptive` | Multi-condition adaptive |
### Strategy Generator (Genetic Programming)
The system evolves strategies via genetic operators:
```python
from malkhut.training.generator import StrategyGenerator, GeneratorConfig
config = GeneratorConfig(
population_size=20, generations=5, tournament_size=3,
mutation_rate=0.15, crossover_rate=0.7,
)
generator = StrategyGenerator(config=config)
population = generator.evolve(baseline_params, scenarios)
# Get successful strategies
successful = generator.get_successful_strategies(min_fitness=0.0)
for genome in successful:
generator.add_to_pool(genome)
```
### Strategy Selector (Market-Adaptive)
The system classifies strategies per market regime and selects the best:
```python
from malkhut.training.selector import StrategySelector, MarketRegime
selector = StrategySelector()
result = selector.select(current_state, available_strategies)
# result.strategy_id → best strategy for current regime
# result.confidence → how confident we are
# result.alternatives → other considered strategies
```
Regimes: `trending_up` , `trending_down` , `high_volatility` , `low_volatility` ,
`mean_reverting` , `momentum` , `choppy` , `liquidity_hole` , `normal`
### Asset Classification (Invariant Taxonomy)
Multi-label asset classification using **only invariant characteristics** — properties
intrinsic to the token that predict price behaviour and order-book performance.
malkhut: three-layer identifier architecture for cross-system asset identification
Layer 1 (canonical identity): symbol, base_asset, name, unified_symbol
(CCXT format), quote_currency
Layer 2 (cross-system): coingecko_id, cmc_id, blockchain, contract_address
Layer 3 (exchange mapping): exchanges tuple
All 13 pre-defined assets migrated with accurate CoinGecko IDs, CMC IDs,
blockchains, and contract addresses (ERC-20 tokens).
7 new query functions: get_asset_by_coingecko_id, get_asset_by_cmc_id,
get_assets_by_base_asset, get_assets_by_blockchain,
get_assets_by_unified_symbol.
33 new tests covering: Layer 1 identity, Layer 2 cross-system identifiers,
identifier query functions, identifier consistency (uniqueness, derivation),
exchange registry.
Total: 1189 tests, 47 files, all green.
Based on research: CCXT BASE/QUOTE is de facto standard, CoinGecko ID
most widely used in crypto-native, ISO 24165 DTI emerging, FIGI for
institutional.
2026-07-12 00:40:14 +02:00
**Three-layer identifier architecture** for cross-system interoperability.
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
**Design principle:** Classify by WHAT an asset IS (fundamental), not how it's trading
right now. Volatility/liquidity appear as long-run statistical profiles, not as
classification axes that change minute-to-minute.
malkhut: three-layer identifier architecture for cross-system asset identification
Layer 1 (canonical identity): symbol, base_asset, name, unified_symbol
(CCXT format), quote_currency
Layer 2 (cross-system): coingecko_id, cmc_id, blockchain, contract_address
Layer 3 (exchange mapping): exchanges tuple
All 13 pre-defined assets migrated with accurate CoinGecko IDs, CMC IDs,
blockchains, and contract addresses (ERC-20 tokens).
7 new query functions: get_asset_by_coingecko_id, get_asset_by_cmc_id,
get_assets_by_base_asset, get_assets_by_blockchain,
get_assets_by_unified_symbol.
33 new tests covering: Layer 1 identity, Layer 2 cross-system identifiers,
identifier query functions, identifier consistency (uniqueness, derivation),
exchange registry.
Total: 1189 tests, 47 files, all green.
Based on research: CCXT BASE/QUOTE is de facto standard, CoinGecko ID
most widely used in crypto-native, ISO 24165 DTI emerging, FIGI for
institutional.
2026-07-12 00:40:14 +02:00
#### Three-Layer Identifier Architecture
| Layer | Fields | Purpose | Example |
|-------|--------|---------|---------|
| **1. Canonical** | `symbol` , `base_asset` , `name` , `unified_symbol` , `quote_currency` | Venue-independent identity | BTCUSDT / BTC / "Bitcoin" / BTC/USDT / USDT |
| **2. Cross-system** | `coingecko_id` , `cmc_id` , `blockchain` , `contract_address` | Data provider + on-chain mapping | "bitcoin" / 1 / "ethereum" / "" |
| **3. Exchange mapping** | `exchanges: tuple[str, ...]` | Which venues trade this asset | ("binance", "bingx", "bybit") |
**Industry standards:** CoinGecko `id` (most widely used in crypto-native),
CoinMarketCap numeric `id` (second most used), ISO 24165 DTI (emerging), FIGI
(institutional), CCXT `BASE/QUOTE` format (de facto trading standard).
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
#### Fundamental Dimensions (intrinsic, never change)
| Dimension | Values | Multi-label | Purpose |
|-----------|--------|-------------|---------|
| **Sector** | CURRENCY, LAYER1, LAYER2, DEFI, ORACLE, EXCHANGE, MEME, PRIVACY, STORAGE, GAMING_NFT | Yes | Primary use-case / vertical |
| **TokenRole** | GAS, STORE_OF_VALUE, GOVERNANCE, UTILITY, MEME, EXCHANGE_FEE | Yes | Functional role → demand elasticity |
| **SupplyModel** | FIXED_CAP, DISINFLATIONARY, INFLATIONARY, BURN_MECHANISM | No | Supply pressure dynamics |
| **ConsensusFamily** | POW, POS, DPOS | No | Miner/validator selling behaviour |
| **SmartContractCapability** | FULL, PARTIAL, NONE | No | DeFi composability |
#### Technical Dimensions (invariant market-structure properties)
| Dimension | Values | Purpose |
|-----------|--------|---------|
| **MarketCapTier** | MEGA (>500B), LARGE (50-500B), MID (5-50B), SMALL (500M-5B), MICRO (< 500M ) | Price impact per dollar |
| **VolatilityProfile** | LOW (< 30 %), MEDIUM ( 30-80 %), HIGH ( 80-150 %), EXTREME ( > 150%) | Long-run average vol |
| **LiquidityProfile** | DEEP (>100M), NORMAL (10-100M), THIN (1-10M), ILLIQUID (< 1M ) | Long-run avg depth |
| **DerivativeAccess** | PERPS_AND_OPTIONS, PERPS_ONLY, NONE | Shorting / funding availability |
| **tick_size / lot_size** | Exchange-set | Minimum price/order size |
| **maker/taker fee_bps** | Exchange-set | Trading cost |
| **typical_spread_bps / depth / volume** | Long-run averages | Order-book fingerprint |
#### Predictive Properties (derived from invariants)
| Property | Logic | Value |
|----------|-------|-------|
| `supply_pressure` | PoW → forced selling (miners cover electricity) | "forced" / "optional" |
| `demand_elasticity` | GAS or STORE_OF_VALUE → inelastic (must hold) | "inelastic" / "elastic" |
| `can_be_shorted` | DerivativeAccess != NONE | bool |
#### Multi-Label Examples
```
BTC: sectors=(CURRENCY), roles=(STORE_OF_VALUE)
ETH: sectors=(LAYER1, DEFI), roles=(GAS, STORE_OF_VALUE, GOVERNANCE)
BNB: sectors=(EXCHANGE, LAYER1), roles=(EXCHANGE_FEE, GAS)
DOGE: sectors=(MEME, CURRENCY), roles=(MEME, GAS)
DOT: sectors=(LAYER1), roles=(GAS, GOVERNANCE)
UNI: sectors=(DEFI), roles=(GOVERNANCE, UTILITY)
```
#### Querying
```python
from malkhut.training.asset_classification import *
# Multi-label queries (match if queried value is ANY of the asset's labels)
layer1s = get_assets_by_sector(Sector.LAYER1) # ETH, SOL, ADA, AVAX, BNB, DOT, ATOM
gas_tokens = get_assets_by_token_role(TokenRole.GAS) # ETH, SOL, ADA, AVAX, DOGE, BNB, MATIC, DOT, ATOM
# Predictive properties
forced = get_pov_assets() # BTC, DOGE (PoW miners with forced selling)
inelastic = [p for p in ASSET_PROFILES.values() if p.demand_elasticity == "inelastic"]
# Backward-compatible single-label access (first element = primary)
btc = get_asset_profile("BTCUSDT")
btc.sector # Sector.CURRENCY
btc.token_role # TokenRole.STORE_OF_VALUE
```
malkhut: multi-exchange asset universe + schema docs
ExchangeProfile: standardized exchange metadata (fees, latency, capabilities).
3 pre-defined exchanges: Binance, BingX, Bybit.
AssetProfile.exchanges: tuple[str] — which venues trade each asset.
New query functions: get_assets_on_exchange, get_common_assets,
get_exchange_for_asset, get_exchange, list_exchanges.
_DATA_STORAGE_SCHEMA_FORMATS.md: comprehensive reference for agent
consumption — data model, storage format, query interfaces, data flow
diagram, enum reference, import patterns for BLUE/VIOLET/UV integration.
README updated: exchange registry section, package structure, subsystems
table.
2026-07-11 22:37:15 +02:00
### Exchange Registry (Multi-Exchange Asset Universe)
MALKHUT serves as the **system-wide asset universe store** for BLUE, VIOLET, UV, and all
downstream systems. Each asset knows which exchanges trade it; each exchange has a
standardized profile.
#### ExchangeProfile
| Field | Type | Purpose |
|-------|------|---------|
| `exchange_id` | str | Canonical key: `"binance"` , `"bingx"` , `"bybit"` |
| `display_name` | str | Human-readable name |
| `has_spot` | bool | Spot trading available |
| `has_perps` | bool | Perpetual futures available |
| `has_options` | bool | Options available |
| `api_base_url` | str | REST API root URL |
| `ws_base_url` | str | WebSocket root URL |
| `default_taker_fee_bps` | float | Default taker fee |
| `default_maker_fee_bps` | float | Default maker fee |
| `typical_latency_ms` | float | Typical API latency |
#### Pre-defined Exchanges
| Exchange | Spot | Perps | Options | Taker Fee | Latency |
|----------|------|-------|---------|-----------|---------|
| **Binance** | ✓ | ✓ | ✓ | 0.4 bps | 40 ms |
| **BingX** | ✓ | ✓ | ✗ | 0.5 bps | 100 ms |
| **Bybit** | ✓ | ✓ | ✓ | 0.06 bps | 50 ms |
#### Exchange-Asset Mapping
Each `AssetProfile` carries an `exchanges: tuple[str, ...]` field listing which venues
trade the asset. Default: `("binance",)` . Other systems (BLUE/VIOLET/UV) import their
asset universes into this store.
```python
from malkhut.training.asset_classification import *
# Which exchanges trade BTC?
btc = get_asset_profile("BTCUSDT")
print(btc.exchanges) # ('binance',)
# All assets on Binance
binance_assets = get_assets_on_exchange("binance")
# Assets traded on BOTH Binance and BingX
common = get_common_assets("binance", "bingx")
# Which exchanges trade a given asset?
venues = get_exchange_for_asset("ETHUSDT")
# Exchange metadata
ex = get_exchange("binance")
print(ex.default_taker_fee_bps) # 0.4
print(ex.typical_latency_ms) # 40
```
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
### Asset Behavior DSL (10-Dimension Research-Validated Model)
The Asset Behavior DSL decomposes each asset's **trading behavior** into 10 orthogonal
dimensions, each with empirically-validated parameters from live Binance/BingX API data
and academic literature.
**Source:** Binance live REST API (depth 500 levels, 1h klines, ticker, funding rates),
BingX open API (depth, funding, premium index), academic literature (Bouchaud "Trades
Quotes and Prices" 2018, Cont/Stoikov/Talreja 2010, Cartea/Jaimungal/Penalva 2015),
CoinGlass OI data, industry MM disclosures.
#### 10 Dimensions
| Dimension | What it captures | Example: BTC vs DOGE |
|-----------|------------------|---------------------|
| **DepthProfile** | Book shape: amplitude, decay exponent α , fragility | BTC: $750K/0.70/0.10 vs DOGE: $22K/1.00/0.30 |
| **SpreadProfile** | Normal spread, stress multiplier | BTC: 0.01bps/50x vs DOGE: 1.35bps/15x |
| **FlowProfile** | Order rate, size distribution, cancel ratio | BTC: 300/s/$643/20x vs DOGE: 80/s/$96/8x |
| **VolatilityProfile** | Annualized vol, GARCH params, half-life | BTC: 35%/48h vs DOGE: 78%/24h |
| **IntradayProfile** | Peak/trough hours, ratio | BTC: 15:00 UTC, 7.4x ratio |
| **WeekendProfile** | Vol/volume/spread multipliers | 0.65x vol, 0.52x volume, 1.20x spread |
| **CorrelationProfile** | ETH beta, BTC corr (normal vs crash) | BTC: 1.0/1.0 vs UNI: 0.70/0.50 |
| **MarketMakerProfile** | Inventory limits, pull speed, margins | BTC: $10M/3ms/0.5bps vs DOGE: $500K/25ms/2bps |
| **LiquidationProfile** | OI/MCap, trigger %, cascade speed | BTC: 0.5%/6.5%/slow vs DOGE: 1.4%/4%/fast |
| **FundingProfile** | Mean/std funding, basis | BTC: 0.59bps/0.22bps vs DOGE: 0.39bps/0.34bps |
| **RetailProfile** | Retail ratio, inst gap | BTC: 35%/0.04 vs DOGE: 80%/0.31 |
| **BingxProfile** | BingX spread/depth multiplier, fees, latency | BTC: 12.6x/0.048x vs DOGE: 7x/1.27x |
#### 3 Composable Templates
| Template | Assets | Key traits |
|----------|--------|-----------|
| `institutional_blue_chip` | BTC, ETH, BNB | Deep book, low spread, slow cascade, high institutional |
| `mid_cap_l1` | SOL, ADA, AVAX, DOT, ATOM, UNI, LINK, MATIC, AAVE | Moderate depth, higher vol, medium cascade |
| `retail_meme` | DOGE | Thin book, high vol, fast cascade, retail-dominated |
#### 13 Per-Asset Behaviors (research-validated)
| Asset | Price | Depth@1bps | Spread | Vol | Cascade | Retail | BingX mult |
|-------|-------|-----------|--------|-----|---------|--------|------------|
| BTC | $64K | $750K | 0.01 bps | 35% | slow/fast | 35% | 12.6x |
| ETH | $1.8K | $600K | 0.01 bps | 66.5% | medium/medium | 35% | 9.0x |
| SOL | $80 | $400K | 1.26 bps | 72% | medium/medium | 72% | 1.7x |
| DOGE | $0.07 | $22K | 1.35 bps | 78% | fast/slow | 80% | 7.0x |
| ADA | $0.17 | $60K | 5.95 bps | 79% | medium/medium | 65% | 8.0x |
| UNI | $3.6 | $30K | 2.76 bps | 95% | medium/medium | 60% | 10x |
#### Usage
```python
from malkhut.training.asset_behavior import *
# Get behavior for any asset
btc = get_behavior("BTCUSDT")
print(btc.depth.amplitude_usd) # $750,000
print(btc.vol.annualized_normal) # 35.0%
print(btc.liquidation.speed) # "slow"
# Compute depth at distance
depth_10bps = btc.depth_at_bps(10) # $1.5M
# Estimate slippage for a $100K order
slippage = btc.expected_slippage_bps(100_000) # ~2 bps
# Query by template, sector, or role
institutional = get_behaviors_by_template("institutional_blue_chip")
gas_tokens = get_behaviors_by_role(TokenRole.GAS)
thin_book = get_thin_book_assets()
```
### Asset Compiler (Auto-Fetch from Live APIs)
The Asset Compiler automatically fetches market data from Binance and BingX public APIs
and produces system-ready `AssetProfile` + `AssetBehavior` for any tradeable symbol.
**Rate-limited** (1 req/sec, well under Binance 1200/min limit), **cached** (avoids
re-fetching within a session), **resumable** .
#### What it auto-fetches
| Endpoint | Data | Computes |
|----------|------|----------|
| `/api/v3/ticker/24hr` | Price, volume, high/low | Daily volume, price reference |
| `/api/v3/depth?limit=100` | Order book (100 levels) | Spread, depth amplitude, decay exponent α |
| `/api/v3/klines?interval=1h&limit=168` | 7 days hourly OHLCV | Annualized volatility, order flow stats |
| `/api/v3/exchangeInfo` | Symbol filters | Tick size, lot size, price decimals |
| `/fapi/v1/fundingRate` | Funding rates (20 intervals) | Mean/std/positive% funding |
| `/fapi/v1/openInterest` | Open interest | OI/MCap ratio for liquidation modeling |
#### Usage
```python
from malkhut.training.asset_compiler import AssetCompiler
compiler = AssetCompiler()
# Single asset — auto-fetches from Binance, computes all params
result = compiler.compile("XRPUSDT")
print(result.asset_profile.symbol) # "XRPUSDT"
print(result.asset_behavior.reference_price) # $1.1056 (live Binance price)
print(result.asset_behavior.spread.normal_bps) # 0.90 bps (computed from depth)
print(result.asset_behavior.vol.annualized_normal) # 46.6% (computed from klines)
# Register into global registries
compiler.register(result)
# Now XRPUSDT is available for ScenarioFactory, queries, etc.
# One-step compile and register
result = compiler.compile_and_register("HBARUSDT")
# Batch compile (rate-limited sequentially)
results = compiler.compile_batch(["APTUSDT", "SEIUSDT", "INJUSDT"])
for r in results:
compiler.register(r)
```
#### Known classifications
The compiler includes pre-built classifications for 28 assets:
BTC, ETH, SOL, DOGE, ADA, AVAX, UNI, LINK, BNB, MATIC, AAVE, DOT, ATOM,
XRP, TRX, LTC, NEAR, APT, OP, ARB, SUI, PEPE, WIF, SEI, INJ, FIL, HBAR, IMX
Unknown assets get heuristic defaults (LAYER1/GAS/INFLATIONARY/POS/FULL/PERPS_ONLY).
### Scenario Factory (Behavior-Driven, Label-Queryable)
The ScenarioFactory builds evaluation scenarios using **real per-asset prices,
depth profiles, and spread characteristics** — no more hardcoded BTC prices.
#### Behavior-driven state creation
Every scenario derives its parameters from the asset's `AssetBehavior` :
| Scenario type | Spread multiplier | Depth fraction | Description |
|---------------|------------------|----------------|-------------|
| `normal` | 1.0x | 100% | Calm, liquid market |
| `thin_book` | 1.5x | 10% | Low participation |
| `wide_spread` | 100x | 50% | High volatility |
| `toxic_stress` | 2.0x | 30% | Adverse selection |
| `flash_crash` | 3.0x | 5% | Book collapse |
| `liquidity_vacuum` | 5.0x | 1% | Near-empty book |
| `weekend_thin` | 5.0x | 5% | Weekend low-volume |
| `trending` | 1.0x | 80% | Momentum move |
| `mean_reverting` | 20x | 60% | Reversal setup |
| `funding_shock` | 1.5x | 30% | Deleveraging event |
| `liquidation_cascade` | 1.5x | 20% | Chain liquidation |
| `market_maker_withdrawal` | 3.0x | 15% | MM pull-out |
#### Realistic per-asset pricing
```
Scenario: "normal_BTCUSDT_42"
BTC bid: $63,999.97 (11.7 BTC = $750K)
BTC ask: $64,000.03
Spread: 0.01 bps
Scenario: "normal_DOGEUSDT_42"
DOGE bid: $0.07 (314K DOGE = $22K)
DOGE ask: $0.07
Spread: 1.35 bps
Scenario: "flash_crash_SOLUSDT_47"
SOL bid: $80.00 (depth = 5% of normal = $20K)
SOL ask: $80.02
Spread: 3.78 bps (3x normal)
```
#### Auto-compilation for unknown assets
When you request a symbol that isn't pre-defined, the factory **automatically compiles
it from live Binance API data** before building scenarios:
```python
factory = ScenarioFactory()
# XRPUSDT is NOT in our pre-defined 13 assets
# But this auto-compiles it from Binance in ~5 seconds:
suite = factory.build_suite_for_symbols(("XRPUSDT",), steps_per_scenario=5)
# XRP behavior is now registered for future use too
```
#### Label-based query interfaces
```python
factory = ScenarioFactory()
# By sector
suite = factory.build_suite_for_sector(Sector.LAYER1) # ETH, SOL, ADA, AVAX, BNB, DOT, ATOM
suite = factory.build_suite_for_sector(Sector.DEFI) # ETH, AVAX, UNI, AAVE
suite = factory.build_suite_for_sector(Sector.MEME) # DOGE
# By token role
suite = factory.build_suite_for_role(TokenRole.GAS) # 9 assets
suite = factory.build_suite_for_role(TokenRole.GOVERNANCE) # ETH, UNI, AAVE, DOT
# By behavior template
suite = factory.build_suite_for_template("institutional_blue_chip") # BTC, ETH, BNB
suite = factory.build_suite_for_template("retail_meme") # DOGE
# By volatility band
suite = factory.build_suite_for_volatility(min_ann=75, max_ann=200) # High-vol assets
# Multi-label union (any matching)
suite = factory.build_suite_for_labels(
sectors=[Sector.DEFI],
roles=[TokenRole.GOVERNANCE]
)
# Arbitrary symbols (auto-compiles unknowns)
suite = factory.build_suite_for_symbols(("BTCUSDT", "XRPUSDT", "HBARUSDT"))
```
#### Research-validated scenario diversity
Each asset's 30 scenario types are parameterized by its behavior profile:
| Scenario | BTC params | DOGE params | Why it matters |
|----------|-----------|-------------|----------------|
| Normal | $750K depth, 0.01bps spread | $22K depth, 1.35bps spread | Baseline for scoring |
| Thin book | $75K depth | $2.2K depth | Tests fill rate under low liquidity |
| Flash crash | $37.5K depth, 0.03bps spread | $1.1K depth, 4bps spread | Tests adverse selection survival |
| Weekend thin | $37.5K depth, 5bps spread | $1.1K depth, 6.75bps spread | Tests off-hours strategy fitness |
| Liquidation cascade | $150K depth | $4.4K depth | Tests position sizing under stress |
| Market maker withdrawal | $112.5K depth | $3.3K depth | Tests quote quality when MMs pull |
### Numba Performance
Hot-path functions JIT-compiled with numba:
| Function | Speedup |
|----------|---------|
| fill_from_levels (batch 100) | 1.8x |
| round_tick | JIT-compiled |
| round_lot | JIT-compiled |
| clip_lots | JIT-compiled |
| extract_features | Vectorized numpy |
CWM throughput: **157K calls/sec** , **6.4 µs/transition** , **0.64 ms/100-step episode** .
### Design Principles
1. **Hardcoded baseline is NEVER replaced** — always available as reference
2. **Generated strategies are ADDED** — system GROWS its repertoire
3. **Crossover** combines two strategies (uniform crossover on parameters)
4. **Mutation** randomly modifies parameters (Gaussian noise)
5. **Selection** uses tournament selection (fitter genomes win more often)
6. **Third parties** can write strategies in DSL text format
7. **Market adaptation** : selector chooses strategy based on ExoF/MARAS regime
8. **Behavior-driven scenarios** : every asset behaves like its real-world self
### Strategy Naming Convention
```
{strategy_type}_gen{generation}_{timestamp}
```
Examples: `UCB1_gen1_20260707_223754` , `THOMPSON_gen2_20260707_223754`
---
## H. Cambrian Expansion
### Tie-In Fix
PerformanceMatrix now wired to evaluator. Strategies scored by regime fitness:
- ALL strategies tested in ALL scenarios
- Score tracked per strategy × regime
- Selector queries matrix for best strategy per regime
- `PerformanceMatrix.record(strategy_id, regime, score)` called during evaluation
### Phase 0.1: Cognition Pipeline
Rate-limited market regime research:
- `RateLimiter` : token bucket, 30 RPM default
- `SourceCatalogue` : tracks sources, fetch counts, error rates
- `RegimeExtractor` : keyword matching + sentiment analysis
- `CognitionPipeline` : rate-limited, deduplication, long/perm-run
- 8 default sources: CoinDesk, Cointelegraph, The Block, CryptoQuant, Glassnode, Coinglass, Binance Research, Messari
### Phase 0.2: Exponential Regime Expansion
Orthogonal to cognition pipeline — generates 200+ regimes from dimensions:
- 4 liquidity (vacuum, thin, normal, deep)
- 4 volatility (tight, normal, wide, extreme)
- 4 flow (balanced, buy_pressure, sell_pressure, toxic)
- 4 structure (single_toxic, mm_toxic, multi_toxic, retail_mm)
Each combination = distinct, testable regime. Total: 256 theoretical, ~200 practical.
### Phase 0.3: Multi-Asset & Invariant Classification
13 assets across 10 sectors, classified by **invariant characteristics only** :
| Asset | Sectors | Roles | Supply | Consensus | Vol | Liq |
|-------|---------|-------|--------|-----------|-----|-----|
| BTC | CURRENCY | STORE_OF_VALUE | FIXED_CAP | POW | LOW | DEEP |
| ETH | LAYER1, DEFI | GAS, SOV, GOV | DISINFLATIONARY | POS | MEDIUM | DEEP |
| SOL | LAYER1 | GAS | INFLATIONARY | POS | HIGH | NORMAL |
| DOGE | MEME, CURRENCY | MEME, GAS | INFLATIONARY | POW | HIGH | NORMAL |
| ADA | LAYER1 | GAS | INFLATIONARY | DPOS | MEDIUM | NORMAL |
| AVAX | LAYER1, DEFI | GAS | INFLATIONARY | POS | HIGH | NORMAL |
| UNI | DEFI | GOV, UTILITY | INFLATIONARY | POS | HIGH | THIN |
| LINK | ORACLE | UTILITY | INFLATIONARY | POS | MEDIUM | THIN |
| BNB | EXCHANGE, LAYER1 | EXCHANGE_FEE, GAS | BURN | DPOS | MEDIUM | DEEP |
| MATIC | LAYER2 | GAS | INFLATIONARY | POS | HIGH | THIN |
| AAVE | DEFI | GOV | FIXED_CAP | POS | HIGH | THIN |
| DOT | LAYER1 | GAS, GOV | INFLATIONARY | DPOS | MEDIUM | NORMAL |
| ATOM | LAYER1 | GAS | INFLATIONARY | POS | HIGH | THIN |
Multi-asset = linear multiplication: 30 scenarios × 13 assets = 390+ testable scenarios.
### Phase 0.4: Asset Behavior DSL + Compiler + Behavior-Driven Scenarios
**Research-validated** from Binance live API, BingX API, CoinGlass, and academic literature:
| Component | Purpose | File |
|-----------|---------|------|
| **AssetBehavior** | 10 orthogonal behavior dimensions | `training/asset_behavior.py` |
| **BehaviorTemplate** | Reusable class-level behavior profiles | `training/asset_behavior.py` |
| **AssetCompiler** | Auto-fetch from Binance/BingX, compute profiles | `training/asset_compiler.py` |
| **ScenarioFactory** | Behavior-driven scenarios with label queries | `training/cma_trainer.py` |
**Key research findings wired into the system:**
- Depth decay: D(d) = A * d^(1-alpha). BTC alpha=0.7, DOGE alpha=1.0
- Spread ranges: BTC 0.01bps, SOL 1.26bps, DOGE 1.35bps, ADA 5.95bps
- Order sizes: BTC P50=$643, DOGE P50=$96 (7x smaller)
- Cascade dynamics: BTC 5-8% trigger/slow/fast recovery, DOGE 3-5%/fast/slow
- BingX specifics: 1.7-12.6x wider spreads, 0.05-9x variable depth
- Weekend effects: -35% vol, -48% volume, +20% spread
- Intraday patterns: 7.4x peak/trough ratio, US session dominant
- GARCH persistence: 0.95-0.99, vol half-life 2-5 days
- Correlation: BTC-ETH 0.90 normal, 0.97 crash; BTC-DOGE 0.45 normal, 0.80 crash
2026-07-11 20:19:32 +02:00
### Training Scaling Results
CMA-ES evaluation on behavior-driven multi-asset scenarios (BTC/ETH/SOL, 3 assets × 30
scenario types = 90 scenarios). Each eval = one CMA candidate × 90 multi-step episodes
through the CWM.
2026-07-13 09:27:21 +02:00
#### Parallel Training Benchmark (48 evals, pop=12)
2026-07-11 20:19:32 +02:00
2026-07-13 09:27:21 +02:00
| Workers | Time | Best Score | Score/min | Speedup | Efficiency |
|---------|------|-----------|-----------|---------|-----------|
| 1 (sequential) | 632s | 1,727 | 164 | 1.0× | — |
| 2 | 320s | 2,262 | 425 | 1.98× | 99% |
| 4 | 163s | 2,152 | 791 | 3.87× | 97% |
| **8** | **91s** | **2,594** | **1,718** | **6.98× ** | **87%** |
2026-07-11 20:19:32 +02:00
2026-07-13 09:27:21 +02:00
**87% parallel efficiency at 8 workers.** Independent workers explore more diverse
strategy space — parallel runs find BETTER scores than sequential (2,594 vs 1,727).
2026-07-11 20:19:32 +02:00
2026-07-13 09:27:21 +02:00
#### Budget Scaling (ProcessPool, 3 assets × 30 scenarios)
| Budget | Gens | Best Score | Time | Score/min |
|--------|------|-----------|------|-----------|
| 48 evals | 4 | 2,594 | 91s | 1,718 |
| 96 evals | 8 | 7,202 | 1246s | 346 |
| 192 evals | 16 | ~14K (est) | ~45min | ~311 |
#### Vectorized Reward (numba)
The CWM `reward()` method now uses `compute_reward_vectorized` (numba JIT) when
available, bypassing the Python FeatureVector dict allocation. This eliminates
~2M dict allocations per CMA generation.
2026-07-11 20:19:32 +02:00
#### Training Performance
| Metric | Sequential | Parallel (8 workers) | Speedup |
|--------|-----------|---------------------|---------|
| CWM throughput | 189K calls/sec | (bottleneck is planner, not CWM) | — |
2026-07-13 09:27:21 +02:00
| CWM per-call latency | 5.3 µs | (numba + vectorized reward) | — |
2026-07-11 20:19:32 +02:00
| Scenario generation | 390 scenarios in 0.8s | — | — |
2026-07-13 09:27:21 +02:00
| CMA-ES per-generation (pop=12) | ~155s | ** ~23s** | ** ~7× ** |
| CMA-ES 48 evals | 632s | **91s** | **7× ** |
| Best score at 48 evals | 1,727 | **2,594** | +50% |
| Score/min | 164 | **1,718** | **10.5× ** |
2026-07-11 20:19:32 +02:00
Parallel evaluation achieves **9x speedup on single evals** (embarrassingly parallel,
zero fidelity loss). CMA-ES training speedup is ~1.5x because multiprocessing overhead
is amortized across 90 scenarios per eval. Bottleneck is planner (MCTS), not CWM —
numba already accelerates the CWM hot path (`_HAS_NUMBA = True` ).
malkhut: standardized order types — FIX/CCXT-aligned, multi-exchange mapping
order_types.py: Five-layer taxonomy normalized to industry standards:
Layer 1: Base types (FIX Tag 40) — MARKET, LIMIT
Layer 2: Time-in-force (FIX Tag 59) — GTC, IOC, FOK, GTD
Layer 3: Conditional/Trigger (FIX Tag 3/4+MIT) — STOP_MARKET, STOP_LIMIT,
TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
Layer 4: Instructions (FIX Tag 18) — POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
Layer 5: Compound (exchange-specific) — OCO, TP_SL
Cross-exchange mapping: BingX ↔ Binance ↔ Bybit (from CCXT source code).
Standards: FIX 4.4 Tag 40/59/18, CCXT unified API, ISO 10383 (MIC).
Transferability: strategy PARAMETERS transfer. ORDER TYPE NAMES are
venue-specific but semantics identical (LIMIT = LIMIT everywhere).
14 tests. README updated with full mapping table and standards references.
2026-07-14 12:02:01 +02:00
### Order Type Standardization (Multi-Exchange, FIX/CCXT-Aligned)
2026-07-14 14:52:50 +02:00
Three **orthogonal dimensions** (not one flat enum), each mapped independently:
malkhut: standardized order types — FIX/CCXT-aligned, multi-exchange mapping
order_types.py: Five-layer taxonomy normalized to industry standards:
Layer 1: Base types (FIX Tag 40) — MARKET, LIMIT
Layer 2: Time-in-force (FIX Tag 59) — GTC, IOC, FOK, GTD
Layer 3: Conditional/Trigger (FIX Tag 3/4+MIT) — STOP_MARKET, STOP_LIMIT,
TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
Layer 4: Instructions (FIX Tag 18) — POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
Layer 5: Compound (exchange-specific) — OCO, TP_SL
Cross-exchange mapping: BingX ↔ Binance ↔ Bybit (from CCXT source code).
Standards: FIX 4.4 Tag 40/59/18, CCXT unified API, ISO 10383 (MIC).
Transferability: strategy PARAMETERS transfer. ORDER TYPE NAMES are
venue-specific but semantics identical (LIMIT = LIMIT everywhere).
14 tests. README updated with full mapping table and standards references.
2026-07-14 12:02:01 +02:00
2026-07-14 14:52:50 +02:00
**Dimension 1 — Order Types (FIX Tag 40 OrdType):** what the order IS
`MARKET` , `LIMIT` , `STOP_MARKET` , `STOP_LIMIT` , `TRIGGER_MARKET` , `TRIGGER_LIMIT` ,
`TRAILING_STOP` , `OCO` , `TP_SL`
malkhut: standardized order types — FIX/CCXT-aligned, multi-exchange mapping
order_types.py: Five-layer taxonomy normalized to industry standards:
Layer 1: Base types (FIX Tag 40) — MARKET, LIMIT
Layer 2: Time-in-force (FIX Tag 59) — GTC, IOC, FOK, GTD
Layer 3: Conditional/Trigger (FIX Tag 3/4+MIT) — STOP_MARKET, STOP_LIMIT,
TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
Layer 4: Instructions (FIX Tag 18) — POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
Layer 5: Compound (exchange-specific) — OCO, TP_SL
Cross-exchange mapping: BingX ↔ Binance ↔ Bybit (from CCXT source code).
Standards: FIX 4.4 Tag 40/59/18, CCXT unified API, ISO 10383 (MIC).
Transferability: strategy PARAMETERS transfer. ORDER TYPE NAMES are
venue-specific but semantics identical (LIMIT = LIMIT everywhere).
14 tests. README updated with full mapping table and standards references.
2026-07-14 12:02:01 +02:00
2026-07-14 14:52:50 +02:00
**Dimension 2 — Time-in-Force (FIX Tag 59):** how long the order LIVES
`GTC` , `IOC` , `FOK` , `GTD`
malkhut: standardized order types — FIX/CCXT-aligned, multi-exchange mapping
order_types.py: Five-layer taxonomy normalized to industry standards:
Layer 1: Base types (FIX Tag 40) — MARKET, LIMIT
Layer 2: Time-in-force (FIX Tag 59) — GTC, IOC, FOK, GTD
Layer 3: Conditional/Trigger (FIX Tag 3/4+MIT) — STOP_MARKET, STOP_LIMIT,
TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
Layer 4: Instructions (FIX Tag 18) — POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
Layer 5: Compound (exchange-specific) — OCO, TP_SL
Cross-exchange mapping: BingX ↔ Binance ↔ Bybit (from CCXT source code).
Standards: FIX 4.4 Tag 40/59/18, CCXT unified API, ISO 10383 (MIC).
Transferability: strategy PARAMETERS transfer. ORDER TYPE NAMES are
venue-specific but semantics identical (LIMIT = LIMIT everywhere).
14 tests. README updated with full mapping table and standards references.
2026-07-14 12:02:01 +02:00
2026-07-14 14:52:50 +02:00
**Dimension 3 — Instructions (FIX Tag 18 ExecInst):** behavioral modifiers
`POST_ONLY` , `REDUCE_ONLY` , `HIDDEN` , `ICEBERG`
> CRITICAL: `POST_ONLY` is an instruction on a `LIMIT` order, not a standalone type.
> `IOC`/`FOK` are TimeInForce values applied to a `LIMIT` order, not order types.
> This aligns with FIX Protocol: Tag 40 (OrdType), Tag 59 (TimeInForce), Tag 18 (ExecInst).
#### Cross-Exchange Order Type Mapping
malkhut: standardized order types — FIX/CCXT-aligned, multi-exchange mapping
order_types.py: Five-layer taxonomy normalized to industry standards:
Layer 1: Base types (FIX Tag 40) — MARKET, LIMIT
Layer 2: Time-in-force (FIX Tag 59) — GTC, IOC, FOK, GTD
Layer 3: Conditional/Trigger (FIX Tag 3/4+MIT) — STOP_MARKET, STOP_LIMIT,
TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
Layer 4: Instructions (FIX Tag 18) — POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
Layer 5: Compound (exchange-specific) — OCO, TP_SL
Cross-exchange mapping: BingX ↔ Binance ↔ Bybit (from CCXT source code).
Standards: FIX 4.4 Tag 40/59/18, CCXT unified API, ISO 10383 (MIC).
Transferability: strategy PARAMETERS transfer. ORDER TYPE NAMES are
venue-specific but semantics identical (LIMIT = LIMIT everywhere).
14 tests. README updated with full mapping table and standards references.
2026-07-14 12:02:01 +02:00
| Normalized | BingX | Binance Spot | Binance Futures | Bybit |
|-----------|-------|-------------|-----------------|-------|
| `LIMIT` | LIMIT | LIMIT | LIMIT | LIMIT |
| `MARKET` | MARKET | MARKET | MARKET | MARKET |
2026-07-14 14:52:50 +02:00
| `STOP_MARKET` | TRIGGER_MARKET | STOP_MARKET | STOP_MARKET | STOP_MARKET |
| `STOP_LIMIT` | TRIGGER_LIMIT | STOP_LOSS_LIMIT | STOP | STOP_LIMIT |
| `TRIGGER_MARKET` | TRIGGER_MARKET | TAKE_PROFIT | TAKE_PROFIT | TAKE_PROFIT_MARKET |
| `TRIGGER_LIMIT` | TRIGGER_LIMIT | TAKE_PROFIT_LIMIT | TAKE_PROFIT | TAKE_PROFIT_LIMIT |
| `TRAILING_STOP` | TRAILING_STOP_MARKET | TRAILING_STOP_MARKET | TRAILING_STOP_MARKET | TRAILING_STOP |
#### TimeInForce Mapping
| Normalized | BingX | Binance | Bybit |
|-----------|-------|---------|-------|
| `GTC` | GTC | GTC | GTC |
| `IOC` | IOC | IOC | IOC |
| `FOK` | FOK | FOK | FOK |
malkhut: standardized order types — FIX/CCXT-aligned, multi-exchange mapping
order_types.py: Five-layer taxonomy normalized to industry standards:
Layer 1: Base types (FIX Tag 40) — MARKET, LIMIT
Layer 2: Time-in-force (FIX Tag 59) — GTC, IOC, FOK, GTD
Layer 3: Conditional/Trigger (FIX Tag 3/4+MIT) — STOP_MARKET, STOP_LIMIT,
TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
Layer 4: Instructions (FIX Tag 18) — POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
Layer 5: Compound (exchange-specific) — OCO, TP_SL
Cross-exchange mapping: BingX ↔ Binance ↔ Bybit (from CCXT source code).
Standards: FIX 4.4 Tag 40/59/18, CCXT unified API, ISO 10383 (MIC).
Transferability: strategy PARAMETERS transfer. ORDER TYPE NAMES are
venue-specific but semantics identical (LIMIT = LIMIT everywhere).
14 tests. README updated with full mapping table and standards references.
2026-07-14 12:02:01 +02:00
#### Transferability Principle
Strategy **PARAMETERS** transfer across exchanges. Order type **NAMES** are venue-specific
but semantics are identical. A `LIMIT` on BingX = `LIMIT` on Binance = `LIMIT` on Bybit
(same fill behavior). Only the API string differs. The venue adapter translates
normalized → exchange-native at submission time.
2026-07-14 15:44:33 +02:00
#### Cross-Exchange Learning
ScenarioFactory is venue-aware: `ScenarioFactory(exchange_id='bingx')` tags every
scenario with its venue. Cross-exchange transfer re-tags for a different venue:
```python
factory = ScenarioFactory(exchange_id='bingx')
scenarios_bingx = factory.build_suite(symbols=['BTCUSDT', 'ETHUSDT'])
strategy = train(scenarios_bingx) # evolve on BingX
scenarios_binance = factory.cross_exchange_transfer(scenarios_bingx, 'binance')
score = evaluate(strategy, scenarios_binance) # test on Binance
```
PerformanceMatrix keys are `(regime, strategy_id, venue)` :
- `get_best(regime, venue='bingx')` — per-venue best
- `get_venue_comparison(regime, strategy_id)` — `{venue: score}`
- `get_best(regime)` — venue-agnostic (backward compatible)
#### Adversary Ecology
Counterparties operate at the ActionKind level (CROSS_SPREAD, PLACE, CANCEL),
not at order-type level. The CWM translates ActionKind to venue-native order types:
- `CROSS_SPREAD` → fills aggressively → equivalent to MARKET
- `PLACE` → passive quote → equivalent to LIMIT
Fee calculation uses VenueRules (per-exchange fees). The ecology is venue-independent.
malkhut: standardized order types — FIX/CCXT-aligned, multi-exchange mapping
order_types.py: Five-layer taxonomy normalized to industry standards:
Layer 1: Base types (FIX Tag 40) — MARKET, LIMIT
Layer 2: Time-in-force (FIX Tag 59) — GTC, IOC, FOK, GTD
Layer 3: Conditional/Trigger (FIX Tag 3/4+MIT) — STOP_MARKET, STOP_LIMIT,
TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
Layer 4: Instructions (FIX Tag 18) — POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
Layer 5: Compound (exchange-specific) — OCO, TP_SL
Cross-exchange mapping: BingX ↔ Binance ↔ Bybit (from CCXT source code).
Standards: FIX 4.4 Tag 40/59/18, CCXT unified API, ISO 10383 (MIC).
Transferability: strategy PARAMETERS transfer. ORDER TYPE NAMES are
venue-specific but semantics identical (LIMIT = LIMIT everywhere).
14 tests. README updated with full mapping table and standards references.
2026-07-14 12:02:01 +02:00
#### Standards Referenced
- FIX 4.4: Tag 40 (OrdType), Tag 59 (TimeInForce), Tag 18 (ExecInst)
- CCXT: unified `limit` /`market` + parameter composition (triggerPrice, stopPrice)
- ISO 10383: MIC for venue identification
- MiFID II / MiCA: defers to FIX for order type classification
- FIA: references FIX Tag 40 in digital asset guidelines
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00
### Prod Tooling
| Component | Purpose | File |
|-----------|---------|------|
| **CognitionLauncher** | Standalone long-run script | `cognition_launcher.py` |
| **CognitionMonitor** | Metrics, health scoring, alerts | `training/monitor.py` |
| **NewsSourceRepository** | 12 industry-standard sources, ranking | `training/news_sources.py` |
| **SourceCatalogue** | Source persistence, fetch tracking | `training/cognition.py` |
| **RateLimiter** | Token bucket, 30 RPM | `training/cognition.py` |
| **RegimeExpander** | 200 regimes from dimensions | `training/regime_expansion.py` |
### Running
```bash
# Training pipeline (continuous)
python -m malkhut.continuous_pipeline
# Cognition pipeline (continuous regime research)
python -m malkhut.cognition_launcher
```