Fill quality is MALKHUT's core aim. Wired end-to-end:
1. FillQuality state (state.py):
- slippage_bps, price_improvement_bps, levels_consumed
- is_maker_fill, rolling_fill_rate, post_fill_adverse_bps
- fill_value_score: composite metric for optimization
- Added to MarketWorldState.fill_quality field
2. HftBacktestCWM.transition() (hft_cwm.py):
- _compute_fill_quality() computes all metrics per transition
- Fill quality now tracked for every CWM step
- Empty book guards added for safety
3. MinimalCryptoLOBCWM.transition() (core.py):
- Same fill quality computation for deterministic fallback
- Empty book guards added
4. Reward function (hft_cwm.py):
- fill_quality_reward = w_fill_probability * fill_value_score (PRIMARY)
- Bonus for maker fills that improve price
- Penalty for adverse selection after fill
- Base reward (PnL, adverse selection, fees) preserved
5. PerformanceMatrix (selector.py):
- RegimeStrategyScore: 4 new fill quality fields
- record(): accepts fill_rate, slippage, price_improvement, fill_value_score
- EMA updates for all fill quality metrics
6. EpisodeResult (cma_trainer.py):
- avg_fill_value_score, avg_price_improvement_bps, avg_post_fill_adverse_bps
- Accumulated per-step during _run_episode
- Recorded to PerformanceMatrix in evaluate_candidate
All 1379+ tests green.
Rolling stats (no episode accumulation), CMA-ES every 10 cycles (3 evals),
gc.collect() after CMA, global try/except for crash safety.
100-opponent swarm, 9 assets, 270 scenarios.
Fixed OOM kill by replacing all_episodes list accumulation with:
- Rolling stats (clear every 20 episodes)
- Only PnL history kept for characterization
- Peak/worst tracking without full episode storage
- Periodic stdout reports from rolling aggregates
100-opponent swarm + 9 assets × 30 scenarios = 270 scenarios per cycle.
CMA-ES every 5 cycles (3 evals). 3-hour target.
- 100 diverse opponents (randomized params within each type)
- Risk gate: empty book guard in _post_only_would_cross
- CMA-ES: only every 5 cycles, 3 evals, robust error handling
- Main loop: try/except prevents silent crashes
- Profiling: 11.7 steps/sec with 100 opponents
Risk gate (risk/gate.py) — 5 stubs implemented:
1. _kill_switch_active(): operator-controlled emergency stop via set_kill_switch()
2. _cancel_rate_would_exceed(): tracks cancel timestamps per symbol in 60s
sliding window, blocks if >= MAX_CANCELS_PER_SYMBOL_PER_MINUTE
3. _would_self_trade(): checks open orders for same symbol+side at same price
(within tick_size), skipping the cancel_order_id for CANCEL_REPLACE
4. _would_exceed_symbol_notional(): sums current open order notional + new
order notional, blocks if > equity * MAX_SYMBOL_NOTIONAL_FRACTION
5. _violates_venue_minima(): checks tick alignment, lot rounding, min_qty,
and min_notional — all float-robust comparisons
ScenarioFactory — 3 remaining hardcoded scenarios converted:
1. _spread_tightening: spread_mult=0.3, depth_fraction=1.0 (was hardcoded BTC)
2. _cross_venue_arb: spread_mult=0.5, depth_fraction=0.5 (was hardcoded BTC)
3. _cross_exchange_arb_stress: spread_mult=0.8, depth_fraction=0.3 (was hardcoded BTC)
All 30 scenarios now use _behavior_state() — zero hardcoded prices remain.
675 tests pass. Zero regressions.
- evaluator: passes scenario.venue to matrix.record(venue=...)
- PerformanceMatrix.record(): accepts venue parameter (default='bingx')
- Enables cross-exchange learnings: same strategy tested on BingX vs Binance
gets separate performance entries per venue
Adversary ecology analysis:
Counterparties operate at ActionKind level (CROSS_SPREAD/PLACE/CANCEL),
not at order-type level. The CWM infers order type from ActionKind:
CROSS_SPREAD → fills aggressively → equivalent to MARKET
PLACE → passive quote → equivalent to LIMIT
This is correct and venue-independent. Fee calculation already uses
VenueRules (per-exchange fees). No adversary changes needed.
ScenarioFactory + CWM + Engine changes:
1. Scenario.venue field (default='bingx') — each scenario tagged with venue
2. ScenarioFactory.exchange_id parameter — controls which exchange scenarios simulate
3. _make_state + _behavior_state: venue propagated to VenueRules.exchange
4. All 34 scenario builders: venue=self.exchange_id
5. cross_exchange_transfer(): re-tag scenarios for different exchange
(strategy evolved on BingX can be re-evaluated on Binance)
6. CWM core.py: is_maker check updated for three-dimensional order model
(POST_ONLY no longer in OrderType; uses post_only flag instead)
Cross-exchange learning flow:
factory_bingx = ScenarioFactory(exchange_id='bingx')
scenarios_bingx = factory_bingx.build_suite(symbols=[...])
strategy = train(scenarios_bingx) # evolve on BingX
factory_binance = ScenarioFactory(exchange_id='binance')
scenarios_binance = factory_bingx.cross_exchange_transfer(
scenarios_bingx, target_exchange='binance')
score = evaluate(strategy, scenarios_binance) # test on Binance
All tests pass. Strategy PARAMETERS transfer; only venue tag + fees + order mapping change.
adapter.py now uses normalize_to_exchange(action.order_type, 'bingx')
to translate normalized order types to BingX-native strings.
Falls back to LIMIT/MARKET/POST_ONLY for backward compatibility.
This is the critical integration point: standardized order types flow
from FulfilmentAction → CWM → VenueAdapter → exchange API.
DaatQuery: 8-feature market state representation
DaatVerdict: KNOWN / MARGINAL / OUT_OF_DISTRIBUTION
daat_classify: cosine RETRIEVE → magnitude GATE → local MODEL
- Cosine finds nearest explored state (directional match)
- Magnitude gate detects out-of-distribution states
- Empty explored set → always OUT_OF_DISTRIBUTION
9 tests covering: known state, OOD, empty explored, marginal, result fields.
No Unicode in code. All tests pass.
Item 1 — Mutation-litmus test (spec §1 item 3):
- test_taker_fee_10x_changes_score: fee change MUST affect score
- test_zero_fees_vs_correct_fees: zero vs 5bps must differ
- BOTH PASS — confirms fees ARE wired into reward function
- If fees were ignored, these tests would go RED
Item 3 — Maker fee verification (spec §1 item 5):
- Added '# UNVERIFIED — no maker fills on record as of 2026-07-13'
to Binance and Bybit exchange profiles
- Maker fee sign (positive on BingX, negative rebate on others)
is correct after fee fix but unverified from actual fills.
Items 2,4-10 remain for implementation.
CMAESTrainer.train() now accepts workers parameter and passes it to
evaluate_candidate(), enabling parallel episode evaluation during
actual training (not just in tests/benchmarks).
Benchmark result: ProcessPoolExecutor is optimal (4.76x speedup).
Ray is slower (0.36x) due to head init + plasma overhead for 90 scenarios.
1. Vectorized reward path (cwm/core.py):
- Wired up existing compute_reward_vectorized from numba_core (was unused!)
- Eliminates FeatureVector dict allocation + Python dict lookups on hot path
- Numba path used when _HAS_NUMBA=True, Python fallback otherwise
- Bit-identical: same math operations, just via numba JIT
2. Ray-based parallel eval (training/ray_eval.py):
- Industrial multi-core execution via Ray (used by OpenAI/Anyscale)
- ray.put() stores params/scenarios in shared object store (no pickle per worker)
- Each worker: own CWM + planner, zero shared state, no races
- Bit-identical: same seed + same params = same results regardless of worker count
- PolicyEvaluator.evaluate_candidate: new use_ray=True parameter
3. VBT post-analysis (training/vbt_analysis.py):
- episodes_to_pnl_array, episodes_to_metrics (Sharpe, Sortino, VaR, win_rate, etc.)
- cross_asset_comparison, parameter_sensitivity
- format_metrics for human-readable output
- Analysis tool only — runs AFTER engine produces results
4. numba_core.py: added missing 'import math' for compute_reward_vectorized
13 new tests: vectorized reward bit-identity, Ray determinism, Ray result fields,
VBT metrics structure, cross-asset comparison, parameter sensitivity, edge cases.
Total: 1178 tests, 50 files, all green, zero regressions.
Architecture: DuckDB for persistence + full in-memory materialization for reads.
All reads served from Python dicts (sub-microsecond). DuckDB only hit on writes.
Performance evolution (get_asset benchmark):
V0 (raw DuckDB): 876µs per call
V1 (LRU cache): 2.3µs per call (380x)
V2 (in-memory): 0.2µs per call (4380x)
All reads now sub-microsecond:
get_asset: 0.2µs (was 876µs)
query(blockian): 6.6µs (was 2.2ms)
query(sector): 6.6µs (was 3.2ms)
exchange lookup: 12.5µs (was 1.5ms)
full scan: 5.9µs (was 1.8ms)
behavior: 0.4µs
Write path: sync_from_profiles batch-inserts all data, then materializes
into Python dicts. Resync: 76ms (was 210ms, 2.8x faster).
Data integrity: DuckDB WAL provides crash recovery. In-memory dicts are
reconstructed from DB on every sync/close-reopen cycle. Zero data loss.