MCTS planner with dynamic book:
Episode 3: slippage=1.125 bps (first aggressive fills)
Episode 20: slippage=0.177 bps (84% reduction)
Episode 30: slippage=0.008 bps (99% reduction!)
The planner learns to:
1. Place passive orders at better offsets
2. Wait for book to move before crossing
3. Use urgency-driven maker/taker decision
4. Reduce slippage through queue position optimization
PnL stays positive throughout (+1687 to +5048 bps).
Fill value improving from -0.319 to -0.000 (less negative = better).
REQUOTE is now distinct from QUOTE:
QUOTE: PLACE new order (no existing to cancel)
REQUOTE: CANCEL_REPLACE existing + place new (immediate)
CHASE: PLACE with short TTL (auto-cancel retry)
CANCEL: Remove existing order
action_menu generates REQUOTE with metadata={'requote': True} for existing orders.
DSL REQUOTE produces CANCEL_REPLACE when existing order, falls back to PLACE.
All 800+ tests pass.
README: Fill Quality section (core optimization target, metrics, reward
function, PerformanceMatrix, CMA-ES integration)
HftBacktestCWM integration doc: FillQuality dataclass, fill_value_score
computation, reward function weighting, PerformanceMatrix tracking
OB microstructure study: Section 13 — Fill Quality Optimization,
per-asset expectations, optimization strategy, connection to OB dynamics
Fill quality is MALKHUT's core aim: the system learns to get better fills
(faster, better-priced, less adverse selection) across regimes and venues.
Fill quality is MALKHUT's core aim. Wired end-to-end:
1. FillQuality state (state.py):
- slippage_bps, price_improvement_bps, levels_consumed
- is_maker_fill, rolling_fill_rate, post_fill_adverse_bps
- fill_value_score: composite metric for optimization
- Added to MarketWorldState.fill_quality field
2. HftBacktestCWM.transition() (hft_cwm.py):
- _compute_fill_quality() computes all metrics per transition
- Fill quality now tracked for every CWM step
- Empty book guards added for safety
3. MinimalCryptoLOBCWM.transition() (core.py):
- Same fill quality computation for deterministic fallback
- Empty book guards added
4. Reward function (hft_cwm.py):
- fill_quality_reward = w_fill_probability * fill_value_score (PRIMARY)
- Bonus for maker fills that improve price
- Penalty for adverse selection after fill
- Base reward (PnL, adverse selection, fees) preserved
5. PerformanceMatrix (selector.py):
- RegimeStrategyScore: 4 new fill quality fields
- record(): accepts fill_rate, slippage, price_improvement, fill_value_score
- EMA updates for all fill quality metrics
6. EpisodeResult (cma_trainer.py):
- avg_fill_value_score, avg_price_improvement_bps, avg_post_fill_adverse_bps
- Accumulated per-step during _run_episode
- Recorded to PerformanceMatrix in evaluate_candidate
All 1379+ tests green.
Rolling stats (no episode accumulation), CMA-ES every 10 cycles (3 evals),
gc.collect() after CMA, global try/except for crash safety.
100-opponent swarm, 9 assets, 270 scenarios.
Fixed OOM kill by replacing all_episodes list accumulation with:
- Rolling stats (clear every 20 episodes)
- Only PnL history kept for characterization
- Peak/worst tracking without full episode storage
- Periodic stdout reports from rolling aggregates
100-opponent swarm + 9 assets × 30 scenarios = 270 scenarios per cycle.
CMA-ES every 5 cycles (3 evals). 3-hour target.
- 100 diverse opponents (randomized params within each type)
- Risk gate: empty book guard in _post_only_would_cross
- CMA-ES: only every 5 cycles, 3 evals, robust error handling
- Main loop: try/except prevents silent crashes
- Profiling: 11.7 steps/sec with 100 opponents
Risk gate (risk/gate.py) — 5 stubs implemented:
1. _kill_switch_active(): operator-controlled emergency stop via set_kill_switch()
2. _cancel_rate_would_exceed(): tracks cancel timestamps per symbol in 60s
sliding window, blocks if >= MAX_CANCELS_PER_SYMBOL_PER_MINUTE
3. _would_self_trade(): checks open orders for same symbol+side at same price
(within tick_size), skipping the cancel_order_id for CANCEL_REPLACE
4. _would_exceed_symbol_notional(): sums current open order notional + new
order notional, blocks if > equity * MAX_SYMBOL_NOTIONAL_FRACTION
5. _violates_venue_minima(): checks tick alignment, lot rounding, min_qty,
and min_notional — all float-robust comparisons
ScenarioFactory — 3 remaining hardcoded scenarios converted:
1. _spread_tightening: spread_mult=0.3, depth_fraction=1.0 (was hardcoded BTC)
2. _cross_venue_arb: spread_mult=0.5, depth_fraction=0.5 (was hardcoded BTC)
3. _cross_exchange_arb_stress: spread_mult=0.8, depth_fraction=0.3 (was hardcoded BTC)
All 30 scenarios now use _behavior_state() — zero hardcoded prices remain.
675 tests pass. Zero regressions.
- evaluator: passes scenario.venue to matrix.record(venue=...)
- PerformanceMatrix.record(): accepts venue parameter (default='bingx')
- Enables cross-exchange learnings: same strategy tested on BingX vs Binance
gets separate performance entries per venue
Adversary ecology analysis:
Counterparties operate at ActionKind level (CROSS_SPREAD/PLACE/CANCEL),
not at order-type level. The CWM infers order type from ActionKind:
CROSS_SPREAD → fills aggressively → equivalent to MARKET
PLACE → passive quote → equivalent to LIMIT
This is correct and venue-independent. Fee calculation already uses
VenueRules (per-exchange fees). No adversary changes needed.
ScenarioFactory + CWM + Engine changes:
1. Scenario.venue field (default='bingx') — each scenario tagged with venue
2. ScenarioFactory.exchange_id parameter — controls which exchange scenarios simulate
3. _make_state + _behavior_state: venue propagated to VenueRules.exchange
4. All 34 scenario builders: venue=self.exchange_id
5. cross_exchange_transfer(): re-tag scenarios for different exchange
(strategy evolved on BingX can be re-evaluated on Binance)
6. CWM core.py: is_maker check updated for three-dimensional order model
(POST_ONLY no longer in OrderType; uses post_only flag instead)
Cross-exchange learning flow:
factory_bingx = ScenarioFactory(exchange_id='bingx')
scenarios_bingx = factory_bingx.build_suite(symbols=[...])
strategy = train(scenarios_bingx) # evolve on BingX
factory_binance = ScenarioFactory(exchange_id='binance')
scenarios_binance = factory_bingx.cross_exchange_transfer(
scenarios_bingx, target_exchange='binance')
score = evaluate(strategy, scenarios_binance) # test on Binance
All tests pass. Strategy PARAMETERS transfer; only venue tag + fees + order mapping change.
Updated README to reflect Fable's corrections:
- OrderType/TimeInForce/Instructions as three orthogonal dimensions
- POST_ONLY/IOC/FOK correctly described as non-types
- BingX trailing_stop -> TRAILING_STOP_MARKET
- Three mapping tables (order type, TIF, instructions)
- Integration status updated
adapter.py now uses normalize_to_exchange(action.order_type, 'bingx')
to translate normalized order types to BingX-native strings.
Falls back to LIMIT/MARKET/POST_ONLY for backward compatibility.
This is the critical integration point: standardized order types flow
from FulfilmentAction → CWM → VenueAdapter → exchange API.
DaatQuery: 8-feature market state representation
DaatVerdict: KNOWN / MARGINAL / OUT_OF_DISTRIBUTION
daat_classify: cosine RETRIEVE → magnitude GATE → local MODEL
- Cosine finds nearest explored state (directional match)
- Magnitude gate detects out-of-distribution states
- Empty explored set → always OUT_OF_DISTRIBUTION
9 tests covering: known state, OOD, empty explored, marginal, result fields.
No Unicode in code. All tests pass.
Item 1 — Mutation-litmus test (spec §1 item 3):
- test_taker_fee_10x_changes_score: fee change MUST affect score
- test_zero_fees_vs_correct_fees: zero vs 5bps must differ
- BOTH PASS — confirms fees ARE wired into reward function
- If fees were ignored, these tests would go RED
Item 3 — Maker fee verification (spec §1 item 5):
- Added '# UNVERIFIED — no maker fills on record as of 2026-07-13'
to Binance and Bybit exchange profiles
- Maker fee sign (positive on BingX, negative rebate on others)
is correct after fee fix but unverified from actual fills.
Items 2,4-10 remain for implementation.