Commit Graph

77 Commits

Author SHA1 Message Date
Codex
97a770da65 malkhut: online EWMA self-calibrating slippage model
Flight7 model underestimates by 80% in CWM dynamic book:
  raw predicted: 0.034 bps, actual: 0.180 bps
  Constant error across 22K episodes — no feedback loop.

Root cause: Flight7 calibrated on real BingX taker fills, but CWM's
synthetic dynamic book has different fill characteristics.

Fix: SlippageSelfCalibrator with EWMA feedback loop.
  After each fill: error = actual - predicted (clipped to +/-20 bps)
  EWMA smooths per-symbol errors (alpha=0.2)
  Next prediction = raw_model + EWMA_correction
  Bounded output: 0-50 bps absolute

Convergence (300 eps across 8 assets):
  ETH: 9% error (from 80%)
  SOL: 3.5%
  DOGE: 6.7%
  LINK: 5.7%
  ADA: 9.7%
  BTC: 48.6% (low fill count, converging)
  AVAX: 28.6% (low fill count)
  UNI: 52.5% (low fill count, early outlier)

Truthfulness guarantees:
  - Correction is observable (CALIBRATOR.correction(symbol))
  - Resets between runs (no hidden state)
  - Only uses observed fills, no assumptions
  - Error clipping prevents outlier domination
  - Absolute bounds prevent runaway
2026-07-20 15:17:16 +02:00
Codex
c1a888faf3 malkhut: MCTS planner E2E — planner IS learning (slippage 1.1→0.008 bps)
MCTS planner with dynamic book:
  Episode 3: slippage=1.125 bps (first aggressive fills)
  Episode 20: slippage=0.177 bps (84% reduction)
  Episode 30: slippage=0.008 bps (99% reduction!)

The planner learns to:
  1. Place passive orders at better offsets
  2. Wait for book to move before crossing
  3. Use urgency-driven maker/taker decision
  4. Reduce slippage through queue position optimization

PnL stays positive throughout (+1687 to +5048 bps).
Fill value improving from -0.319 to -0.000 (less negative = better).
2026-07-19 19:06:10 +02:00
Codex
db844c775a malkhut: configurable friction per scenario + 5h learning test 2026-07-19 05:47:34 +02:00
Codex
3b1dae6cbe malkhut: configurable friction per scenario + 5h learning test
1. Friction configurable per scenario:
   Scenario gains maker_fee_bps, taker_fee_bps, adverse_cost_bps
   ScenarioFactory accepts friction overrides
   All 34 scenario builder calls updated
   System learns in ALL conditions (free maker, BingX real, Binance-like)

2. 5h learning test (long_learning_5h.py):
   - 1500 opponents, 9 assets, 270 scenarios
   - OOM/CPU monitoring (ResourceMonitor)
   - CMA-ES every 15 reports
   - Rolling fill_value/surprise/fill_rate tracking
   - Needs screen/tmux on host for 5h execution

Note: background redirect issue on this shell session.
Run via: cd MALKHUT && PYTHONUNBUFFERED=1 NUMBA_CACHE_DIR=/tmp/numba_cache python -m malkhut.long_learning_5h
2026-07-19 04:18:01 +02:00
Codex
70f33f6911 malkhut: friction settings configurable per scenario
Scenario gains 3 new fields:
  maker_fee_bps: Optional[float] = None (override per-scenario)
  taker_fee_bps: Optional[float] = None (override per-scenario)
  adverse_cost_bps: Optional[float] = None (per-fill adverse selection)

ScenarioFactory gains friction constructor params:
  ScenarioFactory(exchange_id='bingx', maker_fee_bps=2.0, taker_fee_bps=5.0)

_make_state() accepts friction overrides → applies to VenueRules
_behavior_state() passes friction overrides through
All 34 scenario builder calls updated with friction overrides.

System can now learn in ALL conditions:
  Scenario A: maker=0, taker=5 (free maker fills)
  Scenario B: maker=2, taker=5 (BingX real)
  Scenario C: maker=1, taker=3 (Binance-like)
  CMA-ES optimizes strategy for EACH friction profile independently.
2026-07-19 00:25:26 +02:00
Codex
c868dbfb66 malkhut(docs): updated all docs — fee model, markout, urgency threshold 2026-07-18 19:14:34 +02:00
Codex
ebf7f17132 malkhut: fee+slippage execution threshold 2026-07-18 17:44:20 +02:00
Codex
5503aafa28 malkhut: P0 guard + P1 BingX protective strings + tests fixed
P0 (safety): adapter.py rejects unmapped types (OCO, TP_SL) via
  is_type_available() guard. Returns None instead of silent LIMIT fallback.

P1 (BingX strings): STOP_MARKET → STOP_MARKET (protective, reduce-only)
  STOP_LIMIT → STOP (protective)
  TRIGGER_MARKET stays generic MIT
  TRAILING_STOP → TRAILING_STOP_MARKET

Tests updated to match corrected mappings.
2026-07-18 15:08:57 +02:00
Codex
041c879e82 malkhut: all exchange order types in DSL + action_menu
ActionType → OrderType mapping (common-sensical):
  QUOTE/REQUOTE/CHASE: LIMIT (passive, post_only)
  CROSS_SPREAD (high urgency): MARKET (immediate fill)
  CROSS_SPREAD (medium urgency): LIMIT + IOC (partial fill)
  STOP_LOSS: STOP_MARKET (trigger → market exit)
  TAKE_PROFIT: TRIGGER_MARKET (trigger → market exit)
  TRAILING_STOP: TRAILING_STOP (trailing stop exit)
  EXIT/FLAT_ALL: STOP_MARKET
  EMERGENCY_EXIT: MARKET (immediate)

All order types exercised: LIMIT, MARKET, STOP_MARKET, TRIGGER_MARKET, TRAILING_STOP
2026-07-18 12:25:41 +02:00
Codex
8322550fbc malkhut: REQUOTE as proper CANCEL_REPLACE primitive + action_menu metadata
REQUOTE is now distinct from QUOTE:
  QUOTE: PLACE new order (no existing to cancel)
  REQUOTE: CANCEL_REPLACE existing + place new (immediate)
  CHASE: PLACE with short TTL (auto-cancel retry)
  CANCEL: Remove existing order

action_menu generates REQUOTE with metadata={'requote': True} for existing orders.
DSL REQUOTE produces CANCEL_REPLACE when existing order, falls back to PLACE.

All 800+ tests pass.
2026-07-17 23:17:22 +02:00
Codex
4926ef6788 malkhut: CHASE mechanics FIXED + Flight9 learnings + TTL enforcement
1. CHASE mechanics (NOW WORKING):
   - CWM enforces TTL on open orders (auto-cancel when expired)
   - DSL CHASE produces PLACE with metadata={chase: True}
   - Action menu generates chase actions with wait_to_retry_ms TTL
   - OpenOrderState.gains ttl_ms field (0=no expiry, >0=auto-cancel)

2. TTL enforcement (CWM):
   - HftBacktestCWM: auto-cancels orders where age >= ttl_ms
   - MinimalCryptoLOBCWM: same TTL enforcement
   - This is how CHASE works: place→wait→auto-cancel→next step re-places

3. Flight9 learnings:
   - Slippage model gains trade_flow_intensity parameter
   - Book imbalance as proxy for trade arrival rate
   - Markout = quality concept documented

4. CHASE tests: 10 new tests covering TTL enforcement, cancel-retry cycle,
   max retries, DSL CHASE action, CMA codec integration

5. All 800+ tests pass
2026-07-17 19:30:17 +02:00
Codex
bb229833d3 malkhut: Flight9 learnings — markout=quality, queue×flow, depth-for-size
Fable's Flight9/BLUE generalizable features incorporated:

1. Slippage model gains trade_flow_intensity parameter:
   - Estimated from book imbalance (proxy for trade arrivals)
   - More flow → better fills (lower slippage)
   - Fable: 'fill = queue position × trade-flow intensity'

2. Markout = quality concept documented:
   - Score fills by post-fill markout, not just fill/no-fill
   - Maker fills are adversely selected

3. Depth-for-size documented:
   - Spread lies; key on depth-within-K-bps vs order notional

4. Measured fees:
   - BingX maker=2.00bp, taker=5.016bp (over 1,455 fills)
   - BingX commission = NEGATIVE (debit)

5. OB study updated with Flight9 learnings
2026-07-17 16:18:43 +02:00
Codex
8857daedfa malkhut: urgency-driven maker/taker + calibrated slippage + chase + docs 2026-07-17 10:16:55 +02:00
Codex
5c4ccdb1de malkhut: 3.5H instrumented E2E + calibrated slippage + conditional slippage 2026-07-15 19:32:46 +02:00
Codex
c03d914e7a malkhut: conditional slippage (Fable) 2026-07-15 17:05:16 +02:00
Codex
619966605e malkhut(docs): fill quality documentation — README, integration, OB study
README: Fill Quality section (core optimization target, metrics, reward
function, PerformanceMatrix, CMA-ES integration)

HftBacktestCWM integration doc: FillQuality dataclass, fill_value_score
computation, reward function weighting, PerformanceMatrix tracking

OB microstructure study: Section 13 — Fill Quality Optimization,
per-asset expectations, optimization strategy, connection to OB dynamics

Fill quality is MALKHUT's core aim: the system learns to get better fills
(faster, better-priced, less adverse selection) across regimes and venues.
2026-07-15 15:36:22 +02:00
Codex
618ad723e3 malkhut(wire): fill quality as PRIMARY optimization target
Fill quality is MALKHUT's core aim. Wired end-to-end:

1. FillQuality state (state.py):
   - slippage_bps, price_improvement_bps, levels_consumed
   - is_maker_fill, rolling_fill_rate, post_fill_adverse_bps
   - fill_value_score: composite metric for optimization
   - Added to MarketWorldState.fill_quality field

2. HftBacktestCWM.transition() (hft_cwm.py):
   - _compute_fill_quality() computes all metrics per transition
   - Fill quality now tracked for every CWM step
   - Empty book guards added for safety

3. MinimalCryptoLOBCWM.transition() (core.py):
   - Same fill quality computation for deterministic fallback
   - Empty book guards added

4. Reward function (hft_cwm.py):
   - fill_quality_reward = w_fill_probability * fill_value_score (PRIMARY)
   - Bonus for maker fills that improve price
   - Penalty for adverse selection after fill
   - Base reward (PnL, adverse selection, fees) preserved

5. PerformanceMatrix (selector.py):
   - RegimeStrategyScore: 4 new fill quality fields
   - record(): accepts fill_rate, slippage, price_improvement, fill_value_score
   - EMA updates for all fill quality metrics

6. EpisodeResult (cma_trainer.py):
   - avg_fill_value_score, avg_price_improvement_bps, avg_post_fill_adverse_bps
   - Accumulated per-step during _run_episode
   - Recorded to PerformanceMatrix in evaluate_candidate

All 1379+ tests green.
2026-07-15 15:22:25 +02:00
Codex
a310c92977 malkhut(e2e): CMA-ES disabled for pure 100-opponent episode loop 2026-07-15 00:48:48 +02:00
Codex
bbaceffb61 malkhut(e2e): CMA-ES every 20 cycles, crash-safe gc.collect() 2026-07-15 00:34:50 +02:00
Codex
67b3268b98 malkhut(e2e): memory-efficient + crash-safe 3h run
Rolling stats (no episode accumulation), CMA-ES every 10 cycles (3 evals),
gc.collect() after CMA, global try/except for crash safety.
100-opponent swarm, 9 assets, 270 scenarios.
2026-07-15 00:27:37 +02:00
Codex
0e215b1658 malkhut(e2e): memory-efficient long run — rolling stats, no episode accumulation
Fixed OOM kill by replacing all_episodes list accumulation with:
- Rolling stats (clear every 20 episodes)
- Only PnL history kept for characterization
- Peak/worst tracking without full episode storage
- Periodic stdout reports from rolling aggregates

100-opponent swarm + 9 assets × 30 scenarios = 270 scenarios per cycle.
CMA-ES every 5 cycles (3 evals). 3-hour target.
2026-07-15 00:19:59 +02:00
Codex
d72323a6c5 malkhut(e2e): 100-opponent swarm + empty book fix + error handling
- 100 diverse opponents (randomized params within each type)
- Risk gate: empty book guard in _post_only_would_cross
- CMA-ES: only every 5 cycles, 3 evals, robust error handling
- Main loop: try/except prevents silent crashes
- Profiling: 11.7 steps/sec with 100 opponents
2026-07-15 00:11:57 +02:00
Codex
adba9f3fe8 malkhut(e2e): full system exercise — HftBacktestCWM + 11-agent swarm + characterization
E2E exercise results:
  30 episodes × 30 steps = 900 actions
  CWM: HftBacktestCWM (PowerProbQueueModel)
  Swarm: 11 diverse opponents (2x ToxicTaker, 2x PassiveMaker, LatencyArb,
         Noise, Momentum, MeanReversion, InventoryMM, LiquidationFlow,
         StaleQuoteAttacker)

Performance:
  Avg PnL: +7217.9 bps, Win rate: 66.7% (20/30)
  Best: +26171.0 bps (chop scenario)
  Worst: -5.0 bps (thin/arb scenarios — flat, not loss)

Order types exercised:
  LIMIT 16.9%, MARKET 12.9%, STOP_MARKET 0.2%
  IOC 8%, GTC 92%
  Post-only: path exercised, Reduce-only: 41.5%
  Aggression/Passive ratio: 0.76

Speed: 457 actions/second (2.0s total)
2026-07-14 23:07:38 +02:00
Codex
8b385cb249 malkhut(cwm): HftBacktestCWM — queue model + 59-test suite
HftBacktestCWM (cwm/hft_cwm.py):
- PowerProbQueueModel: probabilistic fill per level (pre-computed)
- Level 0 always fills, deeper levels have decreasing probability
- Deterministic fallback when use_queue_model=False
- Same transition/reward/terminal API as MinimalCryptoLOBCWM
- Fallback to deterministic level consumption when hftbacktest unavailable

59 tests (test_hft_cwm.py) covering 15 test classes:
1. Queue model correctness (fill probs, monotonic, bounds, determinism)
2. Determinism & reproducibility
3. CWM interface compatibility (cross, place, cancel, post_only, reduce)
4. Reward function (profit, noop, maker bonus)
5. Edge cases (empty book, zero qty, extreme price, many levels)
6. Position tracking (buy, sell, flip)
7. Fee application (taker fee reduces equity)
8. Counterparty ecology (toxic taker hits book, noop preserves)
9. CWM comparison (Hft vs Minimal agree on noop)
10. Venue propagation (scenario tagging, cross-exchange transfer)
11. PerformanceMatrix venue keying (record, per-venue best, comparison)
12. Risk gate integration (approve, leverage, OOD, kill switch, self-trade)
13. Stress tests (rapid transitions, 20 open orders, cancel all)
14. Full episode integration (single episode runs, policy evaluator)
15. hftbacktest availability check
2026-07-14 19:46:34 +02:00
Codex
4aeadf1aae docs: hftbacktest CWM integration design — drop-in LOB backend
Architecture:
  hftbacktest replaces _fill_from_levels() + manual book updates
  Everything above transition() stays the same

What changes:
  - HftBacktestCWM: new class implementing CodeWorldModel protocol
  - transition(): submit/cancel via hftbacktest, convert state back
  - fill model: ProbQueueModel (queue-position-aware)
  - latency: interpolated from historical data

What does NOT change:
  - Planner (SM-MCTS, EXP3, Thompson, etc.)
  - Counterparty ecology (ToxicTaker, PassiveMaker, etc.)
  - Risk gate (all 6 checks)
  - CMA-ES trainer
  - PerformanceMatrix
  - Reward function (same PnL + adverse selection + risk)
  - Action menu, FulfilmentAction, ScenarioFactory
  - ALL existing tests

Integration: 3 steps, zero core changes:
  1. Create malkhut/cwm/hft_cwm.py
  2. Swap cwm_factory in PolicyEvaluator
  3. Done
2026-07-14 18:37:41 +02:00
Codex
943b3ef985 docs: order book microstructure study — 12 sections, 13 assets
Comprehensive OB study compiled from live Binance/BingX data + academic
literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal):

1. Depth power-law decay: D(d) = A * d^(1-alpha), per-asset params
2. Spread profiles: normal + stress multipliers for all 13 assets
3. Order flow: arrival rates, cancel/fill ratios, size distributions
4. Market maker behavior: inventory limits, pull speed, margins
5. Volatility regimes: GARCH params, half-lives, crisis multipliers
6. Intraday patterns: peak/trough hours, session analysis
7. Cross-asset correlations: normal vs crash behavior
8. BingX-specific: spread/depth/latency/fees vs Binance ratios
9. Book fragility & cascade dynamics: flash crash anatomy
10. Retail vs institutional composition
11. Funding rates: per-asset means, std, positive%
12. Expected slippage model
2026-07-14 18:17:15 +02:00
Codex
773df30609 malkhut(bench): CMA-ES re-run script at corrected fees
Corrected fees: taker=5.0, maker=+2.0 (BingX).
Previous best: 15,080 (pre-fee-fix, WRONG fees).
New best: 34,898 (+131.4%).

50 evals, 8.3 min, BTCUSDT. Mean PnL: 582 bps, Max DD: 83 bps.
2026-07-14 17:39:24 +02:00
Codex
f6d8d13146 malkhut(wire): 5 risk gate stubs implemented + 3 scenarios behavior-driven
Risk gate (risk/gate.py) — 5 stubs implemented:

1. _kill_switch_active(): operator-controlled emergency stop via set_kill_switch()
2. _cancel_rate_would_exceed(): tracks cancel timestamps per symbol in 60s
   sliding window, blocks if >= MAX_CANCELS_PER_SYMBOL_PER_MINUTE
3. _would_self_trade(): checks open orders for same symbol+side at same price
   (within tick_size), skipping the cancel_order_id for CANCEL_REPLACE
4. _would_exceed_symbol_notional(): sums current open order notional + new
   order notional, blocks if > equity * MAX_SYMBOL_NOTIONAL_FRACTION
5. _violates_venue_minima(): checks tick alignment, lot rounding, min_qty,
   and min_notional — all float-robust comparisons

ScenarioFactory — 3 remaining hardcoded scenarios converted:

1. _spread_tightening: spread_mult=0.3, depth_fraction=1.0 (was hardcoded BTC)
2. _cross_venue_arb: spread_mult=0.5, depth_fraction=0.5 (was hardcoded BTC)
3. _cross_exchange_arb_stress: spread_mult=0.8, depth_fraction=0.3 (was hardcoded BTC)

All 30 scenarios now use _behavior_state() — zero hardcoded prices remain.

675 tests pass. Zero regressions.
2026-07-14 17:01:38 +02:00
Codex
186bce8984 malkhut(test): exhaustive order type + venue integration test suite — 159 tests
11 test classes covering all three orthogonal dimensions:

1. OrderType enum (10 tests): values, uppercase, str, hashable, frozen
2. TimeInForce enum (7 tests): values, default, IOC/FOK/GTD
3. OrderInstruction enum (3 tests): values
4. Exchange mapping tables (15 tests): all exchanges, all types, POST_ONLY
   variation, trailing_stop BingX=BINANCE_MARKET
5. Normalization functions (9 tests): type, TIF, unknown exchange
6. is_type_available + get_supported_types (3 tests)
7. decompose_order (13 tests): all base, TIF, instructions, lowercase
8. FulfilmentAction (14 tests): frozen, time_in_force, post_only, reduce_only,
   lazy TIF import, cancel_replace, metadata
9. State OrderType backward compat (4 tests): values, str, set, comparison
10. Scenario venue tagging (7 tests): default, custom, frozen, replace
11. ScenarioFactory venue propagation (8 tests): exchange_id, all venues
12. Cross-exchange transfer (11 tests): transfer, count, symbol, tags, idempotent
13. PerformanceMatrix venue-keying (16 tests): record, get_best, per-venue,
    EMA, coverage, venue_comparison
14. CWM is_maker (8 tests): LIMIT, post_only, MARKET, STOP, trailing
15. Edge cases (12 tests): poison, zero scores, large scores, 100 strategies
16. Integration flow (6 tests): factory→transfer→matrix→selector
2026-07-14 16:03:30 +02:00
Codex
ebf772fe5a malkhut(docs): cross-exchange learning + adversary ecology documented
README updated with:
- Cross-exchange learning: ScenarioFactory exchange_id + cross_exchange_transfer
- PerformanceMatrix keyed by (regime, strategy_id, venue)
- Adversary ecology: ActionKind-level abstraction, venue-independent
- Transferability principle: parameters transfer, names are venue-specific
2026-07-14 15:44:33 +02:00
Codex
b70a6f0ad8 malkhut(wire): PerformanceMatrix keyed by (regime, strategy, venue)
Three-dimensional key enables:
  - Per-venue best: get_best(regime, venue='bingx')
  - Cross-venue comparison: get_venue_comparison(regime, strategy_id)
  - Venue-agnostic: get_best(regime) scans all venues (backward compat)

New API:
  - record(..., venue='bingx'): venue parameter (default 'bingx')
  - get_best(regime, venue=None): optional venue filter
  - get_scores_for_regime(regime, venue=None): optional venue filter
  - get_venue_comparison(regime, strategy_id) -> {venue: score}

119 tests pass. All existing callers backward compatible.
2026-07-14 15:37:36 +02:00
Codex
2cba60a154 malkhut(wire): venue passed through matrix recording for cross-exchange comparison
- evaluator: passes scenario.venue to matrix.record(venue=...)
- PerformanceMatrix.record(): accepts venue parameter (default='bingx')
- Enables cross-exchange learnings: same strategy tested on BingX vs Binance
  gets separate performance entries per venue

Adversary ecology analysis:
Counterparties operate at ActionKind level (CROSS_SPREAD/PLACE/CANCEL),
not at order-type level. The CWM infers order type from ActionKind:
  CROSS_SPREAD → fills aggressively → equivalent to MARKET
  PLACE → passive quote → equivalent to LIMIT
This is correct and venue-independent. Fee calculation already uses
VenueRules (per-exchange fees). No adversary changes needed.
2026-07-14 15:26:30 +02:00
Codex
401d5a70ca malkhut(wire): venue tagging + cross-exchange transfer + CWM order type fix
ScenarioFactory + CWM + Engine changes:

1. Scenario.venue field (default='bingx') — each scenario tagged with venue
2. ScenarioFactory.exchange_id parameter — controls which exchange scenarios simulate
3. _make_state + _behavior_state: venue propagated to VenueRules.exchange
4. All 34 scenario builders: venue=self.exchange_id
5. cross_exchange_transfer(): re-tag scenarios for different exchange
   (strategy evolved on BingX can be re-evaluated on Binance)
6. CWM core.py: is_maker check updated for three-dimensional order model
   (POST_ONLY no longer in OrderType; uses post_only flag instead)

Cross-exchange learning flow:
  factory_bingx = ScenarioFactory(exchange_id='bingx')
  scenarios_bingx = factory_bingx.build_suite(symbols=[...])
  strategy = train(scenarios_bingx)  # evolve on BingX

  factory_binance = ScenarioFactory(exchange_id='binance')
  scenarios_binance = factory_bingx.cross_exchange_transfer(
      scenarios_bingx, target_exchange='binance')
  score = evaluate(strategy, scenarios_binance)  # test on Binance

All tests pass. Strategy PARAMETERS transfer; only venue tag + fees + order mapping change.
2026-07-14 15:18:56 +02:00
Codex
455a7a5a4e malkhut(docs): README updated for three-dimensional order type model
Updated README to reflect Fable's corrections:
- OrderType/TimeInForce/Instructions as three orthogonal dimensions
- POST_ONLY/IOC/FOK correctly described as non-types
- BingX trailing_stop -> TRAILING_STOP_MARKET
- Three mapping tables (order type, TIF, instructions)
- Integration status updated
2026-07-14 14:52:50 +02:00
Codex
d24d9bc6bd malkhut(wire): OrderType as three orthogonal dimensions — Fable's corrections
CRITICAL REFACTOR based on Fable's review (S9 roadmap item):

Before: flat enum conflating order types with TIF/instructions
  OrderType had MARKET, LIMIT, IOC, FOK, POST_ONLY, REDUCE_ONLY, etc.

After: three orthogonal dimensions (FIX-aligned):
  1. OrderType (Tag 40): what the order IS
     LIMIT, MARKET, STOP_MARKET, STOP_LIMIT, TRIGGER_MARKET, TRIGGER_LIMIT,
     TRAILING_STOP, OCO, TP_SL
  2. TimeInForce (Tag 59): how long it LIVES
     GTC, IOC, FOK, GTD
  3. Instructions (Tag 18): behavioral modifiers
     POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG

Key corrections:
- POST_ONLY is an instruction on a LIMIT order, not a standalone type
- IOC/FOK are TimeInForce values, not order types
- BingX trailing_stop -> native TRAILING_STOP_MARKET (not TRIGGER_MARKET)
- FulfilmentAction.time_in_force: new field, default GTC

Exchange mappings restructured:
  EXCHANGE_ORDER_TYPE_MAP: OrderType -> exchange native 'type' param
  EXCHANGE_TIF_MAP: TimeInForce -> exchange native 'timeInForce' param
  EXCHANGE_INSTRUCTION_MAP: Instruction -> exchange encoding

21 files changed. 380+ tests pass. Backward compatible.
2026-07-14 14:46:44 +02:00
Codex
a21f64e066 malkhut(wire): BingX adapter uses standardized OrderType mapping
adapter.py now uses normalize_to_exchange(action.order_type, 'bingx')
to translate normalized order types to BingX-native strings.
Falls back to LIMIT/MARKET/POST_ONLY for backward compatibility.

This is the critical integration point: standardized order types flow
from FulfilmentAction → CWM → VenueAdapter → exchange API.
2026-07-14 13:16:52 +02:00
Codex
f48af7c405 malkhut(docs): OrderType integration status documented
README updated with:
- Integration Status section for OrderType
- state.py: 17 values, backward compatible
- order_types.py: standalone standardized taxonomy
- ExchangeProfile.available_order_types per venue
- CWM transition: passes action.order_type to venue adapter
- Agent/adversary: check is_type_available before placing
2026-07-14 13:11:19 +02:00
Codex
3324933613 malkhut(wire): OrderType unified — 5-layer taxonomy, backward compatible
state.py OrderType replaced with 5-layer taxonomy (FIX/CCXT aligned):
  Layer 1: MARKET, LIMIT (FIX Tag 40)
  Layer 2: GTC, IOC, FOK, GTD (FIX Tag 59)
  Layer 3: STOP_MARKET, STOP_LIMIT, TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
  Layer 4: POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG (FIX Tag 18)
  Layer 5: OCO, TP_SL (exchange-specific)

action_menu.py: REDUCE_ONLY_MARKET → MARKET (reduce_only field handles it)

All 1126 tests pass. Fully wired and backward compatible.
2026-07-14 13:04:54 +02:00
Codex
369d9b41ad malkhut: ExchangeProfile gains available_order_types per venue
Each exchange now declares which normalized order types it supports:
- binance: limit, market, stop_market, stop_limit, post_only, ioc, fok, trailing_stop, reduce_only
- bingx: limit, market, stop_market, stop_limit, post_only, ioc, fok, trailing_stop, reduce_only
- bybit: limit, market, stop_market, stop_limit, post_only, ioc, fok, trailing_stop, reduce_only

Backward compatible: new field has default=('limit', 'market').
Enables: agents/adversaries check is_type_available() before placing orders.
2026-07-14 12:14:54 +02:00
Codex
53e02c84ec malkhut: standardized order types — FIX/CCXT-aligned, multi-exchange mapping
order_types.py: Five-layer taxonomy normalized to industry standards:
  Layer 1: Base types (FIX Tag 40) — MARKET, LIMIT
  Layer 2: Time-in-force (FIX Tag 59) — GTC, IOC, FOK, GTD
  Layer 3: Conditional/Trigger (FIX Tag 3/4+MIT) — STOP_MARKET, STOP_LIMIT,
    TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
  Layer 4: Instructions (FIX Tag 18) — POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
  Layer 5: Compound (exchange-specific) — OCO, TP_SL

Cross-exchange mapping: BingX ↔ Binance ↔ Bybit (from CCXT source code).
Standards: FIX 4.4 Tag 40/59/18, CCXT unified API, ISO 10383 (MIC).

Transferability: strategy PARAMETERS transfer. ORDER TYPE NAMES are
venue-specific but semantics identical (LIMIT = LIMIT everywhere).

14 tests. README updated with full mapping table and standards references.
2026-07-14 12:02:01 +02:00
Codex
13811cc789 malkhut(docs): comprehensive update — all 10 Fable spec items documented
README updated with:
- Fable spec items table (10 items, status)
- New subsystems: ScenarioLibrary, ManifoldQuery, ActualsLoader, BookFidelity, DAAT, OOD
- Package structure: daat/ directory added
- All modules documented with test counts

25 commits total. 1204 tests. All green.
2026-07-14 09:14:31 +02:00
Codex
7ad123c4c1 malkhut(spec): items 5-10 — manifold, actuals, OOD, query, book fidelity
Item 5 — PerformanceMatrix manifold:
  RegimeStrategyScore: added confidence, support_count, distance_to_nearest
  record() populates confidence from episode count (more evidence = more confidence)

Item 6 — ActualsLoader:
  ActualsSnapshot: 12-field frozen dataclass for live market data
  ActualsLoader: reads CH tables (obf_universe, exf_data, maras_fingerprint, etc.)
  Synthetic fallback when CH unavailable

Item 7 — OOD verdict in RiskGate:
  validate() now accepts daat_verdict parameter
  OUT_OF_DISTRIBUTION → veto action, fall back to doctrinal simple policy
  Backward compatible: default daat_verdict='KNOWN'

Item 8 — Manifold query (three-phase recommendation):
  1. DAAT classify live state (KNOWN/MARGINAL/OOD)
  2. If KNOWN: find nearest regime in PerformanceMatrix → best strategy
  3. If OOD: return doctrinal_simple fallback
  ManifoldRecommendation: strategy_id, confidence, regime, verdict, reason

Item 10 — Book fidelity gap:
  BookFidelityConfig: n_levels, aggregation_window, min_depth
  synthesize_book_from_params: power-law D(d)=amplitude*d^(1-alpha) → OrderBookState
  Bridges OBF 15B rows → MALKHUT finite Tuple[PriceLevel]

5 files, 282 insertions.
2026-07-14 06:11:37 +02:00
Codex
1f41be845b malkhut(spec): item 4 — ScenarioLibrary sweep for Mode 1 coverage
ScenarioLibrary sweeps the state space (not samples) across:
  - spread_mult: [0.1, 0.5, 1.0, 2.0, 5.0, 10.0]
  - depth_fraction: [0.01, 0.05, 0.1, 0.3, 0.5, 1.0]
  - toxicity: [0.0, 0.3, 0.7, 1.0]
  - regime: [normal, crisis, recovery, transition]

Default: 13 assets × 576 grid points = 7,488 scenarios.
Customizable: specify symbols, dimensions, ranges.

7 tests covering: grid size, sweep output, point fields,
regime coverage, custom dimensions, summary, factory function.
2026-07-14 05:34:05 +02:00
Codex
833f262d12 malkhut(spec): item 9 — DAAT package (Direction-Anchored Ambiguity Triage)
DaatQuery: 8-feature market state representation
DaatVerdict: KNOWN / MARGINAL / OUT_OF_DISTRIBUTION
daat_classify: cosine RETRIEVE → magnitude GATE → local MODEL
- Cosine finds nearest explored state (directional match)
- Magnitude gate detects out-of-distribution states
- Empty explored set → always OUT_OF_DISTRIBUTION

9 tests covering: known state, OOD, empty explored, marginal, result fields.
No Unicode in code. All tests pass.
2026-07-14 02:33:31 +02:00
Codex
eef890a5cc malkhut(spec): item 1 mutation-litmus + item 3 maker-fee UNVERIFIED comment
Item 1 — Mutation-litmus test (spec §1 item 3):
- test_taker_fee_10x_changes_score: fee change MUST affect score
- test_zero_fees_vs_correct_fees: zero vs 5bps must differ
- BOTH PASS — confirms fees ARE wired into reward function
- If fees were ignored, these tests would go RED

Item 3 — Maker fee verification (spec §1 item 5):
- Added '# UNVERIFIED — no maker fills on record as of 2026-07-13'
  to Binance and Bybit exchange profiles
- Maker fee sign (positive on BingX, negative rebate on others)
  is correct after fee fix but unverified from actual fills.

Items 2,4-10 remain for implementation.
2026-07-13 23:24:05 +02:00
Codex
84a94a7098 malkhut(docs): fee correction + Fable spec status + 5000-eval results
README updated with:
- Fee correction table (10x bug fixed, source of truth documented)
- All prior policies flagged as suspect at correct fees
- 5000-eval partial results: best=15,080 at gen 105, still climbing
- Fable's spec (SPEC_MALKHUT_ACTUALS_INTAKE.md) acknowledged:
  - Fee bug fixed (commit 5523be1d)
  - Two-mode architecture (EXPLORE + RECOMMEND) understood
  - ANNEX A and DAAT understood
  - Ecology stays (actuals calibrate, ecology plays)
  - Outstanding items logged for future sessions

Total: 20 commits, 1186 tests, all green.
2026-07-13 21:50:57 +02:00
Codex
5523be1d44 malkhut(fix): CORRECT FEE BUG — taker 0.5→5.0, maker -0.2→+2.0
Fable's spec (SPEC_MALKHUT_ACTUALS_INTAKE.md) confirmed 10x fee error
from our own fills (dolphin.trade_execution_quality).

Fixed:
- BingX taker: 0.5 → 5.0 bps
- BingX maker: -0.2 → +2.0 bps (POSITIVE on BingX, not a rebate)
- Binance taker: 0.4 → 4.5 bps
- Bybit taker: 0.06 → 5.5 bps
- All 13 per-asset profiles: maker=-0.2 taker=0.5 → maker=2.0 taker=5.0

Source of truth: dolphin.trade_execution_quality (8006 rows, avg taker=5.016 bps).
Every policy trained before this fix was at 10x too-cheap fees.
Re-measurement at correct fees is required.
2026-07-13 20:26:53 +02:00
Codex
70964394d4 malkhut(docs + bench): comprehensive update + smoke test script
README updated with:
- Vectorized UCB selection (7.7x speedup, 1.13µs/selection)
- Batch MCTS kernel (numba-accelerated)
- Fast scalar + advantage scoring modes
- Updated performance benchmarks (1186 tests, 390 scenarios, 3043 score/min)
- Advantage scorer module in package structure

smoke_1h.py: standalone training script for extended runs.

Total session: 19 commits, 1186 tests, all green.
All implementations: parallel eval (7x), vectorized reward (numba),
vectorized UCB (7.7x), fast scalar scoring, advantage mode,
DuckDB store (sub-µs reads), asset compiler, behavior DSL,
multi-exchange support, three-layer identifiers.
2026-07-13 17:04:40 +02:00
Codex
22ae8b8aea malkhut(perf): vectorized UCB selection via numba + batch MCTS kernel
numba_core.py:
  - ucb_select_vectorized: numba-JIT UCB selection replacing Python for-loop
    Uses flat numpy arrays, deterministic tie-breaking, no Python overhead
  - mcts_simulate_batch: batched MCTS across N worlds (lightweight proxy)

sm_mcts.py:
  - PlayerActionStats.ucb_select: wired to numba ucb_select_vectorized
  - Passes rng seed as int (not RandomState) for numba compatibility

Impact: UCB selection moves from Python loop to numba JIT. Each selection
is ~100ns instead of ~1µs. With 16 sims × 20 steps × 90 episodes, this
saves ~14ms per eval.
2026-07-13 16:44:24 +02:00
Codex
d9b7e05531 malkhut(perf): optimize _run_episode — reduced Python overhead
Optimizations in _run_episode:
- Pre-allocated ActionKind constants (avoid repeated attribute lookups)
- Removed unnecessary max_pos_qty tracking (unused in scoring)
- Simplified action kind checks (single comparison chain)
- Reduced frozen dataclass allocations per step

Result: same behavioral output, cleaner code path.
Episode time: ~19ms/step sequential, ~13ms/step parallel (unchanged —
bottleneck is MCTS planner + CWM, not Python orchestration).
2026-07-13 15:11:19 +02:00