Commit Graph

9 Commits

Author SHA1 Message Date
Codex
97a770da65 malkhut: online EWMA self-calibrating slippage model
Flight7 model underestimates by 80% in CWM dynamic book:
  raw predicted: 0.034 bps, actual: 0.180 bps
  Constant error across 22K episodes — no feedback loop.

Root cause: Flight7 calibrated on real BingX taker fills, but CWM's
synthetic dynamic book has different fill characteristics.

Fix: SlippageSelfCalibrator with EWMA feedback loop.
  After each fill: error = actual - predicted (clipped to +/-20 bps)
  EWMA smooths per-symbol errors (alpha=0.2)
  Next prediction = raw_model + EWMA_correction
  Bounded output: 0-50 bps absolute

Convergence (300 eps across 8 assets):
  ETH: 9% error (from 80%)
  SOL: 3.5%
  DOGE: 6.7%
  LINK: 5.7%
  ADA: 9.7%
  BTC: 48.6% (low fill count, converging)
  AVAX: 28.6% (low fill count)
  UNI: 52.5% (low fill count, early outlier)

Truthfulness guarantees:
  - Correction is observable (CALIBRATOR.correction(symbol))
  - Resets between runs (no hidden state)
  - Only uses observed fills, no assumptions
  - Error clipping prevents outlier domination
  - Absolute bounds prevent runaway
2026-07-20 15:17:16 +02:00
Codex
4926ef6788 malkhut: CHASE mechanics FIXED + Flight9 learnings + TTL enforcement
1. CHASE mechanics (NOW WORKING):
   - CWM enforces TTL on open orders (auto-cancel when expired)
   - DSL CHASE produces PLACE with metadata={chase: True}
   - Action menu generates chase actions with wait_to_retry_ms TTL
   - OpenOrderState.gains ttl_ms field (0=no expiry, >0=auto-cancel)

2. TTL enforcement (CWM):
   - HftBacktestCWM: auto-cancels orders where age >= ttl_ms
   - MinimalCryptoLOBCWM: same TTL enforcement
   - This is how CHASE works: place→wait→auto-cancel→next step re-places

3. Flight9 learnings:
   - Slippage model gains trade_flow_intensity parameter
   - Book imbalance as proxy for trade arrival rate
   - Markout = quality concept documented

4. CHASE tests: 10 new tests covering TTL enforcement, cancel-retry cycle,
   max retries, DSL CHASE action, CMA codec integration

5. All 800+ tests pass
2026-07-17 19:30:17 +02:00
Codex
bb229833d3 malkhut: Flight9 learnings — markout=quality, queue×flow, depth-for-size
Fable's Flight9/BLUE generalizable features incorporated:

1. Slippage model gains trade_flow_intensity parameter:
   - Estimated from book imbalance (proxy for trade arrivals)
   - More flow → better fills (lower slippage)
   - Fable: 'fill = queue position × trade-flow intensity'

2. Markout = quality concept documented:
   - Score fills by post-fill markout, not just fill/no-fill
   - Maker fills are adversely selected

3. Depth-for-size documented:
   - Spread lies; key on depth-within-K-bps vs order notional

4. Measured fees:
   - BingX maker=2.00bp, taker=5.016bp (over 1,455 fills)
   - BingX commission = NEGATIVE (debit)

5. OB study updated with Flight9 learnings
2026-07-17 16:18:43 +02:00
Codex
5c4ccdb1de malkhut: 3.5H instrumented E2E + calibrated slippage + conditional slippage 2026-07-15 19:32:46 +02:00
Codex
c03d914e7a malkhut: conditional slippage (Fable) 2026-07-15 17:05:16 +02:00
Codex
618ad723e3 malkhut(wire): fill quality as PRIMARY optimization target
Fill quality is MALKHUT's core aim. Wired end-to-end:

1. FillQuality state (state.py):
   - slippage_bps, price_improvement_bps, levels_consumed
   - is_maker_fill, rolling_fill_rate, post_fill_adverse_bps
   - fill_value_score: composite metric for optimization
   - Added to MarketWorldState.fill_quality field

2. HftBacktestCWM.transition() (hft_cwm.py):
   - _compute_fill_quality() computes all metrics per transition
   - Fill quality now tracked for every CWM step
   - Empty book guards added for safety

3. MinimalCryptoLOBCWM.transition() (core.py):
   - Same fill quality computation for deterministic fallback
   - Empty book guards added

4. Reward function (hft_cwm.py):
   - fill_quality_reward = w_fill_probability * fill_value_score (PRIMARY)
   - Bonus for maker fills that improve price
   - Penalty for adverse selection after fill
   - Base reward (PnL, adverse selection, fees) preserved

5. PerformanceMatrix (selector.py):
   - RegimeStrategyScore: 4 new fill quality fields
   - record(): accepts fill_rate, slippage, price_improvement, fill_value_score
   - EMA updates for all fill quality metrics

6. EpisodeResult (cma_trainer.py):
   - avg_fill_value_score, avg_price_improvement_bps, avg_post_fill_adverse_bps
   - Accumulated per-step during _run_episode
   - Recorded to PerformanceMatrix in evaluate_candidate

All 1379+ tests green.
2026-07-15 15:22:25 +02:00
Codex
401d5a70ca malkhut(wire): venue tagging + cross-exchange transfer + CWM order type fix
ScenarioFactory + CWM + Engine changes:

1. Scenario.venue field (default='bingx') — each scenario tagged with venue
2. ScenarioFactory.exchange_id parameter — controls which exchange scenarios simulate
3. _make_state + _behavior_state: venue propagated to VenueRules.exchange
4. All 34 scenario builders: venue=self.exchange_id
5. cross_exchange_transfer(): re-tag scenarios for different exchange
   (strategy evolved on BingX can be re-evaluated on Binance)
6. CWM core.py: is_maker check updated for three-dimensional order model
   (POST_ONLY no longer in OrderType; uses post_only flag instead)

Cross-exchange learning flow:
  factory_bingx = ScenarioFactory(exchange_id='bingx')
  scenarios_bingx = factory_bingx.build_suite(symbols=[...])
  strategy = train(scenarios_bingx)  # evolve on BingX

  factory_binance = ScenarioFactory(exchange_id='binance')
  scenarios_binance = factory_bingx.cross_exchange_transfer(
      scenarios_bingx, target_exchange='binance')
  score = evaluate(strategy, scenarios_binance)  # test on Binance

All tests pass. Strategy PARAMETERS transfer; only venue tag + fees + order mapping change.
2026-07-14 15:18:56 +02:00
Codex
c6b7a41bb4 malkhut(optim): vectorized reward + Ray parallel eval + VBT post-analysis
1. Vectorized reward path (cwm/core.py):
   - Wired up existing compute_reward_vectorized from numba_core (was unused!)
   - Eliminates FeatureVector dict allocation + Python dict lookups on hot path
   - Numba path used when _HAS_NUMBA=True, Python fallback otherwise
   - Bit-identical: same math operations, just via numba JIT

2. Ray-based parallel eval (training/ray_eval.py):
   - Industrial multi-core execution via Ray (used by OpenAI/Anyscale)
   - ray.put() stores params/scenarios in shared object store (no pickle per worker)
   - Each worker: own CWM + planner, zero shared state, no races
   - Bit-identical: same seed + same params = same results regardless of worker count
   - PolicyEvaluator.evaluate_candidate: new use_ray=True parameter

3. VBT post-analysis (training/vbt_analysis.py):
   - episodes_to_pnl_array, episodes_to_metrics (Sharpe, Sortino, VaR, win_rate, etc.)
   - cross_asset_comparison, parameter_sensitivity
   - format_metrics for human-readable output
   - Analysis tool only — runs AFTER engine produces results

4. numba_core.py: added missing 'import math' for compute_reward_vectorized

13 new tests: vectorized reward bit-identity, Ray determinism, Ray result fields,
VBT metrics structure, cross-asset comparison, parameter sensitivity, edge cases.
Total: 1178 tests, 50 files, all green, zero regressions.
2026-07-12 23:56:16 +02:00
Codex
f943191d56 malkhut(T2): Code World Model — deterministic exchange simulator
CWM core (core.py): price-time priority, sequential level consumption,
partial fills, queue position, latency injection, maker/taker fees.
Numba acceleration (numba_core.py): JIT hot loops, 1.8x fill speedup.
Replay verification (replay_verify.py): binary search, trajectory recording.
Supporting: adverse_selection, correlation, latency_model, multi_level,
queue_model, spread_dynamics, volatility, hftbacktest_validator.
2026-07-11 10:23:44 +02:00