Commit Graph

5 Commits

Author SHA1 Message Date
Codex
618ad723e3 malkhut(wire): fill quality as PRIMARY optimization target
Fill quality is MALKHUT's core aim. Wired end-to-end:

1. FillQuality state (state.py):
   - slippage_bps, price_improvement_bps, levels_consumed
   - is_maker_fill, rolling_fill_rate, post_fill_adverse_bps
   - fill_value_score: composite metric for optimization
   - Added to MarketWorldState.fill_quality field

2. HftBacktestCWM.transition() (hft_cwm.py):
   - _compute_fill_quality() computes all metrics per transition
   - Fill quality now tracked for every CWM step
   - Empty book guards added for safety

3. MinimalCryptoLOBCWM.transition() (core.py):
   - Same fill quality computation for deterministic fallback
   - Empty book guards added

4. Reward function (hft_cwm.py):
   - fill_quality_reward = w_fill_probability * fill_value_score (PRIMARY)
   - Bonus for maker fills that improve price
   - Penalty for adverse selection after fill
   - Base reward (PnL, adverse selection, fees) preserved

5. PerformanceMatrix (selector.py):
   - RegimeStrategyScore: 4 new fill quality fields
   - record(): accepts fill_rate, slippage, price_improvement, fill_value_score
   - EMA updates for all fill quality metrics

6. EpisodeResult (cma_trainer.py):
   - avg_fill_value_score, avg_price_improvement_bps, avg_post_fill_adverse_bps
   - Accumulated per-step during _run_episode
   - Recorded to PerformanceMatrix in evaluate_candidate

All 1379+ tests green.
2026-07-15 15:22:25 +02:00
Codex
b70a6f0ad8 malkhut(wire): PerformanceMatrix keyed by (regime, strategy, venue)
Three-dimensional key enables:
  - Per-venue best: get_best(regime, venue='bingx')
  - Cross-venue comparison: get_venue_comparison(regime, strategy_id)
  - Venue-agnostic: get_best(regime) scans all venues (backward compat)

New API:
  - record(..., venue='bingx'): venue parameter (default 'bingx')
  - get_best(regime, venue=None): optional venue filter
  - get_scores_for_regime(regime, venue=None): optional venue filter
  - get_venue_comparison(regime, strategy_id) -> {venue: score}

119 tests pass. All existing callers backward compatible.
2026-07-14 15:37:36 +02:00
Codex
2cba60a154 malkhut(wire): venue passed through matrix recording for cross-exchange comparison
- evaluator: passes scenario.venue to matrix.record(venue=...)
- PerformanceMatrix.record(): accepts venue parameter (default='bingx')
- Enables cross-exchange learnings: same strategy tested on BingX vs Binance
  gets separate performance entries per venue

Adversary ecology analysis:
Counterparties operate at ActionKind level (CROSS_SPREAD/PLACE/CANCEL),
not at order-type level. The CWM infers order type from ActionKind:
  CROSS_SPREAD → fills aggressively → equivalent to MARKET
  PLACE → passive quote → equivalent to LIMIT
This is correct and venue-independent. Fee calculation already uses
VenueRules (per-exchange fees). No adversary changes needed.
2026-07-14 15:26:30 +02:00
Codex
7ad123c4c1 malkhut(spec): items 5-10 — manifold, actuals, OOD, query, book fidelity
Item 5 — PerformanceMatrix manifold:
  RegimeStrategyScore: added confidence, support_count, distance_to_nearest
  record() populates confidence from episode count (more evidence = more confidence)

Item 6 — ActualsLoader:
  ActualsSnapshot: 12-field frozen dataclass for live market data
  ActualsLoader: reads CH tables (obf_universe, exf_data, maras_fingerprint, etc.)
  Synthetic fallback when CH unavailable

Item 7 — OOD verdict in RiskGate:
  validate() now accepts daat_verdict parameter
  OUT_OF_DISTRIBUTION → veto action, fall back to doctrinal simple policy
  Backward compatible: default daat_verdict='KNOWN'

Item 8 — Manifold query (three-phase recommendation):
  1. DAAT classify live state (KNOWN/MARGINAL/OOD)
  2. If KNOWN: find nearest regime in PerformanceMatrix → best strategy
  3. If OOD: return doctrinal_simple fallback
  ManifoldRecommendation: strategy_id, confidence, regime, verdict, reason

Item 10 — Book fidelity gap:
  BookFidelityConfig: n_levels, aggregation_window, min_depth
  synthesize_book_from_params: power-law D(d)=amplitude*d^(1-alpha) → OrderBookState
  Bridges OBF 15B rows → MALKHUT finite Tuple[PriceLevel]

5 files, 282 insertions.
2026-07-14 06:11:37 +02:00
Codex
ef2f8e8827 malkhut(T6): training core — CMA-ES trainer, registry, pipeline, selector
CMA-ES trainer (cma_trainer.py): self-play pool, bootstrap CI, ScenarioFactory
with behavior-driven scenarios, auto-compile, label query interfaces.
Policy registry (registry.py): CANDIDATE → ACTIVE lifecycle.
Training pipeline (pipeline.py): bounded continuous learning loop + logger.
Strategy selector (selector.py): regime → strategy mapping, performance matrix.
2026-07-11 10:33:56 +02:00