diff --git a/MALKHUT/README.md b/MALKHUT/README.md index 880a40a..a414508 100644 --- a/MALKHUT/README.md +++ b/MALKHUT/README.md @@ -489,6 +489,60 @@ Re-measurement at correct fees is required for production deployment. | **Sync/Async Seams** | `test_sync_async_seams.py` | 7 | Zinc latency, engine budget | | **E2E Integration** | `test_e2e_integration.py` | 2 | Full pipeline: train→register→plan→risk→venue→zinc→ch→reload | +| **HftBacktestCWM** | `cwm/hft_cwm.py` | (new) | Queue-model fills (PowerProb), fill quality tracking, drop-in CWM | + +### Fill Quality — The Core Optimization Target + +MALKHUT is an **execution improvement engine**. Fill quality IS the primary aim — not PnL, not Sharpe, not win rate. The system learns to get **better fills**: faster, better-priced, with less adverse selection. + +#### Fill Quality Metrics (per CWM transition) + +| Metric | Source | What it measures | +|--------|--------|------------------| +| `slippage_bps` | `abs(fill_price - mid) / mid * 10000` | How far from mid did we fill? (aggressive) | +| `price_improvement_bps` | `(best_bid - fill_price) / best_bid * 10000` | How much better than touch? (passive) | +| `levels_consumed` | `fill_qty / avg_level_qty` | Queue depth of fill | +| `is_maker_fill` | `order_type==LIMIT or post_only` | Passive vs aggressive | +| `rolling_fill_rate` | EMA(0.8, 0.2) over recent fills | Recent fill success rate | +| `post_fill_adverse_bps` | `(new_mid - old_mid) / old_mid * 10000` | Price movement after fill | +| `fill_value_score` | `quality - abs(adverse) * 0.5` | **Composite optimization metric** | + +#### Fill Quality in the Reward Function + +``` +reward = w_fill_probability * fill_value_score ← PRIMARY (fill quality) + + w_expected_pnl * pnl ← secondary (PnL) + - w_adverse_selection * toxicity + - w_inventory_risk * inventory_risk + - w_tail_loss * tail_risk + - w_time_decay * time_in_loss + + w_fee_quality * maker_fee_benefit + - spread_cost - taker_fee +``` + +The `fill_value_score` = price_quality - adverse_selection. For maker fills: +`price_quality = price_improvement_bps` (how much better than best bid/ask). +For taker fills: `price_quality = spread_bps - slippage_bps` (how efficiently we crossed). + +#### Fill Quality in the PerformanceMatrix + +`RegimeStrategyScore` now tracks 4 fill quality metrics per (regime, strategy, venue): +- `avg_fill_rate`: rolling fill success rate +- `avg_slippage_bps`: average slippage for aggressive fills +- `avg_price_improvement_bps`: average improvement for passive fills +- `avg_fill_value_score`: composite fill quality metric + +This enables: "Which strategy achieves the best fill quality in regime X on venue Y?" + +#### Fill Quality in CMA-ES + +The CMA-ES optimizer now receives fill quality metrics in each `EpisodeResult`: +- `avg_fill_value_score`: average fill value across the episode +- `avg_price_improvement_bps`: average price improvement +- `avg_post_fill_adverse_bps`: average adverse selection + +The CMA-ES objective is: maximize fill quality (primary) while maintaining positive PnL. + ### Performance Benchmarks | Metric | Value | diff --git a/MALKHUT/docs/HFTBACKTEST_CWM_INTEGRATION.md b/MALKHUT/docs/HFTBACKTEST_CWM_INTEGRATION.md index 032c5d4..8f3f421 100644 --- a/MALKHUT/docs/HFTBACKTEST_CWM_INTEGRATION.md +++ b/MALKHUT/docs/HFTBACKTEST_CWM_INTEGRATION.md @@ -357,3 +357,40 @@ Swapping the implementation behind that protocol is a textbook Strategy pattern. The planner doesn't know or care whether the book is synthesized or hftbacktest. The reward function is pure math on (prev_state, action, next_state) — identical regardless of how next_state was computed. + +## Fill Quality Tracking (CORE Optimization Target) + +Every CWM transition now computes `FillQuality` metrics on the resulting state: + +```python +@dataclass(frozen=True, slots=True) +class FillQuality: + filled: bool # Did this action produce a fill? + fill_qty: float # How much was filled? + fill_price: float # At what price? + slippage_bps: float # Aggressive: distance from mid + price_improvement_bps: float # Passive: improvement over touch + levels_consumed: int # Queue depth consumed + is_maker_fill: bool # Passive (LIMIT) vs aggressive (CROSS) + rolling_fill_rate: float # EMA of recent fill success + post_fill_adverse_bps: float # Price movement after fill + fill_value_score: float # Composite: quality - adverse +``` + +The `fill_value_score` is the PRIMARY optimization target: +- For maker fills: `price_improvement_bps - abs(post_fill_adverse) * 0.5` +- For taker fills: `(spread_bps - slippage_bps) - abs(post_fill_adverse) * 0.5` + +Both `MinimalCryptoLOBCWM` and `HftBacktestCWM` compute these identically. + +The reward function weights fill quality via `w_fill_probability` (default 0.5): +``` +reward = w_fill_probability * fill_value_score ← PRIMARY + + w_expected_pnl * pnl ← secondary + - w_adverse_selection * toxicity + ... +``` + +PerformanceMatrix stores `avg_fill_rate`, `avg_slippage_bps`, `avg_price_improvement_bps`, +`avg_fill_value_score` per (regime, strategy, venue) — enabling: +"Which strategy achieves the best fill quality in regime X on venue Y?" diff --git a/MALKHUT/docs/OB_MICROSTRUCTURE_STUDY.md b/MALKHUT/docs/OB_MICROSTRUCTURE_STUDY.md index c0a7571..cbfd4a5 100644 --- a/MALKHUT/docs/OB_MICROSTRUCTURE_STUDY.md +++ b/MALKHUT/docs/OB_MICROSTRUCTURE_STUDY.md @@ -343,3 +343,52 @@ Example for DOGE ($22K amplitude, alpha=1.00): BTC allows $100K orders with <1 bps slippage. DOGE requires $1K orders for the same. Position sizing must account for the book's capacity, not just the strategy's signal. + +## 13. Fill Quality Optimization (MALKHUT Core) + +MALKHUT is an execution improvement engine. Fill quality IS the primary aim. + +### Fill Quality Metrics + +For each CWM transition, MALKHUT computes: + +- **slippage_bps**: aggressive fills — how far from mid? +- **price_improvement_bps**: passive fills — how much better than touch? +- **levels_consumed**: queue depth of fill +- **post_fill_adverse_bps**: price movement after fill (negative = adverse) +- **fill_value_score**: composite = quality - adverse * 0.5 + +### How This Connects to the OB + +The OB microstructure directly determines fill quality: + +| OB Characteristic | Impact on Fill Quality | +|-------------------|----------------------| +| **Depth at touch** | More depth = more fill opportunities for passive orders | +| **Spread** | Tighter spread = smaller price improvement possible | +| **Depth decay (alpha)** | Steeper decay = fills walk the book faster = higher slippage | +| **Cancel/fill ratio** | Higher ratio = more queue churn = harder to get fills | +| **MM pull speed** | Faster pull = stale quotes less likely = harder to snipe | +| **Book fragility** | During stress, depth drops 70-95% = fills at worse prices | + +### Per-Asset Fill Quality Expectations + +| Asset | Expected Fill Quality | Why | +|-------|----------------------|-----| +| BTC | Excellent | Deep book, tight spread, fast MM re-quote | +| ETH | Good | Similar to BTC, slightly thinner | +| SOL | Moderate | Mid-depth, moderate spread | +| DOGE | Poor | Thin book, wide spread, slow MM | +| ADA | Poor | Thin book, wide spread | + +### Optimization Strategy + +1. **Venue selection**: choose venues with better fill quality (Binance > BingX) +2. **Order type selection**: use POST_ONLY on tight-spread assets, MARKET on thin-spread +3. **Offset optimization**: CMA-ES learns optimal offset per (asset, regime, venue) +4. **Size optimization**: CMA-ES learns optimal size per (asset, regime, venue) +5. **Timing optimization**: CMA-ES learns when to quote vs when to wait + +The PerformanceMatrix tracks `avg_fill_value_score` per (regime, strategy, venue), +enabling the system to learn: "In this regime, on this venue, this strategy +achieves the best fill quality."