malkhut(docs): fill quality documentation — README, integration, OB study

README: Fill Quality section (core optimization target, metrics, reward
function, PerformanceMatrix, CMA-ES integration)

HftBacktestCWM integration doc: FillQuality dataclass, fill_value_score
computation, reward function weighting, PerformanceMatrix tracking

OB microstructure study: Section 13 — Fill Quality Optimization,
per-asset expectations, optimization strategy, connection to OB dynamics

Fill quality is MALKHUT's core aim: the system learns to get better fills
(faster, better-priced, less adverse selection) across regimes and venues.
This commit is contained in:
Codex
2026-07-15 15:36:22 +02:00
parent 618ad723e3
commit 619966605e
3 changed files with 140 additions and 0 deletions

View File

@@ -489,6 +489,60 @@ Re-measurement at correct fees is required for production deployment.
| **Sync/Async Seams** | `test_sync_async_seams.py` | 7 | Zinc latency, engine budget |
| **E2E Integration** | `test_e2e_integration.py` | 2 | Full pipeline: train→register→plan→risk→venue→zinc→ch→reload |
| **HftBacktestCWM** | `cwm/hft_cwm.py` | (new) | Queue-model fills (PowerProb), fill quality tracking, drop-in CWM |
### Fill Quality — The Core Optimization Target
MALKHUT is an **execution improvement engine**. Fill quality IS the primary aim — not PnL, not Sharpe, not win rate. The system learns to get **better fills**: faster, better-priced, with less adverse selection.
#### Fill Quality Metrics (per CWM transition)
| Metric | Source | What it measures |
|--------|--------|------------------|
| `slippage_bps` | `abs(fill_price - mid) / mid * 10000` | How far from mid did we fill? (aggressive) |
| `price_improvement_bps` | `(best_bid - fill_price) / best_bid * 10000` | How much better than touch? (passive) |
| `levels_consumed` | `fill_qty / avg_level_qty` | Queue depth of fill |
| `is_maker_fill` | `order_type==LIMIT or post_only` | Passive vs aggressive |
| `rolling_fill_rate` | EMA(0.8, 0.2) over recent fills | Recent fill success rate |
| `post_fill_adverse_bps` | `(new_mid - old_mid) / old_mid * 10000` | Price movement after fill |
| `fill_value_score` | `quality - abs(adverse) * 0.5` | **Composite optimization metric** |
#### Fill Quality in the Reward Function
```
reward = w_fill_probability * fill_value_score ← PRIMARY (fill quality)
+ w_expected_pnl * pnl ← secondary (PnL)
- w_adverse_selection * toxicity
- w_inventory_risk * inventory_risk
- w_tail_loss * tail_risk
- w_time_decay * time_in_loss
+ w_fee_quality * maker_fee_benefit
- spread_cost - taker_fee
```
The `fill_value_score` = price_quality - adverse_selection. For maker fills:
`price_quality = price_improvement_bps` (how much better than best bid/ask).
For taker fills: `price_quality = spread_bps - slippage_bps` (how efficiently we crossed).
#### Fill Quality in the PerformanceMatrix
`RegimeStrategyScore` now tracks 4 fill quality metrics per (regime, strategy, venue):
- `avg_fill_rate`: rolling fill success rate
- `avg_slippage_bps`: average slippage for aggressive fills
- `avg_price_improvement_bps`: average improvement for passive fills
- `avg_fill_value_score`: composite fill quality metric
This enables: "Which strategy achieves the best fill quality in regime X on venue Y?"
#### Fill Quality in CMA-ES
The CMA-ES optimizer now receives fill quality metrics in each `EpisodeResult`:
- `avg_fill_value_score`: average fill value across the episode
- `avg_price_improvement_bps`: average price improvement
- `avg_post_fill_adverse_bps`: average adverse selection
The CMA-ES objective is: maximize fill quality (primary) while maintaining positive PnL.
### Performance Benchmarks
| Metric | Value |