malkhut(docs): fill quality documentation — README, integration, OB study

README: Fill Quality section (core optimization target, metrics, reward
function, PerformanceMatrix, CMA-ES integration)

HftBacktestCWM integration doc: FillQuality dataclass, fill_value_score
computation, reward function weighting, PerformanceMatrix tracking

OB microstructure study: Section 13 — Fill Quality Optimization,
per-asset expectations, optimization strategy, connection to OB dynamics

Fill quality is MALKHUT's core aim: the system learns to get better fills
(faster, better-priced, less adverse selection) across regimes and venues.
This commit is contained in:
Codex
2026-07-15 15:36:22 +02:00
parent 618ad723e3
commit 619966605e
3 changed files with 140 additions and 0 deletions

View File

@@ -489,6 +489,60 @@ Re-measurement at correct fees is required for production deployment.
| **Sync/Async Seams** | `test_sync_async_seams.py` | 7 | Zinc latency, engine budget | | **Sync/Async Seams** | `test_sync_async_seams.py` | 7 | Zinc latency, engine budget |
| **E2E Integration** | `test_e2e_integration.py` | 2 | Full pipeline: train→register→plan→risk→venue→zinc→ch→reload | | **E2E Integration** | `test_e2e_integration.py` | 2 | Full pipeline: train→register→plan→risk→venue→zinc→ch→reload |
| **HftBacktestCWM** | `cwm/hft_cwm.py` | (new) | Queue-model fills (PowerProb), fill quality tracking, drop-in CWM |
### Fill Quality — The Core Optimization Target
MALKHUT is an **execution improvement engine**. Fill quality IS the primary aim — not PnL, not Sharpe, not win rate. The system learns to get **better fills**: faster, better-priced, with less adverse selection.
#### Fill Quality Metrics (per CWM transition)
| Metric | Source | What it measures |
|--------|--------|------------------|
| `slippage_bps` | `abs(fill_price - mid) / mid * 10000` | How far from mid did we fill? (aggressive) |
| `price_improvement_bps` | `(best_bid - fill_price) / best_bid * 10000` | How much better than touch? (passive) |
| `levels_consumed` | `fill_qty / avg_level_qty` | Queue depth of fill |
| `is_maker_fill` | `order_type==LIMIT or post_only` | Passive vs aggressive |
| `rolling_fill_rate` | EMA(0.8, 0.2) over recent fills | Recent fill success rate |
| `post_fill_adverse_bps` | `(new_mid - old_mid) / old_mid * 10000` | Price movement after fill |
| `fill_value_score` | `quality - abs(adverse) * 0.5` | **Composite optimization metric** |
#### Fill Quality in the Reward Function
```
reward = w_fill_probability * fill_value_score ← PRIMARY (fill quality)
+ w_expected_pnl * pnl ← secondary (PnL)
- w_adverse_selection * toxicity
- w_inventory_risk * inventory_risk
- w_tail_loss * tail_risk
- w_time_decay * time_in_loss
+ w_fee_quality * maker_fee_benefit
- spread_cost - taker_fee
```
The `fill_value_score` = price_quality - adverse_selection. For maker fills:
`price_quality = price_improvement_bps` (how much better than best bid/ask).
For taker fills: `price_quality = spread_bps - slippage_bps` (how efficiently we crossed).
#### Fill Quality in the PerformanceMatrix
`RegimeStrategyScore` now tracks 4 fill quality metrics per (regime, strategy, venue):
- `avg_fill_rate`: rolling fill success rate
- `avg_slippage_bps`: average slippage for aggressive fills
- `avg_price_improvement_bps`: average improvement for passive fills
- `avg_fill_value_score`: composite fill quality metric
This enables: "Which strategy achieves the best fill quality in regime X on venue Y?"
#### Fill Quality in CMA-ES
The CMA-ES optimizer now receives fill quality metrics in each `EpisodeResult`:
- `avg_fill_value_score`: average fill value across the episode
- `avg_price_improvement_bps`: average price improvement
- `avg_post_fill_adverse_bps`: average adverse selection
The CMA-ES objective is: maximize fill quality (primary) while maintaining positive PnL.
### Performance Benchmarks ### Performance Benchmarks
| Metric | Value | | Metric | Value |

View File

@@ -357,3 +357,40 @@ Swapping the implementation behind that protocol is a textbook Strategy pattern.
The planner doesn't know or care whether the book is synthesized or hftbacktest. The planner doesn't know or care whether the book is synthesized or hftbacktest.
The reward function is pure math on (prev_state, action, next_state) — identical The reward function is pure math on (prev_state, action, next_state) — identical
regardless of how next_state was computed. regardless of how next_state was computed.
## Fill Quality Tracking (CORE Optimization Target)
Every CWM transition now computes `FillQuality` metrics on the resulting state:
```python
@dataclass(frozen=True, slots=True)
class FillQuality:
filled: bool # Did this action produce a fill?
fill_qty: float # How much was filled?
fill_price: float # At what price?
slippage_bps: float # Aggressive: distance from mid
price_improvement_bps: float # Passive: improvement over touch
levels_consumed: int # Queue depth consumed
is_maker_fill: bool # Passive (LIMIT) vs aggressive (CROSS)
rolling_fill_rate: float # EMA of recent fill success
post_fill_adverse_bps: float # Price movement after fill
fill_value_score: float # Composite: quality - adverse
```
The `fill_value_score` is the PRIMARY optimization target:
- For maker fills: `price_improvement_bps - abs(post_fill_adverse) * 0.5`
- For taker fills: `(spread_bps - slippage_bps) - abs(post_fill_adverse) * 0.5`
Both `MinimalCryptoLOBCWM` and `HftBacktestCWM` compute these identically.
The reward function weights fill quality via `w_fill_probability` (default 0.5):
```
reward = w_fill_probability * fill_value_score ← PRIMARY
+ w_expected_pnl * pnl ← secondary
- w_adverse_selection * toxicity
...
```
PerformanceMatrix stores `avg_fill_rate`, `avg_slippage_bps`, `avg_price_improvement_bps`,
`avg_fill_value_score` per (regime, strategy, venue) — enabling:
"Which strategy achieves the best fill quality in regime X on venue Y?"

View File

@@ -343,3 +343,52 @@ Example for DOGE ($22K amplitude, alpha=1.00):
BTC allows $100K orders with <1 bps slippage. DOGE requires $1K orders for BTC allows $100K orders with <1 bps slippage. DOGE requires $1K orders for
the same. Position sizing must account for the book's capacity, not just the same. Position sizing must account for the book's capacity, not just
the strategy's signal. the strategy's signal.
## 13. Fill Quality Optimization (MALKHUT Core)
MALKHUT is an execution improvement engine. Fill quality IS the primary aim.
### Fill Quality Metrics
For each CWM transition, MALKHUT computes:
- **slippage_bps**: aggressive fills — how far from mid?
- **price_improvement_bps**: passive fills — how much better than touch?
- **levels_consumed**: queue depth of fill
- **post_fill_adverse_bps**: price movement after fill (negative = adverse)
- **fill_value_score**: composite = quality - adverse * 0.5
### How This Connects to the OB
The OB microstructure directly determines fill quality:
| OB Characteristic | Impact on Fill Quality |
|-------------------|----------------------|
| **Depth at touch** | More depth = more fill opportunities for passive orders |
| **Spread** | Tighter spread = smaller price improvement possible |
| **Depth decay (alpha)** | Steeper decay = fills walk the book faster = higher slippage |
| **Cancel/fill ratio** | Higher ratio = more queue churn = harder to get fills |
| **MM pull speed** | Faster pull = stale quotes less likely = harder to snipe |
| **Book fragility** | During stress, depth drops 70-95% = fills at worse prices |
### Per-Asset Fill Quality Expectations
| Asset | Expected Fill Quality | Why |
|-------|----------------------|-----|
| BTC | Excellent | Deep book, tight spread, fast MM re-quote |
| ETH | Good | Similar to BTC, slightly thinner |
| SOL | Moderate | Mid-depth, moderate spread |
| DOGE | Poor | Thin book, wide spread, slow MM |
| ADA | Poor | Thin book, wide spread |
### Optimization Strategy
1. **Venue selection**: choose venues with better fill quality (Binance > BingX)
2. **Order type selection**: use POST_ONLY on tight-spread assets, MARKET on thin-spread
3. **Offset optimization**: CMA-ES learns optimal offset per (asset, regime, venue)
4. **Size optimization**: CMA-ES learns optimal size per (asset, regime, venue)
5. **Timing optimization**: CMA-ES learns when to quote vs when to wait
The PerformanceMatrix tracks `avg_fill_value_score` per (regime, strategy, venue),
enabling the system to learn: "In this regime, on this venue, this strategy
achieves the best fill quality."