malkhut(docs): fill quality documentation — README, integration, OB study
README: Fill Quality section (core optimization target, metrics, reward function, PerformanceMatrix, CMA-ES integration) HftBacktestCWM integration doc: FillQuality dataclass, fill_value_score computation, reward function weighting, PerformanceMatrix tracking OB microstructure study: Section 13 — Fill Quality Optimization, per-asset expectations, optimization strategy, connection to OB dynamics Fill quality is MALKHUT's core aim: the system learns to get better fills (faster, better-priced, less adverse selection) across regimes and venues.
This commit is contained in:
@@ -357,3 +357,40 @@ Swapping the implementation behind that protocol is a textbook Strategy pattern.
|
||||
The planner doesn't know or care whether the book is synthesized or hftbacktest.
|
||||
The reward function is pure math on (prev_state, action, next_state) — identical
|
||||
regardless of how next_state was computed.
|
||||
|
||||
## Fill Quality Tracking (CORE Optimization Target)
|
||||
|
||||
Every CWM transition now computes `FillQuality` metrics on the resulting state:
|
||||
|
||||
```python
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class FillQuality:
|
||||
filled: bool # Did this action produce a fill?
|
||||
fill_qty: float # How much was filled?
|
||||
fill_price: float # At what price?
|
||||
slippage_bps: float # Aggressive: distance from mid
|
||||
price_improvement_bps: float # Passive: improvement over touch
|
||||
levels_consumed: int # Queue depth consumed
|
||||
is_maker_fill: bool # Passive (LIMIT) vs aggressive (CROSS)
|
||||
rolling_fill_rate: float # EMA of recent fill success
|
||||
post_fill_adverse_bps: float # Price movement after fill
|
||||
fill_value_score: float # Composite: quality - adverse
|
||||
```
|
||||
|
||||
The `fill_value_score` is the PRIMARY optimization target:
|
||||
- For maker fills: `price_improvement_bps - abs(post_fill_adverse) * 0.5`
|
||||
- For taker fills: `(spread_bps - slippage_bps) - abs(post_fill_adverse) * 0.5`
|
||||
|
||||
Both `MinimalCryptoLOBCWM` and `HftBacktestCWM` compute these identically.
|
||||
|
||||
The reward function weights fill quality via `w_fill_probability` (default 0.5):
|
||||
```
|
||||
reward = w_fill_probability * fill_value_score ← PRIMARY
|
||||
+ w_expected_pnl * pnl ← secondary
|
||||
- w_adverse_selection * toxicity
|
||||
...
|
||||
```
|
||||
|
||||
PerformanceMatrix stores `avg_fill_rate`, `avg_slippage_bps`, `avg_price_improvement_bps`,
|
||||
`avg_fill_value_score` per (regime, strategy, venue) — enabling:
|
||||
"Which strategy achieves the best fill quality in regime X on venue Y?"
|
||||
|
||||
Reference in New Issue
Block a user