466 lines
30 KiB
Markdown
466 lines
30 KiB
Markdown
|
|
# SPEC — MALKHUT ACTUALS INTAKE ("game it under actuals; the ecology stays")
|
|||
|
|
|
|||
|
|
**Author:** Claude (Opus 4.8), continuing Fable's review. **Date:** 2026-07-13.
|
|||
|
|
**Addressee:** mimo (mm_, `mm_ob_fill_sim`). **Ordered by:** HJ.
|
|||
|
|
**Status:** SPEC. Companion to `MALKHUT/README.md` (§ADDENDUM) and
|
|||
|
|
`prod/docs/SPEC_UV_SMART_EXEC_MM.md`.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 0. THE ONE-LINE LAW
|
|||
|
|
|
|||
|
|
> **Actuals CALIBRATE the game. The ecology PLAYS it. Real tape grounds the
|
|||
|
|
> adversaries; it NEVER replaces them.**
|
|||
|
|
|
|||
|
|
Every instruction below serves that sentence. If a change would let recorded
|
|||
|
|
history *substitute* for adversarial simulation, it is wrong, no matter how
|
|||
|
|
"realistic" it looks. Replaying our own tape teaches MALKHUT what happened once.
|
|||
|
|
The ecology is what lets it out-play what *could* happen — billions of
|
|||
|
|
strategy × counterparty × regime combinations, in parallel, none of which are in
|
|||
|
|
any tape. **The ecology is the edge. It is not a placeholder for missing data.**
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 0.5 THE TWO MODES (operator, 2026-07-13) — THE ARCHITECTURAL SPINE
|
|||
|
|
|
|||
|
|
MALKHUT is **two engines sharing one CWM**, and every module below belongs to one
|
|||
|
|
of them. This is the frame that makes "actuals" and "ecology" complementary
|
|||
|
|
instead of competing — and it is, structurally, **train vs inference**.
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
┌─ MODE 1: EXPLORE ───────────────────────────────────────────────────┐
|
|||
|
|
│ "Search/play over SYNTHETICS: fast, parallel, acquisitive of BEST" │
|
|||
|
|
│ │
|
|||
|
|
│ ECOLOGY-DOMINANT. Adversaries react. Regimes swept, not sampled. │
|
|||
|
|
│ Billions of strategy × counterparty × regime combinations. │
|
|||
|
|
│ Actuals appear ONLY as calibration (fee tables, latency CDF, │
|
|||
|
|
│ agent base-rates, plausible bounds). NOT as the arena. │
|
|||
|
|
│ │
|
|||
|
|
│ Output: POLICY POOL — a ranked, regime-indexed library of │
|
|||
|
|
│ strategies with known behaviour under known adversary mixes. │
|
|||
|
|
│ Cadence: OFFLINE, hours-to-days, unbounded compute. │
|
|||
|
|
│ Success metric: coverage + robustness, NOT backtest PnL. │
|
|||
|
|
└──────────────────────────────┬───────────────────────────────────────┘
|
|||
|
|
│ policy pool + performance manifold
|
|||
|
|
▼
|
|||
|
|
┌─ MODE 2: RECOMMEND (inference/query) ───────────────────────────────┐
|
|||
|
|
│ "Search over LIVE, ACTUAL (OBF) IRL conditions as INPUTS, to │
|
|||
|
|
│ extrapolate the closest-to-absolute-best strategy → RECOMMEND" │
|
|||
|
|
│ │
|
|||
|
|
│ ACTUALS-DOMINANT. The live book/regime/funding/latency IS the │
|
|||
|
|
│ query. The ecology is now a PRIOR over who is on the other side │
|
|||
|
|
│ right now, not a population to be swept. │
|
|||
|
|
│ │
|
|||
|
|
│ Given: live MarketWorldState (S1 book, S2 funding/dvol, S3 regime, │
|
|||
|
|
│ S4 latency, account, intent) │
|
|||
|
|
│ Ask: which policy in the pool is nearest-optimal for THIS state? │
|
|||
|
|
│ Output: a RECOMMENDATION (policy + its expected behaviour + │
|
|||
|
|
│ confidence + the nearest explored neighbours it interpolates│
|
|||
|
|
│ between). │
|
|||
|
|
│ Cadence: ONLINE, ≤25 ms planning inside ≤100 ms end-to-end. │
|
|||
|
|
│ Success metric: live outcome ≈ predicted outcome (Rule 10). │
|
|||
|
|
└──────────────────────────────────────────────────────────────────────┘
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### Why this is the right shape
|
|||
|
|
|
|||
|
|
- **Mode 1 without Mode 2** is a beautiful simulator nobody trades.
|
|||
|
|
- **Mode 2 without Mode 1** is a lookup table over history — the overfit that
|
|||
|
|
kills quants. It can only recommend what already happened.
|
|||
|
|
- **Together**: Mode 1 explores a space *vastly larger than any tape can contain*
|
|||
|
|
(that is the edge — we out-play participants who only ever fit history), and
|
|||
|
|
Mode 2 *localizes* the live moment inside that explored space and returns the
|
|||
|
|
best-known play, **including for market states we have never actually seen**,
|
|||
|
|
because Mode 1 has *already played them*.
|
|||
|
|
|
|||
|
|
**The word "extrapolate" in the operator's phrasing is load-bearing.** Mode 2's
|
|||
|
|
job is NOT nearest-neighbour lookup. It is: *place the live state inside the
|
|||
|
|
performance manifold Mode 1 built, and interpolate/extrapolate the best play.*
|
|||
|
|
That requires Mode 1 to output a **manifold**, not a leaderboard:
|
|||
|
|
|
|||
|
|
| Mode 1 must emit | Not just |
|
|||
|
|
|---|---|
|
|||
|
|
| policy × (regime, spread, depth, toxicity, latency, funding, inventory) → expected outcome **+ variance** | "policy #7 scored 8,628" |
|
|||
|
|
| the **boundaries** of where each policy was tested (extrapolation beyond = flagged) | a single global champion |
|
|||
|
|
| **why** a policy wins (which adversary it beats, which it loses to) | an opaque score |
|
|||
|
|
|
|||
|
|
### Consequences for the build
|
|||
|
|
|
|||
|
|
1. **`ScenarioLibrary` must SWEEP, not sample.** Mode 1's job is coverage of the
|
|||
|
|
state space, including regions our tape never visited. Tape-derived scenarios
|
|||
|
|
(CLASS-M/X) are the *anchors*; the sweep fills between and beyond them.
|
|||
|
|
2. **The performance matrix in `training/selector.py` becomes the manifold** —
|
|||
|
|
and it must carry **confidence + support count + distance-to-nearest-explored**
|
|||
|
|
per cell. A recommendation from a thinly-explored cell must SAY SO.
|
|||
|
|
3. **Mode 2 must refuse to extrapolate too far.** If the live state is outside
|
|||
|
|
Mode 1's explored envelope (novel regime, latency beyond p99.9, depth collapse
|
|||
|
|
never swept), the honest answer is **"OUT OF DISTRIBUTION — fall back to the
|
|||
|
|
doctrinal simple policy"**, not a confident recommendation. This is the same
|
|||
|
|
law as `INDETERMINATE`: *unknown is not flat, and unknown is not "best guess".*
|
|||
|
|
Wire it as an explicit `RecommendationConfidence.OUT_OF_DISTRIBUTION` verdict
|
|||
|
|
that the risk gate honours.
|
|||
|
|
4. **Mode 2's live search is a LOCALIZATION, not a fresh MCTS from scratch** — the
|
|||
|
|
≤25 ms budget buys you refinement around a pool policy, not global search. The
|
|||
|
|
pool is the compressed product of the billions of Mode-1 trials; Mode 2 spends
|
|||
|
|
its milliseconds *choosing and adapting*, not re-deriving.
|
|||
|
|
5. **Mode 2's own outcomes feed back into Mode 1** as new anchors (live
|
|||
|
|
discrepancies → `live_discrepancies` table → next sweep is denser where we were
|
|||
|
|
wrong). That is the learning loop, and Rule 10 governs it: **where shadow and
|
|||
|
|
live disagree, live is right and the manifold gets corrected.**
|
|||
|
|
|
|||
|
|
### Module ownership by mode
|
|||
|
|
|
|||
|
|
| Module | Mode 1 (EXPLORE) | Mode 2 (RECOMMEND) |
|
|||
|
|
|---|---|---|
|
|||
|
|
| `cwm/core.py` | ✅ the arena | ✅ the local rollout model |
|
|||
|
|
| `counterparties.py` (ecology) | ✅ **swept adversary population** | ✅ **prior over current opponents** |
|
|||
|
|
| `training/cma_trainer.py`, `generator.py` | ✅ | — |
|
|||
|
|
| `ScenarioLibrary` (§4) | ✅ sweeps | — |
|
|||
|
|
| `training/registry.py` (policy pool) | ✅ writes | ✅ reads |
|
|||
|
|
| `training/selector.py` (→ **manifold**) | ✅ builds | ✅ queries |
|
|||
|
|
| `ActualsLoader` (§4) | ⚠️ calibration only | ✅ **the live query itself** |
|
|||
|
|
| `planner/sm_mcts.py` | ✅ deep, unbounded | ✅ **≤25 ms localization** |
|
|||
|
|
| `risk/gate.py` | ✅ constrains training | ✅ hard veto on recommendation |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 1. FINDING #0 — THE FEE BUG IS REAL, AND MEASURED
|
|||
|
|
|
|||
|
|
Fable suspected a 10× unit slip. **Our own fills confirm it.**
|
|||
|
|
|
|||
|
|
```sql
|
|||
|
|
SELECT liquidity_side, order_type, count() n, avg(fee_bps)
|
|||
|
|
FROM dolphin.trade_execution_quality WHERE fee_bps IS NOT NULL GROUP BY 1,2
|
|||
|
|
-- TAKER MARKET 1455 rows avg = 5.016 bps (min 5.000, max 6.649)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
| Where | Code says | Reality (our fills / venue docs) | Factor |
|
|||
|
|
|---|---|---|---|
|
|||
|
|
| `malkhut/training/asset_classification.py:164` (BingX) | `default_taker_fee_bps=0.5` | **5.0** | **10×** |
|
|||
|
|
| `…:154` (Binance) | `default_taker_fee_bps=0.4` | ~4.5 | ~10× |
|
|||
|
|
| `…:174` (Bybit) | `default_taker_fee_bps=0.06` | ~5.5 | ~90× |
|
|||
|
|
| `…:384` (default profile) | `taker_fee_bps=0.5, maker_fee_bps=-0.2` | taker 5.0; maker ≈ +2.0 (BingX perp maker is POSITIVE, not a rebate) | 10× + sign |
|
|||
|
|
|
|||
|
|
**Impact:** these feed `VenueRules.taker_fee_bps` / `maker_fee_bps`
|
|||
|
|
(`malkhut/state.py:109-121`) → the CWM reward → `w_fee_quality`. CMA-ES has been
|
|||
|
|
optimizing against fees an order of magnitude too cheap. **Every policy trained so
|
|||
|
|
far is suspect** — cheap fees reward overtrading, churn, and cross-spread
|
|||
|
|
aggression that real friction annihilates. This is the exact failure the
|
|||
|
|
2026-07-10 venue-friction audit exists to prevent.
|
|||
|
|
|
|||
|
|
**ACTION (do this first, before any other work):**
|
|||
|
|
1. Fix the four sites above. Source of truth = `dolphin.trade_execution_quality`
|
|||
|
|
(`fee_bps`, `liquidity_side`, `order_type`), not vendor marketing pages.
|
|||
|
|
2. Add a **mutation-litmus test**: set `taker_fee_bps` to 0.5 → a test asserting
|
|||
|
|
"policy PnL under realistic friction" must go **RED**. If nothing breaks when
|
|||
|
|
fees change 10×, the reward function isn't actually using them.
|
|||
|
|
3. Re-run every CMA-ES benchmark. The "8,628 / 100.4 bps PnL" headline number is
|
|||
|
|
void until it is re-measured at 5 bps taker.
|
|||
|
|
4. Maker fee sign: verify against a real maker fill before trusting a rebate.
|
|||
|
|
**We have zero MAKER rows in `trade_execution_quality`** (all 1455 are TAKER
|
|||
|
|
MARKET — because UV/F5 only sends MARKET). Until SMART-EXEC produces real
|
|||
|
|
maker fills, the maker fee is an *assumption*: mark it as such in code with a
|
|||
|
|
`# UNVERIFIED — no maker fills on record as of 2026-07-13` comment.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 2. THE DATA SOURCES — WHAT EXISTS, VERIFIED TONIGHT
|
|||
|
|
|
|||
|
|
All queried live on `localhost:8123`. Row counts are real.
|
|||
|
|
|
|||
|
|
| # | Source | Rows | Columns (exact) | State |
|
|||
|
|
|---|---|---|---|---|
|
|||
|
|
| **S1** | `dolphin.obf_universe` | **15,329,163,875** | `ts` DateTime64(3), `symbol`, `spread_bps` f32, `depth_1pct_usd` f64, `depth_quality` f32, `fill_probability` f32, `imbalance` f32, `best_bid` f64, `best_ask` f64, `n_bid_levels` u8, `n_ask_levels` u8 | **LIVE, the motherlode** |
|
|||
|
|
| **S2** | `dolphin.exf_data` | 22,968,094 | `ts` DateTime64(6), `funding_rate` f32, `dvol` f32, `fear_greed` f32, `taker_ratio` f32 | LIVE |
|
|||
|
|
| **S3** | `dolphin.maras_fingerprint` | 1,117,722 | `ts`, `regime` (LowCard), `regime_idx` u8, `confidence`, `final_score`, `conflict_level`, `tier_exf`, `tier_eigen`, `tier_btc`, `tier_esof`, `tier_micro` (+ `_confidence` each), `s_exf_funding_bp`, … | LIVE |
|
|||
|
|
| **S4** | `dolphin.eigen_scans` | 1,529,803 | `ts`, `scan_number` u32, `vel_div` f32, `w50_velocity`, `w750_velocity`, `instability_50`, **`scan_to_fill_ms`**, **`step_bar_ms`**, `scan_uuid` | LIVE — **and it is the latency oracle** |
|
|||
|
|
| **S5** | `dolphin.trade_execution_quality` | 8,006 | `trade_id`, `asset`, `side`, `client_order_id`, `venue_order_id`, `order_type`, **`liquidity_side`**, **`fee_bps`**, `fill_quality_score`, `commission_quote`, `fee_rate`, … | LIVE — **fee/fill ground truth** |
|
|||
|
|
| **S6** | `dolphin.trade_events` | (BLUE's live trades) | `ts`, `trade_id`, `asset`, `side`, `entry_price`, `exit_price`, `pnl`, `pnl_pct`, `exit_reason`, `vel_div_entry`, `boost_at_entry` | LIVE |
|
|||
|
|
| **S7** | `dolphin_uv.exec_journal` | growing | `kind`, `timestamp`, `scan_number`, `intent_id`, `trade_id`, `slot_id`, `asset`, `side`, `action`, `reference_price`, `target_size`, `leverage`, `u_prefix_client_id`, `promo_metadata` (JSON) | LIVE |
|
|||
|
|
| **S8** | `dolphin_uv.tp_exit_ingress` / `max_hold_ingress` | new (2026-07-13) | full per-scan exit diagnostics: `tp_effective_pct`, `tp_mod_factor`, `tp_floor_armed`, `cascade_count`, `imbalance_ma5`, `branch`, … | LIVE as of today |
|
|||
|
|
| **S9** | `dolphin.esof_advisory` | **0** | schema exists (`dow`, `session`, `moon_illumination`, `slot_wr_pct`, …) | ⚠️ **EMPTY** |
|
|||
|
|
| **S10** | `dolphin.obf_fast_intrade` | **0** | schema exists | ⚠️ **EMPTY** |
|
|||
|
|
|
|||
|
|
**Do not spec against S9/S10 as if they were populated.** Either they get a
|
|||
|
|
writer, or they are out of scope. Say so out loud rather than building an intake
|
|||
|
|
for a table that never delivers a row (that is how the missing-CH-table blackout
|
|||
|
|
happened to UV; the fix cost a day).
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 3. FINDING #1 — THE BOOK-FIDELITY GAP (read this twice)
|
|||
|
|
|
|||
|
|
This is the **central technical problem** of the whole integration, and it is not
|
|||
|
|
optional to solve.
|
|||
|
|
|
|||
|
|
**MALKHUT wants a ladder.** `malkhut/state.py:131`:
|
|||
|
|
```python
|
|||
|
|
class OrderBookState:
|
|||
|
|
bids: Tuple[PriceLevel, ...] # price + qty, per level
|
|||
|
|
asks: Tuple[PriceLevel, ...]
|
|||
|
|
```
|
|||
|
|
The CWM's queue-position model, `JOIN_QUEUE`, `LADDER`, `ICEBERG`, queue-ahead
|
|||
|
|
estimates, and adverse-selection accounting all depend on **per-level depth**.
|
|||
|
|
|
|||
|
|
**Our tape has no ladder.** `dolphin.obf_universe` gives, per symbol per ~sub-second
|
|||
|
|
tick: `best_bid`, `best_ask`, `spread_bps`, `depth_1pct_usd` (one aggregate
|
|||
|
|
number), `depth_quality`, `imbalance`, `fill_probability`, `n_bid_levels`,
|
|||
|
|
`n_ask_levels` (**counts**, not the levels themselves).
|
|||
|
|
|
|||
|
|
That is **L1 + shape summary**, not L2. 15.3 billion rows of it — enormously
|
|||
|
|
valuable, but it cannot be *decoded* back into a ladder. Information that was
|
|||
|
|
never recorded cannot be recovered.
|
|||
|
|
|
|||
|
|
### Three honest options — pick explicitly, do not drift
|
|||
|
|
|
|||
|
|
| Option | What it is | Cost | Fidelity |
|
|||
|
|
|---|---|---|---|
|
|||
|
|
| **A. Synthetic ladder w/ declared prior** | Reconstruct levels from `best_bid/ask` + `depth_1pct_usd` + `n_*_levels` + `imbalance` via an explicit shape model (e.g. exponential decay of qty over levels, calibrated so Σqty within 1% = `depth_1pct_usd`, level count = `n_bid_levels`). | LOW | **Approximate.** Queue position becomes a *model*, not a measurement. MUST be labeled as such everywhere it flows. |
|
|||
|
|
| **B. Start recording real L2 now** | New writer → `dolphin_malkhut.book_l2` (or a zinc region + spooler). Depth-20 snapshots at OBF cadence. | MEDIUM (new hose + storage; L2 at 50 symbols × sub-second is BIG — see §7 storage) | **Truth**, but only from the day it starts. Cannot backfill history. |
|
|||
|
|
| **C. hftbacktest with external L2** | Use a third-party L2 feed (Binance archival) for CWM calibration; keep obf_universe for regime/context. | MEDIUM | Truth, but it is *Binance's* book, not BingX's — venue microstructure differs. |
|
|||
|
|
|
|||
|
|
**RECOMMENDATION: A + B in parallel, and say which one a given result came from.**
|
|||
|
|
- **A** unblocks immediate use of 15.3B rows of history for *distributional*
|
|||
|
|
calibration (spread regimes, depth regimes, imbalance dynamics, fill-probability
|
|||
|
|
priors) — where the ladder shape matters less than the aggregate.
|
|||
|
|
- **B** starts the clock on real queue-truth for the queue-sensitive claims
|
|||
|
|
(`JOIN_QUEUE`, maker fills, queue-ahead) — the very claims SMART-EXEC needs.
|
|||
|
|
- **NEVER** let an A-derived queue-position estimate be reported as a measured
|
|||
|
|
fill probability. Tag every artifact: `book_source ∈ {SYNTH_A, REAL_L2_B, EXT_C}`.
|
|||
|
|
Rule 3 (replay correctness before search depth) means exactly this.
|
|||
|
|
|
|||
|
|
**Litmus:** when B has ≥ 1 week of real L2, re-run A-calibrated policies against
|
|||
|
|
B-truth. The delta *is* the measurement of how much the synthetic prior lied. If
|
|||
|
|
that delta is large, every A-era conclusion is downgraded to a hypothesis.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 4. THE PLUG MAP — WHAT GOES WHERE, EXACTLY
|
|||
|
|
|
|||
|
|
`MarketWorldState` (`malkhut/state.py:257`) is the CWM root. It already has the
|
|||
|
|
right holes. Fill them from our sources:
|
|||
|
|
|
|||
|
|
| `MarketWorldState` field | Line | Source | Transform |
|
|||
|
|
|---|---|---|---|
|
|||
|
|
| `book: OrderBookState` | 262 | **S1** `obf_universe` | §3 Option A synth (or B when live). `best_bid`/`best_ask` direct; ladder via declared prior; `symbol` join key. |
|
|||
|
|
| `account: AccountState` | 263 | **DITAv2 ASEx account core** (`asex_account` / `AccountProjectionV2`) — **CONSUMER ONLY** (two-cores-share-nothing law) | For offline training: reconstruct from `dolphin.account_events` / `dolphin_uv.exec_journal`. Never a second writer. |
|
|||
|
|
| `open_orders` | 264 | live: venue adapter; training: **S7** `exec_journal` + **S5** `trade_execution_quality` (client_order_id join) | |
|
|||
|
|
| `trade_path: TradePathState` | 265 | **S8** `tp_exit_ingress` + **S6** `trade_events` | `mae_bps`/`mfe_bps`/`time_in_loss_s` etc. reconstructable from tp_exit_ingress per-scan stamps + trade lifecycle. **`dolphin_regime_score` ← S3; `book_imbalance` ← S1 `imbalance`; `orderflow_toxicity` ← derive (see §5)**. |
|
|||
|
|
| `intent: ExecutionIntent` | 266 | **DITAv2 `KernelIntent`** (binding, README §B.1) | Map ENTER/EXIT + `reference_price`/`target_size`/`leverage`/`metadata.promo_client_id`. `urgency` ← SMART-EXEC urgency class (see §8). |
|
|||
|
|
| **`funding_bps`** | 268 | **S2** `exf_data.funding_rate` | ×10⁴ → bps. Nearest-ts join. |
|
|||
|
|
| **`volatility_state`** | 269 | **S2** `exf_data.dvol` | The sacred vol number. **Same field the gate uses** (`DOLPHIN_VOL_P60_THRESHOLD`, doctrinal 0.00026414). |
|
|||
|
|
| **`market_regime`** | 270 | **S3** `maras_fingerprint.regime` | ⚠️ **taxonomy mismatch — see §6.** |
|
|||
|
|
| **`feed_latency_ms`** | 272 | **S4** `eigen_scans.step_bar_ms` | p50 = **0.08 ms**. |
|
|||
|
|
| **`order_latency_ms`** | 273 | **S4** `eigen_scans.scan_to_fill_ms` | ⚠️ **SAMPLE THE DISTRIBUTION, NOT THE MEAN — see §5.** |
|
|||
|
|
| `venue: VenueRules` | 261 | **S5** for fees (§1); `prod/bingx/` for tick/lot/min_notional | Fee fix is blocking. |
|
|||
|
|
|
|||
|
|
### Where the code changes go
|
|||
|
|
|
|||
|
|
| New module | Path | Job |
|
|||
|
|
|---|---|---|
|
|||
|
|
| `ActualsLoader` | `malkhut/data/actuals.py` **(new pkg `malkhut/data/`)** | CH HTTP reader → typed frames. One method per source S1–S8. Polars, chunked, no full-table loads. |
|
|||
|
|
| `BookSynthesizer` | `malkhut/data/book_synth.py` | §3 Option A. Declared prior, unit-tested against any real L2 we get. Emits `book_source` tag. |
|
|||
|
|
| `LatencyOracle` | `malkhut/data/latency.py` | §5. Empirical CDF sampler, seed-pinned. |
|
|||
|
|
| `EcologyFitter` | `malkhut/data/ecology_fit.py` | §7. Fits counterparty params to tape. **Does not create or delete agent types.** |
|
|||
|
|
| `RegimeBridge` | `malkhut/data/regime_bridge.py` | §6. MARAS ↔ MALKHUT regime mapping, explicit and tested. |
|
|||
|
|
| `ScenarioLibrary` | `malkhut/data/scenarios.py` | §7. CLASS-M/CLASS-X hour selection from tape → adversarial scenario suites. |
|
|||
|
|
|
|||
|
|
Wire them at: `malkhut/training/cma_trainer.py` (scenario source),
|
|||
|
|
`malkhut/cwm/core.py:transition` (latency + venue rules injection),
|
|||
|
|
`malkhut/counterparties.py` (fitted params in, agent classes unchanged).
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 5. FINDING #2 — LATENCY IS NOT A NUMBER, IT IS A MONSTER-TAILED DISTRIBUTION
|
|||
|
|
|
|||
|
|
Measured tonight from `dolphin.eigen_scans` (`scan_to_fill_ms`, n = 1.5M):
|
|||
|
|
|
|||
|
|
| p50 | p95 | p99 | max |
|
|||
|
|
|---|---|---|---|
|
|||
|
|
| **49.9 ms** | **596.9 ms** | **10,793.9 ms** | **450,115.9 ms** (7.5 minutes) |
|
|||
|
|
|
|||
|
|
`step_bar_ms` p50 = **0.08 ms**.
|
|||
|
|
|
|||
|
|
The README's "typical_latency_ms: 100" for BingX is a **fiction that will get us
|
|||
|
|
killed**: it is 2× the median and **0.9% of the mass is beyond 10 SECONDS**. A
|
|||
|
|
policy trained on constant-100ms latency has never met the market that actually
|
|||
|
|
fills our orders. The p99 is what turns a maker quote into an adverse-selected
|
|||
|
|
gift, and it is exactly where the "why did I get filled?" question (Rule 2) gets
|
|||
|
|
its ugliest answer.
|
|||
|
|
|
|||
|
|
**ACTION:** `LatencyOracle` must sample the **empirical CDF**, per-regime where
|
|||
|
|
possible (latency and stress correlate — verify), seed-pinned for determinism
|
|||
|
|
(Rule 9 + our seed-pin doctrine). Constant-latency mode may exist **only** as a
|
|||
|
|
labeled ablation, never as the default. Add a scenario class: **`LATENCY_STORM`**
|
|||
|
|
(sample exclusively from > p95) — if a policy's edge evaporates there, we need to
|
|||
|
|
know before capital does.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 6. FINDING #3 — REGIME TAXONOMY MISMATCH
|
|||
|
|
|
|||
|
|
MARAS actually emits (verified, `dolphin.maras_fingerprint`, 1.1M rows):
|
|||
|
|
|
|||
|
|
| `regime_idx` | `regime` | rows |
|
|||
|
|
|---|---|---|
|
|||
|
|
| 1 | BEARISH | 72,479 |
|
|||
|
|
| 2 | CHOPPY_BEARISH | 516,477 |
|
|||
|
|
| 3 | CHOPPY | 295,366 |
|
|||
|
|
| 4 | SIDEWAYS | 201,358 |
|
|||
|
|
| 5 | CHOPPY_BULLISH | 32,046 |
|
|||
|
|
|
|||
|
|
(idx 0, 6, 7 unobserved in this window — the taxonomy has room MARAS has not used.)
|
|||
|
|
|
|||
|
|
MALKHUT's `StrategySelector` uses a **different, invented** set: `trending_up`,
|
|||
|
|
`trending_down`, `high_volatility`, `low_volatility`, `mean_reverting`,
|
|||
|
|
`momentum`, `choppy`, `liquidity_hole`, `normal`.
|
|||
|
|
|
|||
|
|
**These are two different languages.** A performance matrix keyed on MALKHUT's
|
|||
|
|
regimes cannot be looked up from a MARAS fingerprint without a mapping, and an
|
|||
|
|
implicit/lossy mapping will silently mis-select strategies in production.
|
|||
|
|
|
|||
|
|
**ACTION:** `RegimeBridge` (`malkhut/data/regime_bridge.py`) with an **explicit,
|
|||
|
|
tested, total** mapping MARAS→MALKHUT. Where MALKHUT has an axis MARAS lacks
|
|||
|
|
(e.g. `liquidity_hole`), derive it from S1 (`depth_quality`, `spread_bps`
|
|||
|
|
percentiles) and **name the derivation**. Where MARAS has confidence
|
|||
|
|
(`confidence`, `conflict_level`), **carry it through** — a low-confidence regime
|
|||
|
|
tag must reach the planner as low-confidence, not as a hard label. Prefer
|
|||
|
|
extending MALKHUT to speak MARAS over inventing a third dialect.
|
|||
|
|
|
|||
|
|
**Also plumb the MARAS tiers** (`tier_exf`, `tier_eigen`, `tier_btc`,
|
|||
|
|
`tier_esof`, `tier_micro` + confidences) into the feature vector — that is a
|
|||
|
|
5-tier ensemble view of the market the planner currently cannot see at all, and
|
|||
|
|
it is *already computed and stored*. Free signal.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 7. THE ECOLOGY — HOW ACTUALS FEED IT WITHOUT REPLACING IT
|
|||
|
|
|
|||
|
|
**This section is the point of the whole document. Read it as law.**
|
|||
|
|
|
|||
|
|
`malkhut/counterparties.py` defines the agent ecology (`AgentRole`,
|
|||
|
|
`state.py:61`): PASSIVE_MAKER, TOXIC_TAKER, LATENCY_ARB, MOMENTUM_TAKER,
|
|||
|
|
MEAN_REVERSION_TAKER, INVENTORY_MM, LIQUIDATION_FLOW, NOISE_TRADER,
|
|||
|
|
STALE_QUOTE_ATTACKER. **Nine roles declared, 4 implemented.**
|
|||
|
|
|
|||
|
|
### What actuals DO to the ecology
|
|||
|
|
|
|||
|
|
| Do | How |
|
|||
|
|
|---|---|
|
|||
|
|
| **Fit each agent's parameters** | `EcologyFitter` estimates population parameters from tape: TOXIC_TAKER arrival intensity ← S1 `fill_probability` collapse + adverse post-fill drift in S5/S6; LATENCY_ARB ← S4 latency tails vs price moves; LIQUIDATION_FLOW ← S2 funding extremes + cascade signatures in S8 `cascade_count`; INVENTORY_MM ← S1 `imbalance` mean-reversion; MOMENTUM_TAKER ← S4 `vel_div` / `w50_velocity`. |
|
|||
|
|
| **Set the population MIX per regime** | Which agents dominate under BEARISH vs CHOPPY (S3) is an empirical question our tape can answer. The mix becomes regime-conditional. |
|
|||
|
|
| **Supply the stress scenarios** | CLASS-M / CLASS-X hours (measured hostile windows) → `ScenarioLibrary`. **CLASS-X law: extreme magnitudes are a FEATURE channel, never filtered.** |
|
|||
|
|
| **Bound the plausible** | Tape says what a real book *can* do; adversaries should be allowed to be *worse*, but their base rates must be anchored. |
|
|||
|
|
|
|||
|
|
### What actuals MUST NOT DO
|
|||
|
|
|
|||
|
|
| Never | Why |
|
|||
|
|
|---|---|
|
|||
|
|
| **Replace an agent with a tape replay** | A replayed counterparty cannot react to us. The entire adverse-selection question ("why did I get filled?") requires an opponent that *responds*. A tape is a corpse; the ecology is an opponent. |
|
|||
|
|
| **Delete an agent type for lack of data** | The 5 unimplemented roles are *hypotheses about who is on the other side*. Absence of evidence in our thin tape is not evidence of absence in the market. Implement them; fit what you can; make the rest adversarial priors. |
|
|||
|
|
| **Restrict the game to observed histories** | Billions of combinations, most of which never occurred, is the *product*. Recorded tape is a measure-zero slice of the strategy × counterparty × regime space. Optimizing only over what happened is exactly the overfit that kills quants. |
|
|||
|
|
| **Filter outliers** | CLASS-X law. The 450-second latency tail and the 7.5-minute stop are not noise — they are the market's teeth. |
|
|||
|
|
|
|||
|
|
### The scale ambition (operator's stated aim)
|
|||
|
|
|
|||
|
|
Femtosecond-scale trials, parallel, over **billions** of combinations. Current:
|
|||
|
|
5.3 µs/CWM-transition (numba), 16 ms/episode (parallel), 7× on 8 workers.
|
|||
|
|
Billions of combinations × ~10³ steps at 5.3 µs = ~10¹² × 5.3 µs ≈ **months of
|
|||
|
|
single-box CPU.** So:
|
|||
|
|
|
|||
|
|
1. **Batch/vectorize the counterparty step** (the ecology is the inner loop —
|
|||
|
|
make it a matrix op over the agent population, not a Python loop over agents).
|
|||
|
|
2. **GPU or massively-parallel path** for the rollout kernel (the batch MCTS
|
|||
|
|
kernel already exists — push it further).
|
|||
|
|
3. **Cheap-then-expensive cascade**: screen billions with a cheap surrogate
|
|||
|
|
(vectorized reward, shallow rollout), promote survivors to full CWM depth.
|
|||
|
|
Do NOT run 5.3 µs × depth-3 MCTS on every one of a billion candidates.
|
|||
|
|
4. **Ray path already exists** (`training/ray_eval.py`) — that is the horizontal
|
|||
|
|
scaling seam when the box runs out.
|
|||
|
|
5. Honest note: "femtosecond" is a *metaphor for parallel breadth*, not a
|
|||
|
|
physical claim (a CPU cycle is ~300 picoseconds; a femtosecond is 10⁻¹⁵ s —
|
|||
|
|
light travels 0.3 µm in one). State the real target: **N trials/second at
|
|||
|
|
fidelity F**, and measure it. Otherwise we cannot tell progress from poetry.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 8. INTAKE CONTRACT ALIGNMENT (with SMART-EXEC + UV)
|
|||
|
|
|
|||
|
|
MALKHUT's planner and `SPEC_UV_SMART_EXEC_MM.md` describe the **same seat** at the
|
|||
|
|
venue-adapter boundary. Do not build two doorframes.
|
|||
|
|
|
|||
|
|
- **Input**: DITAv2 `KernelIntent` + `guideline_price` + `urgency_class`
|
|||
|
|
(CATASTROPHIC / PROTECT / HARVEST / ROTATE / ACQUIRE). Map `urgency_class` →
|
|||
|
|
`ExecutionIntent.urgency` (`state.py:246`) and `prefer_maker` (:250).
|
|||
|
|
**CATASTROPHIC ⇒ the planner is BYPASSED. No cleverness on the stop path, ever.**
|
|||
|
|
- **`u-` clientOrderId prefix is LAW** (T9 seam / README §B.2).
|
|||
|
|
- **DUAL-LEVERAGE LAW** (README §B.3): conviction leverage [0.5,9.0] sizes
|
|||
|
|
quantity; venue leverage = `prod/bingx/leverage.py` mapping → int [1,3].
|
|||
|
|
Litmus: 0.5→1, 4.75→2, 8.0→3, 9.0→3. **Note:** we observed a live 3→2 clamp on
|
|||
|
|
2026-07-13 (handover anomaly #2) — reconcile before MALKHUT sizes anything.
|
|||
|
|
- **Execution-truth doctrine is inherited wholesale** (`bdc54fb`): NOT_ATTEMPTED /
|
|||
|
|
REFUSED → rollback sound; **INDETERMINATE → never**. MALKHUT's venue adapter
|
|||
|
|
(`malkhut/venue/bingx/adapter.py`) wraps DITAv2 and therefore inherits the
|
|||
|
|
fences — **verify with a test, do not assume.** See
|
|||
|
|
`COMPREHENSIVE_UV_EDGE_CASE_THEORETICALS.md` §11 M1–M7 for the unaudited seams.
|
|||
|
|
- **No reconcilers.** Bounded, read-only, own-clientOrderId point lookups only.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 9. PERSISTENCE + PROVENANCE
|
|||
|
|
|
|||
|
|
- Namespace `dolphin_malkhut` (approved). **NO TTL** (retention doctrine).
|
|||
|
|
- **DDL ships WITH the code, and the applier's verify-set must require every new
|
|||
|
|
table** — the 2026-07-13 lesson: two UV tables existed as `.sql` files for days
|
|||
|
|
while the runner 404'd every scan and journaled nothing. Pattern to copy:
|
|||
|
|
`prod/clickhouse/uv/apply_uv_ddl.py` (`EXPECTED_TABLES`).
|
|||
|
|
- **Every row carries provenance**: `book_source` (SYNTH_A/REAL_L2_B/EXT_C),
|
|||
|
|
`latency_model` (EMPIRICAL_CDF/CONST), `fee_table_version`, `ecology_fit_version`,
|
|||
|
|
`policy_version`, `seed`. A result whose provenance is unknown is not a result.
|
|||
|
|
- **Cross-reference law**: `fulfilment_decisions.intent_id` + `trade_id` join to
|
|||
|
|
`dolphin_uv.exec_journal` ⇄ venue order history. End-to-end or it is a black box.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 10. ORDER OF WORK (do not reorder)
|
|||
|
|
|
|||
|
|
| # | Task | Gate |
|
|||
|
|
|---|---|---|
|
|||
|
|
| 1 | **Fix the fees** (§1) + mutation-litmus test | A 10× fee change must break a test |
|
|||
|
|
| 2 | **Re-baseline** every CMA-ES number at real friction | Old headline numbers marked VOID |
|
|||
|
|
| 3 | **LatencyOracle** (§5) + `LATENCY_STORM` scenario | Constant-latency is an ablation, not a default |
|
|||
|
|
| 4 | **Decide the book question** (§3) — declare A, B, or C **in writing** | `book_source` tag on every artifact |
|
|||
|
|
| 5 | **RegimeBridge** (§6) — explicit total mapping | Test: every MARAS regime maps; confidence carried |
|
|||
|
|
| 6 | **ActualsLoader** (§4) — S1/S2/S3/S4 into `MarketWorldState` | Replay determinism ×2 (byte-identical minus ts) |
|
|||
|
|
| 7 | **EcologyFitter** (§7) — fit params, **keep all agent types** | Test: fitting cannot reduce the agent-type count |
|
|||
|
|
| 8 | **ScenarioLibrary** — CLASS-M/X hours as suites | CLASS-X magnitudes unfiltered |
|
|||
|
|
| 9 | **Implement the 5 missing agent roles** | Ecology completeness |
|
|||
|
|
| 10 | **Scale the inner loop** (§7) — vectorize ecology, cascade screening | Measured trials/sec, not adjectives |
|
|||
|
|
| 11 | **MODE 1 → manifold** (§0.5): selector emits variance + support + envelope, not a leaderboard | A thin cell must report itself thin |
|
|||
|
|
| 12 | **MODE 2 → localization + OOD verdict** (§0.5): live state query, `OUT_OF_DISTRIBUTION` falls back to doctrinal policy | Test: a novel regime must NOT get a confident recommendation |
|
|||
|
|
| 13 | **Gate-M** (README §F): replay verify → dual-leverage litmus → DARK shadow vs naive → risk-gate mutation litmus → Rule 10 | Money |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 11. THE STANDING GUARDRAILS (Fable's, HJ's, and the house's)
|
|||
|
|
|
|||
|
|
1. **Rule 3 supreme**: replay correctness before search depth. A wrong CWM plus a
|
|||
|
|
deep search is confident nonsense, and confident nonsense is the most expensive
|
|||
|
|
thing we can build.
|
|||
|
|
2. **Rule 10 supreme**: shadow vs live diverges → **trust live**, stand down, file
|
|||
|
|
the discrepancy row.
|
|||
|
|
3. **Mutation litmus everywhere**: 1,186 green tests prove nothing until breaking
|
|||
|
|
the implementation turns one red. Fees, latency, risk-gate constraints, ecology
|
|||
|
|
size — each must have a test that dies when it is sabotaged.
|
|||
|
|
4. **"It looks finished" is the warning sign, not the green light.**
|
|||
|
|
5. **The ecology stays.** If a future refactor makes the counterparties smaller,
|
|||
|
|
fewer, or tamer in the name of realism, it has removed the edge and kept the
|
|||
|
|
costs. That refactor is wrong. This line is the reason this document exists.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
*Written by Claude (Opus 4.8) on Fable's context, at HJ's order, 2026-07-13.
|
|||
|
|
Every path, line number, row count, and measured value herein was verified against
|
|||
|
|
the live system on the night of writing. Where a number was not verified, it says
|
|||
|
|
so.*
|