diff --git a/prod/docs/SPEC_MALKHUT_ACTUALS_INTAKE.md b/prod/docs/SPEC_MALKHUT_ACTUALS_INTAKE.md new file mode 100644 index 0000000..8a2de46 --- /dev/null +++ b/prod/docs/SPEC_MALKHUT_ACTUALS_INTAKE.md @@ -0,0 +1,465 @@ +# SPEC — MALKHUT ACTUALS INTAKE ("game it under actuals; the ecology stays") + +**Author:** Claude (Opus 4.8), continuing Fable's review. **Date:** 2026-07-13. +**Addressee:** mimo (mm_, `mm_ob_fill_sim`). **Ordered by:** HJ. +**Status:** SPEC. Companion to `MALKHUT/README.md` (§ADDENDUM) and +`prod/docs/SPEC_UV_SMART_EXEC_MM.md`. + +--- + +## 0. THE ONE-LINE LAW + +> **Actuals CALIBRATE the game. The ecology PLAYS it. Real tape grounds the +> adversaries; it NEVER replaces them.** + +Every instruction below serves that sentence. If a change would let recorded +history *substitute* for adversarial simulation, it is wrong, no matter how +"realistic" it looks. Replaying our own tape teaches MALKHUT what happened once. +The ecology is what lets it out-play what *could* happen — billions of +strategy × counterparty × regime combinations, in parallel, none of which are in +any tape. **The ecology is the edge. It is not a placeholder for missing data.** + +--- + +## 0.5 THE TWO MODES (operator, 2026-07-13) — THE ARCHITECTURAL SPINE + +MALKHUT is **two engines sharing one CWM**, and every module below belongs to one +of them. This is the frame that makes "actuals" and "ecology" complementary +instead of competing — and it is, structurally, **train vs inference**. + +``` +┌─ MODE 1: EXPLORE ───────────────────────────────────────────────────┐ +│ "Search/play over SYNTHETICS: fast, parallel, acquisitive of BEST" │ +│ │ +│ ECOLOGY-DOMINANT. Adversaries react. Regimes swept, not sampled. │ +│ Billions of strategy × counterparty × regime combinations. │ +│ Actuals appear ONLY as calibration (fee tables, latency CDF, │ +│ agent base-rates, plausible bounds). NOT as the arena. │ +│ │ +│ Output: POLICY POOL — a ranked, regime-indexed library of │ +│ strategies with known behaviour under known adversary mixes. │ +│ Cadence: OFFLINE, hours-to-days, unbounded compute. │ +│ Success metric: coverage + robustness, NOT backtest PnL. │ +└──────────────────────────────┬───────────────────────────────────────┘ + │ policy pool + performance manifold + ▼ +┌─ MODE 2: RECOMMEND (inference/query) ───────────────────────────────┐ +│ "Search over LIVE, ACTUAL (OBF) IRL conditions as INPUTS, to │ +│ extrapolate the closest-to-absolute-best strategy → RECOMMEND" │ +│ │ +│ ACTUALS-DOMINANT. The live book/regime/funding/latency IS the │ +│ query. The ecology is now a PRIOR over who is on the other side │ +│ right now, not a population to be swept. │ +│ │ +│ Given: live MarketWorldState (S1 book, S2 funding/dvol, S3 regime, │ +│ S4 latency, account, intent) │ +│ Ask: which policy in the pool is nearest-optimal for THIS state? │ +│ Output: a RECOMMENDATION (policy + its expected behaviour + │ +│ confidence + the nearest explored neighbours it interpolates│ +│ between). │ +│ Cadence: ONLINE, ≤25 ms planning inside ≤100 ms end-to-end. │ +│ Success metric: live outcome ≈ predicted outcome (Rule 10). │ +└──────────────────────────────────────────────────────────────────────┘ +``` + +### Why this is the right shape + +- **Mode 1 without Mode 2** is a beautiful simulator nobody trades. +- **Mode 2 without Mode 1** is a lookup table over history — the overfit that + kills quants. It can only recommend what already happened. +- **Together**: Mode 1 explores a space *vastly larger than any tape can contain* + (that is the edge — we out-play participants who only ever fit history), and + Mode 2 *localizes* the live moment inside that explored space and returns the + best-known play, **including for market states we have never actually seen**, + because Mode 1 has *already played them*. + +**The word "extrapolate" in the operator's phrasing is load-bearing.** Mode 2's +job is NOT nearest-neighbour lookup. It is: *place the live state inside the +performance manifold Mode 1 built, and interpolate/extrapolate the best play.* +That requires Mode 1 to output a **manifold**, not a leaderboard: + +| Mode 1 must emit | Not just | +|---|---| +| policy × (regime, spread, depth, toxicity, latency, funding, inventory) → expected outcome **+ variance** | "policy #7 scored 8,628" | +| the **boundaries** of where each policy was tested (extrapolation beyond = flagged) | a single global champion | +| **why** a policy wins (which adversary it beats, which it loses to) | an opaque score | + +### Consequences for the build + +1. **`ScenarioLibrary` must SWEEP, not sample.** Mode 1's job is coverage of the + state space, including regions our tape never visited. Tape-derived scenarios + (CLASS-M/X) are the *anchors*; the sweep fills between and beyond them. +2. **The performance matrix in `training/selector.py` becomes the manifold** — + and it must carry **confidence + support count + distance-to-nearest-explored** + per cell. A recommendation from a thinly-explored cell must SAY SO. +3. **Mode 2 must refuse to extrapolate too far.** If the live state is outside + Mode 1's explored envelope (novel regime, latency beyond p99.9, depth collapse + never swept), the honest answer is **"OUT OF DISTRIBUTION — fall back to the + doctrinal simple policy"**, not a confident recommendation. This is the same + law as `INDETERMINATE`: *unknown is not flat, and unknown is not "best guess".* + Wire it as an explicit `RecommendationConfidence.OUT_OF_DISTRIBUTION` verdict + that the risk gate honours. +4. **Mode 2's live search is a LOCALIZATION, not a fresh MCTS from scratch** — the + ≤25 ms budget buys you refinement around a pool policy, not global search. The + pool is the compressed product of the billions of Mode-1 trials; Mode 2 spends + its milliseconds *choosing and adapting*, not re-deriving. +5. **Mode 2's own outcomes feed back into Mode 1** as new anchors (live + discrepancies → `live_discrepancies` table → next sweep is denser where we were + wrong). That is the learning loop, and Rule 10 governs it: **where shadow and + live disagree, live is right and the manifold gets corrected.** + +### Module ownership by mode + +| Module | Mode 1 (EXPLORE) | Mode 2 (RECOMMEND) | +|---|---|---| +| `cwm/core.py` | ✅ the arena | ✅ the local rollout model | +| `counterparties.py` (ecology) | ✅ **swept adversary population** | ✅ **prior over current opponents** | +| `training/cma_trainer.py`, `generator.py` | ✅ | — | +| `ScenarioLibrary` (§4) | ✅ sweeps | — | +| `training/registry.py` (policy pool) | ✅ writes | ✅ reads | +| `training/selector.py` (→ **manifold**) | ✅ builds | ✅ queries | +| `ActualsLoader` (§4) | ⚠️ calibration only | ✅ **the live query itself** | +| `planner/sm_mcts.py` | ✅ deep, unbounded | ✅ **≤25 ms localization** | +| `risk/gate.py` | ✅ constrains training | ✅ hard veto on recommendation | + +--- + +## 1. FINDING #0 — THE FEE BUG IS REAL, AND MEASURED + +Fable suspected a 10× unit slip. **Our own fills confirm it.** + +```sql +SELECT liquidity_side, order_type, count() n, avg(fee_bps) +FROM dolphin.trade_execution_quality WHERE fee_bps IS NOT NULL GROUP BY 1,2 +-- TAKER MARKET 1455 rows avg = 5.016 bps (min 5.000, max 6.649) +``` + +| Where | Code says | Reality (our fills / venue docs) | Factor | +|---|---|---|---| +| `malkhut/training/asset_classification.py:164` (BingX) | `default_taker_fee_bps=0.5` | **5.0** | **10×** | +| `…:154` (Binance) | `default_taker_fee_bps=0.4` | ~4.5 | ~10× | +| `…:174` (Bybit) | `default_taker_fee_bps=0.06` | ~5.5 | ~90× | +| `…:384` (default profile) | `taker_fee_bps=0.5, maker_fee_bps=-0.2` | taker 5.0; maker ≈ +2.0 (BingX perp maker is POSITIVE, not a rebate) | 10× + sign | + +**Impact:** these feed `VenueRules.taker_fee_bps` / `maker_fee_bps` +(`malkhut/state.py:109-121`) → the CWM reward → `w_fee_quality`. CMA-ES has been +optimizing against fees an order of magnitude too cheap. **Every policy trained so +far is suspect** — cheap fees reward overtrading, churn, and cross-spread +aggression that real friction annihilates. This is the exact failure the +2026-07-10 venue-friction audit exists to prevent. + +**ACTION (do this first, before any other work):** +1. Fix the four sites above. Source of truth = `dolphin.trade_execution_quality` + (`fee_bps`, `liquidity_side`, `order_type`), not vendor marketing pages. +2. Add a **mutation-litmus test**: set `taker_fee_bps` to 0.5 → a test asserting + "policy PnL under realistic friction" must go **RED**. If nothing breaks when + fees change 10×, the reward function isn't actually using them. +3. Re-run every CMA-ES benchmark. The "8,628 / 100.4 bps PnL" headline number is + void until it is re-measured at 5 bps taker. +4. Maker fee sign: verify against a real maker fill before trusting a rebate. + **We have zero MAKER rows in `trade_execution_quality`** (all 1455 are TAKER + MARKET — because UV/F5 only sends MARKET). Until SMART-EXEC produces real + maker fills, the maker fee is an *assumption*: mark it as such in code with a + `# UNVERIFIED — no maker fills on record as of 2026-07-13` comment. + +--- + +## 2. THE DATA SOURCES — WHAT EXISTS, VERIFIED TONIGHT + +All queried live on `localhost:8123`. Row counts are real. + +| # | Source | Rows | Columns (exact) | State | +|---|---|---|---|---| +| **S1** | `dolphin.obf_universe` | **15,329,163,875** | `ts` DateTime64(3), `symbol`, `spread_bps` f32, `depth_1pct_usd` f64, `depth_quality` f32, `fill_probability` f32, `imbalance` f32, `best_bid` f64, `best_ask` f64, `n_bid_levels` u8, `n_ask_levels` u8 | **LIVE, the motherlode** | +| **S2** | `dolphin.exf_data` | 22,968,094 | `ts` DateTime64(6), `funding_rate` f32, `dvol` f32, `fear_greed` f32, `taker_ratio` f32 | LIVE | +| **S3** | `dolphin.maras_fingerprint` | 1,117,722 | `ts`, `regime` (LowCard), `regime_idx` u8, `confidence`, `final_score`, `conflict_level`, `tier_exf`, `tier_eigen`, `tier_btc`, `tier_esof`, `tier_micro` (+ `_confidence` each), `s_exf_funding_bp`, … | LIVE | +| **S4** | `dolphin.eigen_scans` | 1,529,803 | `ts`, `scan_number` u32, `vel_div` f32, `w50_velocity`, `w750_velocity`, `instability_50`, **`scan_to_fill_ms`**, **`step_bar_ms`**, `scan_uuid` | LIVE — **and it is the latency oracle** | +| **S5** | `dolphin.trade_execution_quality` | 8,006 | `trade_id`, `asset`, `side`, `client_order_id`, `venue_order_id`, `order_type`, **`liquidity_side`**, **`fee_bps`**, `fill_quality_score`, `commission_quote`, `fee_rate`, … | LIVE — **fee/fill ground truth** | +| **S6** | `dolphin.trade_events` | (BLUE's live trades) | `ts`, `trade_id`, `asset`, `side`, `entry_price`, `exit_price`, `pnl`, `pnl_pct`, `exit_reason`, `vel_div_entry`, `boost_at_entry` | LIVE | +| **S7** | `dolphin_uv.exec_journal` | growing | `kind`, `timestamp`, `scan_number`, `intent_id`, `trade_id`, `slot_id`, `asset`, `side`, `action`, `reference_price`, `target_size`, `leverage`, `u_prefix_client_id`, `promo_metadata` (JSON) | LIVE | +| **S8** | `dolphin_uv.tp_exit_ingress` / `max_hold_ingress` | new (2026-07-13) | full per-scan exit diagnostics: `tp_effective_pct`, `tp_mod_factor`, `tp_floor_armed`, `cascade_count`, `imbalance_ma5`, `branch`, … | LIVE as of today | +| **S9** | `dolphin.esof_advisory` | **0** | schema exists (`dow`, `session`, `moon_illumination`, `slot_wr_pct`, …) | ⚠️ **EMPTY** | +| **S10** | `dolphin.obf_fast_intrade` | **0** | schema exists | ⚠️ **EMPTY** | + +**Do not spec against S9/S10 as if they were populated.** Either they get a +writer, or they are out of scope. Say so out loud rather than building an intake +for a table that never delivers a row (that is how the missing-CH-table blackout +happened to UV; the fix cost a day). + +--- + +## 3. FINDING #1 — THE BOOK-FIDELITY GAP (read this twice) + +This is the **central technical problem** of the whole integration, and it is not +optional to solve. + +**MALKHUT wants a ladder.** `malkhut/state.py:131`: +```python +class OrderBookState: + bids: Tuple[PriceLevel, ...] # price + qty, per level + asks: Tuple[PriceLevel, ...] +``` +The CWM's queue-position model, `JOIN_QUEUE`, `LADDER`, `ICEBERG`, queue-ahead +estimates, and adverse-selection accounting all depend on **per-level depth**. + +**Our tape has no ladder.** `dolphin.obf_universe` gives, per symbol per ~sub-second +tick: `best_bid`, `best_ask`, `spread_bps`, `depth_1pct_usd` (one aggregate +number), `depth_quality`, `imbalance`, `fill_probability`, `n_bid_levels`, +`n_ask_levels` (**counts**, not the levels themselves). + +That is **L1 + shape summary**, not L2. 15.3 billion rows of it — enormously +valuable, but it cannot be *decoded* back into a ladder. Information that was +never recorded cannot be recovered. + +### Three honest options — pick explicitly, do not drift + +| Option | What it is | Cost | Fidelity | +|---|---|---|---| +| **A. Synthetic ladder w/ declared prior** | Reconstruct levels from `best_bid/ask` + `depth_1pct_usd` + `n_*_levels` + `imbalance` via an explicit shape model (e.g. exponential decay of qty over levels, calibrated so Σqty within 1% = `depth_1pct_usd`, level count = `n_bid_levels`). | LOW | **Approximate.** Queue position becomes a *model*, not a measurement. MUST be labeled as such everywhere it flows. | +| **B. Start recording real L2 now** | New writer → `dolphin_malkhut.book_l2` (or a zinc region + spooler). Depth-20 snapshots at OBF cadence. | MEDIUM (new hose + storage; L2 at 50 symbols × sub-second is BIG — see §7 storage) | **Truth**, but only from the day it starts. Cannot backfill history. | +| **C. hftbacktest with external L2** | Use a third-party L2 feed (Binance archival) for CWM calibration; keep obf_universe for regime/context. | MEDIUM | Truth, but it is *Binance's* book, not BingX's — venue microstructure differs. | + +**RECOMMENDATION: A + B in parallel, and say which one a given result came from.** +- **A** unblocks immediate use of 15.3B rows of history for *distributional* + calibration (spread regimes, depth regimes, imbalance dynamics, fill-probability + priors) — where the ladder shape matters less than the aggregate. +- **B** starts the clock on real queue-truth for the queue-sensitive claims + (`JOIN_QUEUE`, maker fills, queue-ahead) — the very claims SMART-EXEC needs. +- **NEVER** let an A-derived queue-position estimate be reported as a measured + fill probability. Tag every artifact: `book_source ∈ {SYNTH_A, REAL_L2_B, EXT_C}`. + Rule 3 (replay correctness before search depth) means exactly this. + +**Litmus:** when B has ≥ 1 week of real L2, re-run A-calibrated policies against +B-truth. The delta *is* the measurement of how much the synthetic prior lied. If +that delta is large, every A-era conclusion is downgraded to a hypothesis. + +--- + +## 4. THE PLUG MAP — WHAT GOES WHERE, EXACTLY + +`MarketWorldState` (`malkhut/state.py:257`) is the CWM root. It already has the +right holes. Fill them from our sources: + +| `MarketWorldState` field | Line | Source | Transform | +|---|---|---|---| +| `book: OrderBookState` | 262 | **S1** `obf_universe` | §3 Option A synth (or B when live). `best_bid`/`best_ask` direct; ladder via declared prior; `symbol` join key. | +| `account: AccountState` | 263 | **DITAv2 ASEx account core** (`asex_account` / `AccountProjectionV2`) — **CONSUMER ONLY** (two-cores-share-nothing law) | For offline training: reconstruct from `dolphin.account_events` / `dolphin_uv.exec_journal`. Never a second writer. | +| `open_orders` | 264 | live: venue adapter; training: **S7** `exec_journal` + **S5** `trade_execution_quality` (client_order_id join) | | +| `trade_path: TradePathState` | 265 | **S8** `tp_exit_ingress` + **S6** `trade_events` | `mae_bps`/`mfe_bps`/`time_in_loss_s` etc. reconstructable from tp_exit_ingress per-scan stamps + trade lifecycle. **`dolphin_regime_score` ← S3; `book_imbalance` ← S1 `imbalance`; `orderflow_toxicity` ← derive (see §5)**. | +| `intent: ExecutionIntent` | 266 | **DITAv2 `KernelIntent`** (binding, README §B.1) | Map ENTER/EXIT + `reference_price`/`target_size`/`leverage`/`metadata.promo_client_id`. `urgency` ← SMART-EXEC urgency class (see §8). | +| **`funding_bps`** | 268 | **S2** `exf_data.funding_rate` | ×10⁴ → bps. Nearest-ts join. | +| **`volatility_state`** | 269 | **S2** `exf_data.dvol` | The sacred vol number. **Same field the gate uses** (`DOLPHIN_VOL_P60_THRESHOLD`, doctrinal 0.00026414). | +| **`market_regime`** | 270 | **S3** `maras_fingerprint.regime` | ⚠️ **taxonomy mismatch — see §6.** | +| **`feed_latency_ms`** | 272 | **S4** `eigen_scans.step_bar_ms` | p50 = **0.08 ms**. | +| **`order_latency_ms`** | 273 | **S4** `eigen_scans.scan_to_fill_ms` | ⚠️ **SAMPLE THE DISTRIBUTION, NOT THE MEAN — see §5.** | +| `venue: VenueRules` | 261 | **S5** for fees (§1); `prod/bingx/` for tick/lot/min_notional | Fee fix is blocking. | + +### Where the code changes go + +| New module | Path | Job | +|---|---|---| +| `ActualsLoader` | `malkhut/data/actuals.py` **(new pkg `malkhut/data/`)** | CH HTTP reader → typed frames. One method per source S1–S8. Polars, chunked, no full-table loads. | +| `BookSynthesizer` | `malkhut/data/book_synth.py` | §3 Option A. Declared prior, unit-tested against any real L2 we get. Emits `book_source` tag. | +| `LatencyOracle` | `malkhut/data/latency.py` | §5. Empirical CDF sampler, seed-pinned. | +| `EcologyFitter` | `malkhut/data/ecology_fit.py` | §7. Fits counterparty params to tape. **Does not create or delete agent types.** | +| `RegimeBridge` | `malkhut/data/regime_bridge.py` | §6. MARAS ↔ MALKHUT regime mapping, explicit and tested. | +| `ScenarioLibrary` | `malkhut/data/scenarios.py` | §7. CLASS-M/CLASS-X hour selection from tape → adversarial scenario suites. | + +Wire them at: `malkhut/training/cma_trainer.py` (scenario source), +`malkhut/cwm/core.py:transition` (latency + venue rules injection), +`malkhut/counterparties.py` (fitted params in, agent classes unchanged). + +--- + +## 5. FINDING #2 — LATENCY IS NOT A NUMBER, IT IS A MONSTER-TAILED DISTRIBUTION + +Measured tonight from `dolphin.eigen_scans` (`scan_to_fill_ms`, n = 1.5M): + +| p50 | p95 | p99 | max | +|---|---|---|---| +| **49.9 ms** | **596.9 ms** | **10,793.9 ms** | **450,115.9 ms** (7.5 minutes) | + +`step_bar_ms` p50 = **0.08 ms**. + +The README's "typical_latency_ms: 100" for BingX is a **fiction that will get us +killed**: it is 2× the median and **0.9% of the mass is beyond 10 SECONDS**. A +policy trained on constant-100ms latency has never met the market that actually +fills our orders. The p99 is what turns a maker quote into an adverse-selected +gift, and it is exactly where the "why did I get filled?" question (Rule 2) gets +its ugliest answer. + +**ACTION:** `LatencyOracle` must sample the **empirical CDF**, per-regime where +possible (latency and stress correlate — verify), seed-pinned for determinism +(Rule 9 + our seed-pin doctrine). Constant-latency mode may exist **only** as a +labeled ablation, never as the default. Add a scenario class: **`LATENCY_STORM`** +(sample exclusively from > p95) — if a policy's edge evaporates there, we need to +know before capital does. + +--- + +## 6. FINDING #3 — REGIME TAXONOMY MISMATCH + +MARAS actually emits (verified, `dolphin.maras_fingerprint`, 1.1M rows): + +| `regime_idx` | `regime` | rows | +|---|---|---| +| 1 | BEARISH | 72,479 | +| 2 | CHOPPY_BEARISH | 516,477 | +| 3 | CHOPPY | 295,366 | +| 4 | SIDEWAYS | 201,358 | +| 5 | CHOPPY_BULLISH | 32,046 | + +(idx 0, 6, 7 unobserved in this window — the taxonomy has room MARAS has not used.) + +MALKHUT's `StrategySelector` uses a **different, invented** set: `trending_up`, +`trending_down`, `high_volatility`, `low_volatility`, `mean_reverting`, +`momentum`, `choppy`, `liquidity_hole`, `normal`. + +**These are two different languages.** A performance matrix keyed on MALKHUT's +regimes cannot be looked up from a MARAS fingerprint without a mapping, and an +implicit/lossy mapping will silently mis-select strategies in production. + +**ACTION:** `RegimeBridge` (`malkhut/data/regime_bridge.py`) with an **explicit, +tested, total** mapping MARAS→MALKHUT. Where MALKHUT has an axis MARAS lacks +(e.g. `liquidity_hole`), derive it from S1 (`depth_quality`, `spread_bps` +percentiles) and **name the derivation**. Where MARAS has confidence +(`confidence`, `conflict_level`), **carry it through** — a low-confidence regime +tag must reach the planner as low-confidence, not as a hard label. Prefer +extending MALKHUT to speak MARAS over inventing a third dialect. + +**Also plumb the MARAS tiers** (`tier_exf`, `tier_eigen`, `tier_btc`, +`tier_esof`, `tier_micro` + confidences) into the feature vector — that is a +5-tier ensemble view of the market the planner currently cannot see at all, and +it is *already computed and stored*. Free signal. + +--- + +## 7. THE ECOLOGY — HOW ACTUALS FEED IT WITHOUT REPLACING IT + +**This section is the point of the whole document. Read it as law.** + +`malkhut/counterparties.py` defines the agent ecology (`AgentRole`, +`state.py:61`): PASSIVE_MAKER, TOXIC_TAKER, LATENCY_ARB, MOMENTUM_TAKER, +MEAN_REVERSION_TAKER, INVENTORY_MM, LIQUIDATION_FLOW, NOISE_TRADER, +STALE_QUOTE_ATTACKER. **Nine roles declared, 4 implemented.** + +### What actuals DO to the ecology + +| Do | How | +|---|---| +| **Fit each agent's parameters** | `EcologyFitter` estimates population parameters from tape: TOXIC_TAKER arrival intensity ← S1 `fill_probability` collapse + adverse post-fill drift in S5/S6; LATENCY_ARB ← S4 latency tails vs price moves; LIQUIDATION_FLOW ← S2 funding extremes + cascade signatures in S8 `cascade_count`; INVENTORY_MM ← S1 `imbalance` mean-reversion; MOMENTUM_TAKER ← S4 `vel_div` / `w50_velocity`. | +| **Set the population MIX per regime** | Which agents dominate under BEARISH vs CHOPPY (S3) is an empirical question our tape can answer. The mix becomes regime-conditional. | +| **Supply the stress scenarios** | CLASS-M / CLASS-X hours (measured hostile windows) → `ScenarioLibrary`. **CLASS-X law: extreme magnitudes are a FEATURE channel, never filtered.** | +| **Bound the plausible** | Tape says what a real book *can* do; adversaries should be allowed to be *worse*, but their base rates must be anchored. | + +### What actuals MUST NOT DO + +| Never | Why | +|---|---| +| **Replace an agent with a tape replay** | A replayed counterparty cannot react to us. The entire adverse-selection question ("why did I get filled?") requires an opponent that *responds*. A tape is a corpse; the ecology is an opponent. | +| **Delete an agent type for lack of data** | The 5 unimplemented roles are *hypotheses about who is on the other side*. Absence of evidence in our thin tape is not evidence of absence in the market. Implement them; fit what you can; make the rest adversarial priors. | +| **Restrict the game to observed histories** | Billions of combinations, most of which never occurred, is the *product*. Recorded tape is a measure-zero slice of the strategy × counterparty × regime space. Optimizing only over what happened is exactly the overfit that kills quants. | +| **Filter outliers** | CLASS-X law. The 450-second latency tail and the 7.5-minute stop are not noise — they are the market's teeth. | + +### The scale ambition (operator's stated aim) + +Femtosecond-scale trials, parallel, over **billions** of combinations. Current: +5.3 µs/CWM-transition (numba), 16 ms/episode (parallel), 7× on 8 workers. +Billions of combinations × ~10³ steps at 5.3 µs = ~10¹² × 5.3 µs ≈ **months of +single-box CPU.** So: + +1. **Batch/vectorize the counterparty step** (the ecology is the inner loop — + make it a matrix op over the agent population, not a Python loop over agents). +2. **GPU or massively-parallel path** for the rollout kernel (the batch MCTS + kernel already exists — push it further). +3. **Cheap-then-expensive cascade**: screen billions with a cheap surrogate + (vectorized reward, shallow rollout), promote survivors to full CWM depth. + Do NOT run 5.3 µs × depth-3 MCTS on every one of a billion candidates. +4. **Ray path already exists** (`training/ray_eval.py`) — that is the horizontal + scaling seam when the box runs out. +5. Honest note: "femtosecond" is a *metaphor for parallel breadth*, not a + physical claim (a CPU cycle is ~300 picoseconds; a femtosecond is 10⁻¹⁵ s — + light travels 0.3 µm in one). State the real target: **N trials/second at + fidelity F**, and measure it. Otherwise we cannot tell progress from poetry. + +--- + +## 8. INTAKE CONTRACT ALIGNMENT (with SMART-EXEC + UV) + +MALKHUT's planner and `SPEC_UV_SMART_EXEC_MM.md` describe the **same seat** at the +venue-adapter boundary. Do not build two doorframes. + +- **Input**: DITAv2 `KernelIntent` + `guideline_price` + `urgency_class` + (CATASTROPHIC / PROTECT / HARVEST / ROTATE / ACQUIRE). Map `urgency_class` → + `ExecutionIntent.urgency` (`state.py:246`) and `prefer_maker` (:250). + **CATASTROPHIC ⇒ the planner is BYPASSED. No cleverness on the stop path, ever.** +- **`u-` clientOrderId prefix is LAW** (T9 seam / README §B.2). +- **DUAL-LEVERAGE LAW** (README §B.3): conviction leverage [0.5,9.0] sizes + quantity; venue leverage = `prod/bingx/leverage.py` mapping → int [1,3]. + Litmus: 0.5→1, 4.75→2, 8.0→3, 9.0→3. **Note:** we observed a live 3→2 clamp on + 2026-07-13 (handover anomaly #2) — reconcile before MALKHUT sizes anything. +- **Execution-truth doctrine is inherited wholesale** (`bdc54fb`): NOT_ATTEMPTED / + REFUSED → rollback sound; **INDETERMINATE → never**. MALKHUT's venue adapter + (`malkhut/venue/bingx/adapter.py`) wraps DITAv2 and therefore inherits the + fences — **verify with a test, do not assume.** See + `COMPREHENSIVE_UV_EDGE_CASE_THEORETICALS.md` §11 M1–M7 for the unaudited seams. +- **No reconcilers.** Bounded, read-only, own-clientOrderId point lookups only. + +--- + +## 9. PERSISTENCE + PROVENANCE + +- Namespace `dolphin_malkhut` (approved). **NO TTL** (retention doctrine). +- **DDL ships WITH the code, and the applier's verify-set must require every new + table** — the 2026-07-13 lesson: two UV tables existed as `.sql` files for days + while the runner 404'd every scan and journaled nothing. Pattern to copy: + `prod/clickhouse/uv/apply_uv_ddl.py` (`EXPECTED_TABLES`). +- **Every row carries provenance**: `book_source` (SYNTH_A/REAL_L2_B/EXT_C), + `latency_model` (EMPIRICAL_CDF/CONST), `fee_table_version`, `ecology_fit_version`, + `policy_version`, `seed`. A result whose provenance is unknown is not a result. +- **Cross-reference law**: `fulfilment_decisions.intent_id` + `trade_id` join to + `dolphin_uv.exec_journal` ⇄ venue order history. End-to-end or it is a black box. + +--- + +## 10. ORDER OF WORK (do not reorder) + +| # | Task | Gate | +|---|---|---| +| 1 | **Fix the fees** (§1) + mutation-litmus test | A 10× fee change must break a test | +| 2 | **Re-baseline** every CMA-ES number at real friction | Old headline numbers marked VOID | +| 3 | **LatencyOracle** (§5) + `LATENCY_STORM` scenario | Constant-latency is an ablation, not a default | +| 4 | **Decide the book question** (§3) — declare A, B, or C **in writing** | `book_source` tag on every artifact | +| 5 | **RegimeBridge** (§6) — explicit total mapping | Test: every MARAS regime maps; confidence carried | +| 6 | **ActualsLoader** (§4) — S1/S2/S3/S4 into `MarketWorldState` | Replay determinism ×2 (byte-identical minus ts) | +| 7 | **EcologyFitter** (§7) — fit params, **keep all agent types** | Test: fitting cannot reduce the agent-type count | +| 8 | **ScenarioLibrary** — CLASS-M/X hours as suites | CLASS-X magnitudes unfiltered | +| 9 | **Implement the 5 missing agent roles** | Ecology completeness | +| 10 | **Scale the inner loop** (§7) — vectorize ecology, cascade screening | Measured trials/sec, not adjectives | +| 11 | **MODE 1 → manifold** (§0.5): selector emits variance + support + envelope, not a leaderboard | A thin cell must report itself thin | +| 12 | **MODE 2 → localization + OOD verdict** (§0.5): live state query, `OUT_OF_DISTRIBUTION` falls back to doctrinal policy | Test: a novel regime must NOT get a confident recommendation | +| 13 | **Gate-M** (README §F): replay verify → dual-leverage litmus → DARK shadow vs naive → risk-gate mutation litmus → Rule 10 | Money | + +--- + +## 11. THE STANDING GUARDRAILS (Fable's, HJ's, and the house's) + +1. **Rule 3 supreme**: replay correctness before search depth. A wrong CWM plus a + deep search is confident nonsense, and confident nonsense is the most expensive + thing we can build. +2. **Rule 10 supreme**: shadow vs live diverges → **trust live**, stand down, file + the discrepancy row. +3. **Mutation litmus everywhere**: 1,186 green tests prove nothing until breaking + the implementation turns one red. Fees, latency, risk-gate constraints, ecology + size — each must have a test that dies when it is sabotaged. +4. **"It looks finished" is the warning sign, not the green light.** +5. **The ecology stays.** If a future refactor makes the counterparties smaller, + fewer, or tamer in the name of realism, it has removed the edge and kept the + costs. That refactor is wrong. This line is the reason this document exists. + +--- + +*Written by Claude (Opus 4.8) on Fable's context, at HJ's order, 2026-07-13. +Every path, line number, row count, and measured value herein was verified against +the live system on the night of writing. Where a number was not verified, it says +so.*