# SPEC — MALKHUT ACTUALS INTAKE ("game it under actuals; the ecology stays") **Author:** Claude (Opus 4.8), continuing Fable's review. **Date:** 2026-07-13. **Addressee:** mimo (mm_, `mm_ob_fill_sim`). **Ordered by:** HJ. **Status:** SPEC. Companion to `MALKHUT/README.md` (§ADDENDUM) and `prod/docs/SPEC_UV_SMART_EXEC_MM.md`. --- ## 0. THE ONE-LINE LAW > **Actuals CALIBRATE the game. The ecology PLAYS it. Real tape grounds the > adversaries; it NEVER replaces them.** Every instruction below serves that sentence. If a change would let recorded history *substitute* for adversarial simulation, it is wrong, no matter how "realistic" it looks. Replaying our own tape teaches MALKHUT what happened once. The ecology is what lets it out-play what *could* happen — billions of strategy × counterparty × regime combinations, in parallel, none of which are in any tape. **The ecology is the edge. It is not a placeholder for missing data.** --- ## 0.5 THE TWO MODES (operator, 2026-07-13) — THE ARCHITECTURAL SPINE MALKHUT is **two engines sharing one CWM**, and every module below belongs to one of them. This is the frame that makes "actuals" and "ecology" complementary instead of competing — and it is, structurally, **train vs inference**. ``` ┌─ MODE 1: EXPLORE ───────────────────────────────────────────────────┐ │ "Search/play over SYNTHETICS: fast, parallel, acquisitive of BEST" │ │ │ │ ECOLOGY-DOMINANT. Adversaries react. Regimes swept, not sampled. │ │ Billions of strategy × counterparty × regime combinations. │ │ Actuals appear ONLY as calibration (fee tables, latency CDF, │ │ agent base-rates, plausible bounds). NOT as the arena. │ │ │ │ Output: POLICY POOL — a ranked, regime-indexed library of │ │ strategies with known behaviour under known adversary mixes. │ │ Cadence: OFFLINE, hours-to-days, unbounded compute. │ │ Success metric: coverage + robustness, NOT backtest PnL. │ └──────────────────────────────┬───────────────────────────────────────┘ │ policy pool + performance manifold ▼ ┌─ MODE 2: RECOMMEND (inference/query) ───────────────────────────────┐ │ "Search over LIVE, ACTUAL (OBF) IRL conditions as INPUTS, to │ │ extrapolate the closest-to-absolute-best strategy → RECOMMEND" │ │ │ │ ACTUALS-DOMINANT. The live book/regime/funding/latency IS the │ │ query. The ecology is now a PRIOR over who is on the other side │ │ right now, not a population to be swept. │ │ │ │ Given: live MarketWorldState (S1 book, S2 funding/dvol, S3 regime, │ │ S4 latency, account, intent) │ │ Ask: which policy in the pool is nearest-optimal for THIS state? │ │ Output: a RECOMMENDATION (policy + its expected behaviour + │ │ confidence + the nearest explored neighbours it interpolates│ │ between). │ │ Cadence: ONLINE, ≤25 ms planning inside ≤100 ms end-to-end. │ │ Success metric: live outcome ≈ predicted outcome (Rule 10). │ └──────────────────────────────────────────────────────────────────────┘ ``` ### Why this is the right shape - **Mode 1 without Mode 2** is a beautiful simulator nobody trades. - **Mode 2 without Mode 1** is a lookup table over history — the overfit that kills quants. It can only recommend what already happened. - **Together**: Mode 1 explores a space *vastly larger than any tape can contain* (that is the edge — we out-play participants who only ever fit history), and Mode 2 *localizes* the live moment inside that explored space and returns the best-known play, **including for market states we have never actually seen**, because Mode 1 has *already played them*. **The word "extrapolate" in the operator's phrasing is load-bearing.** Mode 2's job is NOT nearest-neighbour lookup. It is: *place the live state inside the performance manifold Mode 1 built, and interpolate/extrapolate the best play.* That requires Mode 1 to output a **manifold**, not a leaderboard: | Mode 1 must emit | Not just | |---|---| | policy × (regime, spread, depth, toxicity, latency, funding, inventory) → expected outcome **+ variance** | "policy #7 scored 8,628" | | the **boundaries** of where each policy was tested (extrapolation beyond = flagged) | a single global champion | | **why** a policy wins (which adversary it beats, which it loses to) | an opaque score | ### Consequences for the build 1. **`ScenarioLibrary` must SWEEP, not sample.** Mode 1's job is coverage of the state space, including regions our tape never visited. Tape-derived scenarios (CLASS-M/X) are the *anchors*; the sweep fills between and beyond them. 2. **The performance matrix in `training/selector.py` becomes the manifold** — and it must carry **confidence + support count + distance-to-nearest-explored** per cell. A recommendation from a thinly-explored cell must SAY SO. 3. **Mode 2 must refuse to extrapolate too far.** If the live state is outside Mode 1's explored envelope (novel regime, latency beyond p99.9, depth collapse never swept), the honest answer is **"OUT OF DISTRIBUTION — fall back to the doctrinal simple policy"**, not a confident recommendation. This is the same law as `INDETERMINATE`: *unknown is not flat, and unknown is not "best guess".* Wire it as an explicit `RecommendationConfidence.OUT_OF_DISTRIBUTION` verdict that the risk gate honours. 4. **Mode 2's live search is a LOCALIZATION, not a fresh MCTS from scratch** — the ≤25 ms budget buys you refinement around a pool policy, not global search. The pool is the compressed product of the billions of Mode-1 trials; Mode 2 spends its milliseconds *choosing and adapting*, not re-deriving. 5. **Mode 2's own outcomes feed back into Mode 1** as new anchors (live discrepancies → `live_discrepancies` table → next sweep is denser where we were wrong). That is the learning loop, and Rule 10 governs it: **where shadow and live disagree, live is right and the manifold gets corrected.** ### Module ownership by mode | Module | Mode 1 (EXPLORE) | Mode 2 (RECOMMEND) | |---|---|---| | `cwm/core.py` | ✅ the arena | ✅ the local rollout model | | `counterparties.py` (ecology) | ✅ **swept adversary population** | ✅ **prior over current opponents** | | `training/cma_trainer.py`, `generator.py` | ✅ | — | | `ScenarioLibrary` (§4) | ✅ sweeps | — | | `training/registry.py` (policy pool) | ✅ writes | ✅ reads | | `training/selector.py` (→ **manifold**) | ✅ builds | ✅ queries | | `ActualsLoader` (§4) | ⚠️ calibration only | ✅ **the live query itself** | | `planner/sm_mcts.py` | ✅ deep, unbounded | ✅ **≤25 ms localization** | | `risk/gate.py` | ✅ constrains training | ✅ hard veto on recommendation | --- ## 1. FINDING #0 — THE FEE BUG IS REAL, AND MEASURED Fable suspected a 10× unit slip. **Our own fills confirm it.** ```sql SELECT liquidity_side, order_type, count() n, avg(fee_bps) FROM dolphin.trade_execution_quality WHERE fee_bps IS NOT NULL GROUP BY 1,2 -- TAKER MARKET 1455 rows avg = 5.016 bps (min 5.000, max 6.649) ``` | Where | Code says | Reality (our fills / venue docs) | Factor | |---|---|---|---| | `malkhut/training/asset_classification.py:164` (BingX) | `default_taker_fee_bps=0.5` | **5.0** | **10×** | | `…:154` (Binance) | `default_taker_fee_bps=0.4` | ~4.5 | ~10× | | `…:174` (Bybit) | `default_taker_fee_bps=0.06` | ~5.5 | ~90× | | `…:384` (default profile) | `taker_fee_bps=0.5, maker_fee_bps=-0.2` | taker 5.0; maker ≈ +2.0 (BingX perp maker is POSITIVE, not a rebate) | 10× + sign | **Impact:** these feed `VenueRules.taker_fee_bps` / `maker_fee_bps` (`malkhut/state.py:109-121`) → the CWM reward → `w_fee_quality`. CMA-ES has been optimizing against fees an order of magnitude too cheap. **Every policy trained so far is suspect** — cheap fees reward overtrading, churn, and cross-spread aggression that real friction annihilates. This is the exact failure the 2026-07-10 venue-friction audit exists to prevent. **ACTION (do this first, before any other work):** 1. Fix the four sites above. Source of truth = `dolphin.trade_execution_quality` (`fee_bps`, `liquidity_side`, `order_type`), not vendor marketing pages. 2. Add a **mutation-litmus test**: set `taker_fee_bps` to 0.5 → a test asserting "policy PnL under realistic friction" must go **RED**. If nothing breaks when fees change 10×, the reward function isn't actually using them. 3. Re-run every CMA-ES benchmark. The "8,628 / 100.4 bps PnL" headline number is void until it is re-measured at 5 bps taker. 4. Maker fee sign: verify against a real maker fill before trusting a rebate. **We have zero MAKER rows in `trade_execution_quality`** (all 1455 are TAKER MARKET — because UV/F5 only sends MARKET). Until SMART-EXEC produces real maker fills, the maker fee is an *assumption*: mark it as such in code with a `# UNVERIFIED — no maker fills on record as of 2026-07-13` comment. --- ## 2. THE DATA SOURCES — WHAT EXISTS, VERIFIED TONIGHT All queried live on `localhost:8123`. Row counts are real. | # | Source | Rows | Columns (exact) | State | |---|---|---|---|---| | **S1** | `dolphin.obf_universe` | **15,329,163,875** | `ts` DateTime64(3), `symbol`, `spread_bps` f32, `depth_1pct_usd` f64, `depth_quality` f32, `fill_probability` f32, `imbalance` f32, `best_bid` f64, `best_ask` f64, `n_bid_levels` u8, `n_ask_levels` u8 | **LIVE, the motherlode** | | **S2** | `dolphin.exf_data` | 22,968,094 | `ts` DateTime64(6), `funding_rate` f32, `dvol` f32, `fear_greed` f32, `taker_ratio` f32 | LIVE | | **S3** | `dolphin.maras_fingerprint` | 1,117,722 | `ts`, `regime` (LowCard), `regime_idx` u8, `confidence`, `final_score`, `conflict_level`, `tier_exf`, `tier_eigen`, `tier_btc`, `tier_esof`, `tier_micro` (+ `_confidence` each), `s_exf_funding_bp`, … | LIVE | | **S4** | `dolphin.eigen_scans` | 1,529,803 | `ts`, `scan_number` u32, `vel_div` f32, `w50_velocity`, `w750_velocity`, `instability_50`, **`scan_to_fill_ms`**, **`step_bar_ms`**, `scan_uuid` | LIVE — **and it is the latency oracle** | | **S5** | `dolphin.trade_execution_quality` | 8,006 | `trade_id`, `asset`, `side`, `client_order_id`, `venue_order_id`, `order_type`, **`liquidity_side`**, **`fee_bps`**, `fill_quality_score`, `commission_quote`, `fee_rate`, … | LIVE — **fee/fill ground truth** | | **S6** | `dolphin.trade_events` | (BLUE's live trades) | `ts`, `trade_id`, `asset`, `side`, `entry_price`, `exit_price`, `pnl`, `pnl_pct`, `exit_reason`, `vel_div_entry`, `boost_at_entry` | LIVE | | **S7** | `dolphin_uv.exec_journal` | growing | `kind`, `timestamp`, `scan_number`, `intent_id`, `trade_id`, `slot_id`, `asset`, `side`, `action`, `reference_price`, `target_size`, `leverage`, `u_prefix_client_id`, `promo_metadata` (JSON) | LIVE | | **S8** | `dolphin_uv.tp_exit_ingress` / `max_hold_ingress` | new (2026-07-13) | full per-scan exit diagnostics: `tp_effective_pct`, `tp_mod_factor`, `tp_floor_armed`, `cascade_count`, `imbalance_ma5`, `branch`, … | LIVE as of today | | **S9** | `dolphin.esof_advisory` | **0** | schema exists (`dow`, `session`, `moon_illumination`, `slot_wr_pct`, …) | ⚠️ **EMPTY** | | **S10** | `dolphin.obf_fast_intrade` | **0** | schema exists | ⚠️ **EMPTY** | **Do not spec against S9/S10 as if they were populated.** Either they get a writer, or they are out of scope. Say so out loud rather than building an intake for a table that never delivers a row (that is how the missing-CH-table blackout happened to UV; the fix cost a day). --- ## 3. FINDING #1 — THE BOOK-FIDELITY GAP (read this twice) This is the **central technical problem** of the whole integration, and it is not optional to solve. **MALKHUT wants a ladder.** `malkhut/state.py:131`: ```python class OrderBookState: bids: Tuple[PriceLevel, ...] # price + qty, per level asks: Tuple[PriceLevel, ...] ``` The CWM's queue-position model, `JOIN_QUEUE`, `LADDER`, `ICEBERG`, queue-ahead estimates, and adverse-selection accounting all depend on **per-level depth**. **Our tape has no ladder.** `dolphin.obf_universe` gives, per symbol per ~sub-second tick: `best_bid`, `best_ask`, `spread_bps`, `depth_1pct_usd` (one aggregate number), `depth_quality`, `imbalance`, `fill_probability`, `n_bid_levels`, `n_ask_levels` (**counts**, not the levels themselves). That is **L1 + shape summary**, not L2. 15.3 billion rows of it — enormously valuable, but it cannot be *decoded* back into a ladder. Information that was never recorded cannot be recovered. ### Three honest options — pick explicitly, do not drift | Option | What it is | Cost | Fidelity | |---|---|---|---| | **A. Synthetic ladder w/ declared prior** | Reconstruct levels from `best_bid/ask` + `depth_1pct_usd` + `n_*_levels` + `imbalance` via an explicit shape model (e.g. exponential decay of qty over levels, calibrated so Σqty within 1% = `depth_1pct_usd`, level count = `n_bid_levels`). | LOW | **Approximate.** Queue position becomes a *model*, not a measurement. MUST be labeled as such everywhere it flows. | | **B. Start recording real L2 now** | New writer → `dolphin_malkhut.book_l2` (or a zinc region + spooler). Depth-20 snapshots at OBF cadence. | MEDIUM (new hose + storage; L2 at 50 symbols × sub-second is BIG — see §7 storage) | **Truth**, but only from the day it starts. Cannot backfill history. | | **C. hftbacktest with external L2** | Use a third-party L2 feed (Binance archival) for CWM calibration; keep obf_universe for regime/context. | MEDIUM | Truth, but it is *Binance's* book, not BingX's — venue microstructure differs. | **RECOMMENDATION: A + B in parallel, and say which one a given result came from.** - **A** unblocks immediate use of 15.3B rows of history for *distributional* calibration (spread regimes, depth regimes, imbalance dynamics, fill-probability priors) — where the ladder shape matters less than the aggregate. - **B** starts the clock on real queue-truth for the queue-sensitive claims (`JOIN_QUEUE`, maker fills, queue-ahead) — the very claims SMART-EXEC needs. - **NEVER** let an A-derived queue-position estimate be reported as a measured fill probability. Tag every artifact: `book_source ∈ {SYNTH_A, REAL_L2_B, EXT_C}`. Rule 3 (replay correctness before search depth) means exactly this. **Litmus:** when B has ≥ 1 week of real L2, re-run A-calibrated policies against B-truth. The delta *is* the measurement of how much the synthetic prior lied. If that delta is large, every A-era conclusion is downgraded to a hypothesis. --- ## 4. THE PLUG MAP — WHAT GOES WHERE, EXACTLY `MarketWorldState` (`malkhut/state.py:257`) is the CWM root. It already has the right holes. Fill them from our sources: | `MarketWorldState` field | Line | Source | Transform | |---|---|---|---| | `book: OrderBookState` | 262 | **S1** `obf_universe` | §3 Option A synth (or B when live). `best_bid`/`best_ask` direct; ladder via declared prior; `symbol` join key. | | `account: AccountState` | 263 | **DITAv2 ASEx account core** (`asex_account` / `AccountProjectionV2`) — **CONSUMER ONLY** (two-cores-share-nothing law) | For offline training: reconstruct from `dolphin.account_events` / `dolphin_uv.exec_journal`. Never a second writer. | | `open_orders` | 264 | live: venue adapter; training: **S7** `exec_journal` + **S5** `trade_execution_quality` (client_order_id join) | | | `trade_path: TradePathState` | 265 | **S8** `tp_exit_ingress` + **S6** `trade_events` | `mae_bps`/`mfe_bps`/`time_in_loss_s` etc. reconstructable from tp_exit_ingress per-scan stamps + trade lifecycle. **`dolphin_regime_score` ← S3; `book_imbalance` ← S1 `imbalance`; `orderflow_toxicity` ← derive (see §5)**. | | `intent: ExecutionIntent` | 266 | **DITAv2 `KernelIntent`** (binding, README §B.1) | Map ENTER/EXIT + `reference_price`/`target_size`/`leverage`/`metadata.promo_client_id`. `urgency` ← SMART-EXEC urgency class (see §8). | | **`funding_bps`** | 268 | **S2** `exf_data.funding_rate` | ×10⁴ → bps. Nearest-ts join. | | **`volatility_state`** | 269 | **S2** `exf_data.dvol` | The sacred vol number. **Same field the gate uses** (`DOLPHIN_VOL_P60_THRESHOLD`, doctrinal 0.00026414). | | **`market_regime`** | 270 | **S3** `maras_fingerprint.regime` | ⚠️ **taxonomy mismatch — see §6.** | | **`feed_latency_ms`** | 272 | **S4** `eigen_scans.step_bar_ms` | p50 = **0.08 ms**. | | **`order_latency_ms`** | 273 | **S4** `eigen_scans.scan_to_fill_ms` | ⚠️ **SAMPLE THE DISTRIBUTION, NOT THE MEAN — see §5.** | | `venue: VenueRules` | 261 | **S5** for fees (§1); `prod/bingx/` for tick/lot/min_notional | Fee fix is blocking. | ### Where the code changes go | New module | Path | Job | |---|---|---| | `ActualsLoader` | `malkhut/data/actuals.py` **(new pkg `malkhut/data/`)** | CH HTTP reader → typed frames. One method per source S1–S8. Polars, chunked, no full-table loads. | | `BookSynthesizer` | `malkhut/data/book_synth.py` | §3 Option A. Declared prior, unit-tested against any real L2 we get. Emits `book_source` tag. | | `LatencyOracle` | `malkhut/data/latency.py` | §5. Empirical CDF sampler, seed-pinned. | | `EcologyFitter` | `malkhut/data/ecology_fit.py` | §7. Fits counterparty params to tape. **Does not create or delete agent types.** | | `RegimeBridge` | `malkhut/data/regime_bridge.py` | §6. MARAS ↔ MALKHUT regime mapping, explicit and tested. | | `ScenarioLibrary` | `malkhut/data/scenarios.py` | §7. CLASS-M/CLASS-X hour selection from tape → adversarial scenario suites. | Wire them at: `malkhut/training/cma_trainer.py` (scenario source), `malkhut/cwm/core.py:transition` (latency + venue rules injection), `malkhut/counterparties.py` (fitted params in, agent classes unchanged). --- ## 5. FINDING #2 — LATENCY IS NOT A NUMBER, IT IS A MONSTER-TAILED DISTRIBUTION Measured tonight from `dolphin.eigen_scans` (`scan_to_fill_ms`, n = 1.5M): | p50 | p95 | p99 | max | |---|---|---|---| | **49.9 ms** | **596.9 ms** | **10,793.9 ms** | **450,115.9 ms** (7.5 minutes) | `step_bar_ms` p50 = **0.08 ms**. The README's "typical_latency_ms: 100" for BingX is a **fiction that will get us killed**: it is 2× the median and **0.9% of the mass is beyond 10 SECONDS**. A policy trained on constant-100ms latency has never met the market that actually fills our orders. The p99 is what turns a maker quote into an adverse-selected gift, and it is exactly where the "why did I get filled?" question (Rule 2) gets its ugliest answer. **ACTION:** `LatencyOracle` must sample the **empirical CDF**, per-regime where possible (latency and stress correlate — verify), seed-pinned for determinism (Rule 9 + our seed-pin doctrine). Constant-latency mode may exist **only** as a labeled ablation, never as the default. Add a scenario class: **`LATENCY_STORM`** (sample exclusively from > p95) — if a policy's edge evaporates there, we need to know before capital does. --- ## 6. FINDING #3 — REGIME TAXONOMY MISMATCH MARAS actually emits (verified, `dolphin.maras_fingerprint`, 1.1M rows): | `regime_idx` | `regime` | rows | |---|---|---| | 1 | BEARISH | 72,479 | | 2 | CHOPPY_BEARISH | 516,477 | | 3 | CHOPPY | 295,366 | | 4 | SIDEWAYS | 201,358 | | 5 | CHOPPY_BULLISH | 32,046 | (idx 0, 6, 7 unobserved in this window — the taxonomy has room MARAS has not used.) MALKHUT's `StrategySelector` uses a **different, invented** set: `trending_up`, `trending_down`, `high_volatility`, `low_volatility`, `mean_reverting`, `momentum`, `choppy`, `liquidity_hole`, `normal`. **These are two different languages.** A performance matrix keyed on MALKHUT's regimes cannot be looked up from a MARAS fingerprint without a mapping, and an implicit/lossy mapping will silently mis-select strategies in production. **ACTION:** `RegimeBridge` (`malkhut/data/regime_bridge.py`) with an **explicit, tested, total** mapping MARAS→MALKHUT. Where MALKHUT has an axis MARAS lacks (e.g. `liquidity_hole`), derive it from S1 (`depth_quality`, `spread_bps` percentiles) and **name the derivation**. Where MARAS has confidence (`confidence`, `conflict_level`), **carry it through** — a low-confidence regime tag must reach the planner as low-confidence, not as a hard label. Prefer extending MALKHUT to speak MARAS over inventing a third dialect. **Also plumb the MARAS tiers** (`tier_exf`, `tier_eigen`, `tier_btc`, `tier_esof`, `tier_micro` + confidences) into the feature vector — that is a 5-tier ensemble view of the market the planner currently cannot see at all, and it is *already computed and stored*. Free signal. --- ## 7. THE ECOLOGY — HOW ACTUALS FEED IT WITHOUT REPLACING IT **This section is the point of the whole document. Read it as law.** `malkhut/counterparties.py` defines the agent ecology (`AgentRole`, `state.py:61`): PASSIVE_MAKER, TOXIC_TAKER, LATENCY_ARB, MOMENTUM_TAKER, MEAN_REVERSION_TAKER, INVENTORY_MM, LIQUIDATION_FLOW, NOISE_TRADER, STALE_QUOTE_ATTACKER. **Nine roles declared, 4 implemented.** ### What actuals DO to the ecology | Do | How | |---|---| | **Fit each agent's parameters** | `EcologyFitter` estimates population parameters from tape: TOXIC_TAKER arrival intensity ← S1 `fill_probability` collapse + adverse post-fill drift in S5/S6; LATENCY_ARB ← S4 latency tails vs price moves; LIQUIDATION_FLOW ← S2 funding extremes + cascade signatures in S8 `cascade_count`; INVENTORY_MM ← S1 `imbalance` mean-reversion; MOMENTUM_TAKER ← S4 `vel_div` / `w50_velocity`. | | **Set the population MIX per regime** | Which agents dominate under BEARISH vs CHOPPY (S3) is an empirical question our tape can answer. The mix becomes regime-conditional. | | **Supply the stress scenarios** | CLASS-M / CLASS-X hours (measured hostile windows) → `ScenarioLibrary`. **CLASS-X law: extreme magnitudes are a FEATURE channel, never filtered.** | | **Bound the plausible** | Tape says what a real book *can* do; adversaries should be allowed to be *worse*, but their base rates must be anchored. | ### What actuals MUST NOT DO | Never | Why | |---|---| | **Replace an agent with a tape replay** | A replayed counterparty cannot react to us. The entire adverse-selection question ("why did I get filled?") requires an opponent that *responds*. A tape is a corpse; the ecology is an opponent. | | **Delete an agent type for lack of data** | The 5 unimplemented roles are *hypotheses about who is on the other side*. Absence of evidence in our thin tape is not evidence of absence in the market. Implement them; fit what you can; make the rest adversarial priors. | | **Restrict the game to observed histories** | Billions of combinations, most of which never occurred, is the *product*. Recorded tape is a measure-zero slice of the strategy × counterparty × regime space. Optimizing only over what happened is exactly the overfit that kills quants. | | **Filter outliers** | CLASS-X law. The 450-second latency tail and the 7.5-minute stop are not noise — they are the market's teeth. | ### The scale ambition (operator's stated aim) Femtosecond-scale trials, parallel, over **billions** of combinations. Current: 5.3 µs/CWM-transition (numba), 16 ms/episode (parallel), 7× on 8 workers. Billions of combinations × ~10³ steps at 5.3 µs = ~10¹² × 5.3 µs ≈ **months of single-box CPU.** So: 1. **Batch/vectorize the counterparty step** (the ecology is the inner loop — make it a matrix op over the agent population, not a Python loop over agents). 2. **GPU or massively-parallel path** for the rollout kernel (the batch MCTS kernel already exists — push it further). 3. **Cheap-then-expensive cascade**: screen billions with a cheap surrogate (vectorized reward, shallow rollout), promote survivors to full CWM depth. Do NOT run 5.3 µs × depth-3 MCTS on every one of a billion candidates. 4. **Ray path already exists** (`training/ray_eval.py`) — that is the horizontal scaling seam when the box runs out. 5. Honest note: "femtosecond" is a *metaphor for parallel breadth*, not a physical claim (a CPU cycle is ~300 picoseconds; a femtosecond is 10⁻¹⁵ s — light travels 0.3 µm in one). State the real target: **N trials/second at fidelity F**, and measure it. Otherwise we cannot tell progress from poetry. --- ## 8. INTAKE CONTRACT ALIGNMENT (with SMART-EXEC + UV) MALKHUT's planner and `SPEC_UV_SMART_EXEC_MM.md` describe the **same seat** at the venue-adapter boundary. Do not build two doorframes. - **Input**: DITAv2 `KernelIntent` + `guideline_price` + `urgency_class` (CATASTROPHIC / PROTECT / HARVEST / ROTATE / ACQUIRE). Map `urgency_class` → `ExecutionIntent.urgency` (`state.py:246`) and `prefer_maker` (:250). **CATASTROPHIC ⇒ the planner is BYPASSED. No cleverness on the stop path, ever.** - **`u-` clientOrderId prefix is LAW** (T9 seam / README §B.2). - **DUAL-LEVERAGE LAW** (README §B.3): conviction leverage [0.5,9.0] sizes quantity; venue leverage = `prod/bingx/leverage.py` mapping → int [1,3]. Litmus: 0.5→1, 4.75→2, 8.0→3, 9.0→3. **Note:** we observed a live 3→2 clamp on 2026-07-13 (handover anomaly #2) — reconcile before MALKHUT sizes anything. - **Execution-truth doctrine is inherited wholesale** (`bdc54fb`): NOT_ATTEMPTED / REFUSED → rollback sound; **INDETERMINATE → never**. MALKHUT's venue adapter (`malkhut/venue/bingx/adapter.py`) wraps DITAv2 and therefore inherits the fences — **verify with a test, do not assume.** See `COMPREHENSIVE_UV_EDGE_CASE_THEORETICALS.md` §11 M1–M7 for the unaudited seams. - **No reconcilers.** Bounded, read-only, own-clientOrderId point lookups only. --- ## 9. PERSISTENCE + PROVENANCE - Namespace `dolphin_malkhut` (approved). **NO TTL** (retention doctrine). - **DDL ships WITH the code, and the applier's verify-set must require every new table** — the 2026-07-13 lesson: two UV tables existed as `.sql` files for days while the runner 404'd every scan and journaled nothing. Pattern to copy: `prod/clickhouse/uv/apply_uv_ddl.py` (`EXPECTED_TABLES`). - **Every row carries provenance**: `book_source` (SYNTH_A/REAL_L2_B/EXT_C), `latency_model` (EMPIRICAL_CDF/CONST), `fee_table_version`, `ecology_fit_version`, `policy_version`, `seed`. A result whose provenance is unknown is not a result. - **Cross-reference law**: `fulfilment_decisions.intent_id` + `trade_id` join to `dolphin_uv.exec_journal` ⇄ venue order history. End-to-end or it is a black box. --- ## 10. ORDER OF WORK (do not reorder) | # | Task | Gate | |---|---|---| | 1 | **Fix the fees** (§1) + mutation-litmus test | A 10× fee change must break a test | | 2 | **Re-baseline** every CMA-ES number at real friction | Old headline numbers marked VOID | | 3 | **LatencyOracle** (§5) + `LATENCY_STORM` scenario | Constant-latency is an ablation, not a default | | 4 | **Decide the book question** (§3) — declare A, B, or C **in writing** | `book_source` tag on every artifact | | 5 | **RegimeBridge** (§6) — explicit total mapping | Test: every MARAS regime maps; confidence carried | | 6 | **ActualsLoader** (§4) — S1/S2/S3/S4 into `MarketWorldState` | Replay determinism ×2 (byte-identical minus ts) | | 7 | **EcologyFitter** (§7) — fit params, **keep all agent types** | Test: fitting cannot reduce the agent-type count | | 8 | **ScenarioLibrary** — CLASS-M/X hours as suites | CLASS-X magnitudes unfiltered | | 9 | **Implement the 5 missing agent roles** | Ecology completeness | | 10 | **Scale the inner loop** (§7) — vectorize ecology, cascade screening | Measured trials/sec, not adjectives | | 11 | **MODE 1 → manifold** (§0.5): selector emits variance + support + envelope, not a leaderboard | A thin cell must report itself thin | | 12 | **MODE 2 → localization + OOD verdict** (§0.5): live state query, `OUT_OF_DISTRIBUTION` falls back to doctrinal policy | Test: a novel regime must NOT get a confident recommendation | | 13 | **Gate-M** (README §F): replay verify → dual-leverage litmus → DARK shadow vs naive → risk-gate mutation litmus → Rule 10 | Money | --- ## 11. THE STANDING GUARDRAILS (Fable's, HJ's, and the house's) 1. **Rule 3 supreme**: replay correctness before search depth. A wrong CWM plus a deep search is confident nonsense, and confident nonsense is the most expensive thing we can build. 2. **Rule 10 supreme**: shadow vs live diverges → **trust live**, stand down, file the discrepancy row. 3. **Mutation litmus everywhere**: 1,186 green tests prove nothing until breaking the implementation turns one red. Fees, latency, risk-gate constraints, ecology size — each must have a test that dies when it is sabotaged. 4. **"It looks finished" is the warning sign, not the green light.** 5. **The ecology stays.** If a future refactor makes the counterparties smaller, fewer, or tamer in the name of realism, it has removed the edge and kept the costs. That refactor is wrong. This line is the reason this document exists. --- *Written by Claude (Opus 4.8) on Fable's context, at HJ's order, 2026-07-13. Every path, line number, row count, and measured value herein was verified against the live system on the night of writing. Where a number was not verified, it says so.*