Files
sentiment-engine/prod/docs/SPEC_MALKHUT_ACTUALS_INTAKE.md
Codex 317e9d7a13 docs(spec): MALKHUT actuals-intake — two-mode architecture, verified plug map, 3 findings
Two modes (operator): EXPLORE (synthetics, ecology-dominant, billions of combos,
emits a MANIFOLD) -> RECOMMEND (live OBF as query, localize + extrapolate, OOD
verdict falls back to doctrinal). Ecology stays: actuals calibrate, ecology plays.

Findings verified live tonight:
- FEE BUG CONFIRMED: trade_execution_quality says 5.016 bps taker; code says 0.5
  (asset_classification.py:154/164/174/384). Every CMA-ES number is void until re-baselined.
- BOOK GAP: obf_universe (15.3B rows) is L1+aggregates, NOT a ladder. Three options,
  must declare which; book_source provenance tag on every artifact.
- LATENCY: p50 49.9ms / p95 597ms / p99 10.8s / max 450s. Constant-100ms is fiction.
- REGIME MISMATCH: MARAS emits 5 regimes; MALKHUT selector invented 9. RegimeBridge needed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 18:32:54 +02:00

30 KiB
Raw Blame History

SPEC — MALKHUT ACTUALS INTAKE ("game it under actuals; the ecology stays")

Author: Claude (Opus 4.8), continuing Fable's review. Date: 2026-07-13. Addressee: mimo (mm_, mm_ob_fill_sim). Ordered by: HJ. Status: SPEC. Companion to MALKHUT/README.md (§ADDENDUM) and prod/docs/SPEC_UV_SMART_EXEC_MM.md.


0. THE ONE-LINE LAW

Actuals CALIBRATE the game. The ecology PLAYS it. Real tape grounds the adversaries; it NEVER replaces them.

Every instruction below serves that sentence. If a change would let recorded history substitute for adversarial simulation, it is wrong, no matter how "realistic" it looks. Replaying our own tape teaches MALKHUT what happened once. The ecology is what lets it out-play what could happen — billions of strategy × counterparty × regime combinations, in parallel, none of which are in any tape. The ecology is the edge. It is not a placeholder for missing data.


0.5 THE TWO MODES (operator, 2026-07-13) — THE ARCHITECTURAL SPINE

MALKHUT is two engines sharing one CWM, and every module below belongs to one of them. This is the frame that makes "actuals" and "ecology" complementary instead of competing — and it is, structurally, train vs inference.

┌─ MODE 1: EXPLORE ───────────────────────────────────────────────────┐
│  "Search/play over SYNTHETICS: fast, parallel, acquisitive of BEST"  │
│                                                                      │
│  ECOLOGY-DOMINANT. Adversaries react. Regimes swept, not sampled.    │
│  Billions of strategy × counterparty × regime combinations.          │
│  Actuals appear ONLY as calibration (fee tables, latency CDF,        │
│  agent base-rates, plausible bounds). NOT as the arena.              │
│                                                                      │
│  Output: POLICY POOL — a ranked, regime-indexed library of           │
│  strategies with known behaviour under known adversary mixes.        │
│  Cadence: OFFLINE, hours-to-days, unbounded compute.                 │
│  Success metric: coverage + robustness, NOT backtest PnL.            │
└──────────────────────────────┬───────────────────────────────────────┘
                               │  policy pool + performance manifold
                               ▼
┌─ MODE 2: RECOMMEND (inference/query) ───────────────────────────────┐
│  "Search over LIVE, ACTUAL (OBF) IRL conditions as INPUTS, to       │
│   extrapolate the closest-to-absolute-best strategy → RECOMMEND"     │
│                                                                      │
│  ACTUALS-DOMINANT. The live book/regime/funding/latency IS the       │
│  query. The ecology is now a PRIOR over who is on the other side     │
│  right now, not a population to be swept.                            │
│                                                                      │
│  Given: live MarketWorldState (S1 book, S2 funding/dvol, S3 regime,  │
│         S4 latency, account, intent)                                 │
│  Ask:   which policy in the pool is nearest-optimal for THIS state?  │
│  Output: a RECOMMENDATION (policy + its expected behaviour +         │
│          confidence + the nearest explored neighbours it interpolates│
│          between).                                                   │
│  Cadence: ONLINE, ≤25 ms planning inside ≤100 ms end-to-end.         │
│  Success metric: live outcome ≈ predicted outcome (Rule 10).         │
└──────────────────────────────────────────────────────────────────────┘

Why this is the right shape

  • Mode 1 without Mode 2 is a beautiful simulator nobody trades.
  • Mode 2 without Mode 1 is a lookup table over history — the overfit that kills quants. It can only recommend what already happened.
  • Together: Mode 1 explores a space vastly larger than any tape can contain (that is the edge — we out-play participants who only ever fit history), and Mode 2 localizes the live moment inside that explored space and returns the best-known play, including for market states we have never actually seen, because Mode 1 has already played them.

The word "extrapolate" in the operator's phrasing is load-bearing. Mode 2's job is NOT nearest-neighbour lookup. It is: place the live state inside the performance manifold Mode 1 built, and interpolate/extrapolate the best play. That requires Mode 1 to output a manifold, not a leaderboard:

Mode 1 must emit Not just
policy × (regime, spread, depth, toxicity, latency, funding, inventory) → expected outcome + variance "policy #7 scored 8,628"
the boundaries of where each policy was tested (extrapolation beyond = flagged) a single global champion
why a policy wins (which adversary it beats, which it loses to) an opaque score

Consequences for the build

  1. ScenarioLibrary must SWEEP, not sample. Mode 1's job is coverage of the state space, including regions our tape never visited. Tape-derived scenarios (CLASS-M/X) are the anchors; the sweep fills between and beyond them.
  2. The performance matrix in training/selector.py becomes the manifold — and it must carry confidence + support count + distance-to-nearest-explored per cell. A recommendation from a thinly-explored cell must SAY SO.
  3. Mode 2 must refuse to extrapolate too far. If the live state is outside Mode 1's explored envelope (novel regime, latency beyond p99.9, depth collapse never swept), the honest answer is "OUT OF DISTRIBUTION — fall back to the doctrinal simple policy", not a confident recommendation. This is the same law as INDETERMINATE: unknown is not flat, and unknown is not "best guess". Wire it as an explicit RecommendationConfidence.OUT_OF_DISTRIBUTION verdict that the risk gate honours.
  4. Mode 2's live search is a LOCALIZATION, not a fresh MCTS from scratch — the ≤25 ms budget buys you refinement around a pool policy, not global search. The pool is the compressed product of the billions of Mode-1 trials; Mode 2 spends its milliseconds choosing and adapting, not re-deriving.
  5. Mode 2's own outcomes feed back into Mode 1 as new anchors (live discrepancies → live_discrepancies table → next sweep is denser where we were wrong). That is the learning loop, and Rule 10 governs it: where shadow and live disagree, live is right and the manifold gets corrected.

Module ownership by mode

Module Mode 1 (EXPLORE) Mode 2 (RECOMMEND)
cwm/core.py ✅ the arena ✅ the local rollout model
counterparties.py (ecology) ✅ swept adversary population ✅ prior over current opponents
training/cma_trainer.py, generator.py ✅ —
ScenarioLibrary (§4) ✅ sweeps —
training/registry.py (policy pool) ✅ writes ✅ reads
training/selector.py (→ manifold) ✅ builds ✅ queries
ActualsLoader (§4) ⚠️ calibration only ✅ the live query itself
planner/sm_mcts.py ✅ deep, unbounded ✅ ≤25 ms localization
risk/gate.py ✅ constrains training ✅ hard veto on recommendation

1. FINDING #0 — THE FEE BUG IS REAL, AND MEASURED

Fable suspected a 10× unit slip. Our own fills confirm it.

SELECT liquidity_side, order_type, count() n, avg(fee_bps)
FROM dolphin.trade_execution_quality WHERE fee_bps IS NOT NULL GROUP BY 1,2
-- TAKER  MARKET  1455 rows  avg = 5.016 bps  (min 5.000, max 6.649)
Where Code says Reality (our fills / venue docs) Factor
malkhut/training/asset_classification.py:164 (BingX) default_taker_fee_bps=0.5 5.0 10×
…:154 (Binance) default_taker_fee_bps=0.4 ~4.5 ~10×
…:174 (Bybit) default_taker_fee_bps=0.06 ~5.5 ~90×
…:384 (default profile) taker_fee_bps=0.5, maker_fee_bps=-0.2 taker 5.0; maker ≈ +2.0 (BingX perp maker is POSITIVE, not a rebate) 10× + sign

Impact: these feed VenueRules.taker_fee_bps / maker_fee_bps (malkhut/state.py:109-121) → the CWM reward → w_fee_quality. CMA-ES has been optimizing against fees an order of magnitude too cheap. Every policy trained so far is suspect — cheap fees reward overtrading, churn, and cross-spread aggression that real friction annihilates. This is the exact failure the 2026-07-10 venue-friction audit exists to prevent.

ACTION (do this first, before any other work):

  1. Fix the four sites above. Source of truth = dolphin.trade_execution_quality (fee_bps, liquidity_side, order_type), not vendor marketing pages.
  2. Add a mutation-litmus test: set taker_fee_bps to 0.5 → a test asserting "policy PnL under realistic friction" must go RED. If nothing breaks when fees change 10×, the reward function isn't actually using them.
  3. Re-run every CMA-ES benchmark. The "8,628 / 100.4 bps PnL" headline number is void until it is re-measured at 5 bps taker.
  4. Maker fee sign: verify against a real maker fill before trusting a rebate. We have zero MAKER rows in trade_execution_quality (all 1455 are TAKER MARKET — because UV/F5 only sends MARKET). Until SMART-EXEC produces real maker fills, the maker fee is an assumption: mark it as such in code with a # UNVERIFIED — no maker fills on record as of 2026-07-13 comment.

2. THE DATA SOURCES — WHAT EXISTS, VERIFIED TONIGHT

All queried live on localhost:8123. Row counts are real.

# Source Rows Columns (exact) State
S1 dolphin.obf_universe 15,329,163,875 ts DateTime64(3), symbol, spread_bps f32, depth_1pct_usd f64, depth_quality f32, fill_probability f32, imbalance f32, best_bid f64, best_ask f64, n_bid_levels u8, n_ask_levels u8 LIVE, the motherlode
S2 dolphin.exf_data 22,968,094 ts DateTime64(6), funding_rate f32, dvol f32, fear_greed f32, taker_ratio f32 LIVE
S3 dolphin.maras_fingerprint 1,117,722 ts, regime (LowCard), regime_idx u8, confidence, final_score, conflict_level, tier_exf, tier_eigen, tier_btc, tier_esof, tier_micro (+ _confidence each), s_exf_funding_bp, … LIVE
S4 dolphin.eigen_scans 1,529,803 ts, scan_number u32, vel_div f32, w50_velocity, w750_velocity, instability_50, scan_to_fill_ms, step_bar_ms, scan_uuid LIVE — and it is the latency oracle
S5 dolphin.trade_execution_quality 8,006 trade_id, asset, side, client_order_id, venue_order_id, order_type, liquidity_side, fee_bps, fill_quality_score, commission_quote, fee_rate, … LIVE — fee/fill ground truth
S6 dolphin.trade_events (BLUE's live trades) ts, trade_id, asset, side, entry_price, exit_price, pnl, pnl_pct, exit_reason, vel_div_entry, boost_at_entry LIVE
S7 dolphin_uv.exec_journal growing kind, timestamp, scan_number, intent_id, trade_id, slot_id, asset, side, action, reference_price, target_size, leverage, u_prefix_client_id, promo_metadata (JSON) LIVE
S8 dolphin_uv.tp_exit_ingress / max_hold_ingress new (2026-07-13) full per-scan exit diagnostics: tp_effective_pct, tp_mod_factor, tp_floor_armed, cascade_count, imbalance_ma5, branch, … LIVE as of today
S9 dolphin.esof_advisory 0 schema exists (dow, session, moon_illumination, slot_wr_pct, …) ⚠️ EMPTY
S10 dolphin.obf_fast_intrade 0 schema exists ⚠️ EMPTY

Do not spec against S9/S10 as if they were populated. Either they get a writer, or they are out of scope. Say so out loud rather than building an intake for a table that never delivers a row (that is how the missing-CH-table blackout happened to UV; the fix cost a day).


3. FINDING #1 — THE BOOK-FIDELITY GAP (read this twice)

This is the central technical problem of the whole integration, and it is not optional to solve.

MALKHUT wants a ladder. malkhut/state.py:131:

class OrderBookState:
    bids: Tuple[PriceLevel, ...]   # price + qty, per level
    asks: Tuple[PriceLevel, ...]

The CWM's queue-position model, JOIN_QUEUE, LADDER, ICEBERG, queue-ahead estimates, and adverse-selection accounting all depend on per-level depth.

Our tape has no ladder. dolphin.obf_universe gives, per symbol per ~sub-second tick: best_bid, best_ask, spread_bps, depth_1pct_usd (one aggregate number), depth_quality, imbalance, fill_probability, n_bid_levels, n_ask_levels (counts, not the levels themselves).

That is L1 + shape summary, not L2. 15.3 billion rows of it — enormously valuable, but it cannot be decoded back into a ladder. Information that was never recorded cannot be recovered.

Three honest options — pick explicitly, do not drift

Option What it is Cost Fidelity
A. Synthetic ladder w/ declared prior Reconstruct levels from best_bid/ask + depth_1pct_usd + n_*_levels + imbalance via an explicit shape model (e.g. exponential decay of qty over levels, calibrated so Σqty within 1% = depth_1pct_usd, level count = n_bid_levels). LOW Approximate. Queue position becomes a model, not a measurement. MUST be labeled as such everywhere it flows.
B. Start recording real L2 now New writer → dolphin_malkhut.book_l2 (or a zinc region + spooler). Depth-20 snapshots at OBF cadence. MEDIUM (new hose + storage; L2 at 50 symbols × sub-second is BIG — see §7 storage) Truth, but only from the day it starts. Cannot backfill history.
C. hftbacktest with external L2 Use a third-party L2 feed (Binance archival) for CWM calibration; keep obf_universe for regime/context. MEDIUM Truth, but it is Binance's book, not BingX's — venue microstructure differs.

RECOMMENDATION: A + B in parallel, and say which one a given result came from.

  • A unblocks immediate use of 15.3B rows of history for distributional calibration (spread regimes, depth regimes, imbalance dynamics, fill-probability priors) — where the ladder shape matters less than the aggregate.
  • B starts the clock on real queue-truth for the queue-sensitive claims (JOIN_QUEUE, maker fills, queue-ahead) — the very claims SMART-EXEC needs.
  • NEVER let an A-derived queue-position estimate be reported as a measured fill probability. Tag every artifact: book_source ∈ {SYNTH_A, REAL_L2_B, EXT_C}. Rule 3 (replay correctness before search depth) means exactly this.

Litmus: when B has ≥ 1 week of real L2, re-run A-calibrated policies against B-truth. The delta is the measurement of how much the synthetic prior lied. If that delta is large, every A-era conclusion is downgraded to a hypothesis.


4. THE PLUG MAP — WHAT GOES WHERE, EXACTLY

MarketWorldState (malkhut/state.py:257) is the CWM root. It already has the right holes. Fill them from our sources:

MarketWorldState field Line Source Transform
book: OrderBookState 262 S1 obf_universe §3 Option A synth (or B when live). best_bid/best_ask direct; ladder via declared prior; symbol join key.
account: AccountState 263 DITAv2 ASEx account core (asex_account / AccountProjectionV2) — CONSUMER ONLY (two-cores-share-nothing law) For offline training: reconstruct from dolphin.account_events / dolphin_uv.exec_journal. Never a second writer.
open_orders 264 live: venue adapter; training: S7 exec_journal + S5 trade_execution_quality (client_order_id join)
trade_path: TradePathState 265 S8 tp_exit_ingress + S6 trade_events mae_bps/mfe_bps/time_in_loss_s etc. reconstructable from tp_exit_ingress per-scan stamps + trade lifecycle. dolphin_regime_score ← S3; book_imbalance ← S1 imbalance; orderflow_toxicity ← derive (see §5).
intent: ExecutionIntent 266 DITAv2 KernelIntent (binding, README §B.1) Map ENTER/EXIT + reference_price/target_size/leverage/metadata.promo_client_id. urgency ← SMART-EXEC urgency class (see §8).
funding_bps 268 S2 exf_data.funding_rate ×10⁴ → bps. Nearest-ts join.
volatility_state 269 S2 exf_data.dvol The sacred vol number. Same field the gate uses (DOLPHIN_VOL_P60_THRESHOLD, doctrinal 0.00026414).
market_regime 270 S3 maras_fingerprint.regime ⚠️ taxonomy mismatch — see §6.
feed_latency_ms 272 S4 eigen_scans.step_bar_ms p50 = 0.08 ms.
order_latency_ms 273 S4 eigen_scans.scan_to_fill_ms ⚠️ SAMPLE THE DISTRIBUTION, NOT THE MEAN — see §5.
venue: VenueRules 261 S5 for fees (§1); prod/bingx/ for tick/lot/min_notional Fee fix is blocking.

Where the code changes go

New module Path Job
ActualsLoader malkhut/data/actuals.py (new pkg malkhut/data/) CH HTTP reader → typed frames. One method per source S1–S8. Polars, chunked, no full-table loads.
BookSynthesizer malkhut/data/book_synth.py §3 Option A. Declared prior, unit-tested against any real L2 we get. Emits book_source tag.
LatencyOracle malkhut/data/latency.py §5. Empirical CDF sampler, seed-pinned.
EcologyFitter malkhut/data/ecology_fit.py §7. Fits counterparty params to tape. Does not create or delete agent types.
RegimeBridge malkhut/data/regime_bridge.py §6. MARAS ↔ MALKHUT regime mapping, explicit and tested.
ScenarioLibrary malkhut/data/scenarios.py §7. CLASS-M/CLASS-X hour selection from tape → adversarial scenario suites.

Wire them at: malkhut/training/cma_trainer.py (scenario source), malkhut/cwm/core.py:transition (latency + venue rules injection), malkhut/counterparties.py (fitted params in, agent classes unchanged).


5. FINDING #2 — LATENCY IS NOT A NUMBER, IT IS A MONSTER-TAILED DISTRIBUTION

Measured tonight from dolphin.eigen_scans (scan_to_fill_ms, n = 1.5M):

p50 p95 p99 max
49.9 ms 596.9 ms 10,793.9 ms 450,115.9 ms (7.5 minutes)

step_bar_ms p50 = 0.08 ms.

The README's "typical_latency_ms: 100" for BingX is a fiction that will get us killed: it is 2× the median and 0.9% of the mass is beyond 10 SECONDS. A policy trained on constant-100ms latency has never met the market that actually fills our orders. The p99 is what turns a maker quote into an adverse-selected gift, and it is exactly where the "why did I get filled?" question (Rule 2) gets its ugliest answer.

ACTION: LatencyOracle must sample the empirical CDF, per-regime where possible (latency and stress correlate — verify), seed-pinned for determinism (Rule 9 + our seed-pin doctrine). Constant-latency mode may exist only as a labeled ablation, never as the default. Add a scenario class: LATENCY_STORM (sample exclusively from > p95) — if a policy's edge evaporates there, we need to know before capital does.


6. FINDING #3 — REGIME TAXONOMY MISMATCH

MARAS actually emits (verified, dolphin.maras_fingerprint, 1.1M rows):

regime_idx regime rows
1 BEARISH 72,479
2 CHOPPY_BEARISH 516,477
3 CHOPPY 295,366
4 SIDEWAYS 201,358
5 CHOPPY_BULLISH 32,046

(idx 0, 6, 7 unobserved in this window — the taxonomy has room MARAS has not used.)

MALKHUT's StrategySelector uses a different, invented set: trending_up, trending_down, high_volatility, low_volatility, mean_reverting, momentum, choppy, liquidity_hole, normal.

These are two different languages. A performance matrix keyed on MALKHUT's regimes cannot be looked up from a MARAS fingerprint without a mapping, and an implicit/lossy mapping will silently mis-select strategies in production.

ACTION: RegimeBridge (malkhut/data/regime_bridge.py) with an explicit, tested, total mapping MARAS→MALKHUT. Where MALKHUT has an axis MARAS lacks (e.g. liquidity_hole), derive it from S1 (depth_quality, spread_bps percentiles) and name the derivation. Where MARAS has confidence (confidence, conflict_level), carry it through — a low-confidence regime tag must reach the planner as low-confidence, not as a hard label. Prefer extending MALKHUT to speak MARAS over inventing a third dialect.

Also plumb the MARAS tiers (tier_exf, tier_eigen, tier_btc, tier_esof, tier_micro + confidences) into the feature vector — that is a 5-tier ensemble view of the market the planner currently cannot see at all, and it is already computed and stored. Free signal.


7. THE ECOLOGY — HOW ACTUALS FEED IT WITHOUT REPLACING IT

This section is the point of the whole document. Read it as law.

malkhut/counterparties.py defines the agent ecology (AgentRole, state.py:61): PASSIVE_MAKER, TOXIC_TAKER, LATENCY_ARB, MOMENTUM_TAKER, MEAN_REVERSION_TAKER, INVENTORY_MM, LIQUIDATION_FLOW, NOISE_TRADER, STALE_QUOTE_ATTACKER. Nine roles declared, 4 implemented.

What actuals DO to the ecology

Do How
Fit each agent's parameters EcologyFitter estimates population parameters from tape: TOXIC_TAKER arrival intensity ← S1 fill_probability collapse + adverse post-fill drift in S5/S6; LATENCY_ARB ← S4 latency tails vs price moves; LIQUIDATION_FLOW ← S2 funding extremes + cascade signatures in S8 cascade_count; INVENTORY_MM ← S1 imbalance mean-reversion; MOMENTUM_TAKER ← S4 vel_div / w50_velocity.
Set the population MIX per regime Which agents dominate under BEARISH vs CHOPPY (S3) is an empirical question our tape can answer. The mix becomes regime-conditional.
Supply the stress scenarios CLASS-M / CLASS-X hours (measured hostile windows) → ScenarioLibrary. CLASS-X law: extreme magnitudes are a FEATURE channel, never filtered.
Bound the plausible Tape says what a real book can do; adversaries should be allowed to be worse, but their base rates must be anchored.

What actuals MUST NOT DO

Never Why
Replace an agent with a tape replay A replayed counterparty cannot react to us. The entire adverse-selection question ("why did I get filled?") requires an opponent that responds. A tape is a corpse; the ecology is an opponent.
Delete an agent type for lack of data The 5 unimplemented roles are hypotheses about who is on the other side. Absence of evidence in our thin tape is not evidence of absence in the market. Implement them; fit what you can; make the rest adversarial priors.
Restrict the game to observed histories Billions of combinations, most of which never occurred, is the product. Recorded tape is a measure-zero slice of the strategy × counterparty × regime space. Optimizing only over what happened is exactly the overfit that kills quants.
Filter outliers CLASS-X law. The 450-second latency tail and the 7.5-minute stop are not noise — they are the market's teeth.

The scale ambition (operator's stated aim)

Femtosecond-scale trials, parallel, over billions of combinations. Current: 5.3 µs/CWM-transition (numba), 16 ms/episode (parallel), 7× on 8 workers. Billions of combinations × ~10³ steps at 5.3 µs = ~10¹² × 5.3 µs ≈ months of single-box CPU. So:

  1. Batch/vectorize the counterparty step (the ecology is the inner loop — make it a matrix op over the agent population, not a Python loop over agents).
  2. GPU or massively-parallel path for the rollout kernel (the batch MCTS kernel already exists — push it further).
  3. Cheap-then-expensive cascade: screen billions with a cheap surrogate (vectorized reward, shallow rollout), promote survivors to full CWM depth. Do NOT run 5.3 µs × depth-3 MCTS on every one of a billion candidates.
  4. Ray path already exists (training/ray_eval.py) — that is the horizontal scaling seam when the box runs out.
  5. Honest note: "femtosecond" is a metaphor for parallel breadth, not a physical claim (a CPU cycle is ~300 picoseconds; a femtosecond is 10⁻¹⁵ s — light travels 0.3 µm in one). State the real target: N trials/second at fidelity F, and measure it. Otherwise we cannot tell progress from poetry.

8. INTAKE CONTRACT ALIGNMENT (with SMART-EXEC + UV)

MALKHUT's planner and SPEC_UV_SMART_EXEC_MM.md describe the same seat at the venue-adapter boundary. Do not build two doorframes.

  • Input: DITAv2 KernelIntent + guideline_price + urgency_class (CATASTROPHIC / PROTECT / HARVEST / ROTATE / ACQUIRE). Map urgency_class → ExecutionIntent.urgency (state.py:246) and prefer_maker (:250). CATASTROPHIC ⇒ the planner is BYPASSED. No cleverness on the stop path, ever.
  • u- clientOrderId prefix is LAW (T9 seam / README §B.2).
  • DUAL-LEVERAGE LAW (README §B.3): conviction leverage [0.5,9.0] sizes quantity; venue leverage = prod/bingx/leverage.py mapping → int [1,3]. Litmus: 0.5→1, 4.75→2, 8.0→3, 9.0→3. Note: we observed a live 3→2 clamp on 2026-07-13 (handover anomaly #2) — reconcile before MALKHUT sizes anything.
  • Execution-truth doctrine is inherited wholesale (bdc54fb): NOT_ATTEMPTED / REFUSED → rollback sound; INDETERMINATE → never. MALKHUT's venue adapter (malkhut/venue/bingx/adapter.py) wraps DITAv2 and therefore inherits the fences — verify with a test, do not assume. See COMPREHENSIVE_UV_EDGE_CASE_THEORETICALS.md §11 M1–M7 for the unaudited seams.
  • No reconcilers. Bounded, read-only, own-clientOrderId point lookups only.

9. PERSISTENCE + PROVENANCE

  • Namespace dolphin_malkhut (approved). NO TTL (retention doctrine).
  • DDL ships WITH the code, and the applier's verify-set must require every new table — the 2026-07-13 lesson: two UV tables existed as .sql files for days while the runner 404'd every scan and journaled nothing. Pattern to copy: prod/clickhouse/uv/apply_uv_ddl.py (EXPECTED_TABLES).
  • Every row carries provenance: book_source (SYNTH_A/REAL_L2_B/EXT_C), latency_model (EMPIRICAL_CDF/CONST), fee_table_version, ecology_fit_version, policy_version, seed. A result whose provenance is unknown is not a result.
  • Cross-reference law: fulfilment_decisions.intent_id + trade_id join to dolphin_uv.exec_journal ⇄ venue order history. End-to-end or it is a black box.

10. ORDER OF WORK (do not reorder)

# Task Gate
1 Fix the fees (§1) + mutation-litmus test A 10× fee change must break a test
2 Re-baseline every CMA-ES number at real friction Old headline numbers marked VOID
3 LatencyOracle (§5) + LATENCY_STORM scenario Constant-latency is an ablation, not a default
4 Decide the book question (§3) — declare A, B, or C in writing book_source tag on every artifact
5 RegimeBridge (§6) — explicit total mapping Test: every MARAS regime maps; confidence carried
6 ActualsLoader (§4) — S1/S2/S3/S4 into MarketWorldState Replay determinism ×2 (byte-identical minus ts)
7 EcologyFitter (§7) — fit params, keep all agent types Test: fitting cannot reduce the agent-type count
8 ScenarioLibrary — CLASS-M/X hours as suites CLASS-X magnitudes unfiltered
9 Implement the 5 missing agent roles Ecology completeness
10 Scale the inner loop (§7) — vectorize ecology, cascade screening Measured trials/sec, not adjectives
11 MODE 1 → manifold (§0.5): selector emits variance + support + envelope, not a leaderboard A thin cell must report itself thin
12 MODE 2 → localization + OOD verdict (§0.5): live state query, OUT_OF_DISTRIBUTION falls back to doctrinal policy Test: a novel regime must NOT get a confident recommendation
13 Gate-M (README §F): replay verify → dual-leverage litmus → DARK shadow vs naive → risk-gate mutation litmus → Rule 10 Money

11. THE STANDING GUARDRAILS (Fable's, HJ's, and the house's)

  1. Rule 3 supreme: replay correctness before search depth. A wrong CWM plus a deep search is confident nonsense, and confident nonsense is the most expensive thing we can build.
  2. Rule 10 supreme: shadow vs live diverges → trust live, stand down, file the discrepancy row.
  3. Mutation litmus everywhere: 1,186 green tests prove nothing until breaking the implementation turns one red. Fees, latency, risk-gate constraints, ecology size — each must have a test that dies when it is sabotaged.
  4. "It looks finished" is the warning sign, not the green light.
  5. The ecology stays. If a future refactor makes the counterparties smaller, fewer, or tamer in the name of realism, it has removed the edge and kept the costs. That refactor is wrong. This line is the reason this document exists.

Written by Claude (Opus 4.8) on Fable's context, at HJ's order, 2026-07-13. Every path, line number, row count, and measured value herein was verified against the live system on the night of writing. Where a number was not verified, it says so.