Two modes (operator): EXPLORE (synthetics, ecology-dominant, billions of combos, emits a MANIFOLD) -> RECOMMEND (live OBF as query, localize + extrapolate, OOD verdict falls back to doctrinal). Ecology stays: actuals calibrate, ecology plays. Findings verified live tonight: - FEE BUG CONFIRMED: trade_execution_quality says 5.016 bps taker; code says 0.5 (asset_classification.py:154/164/174/384). Every CMA-ES number is void until re-baselined. - BOOK GAP: obf_universe (15.3B rows) is L1+aggregates, NOT a ladder. Three options, must declare which; book_source provenance tag on every artifact. - LATENCY: p50 49.9ms / p95 597ms / p99 10.8s / max 450s. Constant-100ms is fiction. - REGIME MISMATCH: MARAS emits 5 regimes; MALKHUT selector invented 9. RegimeBridge needed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
30 KiB
SPEC — MALKHUT ACTUALS INTAKE ("game it under actuals; the ecology stays")
Author: Claude (Opus 4.8), continuing Fable's review. Date: 2026-07-13.
Addressee: mimo (mm_, mm_ob_fill_sim). Ordered by: HJ.
Status: SPEC. Companion to MALKHUT/README.md (§ADDENDUM) and
prod/docs/SPEC_UV_SMART_EXEC_MM.md.
0. THE ONE-LINE LAW
Actuals CALIBRATE the game. The ecology PLAYS it. Real tape grounds the adversaries; it NEVER replaces them.
Every instruction below serves that sentence. If a change would let recorded history substitute for adversarial simulation, it is wrong, no matter how "realistic" it looks. Replaying our own tape teaches MALKHUT what happened once. The ecology is what lets it out-play what could happen — billions of strategy × counterparty × regime combinations, in parallel, none of which are in any tape. The ecology is the edge. It is not a placeholder for missing data.
0.5 THE TWO MODES (operator, 2026-07-13) — THE ARCHITECTURAL SPINE
MALKHUT is two engines sharing one CWM, and every module below belongs to one of them. This is the frame that makes "actuals" and "ecology" complementary instead of competing — and it is, structurally, train vs inference.
┌─ MODE 1: EXPLORE ───────────────────────────────────────────────────┐
│ "Search/play over SYNTHETICS: fast, parallel, acquisitive of BEST" │
│ │
│ ECOLOGY-DOMINANT. Adversaries react. Regimes swept, not sampled. │
│ Billions of strategy × counterparty × regime combinations. │
│ Actuals appear ONLY as calibration (fee tables, latency CDF, │
│ agent base-rates, plausible bounds). NOT as the arena. │
│ │
│ Output: POLICY POOL — a ranked, regime-indexed library of │
│ strategies with known behaviour under known adversary mixes. │
│ Cadence: OFFLINE, hours-to-days, unbounded compute. │
│ Success metric: coverage + robustness, NOT backtest PnL. │
└──────────────────────────────┬───────────────────────────────────────┘
│ policy pool + performance manifold
▼
┌─ MODE 2: RECOMMEND (inference/query) ───────────────────────────────┐
│ "Search over LIVE, ACTUAL (OBF) IRL conditions as INPUTS, to │
│ extrapolate the closest-to-absolute-best strategy → RECOMMEND" │
│ │
│ ACTUALS-DOMINANT. The live book/regime/funding/latency IS the │
│ query. The ecology is now a PRIOR over who is on the other side │
│ right now, not a population to be swept. │
│ │
│ Given: live MarketWorldState (S1 book, S2 funding/dvol, S3 regime, │
│ S4 latency, account, intent) │
│ Ask: which policy in the pool is nearest-optimal for THIS state? │
│ Output: a RECOMMENDATION (policy + its expected behaviour + │
│ confidence + the nearest explored neighbours it interpolates│
│ between). │
│ Cadence: ONLINE, ≤25 ms planning inside ≤100 ms end-to-end. │
│ Success metric: live outcome ≈ predicted outcome (Rule 10). │
└──────────────────────────────────────────────────────────────────────┘
Why this is the right shape
- Mode 1 without Mode 2 is a beautiful simulator nobody trades.
- Mode 2 without Mode 1 is a lookup table over history — the overfit that kills quants. It can only recommend what already happened.
- Together: Mode 1 explores a space vastly larger than any tape can contain (that is the edge — we out-play participants who only ever fit history), and Mode 2 localizes the live moment inside that explored space and returns the best-known play, including for market states we have never actually seen, because Mode 1 has already played them.
The word "extrapolate" in the operator's phrasing is load-bearing. Mode 2's job is NOT nearest-neighbour lookup. It is: place the live state inside the performance manifold Mode 1 built, and interpolate/extrapolate the best play. That requires Mode 1 to output a manifold, not a leaderboard:
| Mode 1 must emit | Not just |
|---|---|
| policy × (regime, spread, depth, toxicity, latency, funding, inventory) → expected outcome + variance | "policy #7 scored 8,628" |
| the boundaries of where each policy was tested (extrapolation beyond = flagged) | a single global champion |
| why a policy wins (which adversary it beats, which it loses to) | an opaque score |
Consequences for the build
ScenarioLibrarymust SWEEP, not sample. Mode 1's job is coverage of the state space, including regions our tape never visited. Tape-derived scenarios (CLASS-M/X) are the anchors; the sweep fills between and beyond them.- The performance matrix in
training/selector.pybecomes the manifold — and it must carry confidence + support count + distance-to-nearest-explored per cell. A recommendation from a thinly-explored cell must SAY SO. - Mode 2 must refuse to extrapolate too far. If the live state is outside
Mode 1's explored envelope (novel regime, latency beyond p99.9, depth collapse
never swept), the honest answer is "OUT OF DISTRIBUTION — fall back to the
doctrinal simple policy", not a confident recommendation. This is the same
law as
INDETERMINATE: unknown is not flat, and unknown is not "best guess". Wire it as an explicitRecommendationConfidence.OUT_OF_DISTRIBUTIONverdict that the risk gate honours. - Mode 2's live search is a LOCALIZATION, not a fresh MCTS from scratch — the ≤25 ms budget buys you refinement around a pool policy, not global search. The pool is the compressed product of the billions of Mode-1 trials; Mode 2 spends its milliseconds choosing and adapting, not re-deriving.
- Mode 2's own outcomes feed back into Mode 1 as new anchors (live
discrepancies →
live_discrepanciestable → next sweep is denser where we were wrong). That is the learning loop, and Rule 10 governs it: where shadow and live disagree, live is right and the manifold gets corrected.
Module ownership by mode
| Module | Mode 1 (EXPLORE) | Mode 2 (RECOMMEND) |
|---|---|---|
cwm/core.py |
✅ the arena | ✅ the local rollout model |
counterparties.py (ecology) |
✅ swept adversary population | ✅ prior over current opponents |
training/cma_trainer.py, generator.py |
✅ | — |
ScenarioLibrary (§4) |
✅ sweeps | — |
training/registry.py (policy pool) |
✅ writes | ✅ reads |
training/selector.py (→ manifold) |
✅ builds | ✅ queries |
ActualsLoader (§4) |
⚠️ calibration only | ✅ the live query itself |
planner/sm_mcts.py |
✅ deep, unbounded | ✅ ≤25 ms localization |
risk/gate.py |
✅ constrains training | ✅ hard veto on recommendation |
1. FINDING #0 — THE FEE BUG IS REAL, AND MEASURED
Fable suspected a 10× unit slip. Our own fills confirm it.
SELECT liquidity_side, order_type, count() n, avg(fee_bps)
FROM dolphin.trade_execution_quality WHERE fee_bps IS NOT NULL GROUP BY 1,2
-- TAKER MARKET 1455 rows avg = 5.016 bps (min 5.000, max 6.649)
| Where | Code says | Reality (our fills / venue docs) | Factor |
|---|---|---|---|
malkhut/training/asset_classification.py:164 (BingX) |
default_taker_fee_bps=0.5 |
5.0 | 10× |
…:154 (Binance) |
default_taker_fee_bps=0.4 |
~4.5 | ~10× |
…:174 (Bybit) |
default_taker_fee_bps=0.06 |
~5.5 | ~90× |
…:384 (default profile) |
taker_fee_bps=0.5, maker_fee_bps=-0.2 |
taker 5.0; maker ≈ +2.0 (BingX perp maker is POSITIVE, not a rebate) | 10× + sign |
Impact: these feed VenueRules.taker_fee_bps / maker_fee_bps
(malkhut/state.py:109-121) → the CWM reward → w_fee_quality. CMA-ES has been
optimizing against fees an order of magnitude too cheap. Every policy trained so
far is suspect — cheap fees reward overtrading, churn, and cross-spread
aggression that real friction annihilates. This is the exact failure the
2026-07-10 venue-friction audit exists to prevent.
ACTION (do this first, before any other work):
- Fix the four sites above. Source of truth =
dolphin.trade_execution_quality(fee_bps,liquidity_side,order_type), not vendor marketing pages. - Add a mutation-litmus test: set
taker_fee_bpsto 0.5 → a test asserting "policy PnL under realistic friction" must go RED. If nothing breaks when fees change 10×, the reward function isn't actually using them. - Re-run every CMA-ES benchmark. The "8,628 / 100.4 bps PnL" headline number is void until it is re-measured at 5 bps taker.
- Maker fee sign: verify against a real maker fill before trusting a rebate.
We have zero MAKER rows in
trade_execution_quality(all 1455 are TAKER MARKET — because UV/F5 only sends MARKET). Until SMART-EXEC produces real maker fills, the maker fee is an assumption: mark it as such in code with a# UNVERIFIED — no maker fills on record as of 2026-07-13comment.
2. THE DATA SOURCES — WHAT EXISTS, VERIFIED TONIGHT
All queried live on localhost:8123. Row counts are real.
| # | Source | Rows | Columns (exact) | State |
|---|---|---|---|---|
| S1 | dolphin.obf_universe |
15,329,163,875 | ts DateTime64(3), symbol, spread_bps f32, depth_1pct_usd f64, depth_quality f32, fill_probability f32, imbalance f32, best_bid f64, best_ask f64, n_bid_levels u8, n_ask_levels u8 |
LIVE, the motherlode |
| S2 | dolphin.exf_data |
22,968,094 | ts DateTime64(6), funding_rate f32, dvol f32, fear_greed f32, taker_ratio f32 |
LIVE |
| S3 | dolphin.maras_fingerprint |
1,117,722 | ts, regime (LowCard), regime_idx u8, confidence, final_score, conflict_level, tier_exf, tier_eigen, tier_btc, tier_esof, tier_micro (+ _confidence each), s_exf_funding_bp, … |
LIVE |
| S4 | dolphin.eigen_scans |
1,529,803 | ts, scan_number u32, vel_div f32, w50_velocity, w750_velocity, instability_50, scan_to_fill_ms, step_bar_ms, scan_uuid |
LIVE — and it is the latency oracle |
| S5 | dolphin.trade_execution_quality |
8,006 | trade_id, asset, side, client_order_id, venue_order_id, order_type, liquidity_side, fee_bps, fill_quality_score, commission_quote, fee_rate, … |
LIVE — fee/fill ground truth |
| S6 | dolphin.trade_events |
(BLUE's live trades) | ts, trade_id, asset, side, entry_price, exit_price, pnl, pnl_pct, exit_reason, vel_div_entry, boost_at_entry |
LIVE |
| S7 | dolphin_uv.exec_journal |
growing | kind, timestamp, scan_number, intent_id, trade_id, slot_id, asset, side, action, reference_price, target_size, leverage, u_prefix_client_id, promo_metadata (JSON) |
LIVE |
| S8 | dolphin_uv.tp_exit_ingress / max_hold_ingress |
new (2026-07-13) | full per-scan exit diagnostics: tp_effective_pct, tp_mod_factor, tp_floor_armed, cascade_count, imbalance_ma5, branch, … |
LIVE as of today |
| S9 | dolphin.esof_advisory |
0 | schema exists (dow, session, moon_illumination, slot_wr_pct, …) |
⚠️ EMPTY |
| S10 | dolphin.obf_fast_intrade |
0 | schema exists | ⚠️ EMPTY |
Do not spec against S9/S10 as if they were populated. Either they get a writer, or they are out of scope. Say so out loud rather than building an intake for a table that never delivers a row (that is how the missing-CH-table blackout happened to UV; the fix cost a day).
3. FINDING #1 — THE BOOK-FIDELITY GAP (read this twice)
This is the central technical problem of the whole integration, and it is not optional to solve.
MALKHUT wants a ladder. malkhut/state.py:131:
class OrderBookState:
bids: Tuple[PriceLevel, ...] # price + qty, per level
asks: Tuple[PriceLevel, ...]
The CWM's queue-position model, JOIN_QUEUE, LADDER, ICEBERG, queue-ahead
estimates, and adverse-selection accounting all depend on per-level depth.
Our tape has no ladder. dolphin.obf_universe gives, per symbol per ~sub-second
tick: best_bid, best_ask, spread_bps, depth_1pct_usd (one aggregate
number), depth_quality, imbalance, fill_probability, n_bid_levels,
n_ask_levels (counts, not the levels themselves).
That is L1 + shape summary, not L2. 15.3 billion rows of it — enormously valuable, but it cannot be decoded back into a ladder. Information that was never recorded cannot be recovered.
Three honest options — pick explicitly, do not drift
| Option | What it is | Cost | Fidelity |
|---|---|---|---|
| A. Synthetic ladder w/ declared prior | Reconstruct levels from best_bid/ask + depth_1pct_usd + n_*_levels + imbalance via an explicit shape model (e.g. exponential decay of qty over levels, calibrated so Σqty within 1% = depth_1pct_usd, level count = n_bid_levels). |
LOW | Approximate. Queue position becomes a model, not a measurement. MUST be labeled as such everywhere it flows. |
| B. Start recording real L2 now | New writer → dolphin_malkhut.book_l2 (or a zinc region + spooler). Depth-20 snapshots at OBF cadence. |
MEDIUM (new hose + storage; L2 at 50 symbols × sub-second is BIG — see §7 storage) | Truth, but only from the day it starts. Cannot backfill history. |
| C. hftbacktest with external L2 | Use a third-party L2 feed (Binance archival) for CWM calibration; keep obf_universe for regime/context. | MEDIUM | Truth, but it is Binance's book, not BingX's — venue microstructure differs. |
RECOMMENDATION: A + B in parallel, and say which one a given result came from.
- A unblocks immediate use of 15.3B rows of history for distributional calibration (spread regimes, depth regimes, imbalance dynamics, fill-probability priors) — where the ladder shape matters less than the aggregate.
- B starts the clock on real queue-truth for the queue-sensitive claims
(
JOIN_QUEUE, maker fills, queue-ahead) — the very claims SMART-EXEC needs. - NEVER let an A-derived queue-position estimate be reported as a measured
fill probability. Tag every artifact:
book_source ∈ {SYNTH_A, REAL_L2_B, EXT_C}. Rule 3 (replay correctness before search depth) means exactly this.
Litmus: when B has ≥ 1 week of real L2, re-run A-calibrated policies against B-truth. The delta is the measurement of how much the synthetic prior lied. If that delta is large, every A-era conclusion is downgraded to a hypothesis.
4. THE PLUG MAP — WHAT GOES WHERE, EXACTLY
MarketWorldState (malkhut/state.py:257) is the CWM root. It already has the
right holes. Fill them from our sources:
MarketWorldState field |
Line | Source | Transform |
|---|---|---|---|
book: OrderBookState |
262 | S1 obf_universe |
§3 Option A synth (or B when live). best_bid/best_ask direct; ladder via declared prior; symbol join key. |
account: AccountState |
263 | DITAv2 ASEx account core (asex_account / AccountProjectionV2) — CONSUMER ONLY (two-cores-share-nothing law) |
For offline training: reconstruct from dolphin.account_events / dolphin_uv.exec_journal. Never a second writer. |
open_orders |
264 | live: venue adapter; training: S7 exec_journal + S5 trade_execution_quality (client_order_id join) |
|
trade_path: TradePathState |
265 | S8 tp_exit_ingress + S6 trade_events |
mae_bps/mfe_bps/time_in_loss_s etc. reconstructable from tp_exit_ingress per-scan stamps + trade lifecycle. dolphin_regime_score ← S3; book_imbalance ← S1 imbalance; orderflow_toxicity ← derive (see §5). |
intent: ExecutionIntent |
266 | DITAv2 KernelIntent (binding, README §B.1) |
Map ENTER/EXIT + reference_price/target_size/leverage/metadata.promo_client_id. urgency ← SMART-EXEC urgency class (see §8). |
funding_bps |
268 | S2 exf_data.funding_rate |
×10⁴ → bps. Nearest-ts join. |
volatility_state |
269 | S2 exf_data.dvol |
The sacred vol number. Same field the gate uses (DOLPHIN_VOL_P60_THRESHOLD, doctrinal 0.00026414). |
market_regime |
270 | S3 maras_fingerprint.regime |
⚠️ taxonomy mismatch — see §6. |
feed_latency_ms |
272 | S4 eigen_scans.step_bar_ms |
p50 = 0.08 ms. |
order_latency_ms |
273 | S4 eigen_scans.scan_to_fill_ms |
⚠️ SAMPLE THE DISTRIBUTION, NOT THE MEAN — see §5. |
venue: VenueRules |
261 | S5 for fees (§1); prod/bingx/ for tick/lot/min_notional |
Fee fix is blocking. |
Where the code changes go
| New module | Path | Job |
|---|---|---|
ActualsLoader |
malkhut/data/actuals.py (new pkg malkhut/data/) |
CH HTTP reader → typed frames. One method per source S1–S8. Polars, chunked, no full-table loads. |
BookSynthesizer |
malkhut/data/book_synth.py |
§3 Option A. Declared prior, unit-tested against any real L2 we get. Emits book_source tag. |
LatencyOracle |
malkhut/data/latency.py |
§5. Empirical CDF sampler, seed-pinned. |
EcologyFitter |
malkhut/data/ecology_fit.py |
§7. Fits counterparty params to tape. Does not create or delete agent types. |
RegimeBridge |
malkhut/data/regime_bridge.py |
§6. MARAS ↔ MALKHUT regime mapping, explicit and tested. |
ScenarioLibrary |
malkhut/data/scenarios.py |
§7. CLASS-M/CLASS-X hour selection from tape → adversarial scenario suites. |
Wire them at: malkhut/training/cma_trainer.py (scenario source),
malkhut/cwm/core.py:transition (latency + venue rules injection),
malkhut/counterparties.py (fitted params in, agent classes unchanged).
5. FINDING #2 — LATENCY IS NOT A NUMBER, IT IS A MONSTER-TAILED DISTRIBUTION
Measured tonight from dolphin.eigen_scans (scan_to_fill_ms, n = 1.5M):
| p50 | p95 | p99 | max |
|---|---|---|---|
| 49.9 ms | 596.9 ms | 10,793.9 ms | 450,115.9 ms (7.5 minutes) |
step_bar_ms p50 = 0.08 ms.
The README's "typical_latency_ms: 100" for BingX is a fiction that will get us killed: it is 2× the median and 0.9% of the mass is beyond 10 SECONDS. A policy trained on constant-100ms latency has never met the market that actually fills our orders. The p99 is what turns a maker quote into an adverse-selected gift, and it is exactly where the "why did I get filled?" question (Rule 2) gets its ugliest answer.
ACTION: LatencyOracle must sample the empirical CDF, per-regime where
possible (latency and stress correlate — verify), seed-pinned for determinism
(Rule 9 + our seed-pin doctrine). Constant-latency mode may exist only as a
labeled ablation, never as the default. Add a scenario class: LATENCY_STORM
(sample exclusively from > p95) — if a policy's edge evaporates there, we need to
know before capital does.
6. FINDING #3 — REGIME TAXONOMY MISMATCH
MARAS actually emits (verified, dolphin.maras_fingerprint, 1.1M rows):
regime_idx |
regime |
rows |
|---|---|---|
| 1 | BEARISH | 72,479 |
| 2 | CHOPPY_BEARISH | 516,477 |
| 3 | CHOPPY | 295,366 |
| 4 | SIDEWAYS | 201,358 |
| 5 | CHOPPY_BULLISH | 32,046 |
(idx 0, 6, 7 unobserved in this window — the taxonomy has room MARAS has not used.)
MALKHUT's StrategySelector uses a different, invented set: trending_up,
trending_down, high_volatility, low_volatility, mean_reverting,
momentum, choppy, liquidity_hole, normal.
These are two different languages. A performance matrix keyed on MALKHUT's regimes cannot be looked up from a MARAS fingerprint without a mapping, and an implicit/lossy mapping will silently mis-select strategies in production.
ACTION: RegimeBridge (malkhut/data/regime_bridge.py) with an explicit,
tested, total mapping MARAS→MALKHUT. Where MALKHUT has an axis MARAS lacks
(e.g. liquidity_hole), derive it from S1 (depth_quality, spread_bps
percentiles) and name the derivation. Where MARAS has confidence
(confidence, conflict_level), carry it through — a low-confidence regime
tag must reach the planner as low-confidence, not as a hard label. Prefer
extending MALKHUT to speak MARAS over inventing a third dialect.
Also plumb the MARAS tiers (tier_exf, tier_eigen, tier_btc,
tier_esof, tier_micro + confidences) into the feature vector — that is a
5-tier ensemble view of the market the planner currently cannot see at all, and
it is already computed and stored. Free signal.
7. THE ECOLOGY — HOW ACTUALS FEED IT WITHOUT REPLACING IT
This section is the point of the whole document. Read it as law.
malkhut/counterparties.py defines the agent ecology (AgentRole,
state.py:61): PASSIVE_MAKER, TOXIC_TAKER, LATENCY_ARB, MOMENTUM_TAKER,
MEAN_REVERSION_TAKER, INVENTORY_MM, LIQUIDATION_FLOW, NOISE_TRADER,
STALE_QUOTE_ATTACKER. Nine roles declared, 4 implemented.
What actuals DO to the ecology
| Do | How |
|---|---|
| Fit each agent's parameters | EcologyFitter estimates population parameters from tape: TOXIC_TAKER arrival intensity ← S1 fill_probability collapse + adverse post-fill drift in S5/S6; LATENCY_ARB ← S4 latency tails vs price moves; LIQUIDATION_FLOW ← S2 funding extremes + cascade signatures in S8 cascade_count; INVENTORY_MM ← S1 imbalance mean-reversion; MOMENTUM_TAKER ← S4 vel_div / w50_velocity. |
| Set the population MIX per regime | Which agents dominate under BEARISH vs CHOPPY (S3) is an empirical question our tape can answer. The mix becomes regime-conditional. |
| Supply the stress scenarios | CLASS-M / CLASS-X hours (measured hostile windows) → ScenarioLibrary. CLASS-X law: extreme magnitudes are a FEATURE channel, never filtered. |
| Bound the plausible | Tape says what a real book can do; adversaries should be allowed to be worse, but their base rates must be anchored. |
What actuals MUST NOT DO
| Never | Why |
|---|---|
| Replace an agent with a tape replay | A replayed counterparty cannot react to us. The entire adverse-selection question ("why did I get filled?") requires an opponent that responds. A tape is a corpse; the ecology is an opponent. |
| Delete an agent type for lack of data | The 5 unimplemented roles are hypotheses about who is on the other side. Absence of evidence in our thin tape is not evidence of absence in the market. Implement them; fit what you can; make the rest adversarial priors. |
| Restrict the game to observed histories | Billions of combinations, most of which never occurred, is the product. Recorded tape is a measure-zero slice of the strategy × counterparty × regime space. Optimizing only over what happened is exactly the overfit that kills quants. |
| Filter outliers | CLASS-X law. The 450-second latency tail and the 7.5-minute stop are not noise — they are the market's teeth. |
The scale ambition (operator's stated aim)
Femtosecond-scale trials, parallel, over billions of combinations. Current: 5.3 µs/CWM-transition (numba), 16 ms/episode (parallel), 7× on 8 workers. Billions of combinations × ~10³ steps at 5.3 µs = ~10¹² × 5.3 µs ≈ months of single-box CPU. So:
- Batch/vectorize the counterparty step (the ecology is the inner loop — make it a matrix op over the agent population, not a Python loop over agents).
- GPU or massively-parallel path for the rollout kernel (the batch MCTS kernel already exists — push it further).
- Cheap-then-expensive cascade: screen billions with a cheap surrogate (vectorized reward, shallow rollout), promote survivors to full CWM depth. Do NOT run 5.3 µs × depth-3 MCTS on every one of a billion candidates.
- Ray path already exists (
training/ray_eval.py) — that is the horizontal scaling seam when the box runs out. - Honest note: "femtosecond" is a metaphor for parallel breadth, not a physical claim (a CPU cycle is ~300 picoseconds; a femtosecond is 10⁻¹⁵ s — light travels 0.3 µm in one). State the real target: N trials/second at fidelity F, and measure it. Otherwise we cannot tell progress from poetry.
8. INTAKE CONTRACT ALIGNMENT (with SMART-EXEC + UV)
MALKHUT's planner and SPEC_UV_SMART_EXEC_MM.md describe the same seat at the
venue-adapter boundary. Do not build two doorframes.
- Input: DITAv2
KernelIntent+guideline_price+urgency_class(CATASTROPHIC / PROTECT / HARVEST / ROTATE / ACQUIRE). Mapurgency_class→ExecutionIntent.urgency(state.py:246) andprefer_maker(:250). CATASTROPHIC ⇒ the planner is BYPASSED. No cleverness on the stop path, ever. u-clientOrderId prefix is LAW (T9 seam / README §B.2).- DUAL-LEVERAGE LAW (README §B.3): conviction leverage [0.5,9.0] sizes
quantity; venue leverage =
prod/bingx/leverage.pymapping → int [1,3]. Litmus: 0.5→1, 4.75→2, 8.0→3, 9.0→3. Note: we observed a live 3→2 clamp on 2026-07-13 (handover anomaly #2) — reconcile before MALKHUT sizes anything. - Execution-truth doctrine is inherited wholesale (
bdc54fb): NOT_ATTEMPTED / REFUSED → rollback sound; INDETERMINATE → never. MALKHUT's venue adapter (malkhut/venue/bingx/adapter.py) wraps DITAv2 and therefore inherits the fences — verify with a test, do not assume. SeeCOMPREHENSIVE_UV_EDGE_CASE_THEORETICALS.md§11 M1–M7 for the unaudited seams. - No reconcilers. Bounded, read-only, own-clientOrderId point lookups only.
9. PERSISTENCE + PROVENANCE
- Namespace
dolphin_malkhut(approved). NO TTL (retention doctrine). - DDL ships WITH the code, and the applier's verify-set must require every new
table — the 2026-07-13 lesson: two UV tables existed as
.sqlfiles for days while the runner 404'd every scan and journaled nothing. Pattern to copy:prod/clickhouse/uv/apply_uv_ddl.py(EXPECTED_TABLES). - Every row carries provenance:
book_source(SYNTH_A/REAL_L2_B/EXT_C),latency_model(EMPIRICAL_CDF/CONST),fee_table_version,ecology_fit_version,policy_version,seed. A result whose provenance is unknown is not a result. - Cross-reference law:
fulfilment_decisions.intent_id+trade_idjoin todolphin_uv.exec_journal⇄ venue order history. End-to-end or it is a black box.
10. ORDER OF WORK (do not reorder)
| # | Task | Gate |
|---|---|---|
| 1 | Fix the fees (§1) + mutation-litmus test | A 10× fee change must break a test |
| 2 | Re-baseline every CMA-ES number at real friction | Old headline numbers marked VOID |
| 3 | LatencyOracle (§5) + LATENCY_STORM scenario |
Constant-latency is an ablation, not a default |
| 4 | Decide the book question (§3) — declare A, B, or C in writing | book_source tag on every artifact |
| 5 | RegimeBridge (§6) — explicit total mapping | Test: every MARAS regime maps; confidence carried |
| 6 | ActualsLoader (§4) — S1/S2/S3/S4 into MarketWorldState |
Replay determinism ×2 (byte-identical minus ts) |
| 7 | EcologyFitter (§7) — fit params, keep all agent types | Test: fitting cannot reduce the agent-type count |
| 8 | ScenarioLibrary — CLASS-M/X hours as suites | CLASS-X magnitudes unfiltered |
| 9 | Implement the 5 missing agent roles | Ecology completeness |
| 10 | Scale the inner loop (§7) — vectorize ecology, cascade screening | Measured trials/sec, not adjectives |
| 11 | MODE 1 → manifold (§0.5): selector emits variance + support + envelope, not a leaderboard | A thin cell must report itself thin |
| 12 | MODE 2 → localization + OOD verdict (§0.5): live state query, OUT_OF_DISTRIBUTION falls back to doctrinal policy |
Test: a novel regime must NOT get a confident recommendation |
| 13 | Gate-M (README §F): replay verify → dual-leverage litmus → DARK shadow vs naive → risk-gate mutation litmus → Rule 10 | Money |
11. THE STANDING GUARDRAILS (Fable's, HJ's, and the house's)
- Rule 3 supreme: replay correctness before search depth. A wrong CWM plus a deep search is confident nonsense, and confident nonsense is the most expensive thing we can build.
- Rule 10 supreme: shadow vs live diverges → trust live, stand down, file the discrepancy row.
- Mutation litmus everywhere: 1,186 green tests prove nothing until breaking the implementation turns one red. Fees, latency, risk-gate constraints, ecology size — each must have a test that dies when it is sabotaged.
- "It looks finished" is the warning sign, not the green light.
- The ecology stays. If a future refactor makes the counterparties smaller, fewer, or tamer in the name of realism, it has removed the edge and kept the costs. That refactor is wrong. This line is the reason this document exists.
Written by Claude (Opus 4.8) on Fable's context, at HJ's order, 2026-07-13. Every path, line number, row count, and measured value herein was verified against the live system on the night of writing. Where a number was not verified, it says so.