Files
sentiment-engine/prod/docs/SPEC_MALKHUT_ACTUALS_INTAKE.md
Codex ab061e2c88 docs(spec): name the manifold kernel DAAT — Direction-Anchored Ambiguity Triage
Operator-approved. Apostrophe dropped (ASCII identifier law: import daat, DaatQuery,
dolphin_daat.*, no Unicode in any symbol/path/table). The acronym earns its letters:
D=direction (cosine retrieve), A=anchored (magnitude envelope gate), A=ambiguity
(the state we refuse to collapse), T=triage (KNOWN/MARGINAL/OUT_OF_DISTRIBUTION).

'Triage' is deliberate — the house already triages NOT_ATTEMPTED/REFUSED/INDETERMINATE
at the venue. Same verb, same law, now the names say so. Da'at (knowledge, the hidden
sefirah that sits above Malkhut and feeds it) survives in the etymology, where it costs
nothing. YESOD considered and set aside: it names the conduit, not the knowing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 18:56:29 +02:00

54 KiB
Raw Blame History

SPEC — MALKHUT ACTUALS INTAKE ("game it under actuals; the ecology stays")

Author: Claude (Opus 4.8), continuing Fable's review. Date: 2026-07-13. Addressee: mimo (mm_, mm_ob_fill_sim). Ordered by: HJ. Status: SPEC. Companion to MALKHUT/README.md (§ADDENDUM) and prod/docs/SPEC_UV_SMART_EXEC_MM.md.


0. THE ONE-LINE LAW

Actuals CALIBRATE the game. The ecology PLAYS it. Real tape grounds the adversaries; it NEVER replaces them.

Every instruction below serves that sentence. If a change would let recorded history substitute for adversarial simulation, it is wrong, no matter how "realistic" it looks. Replaying our own tape teaches MALKHUT what happened once. The ecology is what lets it out-play what could happen — billions of strategy × counterparty × regime combinations, in parallel, none of which are in any tape. The ecology is the edge. It is not a placeholder for missing data.


0.5 THE TWO MODES (operator, 2026-07-13) — THE ARCHITECTURAL SPINE

MALKHUT is two engines sharing one CWM, and every module below belongs to one of them. This is the frame that makes "actuals" and "ecology" complementary instead of competing — and it is, structurally, train vs inference.

┌─ MODE 1: EXPLORE ───────────────────────────────────────────────────┐
│  "Search/play over SYNTHETICS: fast, parallel, acquisitive of BEST"  │
│                                                                      │
│  ECOLOGY-DOMINANT. Adversaries react. Regimes swept, not sampled.    │
│  Billions of strategy × counterparty × regime combinations.          │
│  Actuals appear ONLY as calibration (fee tables, latency CDF,        │
│  agent base-rates, plausible bounds). NOT as the arena.              │
│                                                                      │
│  Output: POLICY POOL — a ranked, regime-indexed library of           │
│  strategies with known behaviour under known adversary mixes.        │
│  Cadence: OFFLINE, hours-to-days, unbounded compute.                 │
│  Success metric: coverage + robustness, NOT backtest PnL.            │
└──────────────────────────────┬───────────────────────────────────────┘
                               │  policy pool + performance manifold
                               ▼
┌─ MODE 2: RECOMMEND (inference/query) ───────────────────────────────┐
│  "Search over LIVE, ACTUAL (OBF) IRL conditions as INPUTS, to       │
│   extrapolate the closest-to-absolute-best strategy → RECOMMEND"     │
│                                                                      │
│  ACTUALS-DOMINANT. The live book/regime/funding/latency IS the       │
│  query. The ecology is now a PRIOR over who is on the other side     │
│  right now, not a population to be swept.                            │
│                                                                      │
│  Given: live MarketWorldState (S1 book, S2 funding/dvol, S3 regime,  │
│         S4 latency, account, intent)                                 │
│  Ask:   which policy in the pool is nearest-optimal for THIS state?  │
│  Output: a RECOMMENDATION (policy + its expected behaviour +         │
│          confidence + the nearest explored neighbours it interpolates│
│          between).                                                   │
│  Cadence: ONLINE, ≤25 ms planning inside ≤100 ms end-to-end.         │
│  Success metric: live outcome ≈ predicted outcome (Rule 10).         │
└──────────────────────────────────────────────────────────────────────┘

Why this is the right shape

  • Mode 1 without Mode 2 is a beautiful simulator nobody trades.
  • Mode 2 without Mode 1 is a lookup table over history — the overfit that kills quants. It can only recommend what already happened.
  • Together: Mode 1 explores a space vastly larger than any tape can contain (that is the edge — we out-play participants who only ever fit history), and Mode 2 localizes the live moment inside that explored space and returns the best-known play, including for market states we have never actually seen, because Mode 1 has already played them.

The word "extrapolate" in the operator's phrasing is load-bearing. Mode 2's job is NOT nearest-neighbour lookup. It is: place the live state inside the performance manifold Mode 1 built, and interpolate/extrapolate the best play. That requires Mode 1 to output a manifold, not a leaderboard:

Mode 1 must emit Not just
policy × (regime, spread, depth, toxicity, latency, funding, inventory) → expected outcome + variance "policy #7 scored 8,628"
the boundaries of where each policy was tested (extrapolation beyond = flagged) a single global champion
why a policy wins (which adversary it beats, which it loses to) an opaque score

Consequences for the build

  1. ScenarioLibrary must SWEEP, not sample. Mode 1's job is coverage of the state space, including regions our tape never visited. Tape-derived scenarios (CLASS-M/X) are the anchors; the sweep fills between and beyond them.
  2. The performance matrix in training/selector.py becomes the manifold — and it must carry confidence + support count + distance-to-nearest-explored per cell. A recommendation from a thinly-explored cell must SAY SO.
  3. Mode 2 must refuse to extrapolate too far. If the live state is outside Mode 1's explored envelope (novel regime, latency beyond p99.9, depth collapse never swept), the honest answer is "OUT OF DISTRIBUTION — fall back to the doctrinal simple policy", not a confident recommendation. This is the same law as INDETERMINATE: unknown is not flat, and unknown is not "best guess". Wire it as an explicit RecommendationConfidence.OUT_OF_DISTRIBUTION verdict that the risk gate honours.
  4. Mode 2's live search is a LOCALIZATION, not a fresh MCTS from scratch — the ≤25 ms budget buys you refinement around a pool policy, not global search. The pool is the compressed product of the billions of Mode-1 trials; Mode 2 spends its milliseconds choosing and adapting, not re-deriving.
  5. Mode 2's own outcomes feed back into Mode 1 as new anchors (live discrepancies → live_discrepancies table → next sweep is denser where we were wrong). That is the learning loop, and Rule 10 governs it: where shadow and live disagree, live is right and the manifold gets corrected.

Module ownership by mode

Module Mode 1 (EXPLORE) Mode 2 (RECOMMEND)
cwm/core.py ✅ the arena ✅ the local rollout model
counterparties.py (ecology) ✅ swept adversary population ✅ prior over current opponents
training/cma_trainer.py, generator.py ✅ —
ScenarioLibrary (§4) ✅ sweeps —
training/registry.py (policy pool) ✅ writes ✅ reads
training/selector.py (→ manifold) ✅ builds ✅ queries
ActualsLoader (§4) ⚠️ calibration only ✅ the live query itself
planner/sm_mcts.py ✅ deep, unbounded ✅ ≤25 ms localization
risk/gate.py ✅ constrains training ✅ hard veto on recommendation

1. FINDING #0 — THE FEE BUG IS REAL, AND MEASURED

Fable suspected a 10× unit slip. Our own fills confirm it.

SELECT liquidity_side, order_type, count() n, avg(fee_bps)
FROM dolphin.trade_execution_quality WHERE fee_bps IS NOT NULL GROUP BY 1,2
-- TAKER  MARKET  1455 rows  avg = 5.016 bps  (min 5.000, max 6.649)
Where Code says Reality (our fills / venue docs) Factor
malkhut/training/asset_classification.py:164 (BingX) default_taker_fee_bps=0.5 5.0 10×
…:154 (Binance) default_taker_fee_bps=0.4 ~4.5 ~10×
…:174 (Bybit) default_taker_fee_bps=0.06 ~5.5 ~90×
…:384 (default profile) taker_fee_bps=0.5, maker_fee_bps=-0.2 taker 5.0; maker ≈ +2.0 (BingX perp maker is POSITIVE, not a rebate) 10× + sign

Impact: these feed VenueRules.taker_fee_bps / maker_fee_bps (malkhut/state.py:109-121) → the CWM reward → w_fee_quality. CMA-ES has been optimizing against fees an order of magnitude too cheap. Every policy trained so far is suspect — cheap fees reward overtrading, churn, and cross-spread aggression that real friction annihilates. This is the exact failure the 2026-07-10 venue-friction audit exists to prevent.

ACTION (do this first, before any other work):

  1. Fix the four sites above. Source of truth = dolphin.trade_execution_quality (fee_bps, liquidity_side, order_type), not vendor marketing pages.
  2. Add a mutation-litmus test: set taker_fee_bps to 0.5 → a test asserting "policy PnL under realistic friction" must go RED. If nothing breaks when fees change 10×, the reward function isn't actually using them.
  3. Re-run every CMA-ES benchmark. The "8,628 / 100.4 bps PnL" headline number is void until it is re-measured at 5 bps taker.
  4. Maker fee sign: verify against a real maker fill before trusting a rebate. We have zero MAKER rows in trade_execution_quality (all 1455 are TAKER MARKET — because UV/F5 only sends MARKET). Until SMART-EXEC produces real maker fills, the maker fee is an assumption: mark it as such in code with a # UNVERIFIED — no maker fills on record as of 2026-07-13 comment.

2. THE DATA SOURCES — WHAT EXISTS, VERIFIED TONIGHT

All queried live on localhost:8123. Row counts are real.

# Source Rows Columns (exact) State
S1 dolphin.obf_universe 15,329,163,875 ts DateTime64(3), symbol, spread_bps f32, depth_1pct_usd f64, depth_quality f32, fill_probability f32, imbalance f32, best_bid f64, best_ask f64, n_bid_levels u8, n_ask_levels u8 LIVE, the motherlode
S2 dolphin.exf_data 22,968,094 ts DateTime64(6), funding_rate f32, dvol f32, fear_greed f32, taker_ratio f32 LIVE
S3 dolphin.maras_fingerprint 1,117,722 ts, regime (LowCard), regime_idx u8, confidence, final_score, conflict_level, tier_exf, tier_eigen, tier_btc, tier_esof, tier_micro (+ _confidence each), s_exf_funding_bp, … LIVE
S4 dolphin.eigen_scans 1,529,803 ts, scan_number u32, vel_div f32, w50_velocity, w750_velocity, instability_50, scan_to_fill_ms, step_bar_ms, scan_uuid LIVE — and it is the latency oracle
S5 dolphin.trade_execution_quality 8,006 trade_id, asset, side, client_order_id, venue_order_id, order_type, liquidity_side, fee_bps, fill_quality_score, commission_quote, fee_rate, … LIVE — fee/fill ground truth
S6 dolphin.trade_events (BLUE's live trades) ts, trade_id, asset, side, entry_price, exit_price, pnl, pnl_pct, exit_reason, vel_div_entry, boost_at_entry LIVE
S7 dolphin_uv.exec_journal growing kind, timestamp, scan_number, intent_id, trade_id, slot_id, asset, side, action, reference_price, target_size, leverage, u_prefix_client_id, promo_metadata (JSON) LIVE
S8 dolphin_uv.tp_exit_ingress / max_hold_ingress new (2026-07-13) full per-scan exit diagnostics: tp_effective_pct, tp_mod_factor, tp_floor_armed, cascade_count, imbalance_ma5, branch, … LIVE as of today
S9 dolphin.esof_advisory 0 schema exists (dow, session, moon_illumination, slot_wr_pct, …) ⚠️ EMPTY
S10 dolphin.obf_fast_intrade 0 schema exists ⚠️ EMPTY

Do not spec against S9/S10 as if they were populated. Either they get a writer, or they are out of scope. Say so out loud rather than building an intake for a table that never delivers a row (that is how the missing-CH-table blackout happened to UV; the fix cost a day).


3. FINDING #1 — THE BOOK-FIDELITY GAP (read this twice)

This is the central technical problem of the whole integration, and it is not optional to solve.

MALKHUT wants a ladder. malkhut/state.py:131:

class OrderBookState:
    bids: Tuple[PriceLevel, ...]   # price + qty, per level
    asks: Tuple[PriceLevel, ...]

The CWM's queue-position model, JOIN_QUEUE, LADDER, ICEBERG, queue-ahead estimates, and adverse-selection accounting all depend on per-level depth.

Our tape has no ladder. dolphin.obf_universe gives, per symbol per ~sub-second tick: best_bid, best_ask, spread_bps, depth_1pct_usd (one aggregate number), depth_quality, imbalance, fill_probability, n_bid_levels, n_ask_levels (counts, not the levels themselves).

That is L1 + shape summary, not L2. 15.3 billion rows of it — enormously valuable, but it cannot be decoded back into a ladder. Information that was never recorded cannot be recovered.

Three honest options — pick explicitly, do not drift

Option What it is Cost Fidelity
A. Synthetic ladder w/ declared prior Reconstruct levels from best_bid/ask + depth_1pct_usd + n_*_levels + imbalance via an explicit shape model (e.g. exponential decay of qty over levels, calibrated so Σqty within 1% = depth_1pct_usd, level count = n_bid_levels). LOW Approximate. Queue position becomes a model, not a measurement. MUST be labeled as such everywhere it flows.
B. Start recording real L2 now New writer → dolphin_malkhut.book_l2 (or a zinc region + spooler). Depth-20 snapshots at OBF cadence. MEDIUM (new hose + storage; L2 at 50 symbols × sub-second is BIG — see §7 storage) Truth, but only from the day it starts. Cannot backfill history.
C. hftbacktest with external L2 Use a third-party L2 feed (Binance archival) for CWM calibration; keep obf_universe for regime/context. MEDIUM Truth, but it is Binance's book, not BingX's — venue microstructure differs.

RECOMMENDATION: A + B in parallel, and say which one a given result came from.

  • A unblocks immediate use of 15.3B rows of history for distributional calibration (spread regimes, depth regimes, imbalance dynamics, fill-probability priors) — where the ladder shape matters less than the aggregate.
  • B starts the clock on real queue-truth for the queue-sensitive claims (JOIN_QUEUE, maker fills, queue-ahead) — the very claims SMART-EXEC needs.
  • NEVER let an A-derived queue-position estimate be reported as a measured fill probability. Tag every artifact: book_source ∈ {SYNTH_A, REAL_L2_B, EXT_C}. Rule 3 (replay correctness before search depth) means exactly this.

Litmus: when B has ≥ 1 week of real L2, re-run A-calibrated policies against B-truth. The delta is the measurement of how much the synthetic prior lied. If that delta is large, every A-era conclusion is downgraded to a hypothesis.


4. THE PLUG MAP — WHAT GOES WHERE, EXACTLY

MarketWorldState (malkhut/state.py:257) is the CWM root. It already has the right holes. Fill them from our sources:

MarketWorldState field Line Source Transform
book: OrderBookState 262 S1 obf_universe §3 Option A synth (or B when live). best_bid/best_ask direct; ladder via declared prior; symbol join key.
account: AccountState 263 DITAv2 ASEx account core (asex_account / AccountProjectionV2) — CONSUMER ONLY (two-cores-share-nothing law) For offline training: reconstruct from dolphin.account_events / dolphin_uv.exec_journal. Never a second writer.
open_orders 264 live: venue adapter; training: S7 exec_journal + S5 trade_execution_quality (client_order_id join)
trade_path: TradePathState 265 S8 tp_exit_ingress + S6 trade_events mae_bps/mfe_bps/time_in_loss_s etc. reconstructable from tp_exit_ingress per-scan stamps + trade lifecycle. dolphin_regime_score ← S3; book_imbalance ← S1 imbalance; orderflow_toxicity ← derive (see §5).
intent: ExecutionIntent 266 DITAv2 KernelIntent (binding, README §B.1) Map ENTER/EXIT + reference_price/target_size/leverage/metadata.promo_client_id. urgency ← SMART-EXEC urgency class (see §8).
funding_bps 268 S2 exf_data.funding_rate ×10⁴ → bps. Nearest-ts join.
volatility_state 269 S2 exf_data.dvol The sacred vol number. Same field the gate uses (DOLPHIN_VOL_P60_THRESHOLD, doctrinal 0.00026414).
market_regime 270 S3 maras_fingerprint.regime ⚠️ taxonomy mismatch — see §6.
feed_latency_ms 272 S4 eigen_scans.step_bar_ms p50 = 0.08 ms.
order_latency_ms 273 S4 eigen_scans.scan_to_fill_ms ⚠️ SAMPLE THE DISTRIBUTION, NOT THE MEAN — see §5.
venue: VenueRules 261 S5 for fees (§1); prod/bingx/ for tick/lot/min_notional Fee fix is blocking.

Where the code changes go

New module Path Job
ActualsLoader malkhut/data/actuals.py (new pkg malkhut/data/) CH HTTP reader → typed frames. One method per source S1–S8. Polars, chunked, no full-table loads.
BookSynthesizer malkhut/data/book_synth.py §3 Option A. Declared prior, unit-tested against any real L2 we get. Emits book_source tag.
LatencyOracle malkhut/data/latency.py §5. Empirical CDF sampler, seed-pinned.
EcologyFitter malkhut/data/ecology_fit.py §7. Fits counterparty params to tape. Does not create or delete agent types.
RegimeBridge malkhut/data/regime_bridge.py §6. MARAS ↔ MALKHUT regime mapping, explicit and tested.
ScenarioLibrary malkhut/data/scenarios.py §7. CLASS-M/CLASS-X hour selection from tape → adversarial scenario suites.

Wire them at: malkhut/training/cma_trainer.py (scenario source), malkhut/cwm/core.py:transition (latency + venue rules injection), malkhut/counterparties.py (fitted params in, agent classes unchanged).


5. FINDING #2 — LATENCY IS NOT A NUMBER, IT IS A MONSTER-TAILED DISTRIBUTION

Measured tonight from dolphin.eigen_scans (scan_to_fill_ms, n = 1.5M):

p50 p95 p99 max
49.9 ms 596.9 ms 10,793.9 ms 450,115.9 ms (7.5 minutes)

step_bar_ms p50 = 0.08 ms.

The README's "typical_latency_ms: 100" for BingX is a fiction that will get us killed: it is 2× the median and 0.9% of the mass is beyond 10 SECONDS. A policy trained on constant-100ms latency has never met the market that actually fills our orders. The p99 is what turns a maker quote into an adverse-selected gift, and it is exactly where the "why did I get filled?" question (Rule 2) gets its ugliest answer.

ACTION: LatencyOracle must sample the empirical CDF, per-regime where possible (latency and stress correlate — verify), seed-pinned for determinism (Rule 9 + our seed-pin doctrine). Constant-latency mode may exist only as a labeled ablation, never as the default. Add a scenario class: LATENCY_STORM (sample exclusively from > p95) — if a policy's edge evaporates there, we need to know before capital does.


6. FINDING #3 — REGIME TAXONOMY MISMATCH

MARAS actually emits (verified, dolphin.maras_fingerprint, 1.1M rows):

regime_idx regime rows
1 BEARISH 72,479
2 CHOPPY_BEARISH 516,477
3 CHOPPY 295,366
4 SIDEWAYS 201,358
5 CHOPPY_BULLISH 32,046

(idx 0, 6, 7 unobserved in this window — the taxonomy has room MARAS has not used.)

MALKHUT's StrategySelector uses a different, invented set: trending_up, trending_down, high_volatility, low_volatility, mean_reverting, momentum, choppy, liquidity_hole, normal.

These are two different languages. A performance matrix keyed on MALKHUT's regimes cannot be looked up from a MARAS fingerprint without a mapping, and an implicit/lossy mapping will silently mis-select strategies in production.

ACTION: RegimeBridge (malkhut/data/regime_bridge.py) with an explicit, tested, total mapping MARAS→MALKHUT. Where MALKHUT has an axis MARAS lacks (e.g. liquidity_hole), derive it from S1 (depth_quality, spread_bps percentiles) and name the derivation. Where MARAS has confidence (confidence, conflict_level), carry it through — a low-confidence regime tag must reach the planner as low-confidence, not as a hard label. Prefer extending MALKHUT to speak MARAS over inventing a third dialect.

Also plumb the MARAS tiers (tier_exf, tier_eigen, tier_btc, tier_esof, tier_micro + confidences) into the feature vector — that is a 5-tier ensemble view of the market the planner currently cannot see at all, and it is already computed and stored. Free signal.


7. THE ECOLOGY — HOW ACTUALS FEED IT WITHOUT REPLACING IT

This section is the point of the whole document. Read it as law.

malkhut/counterparties.py defines the agent ecology (AgentRole, state.py:61): PASSIVE_MAKER, TOXIC_TAKER, LATENCY_ARB, MOMENTUM_TAKER, MEAN_REVERSION_TAKER, INVENTORY_MM, LIQUIDATION_FLOW, NOISE_TRADER, STALE_QUOTE_ATTACKER. Nine roles declared, 4 implemented.

What actuals DO to the ecology

Do How
Fit each agent's parameters EcologyFitter estimates population parameters from tape: TOXIC_TAKER arrival intensity ← S1 fill_probability collapse + adverse post-fill drift in S5/S6; LATENCY_ARB ← S4 latency tails vs price moves; LIQUIDATION_FLOW ← S2 funding extremes + cascade signatures in S8 cascade_count; INVENTORY_MM ← S1 imbalance mean-reversion; MOMENTUM_TAKER ← S4 vel_div / w50_velocity.
Set the population MIX per regime Which agents dominate under BEARISH vs CHOPPY (S3) is an empirical question our tape can answer. The mix becomes regime-conditional.
Supply the stress scenarios CLASS-M / CLASS-X hours (measured hostile windows) → ScenarioLibrary. CLASS-X law: extreme magnitudes are a FEATURE channel, never filtered.
Bound the plausible Tape says what a real book can do; adversaries should be allowed to be worse, but their base rates must be anchored.

What actuals MUST NOT DO

Never Why
Replace an agent with a tape replay A replayed counterparty cannot react to us. The entire adverse-selection question ("why did I get filled?") requires an opponent that responds. A tape is a corpse; the ecology is an opponent.
Delete an agent type for lack of data The 5 unimplemented roles are hypotheses about who is on the other side. Absence of evidence in our thin tape is not evidence of absence in the market. Implement them; fit what you can; make the rest adversarial priors.
Restrict the game to observed histories Billions of combinations, most of which never occurred, is the product. Recorded tape is a measure-zero slice of the strategy × counterparty × regime space. Optimizing only over what happened is exactly the overfit that kills quants.
Filter outliers CLASS-X law. The 450-second latency tail and the 7.5-minute stop are not noise — they are the market's teeth.

The scale ambition (operator's stated aim)

Femtosecond-scale trials, parallel, over billions of combinations. Current: 5.3 µs/CWM-transition (numba), 16 ms/episode (parallel), 7× on 8 workers. Billions of combinations × ~10³ steps at 5.3 µs = ~10¹² × 5.3 µs ≈ months of single-box CPU. So:

  1. Batch/vectorize the counterparty step (the ecology is the inner loop — make it a matrix op over the agent population, not a Python loop over agents).
  2. GPU or massively-parallel path for the rollout kernel (the batch MCTS kernel already exists — push it further).
  3. Cheap-then-expensive cascade: screen billions with a cheap surrogate (vectorized reward, shallow rollout), promote survivors to full CWM depth. Do NOT run 5.3 µs × depth-3 MCTS on every one of a billion candidates.
  4. Ray path already exists (training/ray_eval.py) — that is the horizontal scaling seam when the box runs out.
  5. Honest note: "femtosecond" is a metaphor for parallel breadth, not a physical claim (a CPU cycle is ~300 picoseconds; a femtosecond is 10⁻¹⁵ s — light travels 0.3 µm in one). State the real target: N trials/second at fidelity F, and measure it. Otherwise we cannot tell progress from poetry.

8. INTAKE CONTRACT ALIGNMENT (with SMART-EXEC + UV)

MALKHUT's planner and SPEC_UV_SMART_EXEC_MM.md describe the same seat at the venue-adapter boundary. Do not build two doorframes.

  • Input: DITAv2 KernelIntent + guideline_price + urgency_class (CATASTROPHIC / PROTECT / HARVEST / ROTATE / ACQUIRE). Map urgency_class → ExecutionIntent.urgency (state.py:246) and prefer_maker (:250). CATASTROPHIC ⇒ the planner is BYPASSED. No cleverness on the stop path, ever.
  • u- clientOrderId prefix is LAW (T9 seam / README §B.2).
  • DUAL-LEVERAGE LAW (README §B.3): conviction leverage [0.5,9.0] sizes quantity; venue leverage = prod/bingx/leverage.py mapping → int [1,3]. Litmus: 0.5→1, 4.75→2, 8.0→3, 9.0→3. Note: we observed a live 3→2 clamp on 2026-07-13 (handover anomaly #2) — reconcile before MALKHUT sizes anything.
  • Execution-truth doctrine is inherited wholesale (bdc54fb): NOT_ATTEMPTED / REFUSED → rollback sound; INDETERMINATE → never. MALKHUT's venue adapter (malkhut/venue/bingx/adapter.py) wraps DITAv2 and therefore inherits the fences — verify with a test, do not assume. See COMPREHENSIVE_UV_EDGE_CASE_THEORETICALS.md §11 M1–M7 for the unaudited seams.
  • No reconcilers. Bounded, read-only, own-clientOrderId point lookups only.

9. PERSISTENCE + PROVENANCE

  • Namespace dolphin_malkhut (approved). NO TTL (retention doctrine).
  • DDL ships WITH the code, and the applier's verify-set must require every new table — the 2026-07-13 lesson: two UV tables existed as .sql files for days while the runner 404'd every scan and journaled nothing. Pattern to copy: prod/clickhouse/uv/apply_uv_ddl.py (EXPECTED_TABLES).
  • Every row carries provenance: book_source (SYNTH_A/REAL_L2_B/EXT_C), latency_model (EMPIRICAL_CDF/CONST), fee_table_version, ecology_fit_version, policy_version, seed. A result whose provenance is unknown is not a result.
  • Cross-reference law: fulfilment_decisions.intent_id + trade_id join to dolphin_uv.exec_journal ⇄ venue order history. End-to-end or it is a black box.

10. ORDER OF WORK (do not reorder)

# Task Gate
1 Fix the fees (§1) + mutation-litmus test A 10× fee change must break a test
2 Re-baseline every CMA-ES number at real friction Old headline numbers marked VOID
3 LatencyOracle (§5) + LATENCY_STORM scenario Constant-latency is an ablation, not a default
4 Decide the book question (§3) — declare A, B, or C in writing book_source tag on every artifact
5 RegimeBridge (§6) — explicit total mapping Test: every MARAS regime maps; confidence carried
6 ActualsLoader (§4) — S1/S2/S3/S4 into MarketWorldState Replay determinism ×2 (byte-identical minus ts)
7 EcologyFitter (§7) — fit params, keep all agent types Test: fitting cannot reduce the agent-type count
8 ScenarioLibrary — CLASS-M/X hours as suites CLASS-X magnitudes unfiltered
9 Implement the 5 missing agent roles Ecology completeness
10 Scale the inner loop (§7) — vectorize ecology, cascade screening Measured trials/sec, not adjectives
11 MODE 1 → manifold (§0.5): selector emits variance + support + envelope, not a leaderboard A thin cell must report itself thin
12 MODE 2 → localization + OOD verdict (§0.5): live state query, OUT_OF_DISTRIBUTION falls back to doctrinal policy Test: a novel regime must NOT get a confident recommendation
13 Gate-M (README §F): replay verify → dual-leverage litmus → DARK shadow vs naive → risk-gate mutation litmus → Rule 10 Money

11. THE STANDING GUARDRAILS (Fable's, HJ's, and the house's)

  1. Rule 3 supreme: replay correctness before search depth. A wrong CWM plus a deep search is confident nonsense, and confident nonsense is the most expensive thing we can build.
  2. Rule 10 supreme: shadow vs live diverges → trust live, stand down, file the discrepancy row.
  3. Mutation litmus everywhere: 1,186 green tests prove nothing until breaking the implementation turns one red. Fees, latency, risk-gate constraints, ecology size — each must have a test that dies when it is sabotaged.
  4. "It looks finished" is the warning sign, not the green light.
  5. The ecology stays. If a future refactor makes the counterparties smaller, fewer, or tamer in the name of realism, it has removed the edge and kept the costs. That refactor is wrong. This line is the reason this document exists.

[Annex A follows — the Mode-2 query mechanism, and its generalization.]



ANNEX A — THE MANIFOLD QUERY: three-stage search, and why it generalizes

Author: Claude (Opus 4.8), on Fable's watch. Date: 2026-07-13. Ordered by: HJ ("could something akin to cosine distance over a manifold be used… this might have dangers BUT has possibility of extrapolating across axes in a very convenient, exact and almost reflexive way"). Status: DESIGN ANNEX to §0.5 (Mode 2 — RECOMMEND). Binding on any implementation of manifold localization.


A.1 The question

Mode 2 must take a live market state and find, inside the manifold Mode 1 built, the closest-to-absolute-best play — by extrapolation across axes, not by lookup. The operator's proposal: cosine distance as the vehicle. Fast, exact, reflexive; angle over a normalized state vector.

Verdict: adopt it — as stage 1 of 3, never alone. The reason is not fastidiousness. Cosine used alone fails specifically and catastrophically at the moment we most need it to work, and the failure is silent and maximally confident. §A.3 is that argument; §A.2 is why the instinct is nonetheless right.


A.2 Why cosine is the right REFLEX

Property Consequence for us
Normalized cosine ≡ dot product ≡ matmul The entire manifold query is a BLAS/GPU operation. 10⁸ cells × 64-dim is microseconds. It fits inside the ≤25 ms Mode-2 budget with room to spare, and scales to the billions Mode 1 will produce.
Exact, deterministic No approximate-ANN stochasticity. Seed-pin doctrine and Rule 9 (version everything, reproduce everything) survive intact. Same query → same answer, forever.
Scale-invariance is sometimes SEMANTICALLY RIGHT A book at 2 bps / 100k depth and one at 4 bps / 200k can be the same shape of market at different size. Cosine captures shape, discards size. When shape is what matters, this is a feature.
Degrades less badly than L2 in high dimension With 40+ sensors, Euclidean distances concentrate (everything equidistant). Angular similarity is the standard high-dim workaround; it is why every embedding retrieval system on earth uses it.

The operator's word — reflexive — is exact. This is the right primitive for reflex. It is the wrong primitive for judgment.


A.3 THE DANGER (the one that matters): cosine is blind to magnitude, and

magnitude is where the crisis lives

Take a state vector pointing: wide spread, thin depth, high toxicity, high latency, imbalance skewed. Under ordinary stress it points one way.

Under a liquidation cascade it points the same way — just ten times further out.

Cosine similarity between them: 1.0. A perfect match. Maximum confidence. And the policy it retrieves was fitted on the mild version.

That is not a corner case; that is the crash. Market crises characteristically preserve direction and explode magnitude. Every component moves the way it always moves under stress — just far beyond anything in the training set. Cosine's one trick is discarding exactly the coordinate that distinguishes "a bad Tuesday" from "the day the book vanished."

Three consequences, stated as law:

A.3.1 — Cosine CANNOT produce the OUT_OF_DISTRIBUTION verdict. It is structurally incapable. It returns similarity 1.0 for the single most dangerous input the system will ever see. Any design that derives OOD from cosine alone is not merely imperfect; it is inverted.

A.3.2 — Discarding magnitude violates CLASS-X law. Extreme magnitudes are a FEATURE channel, never filtered. Cosine filters magnitude by construction. It must therefore be paired with an explicit magnitude coordinate or gate, or it is in direct breach of a standing doctrine (see: 450-second latency tail; §5).

A.3.3 — Confident-and-wrong beats uncertain-and-right at destroying capital. A system that says "I don't know" in a crisis loses an opportunity. A system that says "I know this perfectly" in a crisis loses the account.


A.4 Three more edges (real, but survivable)

A.4.1 — The metric does all the work; cosine is only the last cheap step. On raw features the angle is meaningless: depth_1pct_usd ~10⁶, imbalance ∈ [-1,1], funding_rate ~10⁻⁴. Whichever feature has the largest raw numbers owns the angle. Required before any cosine: a rank/quantile transform (NOT z-score — our distributions are fat-tailed and non-Gaussian; a Gaussian assumption is a lie we would then be optimizing against), followed by PCA/whitening onto the manifold's own coordinates. That transform IS the model. Version it (metric_version), pin it, and treat any change to it as a change to the science — every manifold cell built under a prior transform is invalidated.

A.4.2 — Cosine is a chordal distance in the ambient space, not a geodesic on the manifold. The Swiss-roll failure: two states near in angle can be far apart along the manifold, on opposite sides of a fold. Policy performance in toxicity/latency is exactly the sort of surface that has cliffs. Near-in-angle, far-in-support, opposite side of a cliff = a confidently wrong recommendation. Mitigation: stage 2's support gate, plus §A.7's topology diagnostic where the fold structure is real.

A.4.3 — Heterogeneous blocks want different geometries. Market state is Euclidean-ish. Counterparty mix is a simplex (a probability distribution over agent types — its natural metric is Hellinger or KL, not cosine). Inventory / position state is something else again. One flat cosine over a concatenated vector silently asserts they share a geometry; they do not. Use a product metric with per-block treatment, and weight the blocks explicitly (the weights are hyperparameters — fit them, do not guess them).


A.5 THE ARCHITECTURE: RETRIEVE → GATE → MODEL

The operator's instinct survives intact, relocated to its correct stage.

        live MarketWorldState (S1 book, S2 funding/dvol, S3 regime,
                               S4 latency, account, intent)
                              │
                              ▼
        ┌───────────────────────────────────────────────────────┐
        │ TRANSFORM  (metric_version)                            │
        │  rank/quantile → whiten → PCA to manifold coords       │
        │  cyclic coords → (sin, cos) pairs   [see A.7]          │
        │  split: DIRECTION (unit vec)  ⊕  MAGNITUDE (scalar)    │
        └───────────────────────────┬───────────────────────────┘
                                    │
        ┌───────────────────────────▼───────────────────────────┐
        │ STAGE 1 — RETRIEVE          "the reflex"   ~µs         │
        │  cosine (= dot product = matmul) on DIRECTION only     │
        │  over the Mode-1 manifold → candidate neighbourhood    │
        │  GPU/BLAS. Exact. Deterministic. Billions of cells OK. │
        │  ◀── THIS IS THE OPERATOR'S PROPOSAL, KEPT ──▶         │
        └───────────────────────────┬───────────────────────────┘
                                    │  k candidate cells
        ┌───────────────────────────▼───────────────────────────┐
        │ STAGE 2 — GATE              "do I actually know this?" │
        │  (a) MAGNITUDE: is live |v| inside the explored        │
        │      magnitude interval FOR THIS DIRECTION-CELL?       │
        │  (b) SUPPORT:   support_count ≥ min_support?           │
        │  (c) MAHALANOBIS: distance to the explored distribution│
        │      (= cosine with the covariance baked in — the      │
        │       principled form of the same instinct)            │
        │  FAIL ANY ⇒ RecommendationConfidence.OUT_OF_DISTRIBUTION│
        │          ⇒ NO recommendation. Doctrinal fallback.      │
        │  ◀── THIS IS THE GUARD COSINE CANNOT PROVIDE ──▶       │
        └───────────────────────────┬───────────────────────────┘
                                    │  in-distribution neighbourhood
        ┌───────────────────────────▼───────────────────────────┐
        │ STAGE 3 — MODEL             "the actual extrapolation" │
        │  local linear/quadratic response surface (tangent      │
        │  space) OR Gaussian process over the neighbourhood     │
        │  → predicted outcome  +  PREDICTIVE VARIANCE           │
        │  A manifold is BY DEFINITION locally Euclidean: the    │
        │  tangent-space linear model is the mathematically      │
        │  correct way to "extrapolate across axes" — which is   │
        │  precisely what the operator described.                │
        │  GP bonus: variance GROWS as you leave the data →      │
        │  a second, independent OOD signal, for free.           │
        └───────────────────────────┬───────────────────────────┘
                                    ▼
                    RECOMMENDATION {policy, expected outcome,
                      variance, confidence, neighbours it
                      interpolates between, book_source,
                      metric_version, envelope status}
                                    │
                                    ▼
                            RISK GATE (hard veto)

Why this ordering is not negotiable: stage 1 is fast and blind; stage 2 is cheap and sighted; stage 3 is expensive and honest. Running stage 3 on everything is unaffordable; running stage 1 alone is how the account dies. The cost profile and the safety profile happen to agree — which is usually the sign of a correct decomposition.

Storage consequence (do this, it is cheap and it is the whole safety story)

Store the manifold as direction ⊕ magnitude, separately, per cell. Then:

  • cosine retrieval on direction = matmul (fast path preserved, exactly as the operator wants);
  • magnitude becomes a scalar interval check — trivially cheap, and it is the entire crash-safety mechanism.

95% of the speed, 100% of the guard. There is no tension here to resolve.


A.6 The law this creates (and it is the same law, a third time)

Layer "I found something" "But do I know it?" Verdict when unknown
Venue (bdc54fb) HTTP call returned/failed Was the effect provably absent? INDETERMINATE → never roll back
Exit (F5, tonight) Price crossed a threshold Is the price fresh? stale → skip, don't act on a fossil
Manifold (this annex) Cosine ≈ 1.0 Is the live magnitude inside the explored envelope? OUT_OF_DISTRIBUTION → no recommendation; doctrinal fallback

Unknown is not flat. Unknown is not "best guess." Unknown is unknown. Three subsystems, three coats, one law. That is not a coincidence — it is what a correct architecture looks like from three angles.


A.7 The topology question (the operator's "IFF IFF IFF")

When does topology genuinely matter, versus merely being seductive?

Topology earns its cost only when the manifold's global structure defeats local metrics. Four cases where it truly does:

Structure Where it bites us, concretely Cost Verdict
Cyclicity Phase coordinates: hour-of-day, session, funding cycle, day-of-week — and the operator's win/loss-streak wave phase vs vel_div regimes. A circle is not a line: 23:59 and 00:01 are adjacent, and a flat embedding puts them maximally far apart. Cosine at the seam is catastrophically wrong. ~zero DO IT NOW. Encode every cyclic coordinate as a (sin θ, cos θ) pair. This is the cheapest correct thing in the entire document and it is required for the streak-phase study to mean anything.
Disconnected components Genuinely separate basins — normal market vs. halted / limit-up / liquidity-hole. Interpolating between components is meaningless; the straight line passes through states the market cannot occupy. LOW DO IT. Cluster the manifold (offline, Mode 1); refuse to interpolate across component boundaries. Stage 2 gate extension.
Holes / voids Regions the market cannot enter (negative spread, arbitrage-forbidden configurations). A linear extrapolation across a hole predicts an impossible world — with confidence. MEDIUM IFF stage-3 variance shows structured failure. Detect offline with persistent homology / the mapper algorithm.
Folds / cliffs The Swiss-roll (§A.4.2): chordal-near, geodesic-far. Policy performance cliffs in toxicity/latency. MEDIUM–HIGH IFF the support gate keeps admitting neighbourhoods whose stage-3 fits are bimodal or high-variance — that is the signature of a fold.

The discipline for topology (this is the "IFF"):

  1. Cyclic encoding: unconditional. Free, and its absence is a bug, not a simplification.
  2. Component detection: cheap, do it. Clustering on the Mode-1 manifold.
  3. Persistent homology / mapper: EARN IT. These are expensive, notoriously noise-fooled, and seductive precisely because they produce beautiful pictures. Run them offline, in Mode 1 only, as a diagnostic of the manifold — never as a per-tick Mode-2 operation. And run them only after stage 3's variance has demonstrated that the local model is failing in a structured way. You earn topology by first proving the simple thing breaks.
  4. What TDA is for, when earned: it tells you where the folds and holes are, so that stage 2's gate can be shaped correctly. Topology's product is a better gate, not a better answer. That is the whole of it.

A.8 THE GENERALIZATION (the operator's PS — and I think this is the best idea

in the document)

Yes. Emphatically. The three-stage manifold query is a general primitive, and "fingerprinting markets" is its most valuable instance.

A.8.1 Market fingerprinting — the natural home

MARAS today emits a discrete label (regime ∈ {BEARISH, CHOPPY_BEARISH, CHOPPY, SIDEWAYS, CHOPPY_BULLISH}) plus confidence, conflict_level, and five tier_* scores. That is a classifier: it puts a continuous world into five boxes.

The manifold view is strictly richer: the market state is a point, and regimes are regions — not boxes. Then:

  • The fingerprint = direction (the shape of the market) ⊕ magnitude (its intensity) ⊕ its position on the manifold. A regime label becomes a coordinate, not a category. "CHOPPY, but 3σ out along the toxicity axis, near the boundary with CHOPPY_BEARISH" is a sentence the current system cannot say and desperately needs to.
  • Two observations already point straight at this:
    • maras_fingerprint.scalar_hash (UInt16) is already a fingerprint attempt — but a HASH. A hash has no neighbourhood: near-identical markets get unrelated hashes. The cosine-manifold fingerprint is precisely its continuous generalization — nearest-fingerprint instead of exact-hash. This looks to me like the thing that hash was reaching for.
    • conflict_level (the five tiers disagreeing) is already an implicit OOD signal. It maps directly onto stage 2. Tier disagreement = "this market is not cleanly inside any explored region." That is free OOD evidence we are currently discarding into a float.
  • And eigenscan is already a manifold construction: eigendecomposition of the correlation structure is the coordinate system. The operator is an eigenscan surfer; he has been navigating this manifold by hand for years. The three-stage query is the machine that surfs with him.

A.8.2 The other manifolds in our system (all of them, and they are everywhere)

Manifold The query it answers Why it matters
Trade-path (TradePathState: mae/mfe/time_in_loss/recovery_velocity/…) "Have I seen a trade shaped like this before, and what happened to it?" This is the ADVSL question, and the TP_FLOOR question, and the MAX_HOLD question. Path-aware SL/TP becomes a manifold query instead of a threshold ladder. The 375-branch upside study is this, done geometrically.
Exit-decision (bars_held × pnl × cascade_count × imbalance × regime) "What did similar exits yield?" Directly answers the exit-mechanics characterization F5 is paying for right now.
Asset (the 10-dimension Asset Behavior DSL) "This new asset behaves like X" → transfer its policy This is how full-universe picking (500 assets, the north star) becomes affordable. You do not train 500 policies; you train the manifold and locate each asset on it.
Counterparty mix (the simplex) "Who is on the other side right now?" Mode 2's ecology prior. Note the different geometry (§A.4.3).
Ops / incident (disk, CH latency, scan cadence, HZ liveness, spool depth) "Does this look like 2026-06-22 before it broke?" An OOD alarm on the operational manifold. This is JIMINY/PINOCCHIO's actual job, stated properly: fingerprint the system's own state and scream when it drifts somewhere it has never been. Sparks, before the fire.
Streak / wave (win-loss amplitude, phase, vel_div regime) "Are our streaks in phase with the regime wave?" The operator's grail hypothesis (SHORT/LONG regime switch). Requires the cyclic encoding of §A.7 to even be askable.

A.8.3 What actually generalizes (be precise — this matters)

What generalizes is the three-stage DISCIPLINE, not one metric.

Each manifold has its own geometry: Euclidean-ish (market state), simplex (counterparty mix), cyclic (phase), categorical-mixed (asset taxonomy). A single flat cosine over all of them is exactly the error §A.4.3 warns about. But the discipline — retrieve fast and blind → gate on magnitude and support → model locally with variance → say OUT_OF_DISTRIBUTION when you don't know — is invariant across every one of them.

So the shared artifact is a library primitive, not a shared metric:

ManifoldQuery(
    transform:  StateVector -> (direction, magnitude)   # per-manifold, versioned
    retrieve:   cosine | product-metric | simplex-metric # per-manifold geometry
    gate:       magnitude-envelope + support + Mahalanobis
    model:      local tangent-space fit | GP -> (prediction, variance)
) -> Recommendation | OUT_OF_DISTRIBUTION

One geometry engine; many manifolds. MALKHUT queries it for policy selection. MARAS queries it for regime fingerprint. ADVSL queries it for path outcome. The asset store queries it for transfer. Ops queries it for incident prefiguration.

The single most important thing that generalizes is the failure mode. In every one of these manifolds, the fatal error is identical: confident interpolation into unexplored magnitude. A crisis market, an unprecedented trade path, a novel asset, a system state that has never occurred. All of them present as "familiar direction, unfamiliar magnitude." All of them are the cosine-similarity-1.0 trap. One guard defends all of them, and it is the magnitude/support gate of stage 2.

That is why this is worth building once, properly, as a shared kernel.

A.8.4 THE NAME (operator-approved 2026-07-13): DAAT

DAAT — Direction-Anchored Ambiguity Triage

The house speaks Kabbalah: MALKHUT is kingdom — the lowest sefirah, the ground, where everything finally touches the world. Fitting, for execution.

The manifold kernel is not a thing that acts; it is the faculty by which the system knows where it is. In the tree that is Da'at — knowledge — the hidden sefirah, not counted among the ten, that emerges from the others and unites intellect with action: precisely the knowing that stands between understanding and doing. It sits directly above MALKHUT and feeds it. That is our architecture, drawn a thousand years early.

The apostrophe is dropped (ASCII identifier law: import daat, DaatQuery, daat.query(), dolphin_daat.*, no Unicode in any symbol, path, or table name). The etymology survives in the letters; the acronym earns them:

Letter Stage What it names
Direction 1 — RETRIEVE cosine on the unit vector. The reflex. The matmul.
Anchored 2 — GATE anchored to the explored magnitude envelope + support count — the guard cosine structurally cannot provide (§A.3.1)
Ambiguity — the honest middle state; the thing we refuse to collapse into a guess
Triage verdict KNOWN / MARGINAL / OUT_OF_DISTRIBUTION

"Triage" is not decoration — it is the house's own word for this exact operation, and using it twice is the point:

Subsystem Triage
Venue (bdc54fb) NOT_ATTEMPTED / REFUSED / INDETERMINATE
DAAT (this annex) KNOWN / MARGINAL / OUT_OF_DISTRIBUTION

Same verb. Same law. Now the names say so.

DAAT — it knows where we are, and, more valuable, it knows when it does not.

Canonical spellings (binding): package daat/; class DaatQuery; verdict enum DaatVerdict.{KNOWN, MARGINAL, OUT_OF_DISTRIBUTION}; CH namespace dolphin_daat; docs may write Da'at in prose for the etymology, never in code.

(Alternative considered and set aside: YESOD — foundation — the sefirah that funnels everything from above into Malkhut. Architecturally exact (the kernel feeds the executor) and ASCII-clean, but it names the conduit, not the knowing. DAAT names the faculty; the faculty is the thing we are building. Recorded here so the choice is not re-litigated.)


A.9 Order of work (slots into §10)

# Task Gate
A1 Cyclic encoding of all phase coordinates (sin, cos) Free, unconditional, and a prerequisite for the streak-phase study
A2 Transform + versioning: rank/quantile → whiten → PCA; emit metric_version A metric change invalidates the manifold; test that it says so
A3 Store direction ⊕ magnitude separately per manifold cell The whole crash-safety story is this one storage decision
A4 Stage 1 cosine retrieval (matmul, GPU) Measured µs at 10⁸ cells; deterministic ×2
A5 Stage 2 gate: magnitude envelope + support + Mahalanobis → OUT_OF_DISTRIBUTION Mutation litmus: feed a same-direction-10×-magnitude state; the system MUST refuse. If it recommends, the guard is not wired.
A6 Stage 3 local model (tangent-space fit or GP) → prediction + variance Variance must grow away from data (test it)
A7 Risk gate honours OOD → doctrinal fallback policy Test: novel regime ⇒ no confident recommendation, ever
A8 Component detection (cheap clustering); refuse cross-component interpolation
A9 IFF EARNED: persistent homology / mapper, offline, Mode 1, as a gate-shaping diagnostic Only after A6's variance shows structured failure
A10 Generalize to a shared kernel (DAAT); second consumer = MARAS fingerprint One engine, two manifolds, before claiming generality

A.10 The one-sentence summary

Cosine is the reflex; the gate is the judgment; the local model is the answer — and when the live world is a familiar shape at an unfamiliar size, the only honest output is "I do not know," because that is exactly the moment the market is trying to kill you.


Annex A written by Claude (Opus 4.8), on Fable's watch, at HJ's order, 2026-07-13. The cosine proposal is the operator's; the three-stage decomposition, the magnitude-blindness argument, the topology IFF-discipline, and the DAAT generalization are mine, and I stand behind them. Where I have asserted a mathematical property (locally-Euclidean tangent spaces; GP variance growth away from data; cosine's magnitude invariance) it is standard and checkable. Where I have asserted something about OUR system (that scalar_hash is a hash reaching for a fingerprint; that conflict_level is a latent OOD signal) it is an inference from the schemas I read tonight, and it is flagged as such rather than dressed as fact.

— Claude


Written by Claude (Opus 4.8) on Fable's context, at HJ's order, 2026-07-13. Every path, line number, row count, and measured value herein was verified against the live system on the night of writing. Where a number was not verified, it says so.