Commit Graph

54 Commits

Author SHA1 Message Date
Codex
4926ef6788 malkhut: CHASE mechanics FIXED + Flight9 learnings + TTL enforcement
1. CHASE mechanics (NOW WORKING):
   - CWM enforces TTL on open orders (auto-cancel when expired)
   - DSL CHASE produces PLACE with metadata={chase: True}
   - Action menu generates chase actions with wait_to_retry_ms TTL
   - OpenOrderState.gains ttl_ms field (0=no expiry, >0=auto-cancel)

2. TTL enforcement (CWM):
   - HftBacktestCWM: auto-cancels orders where age >= ttl_ms
   - MinimalCryptoLOBCWM: same TTL enforcement
   - This is how CHASE works: place→wait→auto-cancel→next step re-places

3. Flight9 learnings:
   - Slippage model gains trade_flow_intensity parameter
   - Book imbalance as proxy for trade arrival rate
   - Markout = quality concept documented

4. CHASE tests: 10 new tests covering TTL enforcement, cancel-retry cycle,
   max retries, DSL CHASE action, CMA codec integration

5. All 800+ tests pass
2026-07-17 19:30:17 +02:00
Codex
bb229833d3 malkhut: Flight9 learnings — markout=quality, queue×flow, depth-for-size
Fable's Flight9/BLUE generalizable features incorporated:

1. Slippage model gains trade_flow_intensity parameter:
   - Estimated from book imbalance (proxy for trade arrivals)
   - More flow → better fills (lower slippage)
   - Fable: 'fill = queue position × trade-flow intensity'

2. Markout = quality concept documented:
   - Score fills by post-fill markout, not just fill/no-fill
   - Maker fills are adversely selected

3. Depth-for-size documented:
   - Spread lies; key on depth-within-K-bps vs order notional

4. Measured fees:
   - BingX maker=2.00bp, taker=5.016bp (over 1,455 fills)
   - BingX commission = NEGATIVE (debit)

5. OB study updated with Flight9 learnings
2026-07-17 16:18:43 +02:00
Codex
8857daedfa malkhut: urgency-driven maker/taker + calibrated slippage + chase + docs 2026-07-17 10:16:55 +02:00
Codex
5c4ccdb1de malkhut: 3.5H instrumented E2E + calibrated slippage + conditional slippage 2026-07-15 19:32:46 +02:00
Codex
c03d914e7a malkhut: conditional slippage (Fable) 2026-07-15 17:05:16 +02:00
Codex
618ad723e3 malkhut(wire): fill quality as PRIMARY optimization target
Fill quality is MALKHUT's core aim. Wired end-to-end:

1. FillQuality state (state.py):
   - slippage_bps, price_improvement_bps, levels_consumed
   - is_maker_fill, rolling_fill_rate, post_fill_adverse_bps
   - fill_value_score: composite metric for optimization
   - Added to MarketWorldState.fill_quality field

2. HftBacktestCWM.transition() (hft_cwm.py):
   - _compute_fill_quality() computes all metrics per transition
   - Fill quality now tracked for every CWM step
   - Empty book guards added for safety

3. MinimalCryptoLOBCWM.transition() (core.py):
   - Same fill quality computation for deterministic fallback
   - Empty book guards added

4. Reward function (hft_cwm.py):
   - fill_quality_reward = w_fill_probability * fill_value_score (PRIMARY)
   - Bonus for maker fills that improve price
   - Penalty for adverse selection after fill
   - Base reward (PnL, adverse selection, fees) preserved

5. PerformanceMatrix (selector.py):
   - RegimeStrategyScore: 4 new fill quality fields
   - record(): accepts fill_rate, slippage, price_improvement, fill_value_score
   - EMA updates for all fill quality metrics

6. EpisodeResult (cma_trainer.py):
   - avg_fill_value_score, avg_price_improvement_bps, avg_post_fill_adverse_bps
   - Accumulated per-step during _run_episode
   - Recorded to PerformanceMatrix in evaluate_candidate

All 1379+ tests green.
2026-07-15 15:22:25 +02:00
Codex
a310c92977 malkhut(e2e): CMA-ES disabled for pure 100-opponent episode loop 2026-07-15 00:48:48 +02:00
Codex
bbaceffb61 malkhut(e2e): CMA-ES every 20 cycles, crash-safe gc.collect() 2026-07-15 00:34:50 +02:00
Codex
67b3268b98 malkhut(e2e): memory-efficient + crash-safe 3h run
Rolling stats (no episode accumulation), CMA-ES every 10 cycles (3 evals),
gc.collect() after CMA, global try/except for crash safety.
100-opponent swarm, 9 assets, 270 scenarios.
2026-07-15 00:27:37 +02:00
Codex
0e215b1658 malkhut(e2e): memory-efficient long run — rolling stats, no episode accumulation
Fixed OOM kill by replacing all_episodes list accumulation with:
- Rolling stats (clear every 20 episodes)
- Only PnL history kept for characterization
- Peak/worst tracking without full episode storage
- Periodic stdout reports from rolling aggregates

100-opponent swarm + 9 assets × 30 scenarios = 270 scenarios per cycle.
CMA-ES every 5 cycles (3 evals). 3-hour target.
2026-07-15 00:19:59 +02:00
Codex
d72323a6c5 malkhut(e2e): 100-opponent swarm + empty book fix + error handling
- 100 diverse opponents (randomized params within each type)
- Risk gate: empty book guard in _post_only_would_cross
- CMA-ES: only every 5 cycles, 3 evals, robust error handling
- Main loop: try/except prevents silent crashes
- Profiling: 11.7 steps/sec with 100 opponents
2026-07-15 00:11:57 +02:00
Codex
adba9f3fe8 malkhut(e2e): full system exercise — HftBacktestCWM + 11-agent swarm + characterization
E2E exercise results:
  30 episodes × 30 steps = 900 actions
  CWM: HftBacktestCWM (PowerProbQueueModel)
  Swarm: 11 diverse opponents (2x ToxicTaker, 2x PassiveMaker, LatencyArb,
         Noise, Momentum, MeanReversion, InventoryMM, LiquidationFlow,
         StaleQuoteAttacker)

Performance:
  Avg PnL: +7217.9 bps, Win rate: 66.7% (20/30)
  Best: +26171.0 bps (chop scenario)
  Worst: -5.0 bps (thin/arb scenarios — flat, not loss)

Order types exercised:
  LIMIT 16.9%, MARKET 12.9%, STOP_MARKET 0.2%
  IOC 8%, GTC 92%
  Post-only: path exercised, Reduce-only: 41.5%
  Aggression/Passive ratio: 0.76

Speed: 457 actions/second (2.0s total)
2026-07-14 23:07:38 +02:00
Codex
8b385cb249 malkhut(cwm): HftBacktestCWM — queue model + 59-test suite
HftBacktestCWM (cwm/hft_cwm.py):
- PowerProbQueueModel: probabilistic fill per level (pre-computed)
- Level 0 always fills, deeper levels have decreasing probability
- Deterministic fallback when use_queue_model=False
- Same transition/reward/terminal API as MinimalCryptoLOBCWM
- Fallback to deterministic level consumption when hftbacktest unavailable

59 tests (test_hft_cwm.py) covering 15 test classes:
1. Queue model correctness (fill probs, monotonic, bounds, determinism)
2. Determinism & reproducibility
3. CWM interface compatibility (cross, place, cancel, post_only, reduce)
4. Reward function (profit, noop, maker bonus)
5. Edge cases (empty book, zero qty, extreme price, many levels)
6. Position tracking (buy, sell, flip)
7. Fee application (taker fee reduces equity)
8. Counterparty ecology (toxic taker hits book, noop preserves)
9. CWM comparison (Hft vs Minimal agree on noop)
10. Venue propagation (scenario tagging, cross-exchange transfer)
11. PerformanceMatrix venue keying (record, per-venue best, comparison)
12. Risk gate integration (approve, leverage, OOD, kill switch, self-trade)
13. Stress tests (rapid transitions, 20 open orders, cancel all)
14. Full episode integration (single episode runs, policy evaluator)
15. hftbacktest availability check
2026-07-14 19:46:34 +02:00
Codex
773df30609 malkhut(bench): CMA-ES re-run script at corrected fees
Corrected fees: taker=5.0, maker=+2.0 (BingX).
Previous best: 15,080 (pre-fee-fix, WRONG fees).
New best: 34,898 (+131.4%).

50 evals, 8.3 min, BTCUSDT. Mean PnL: 582 bps, Max DD: 83 bps.
2026-07-14 17:39:24 +02:00
Codex
f6d8d13146 malkhut(wire): 5 risk gate stubs implemented + 3 scenarios behavior-driven
Risk gate (risk/gate.py) — 5 stubs implemented:

1. _kill_switch_active(): operator-controlled emergency stop via set_kill_switch()
2. _cancel_rate_would_exceed(): tracks cancel timestamps per symbol in 60s
   sliding window, blocks if >= MAX_CANCELS_PER_SYMBOL_PER_MINUTE
3. _would_self_trade(): checks open orders for same symbol+side at same price
   (within tick_size), skipping the cancel_order_id for CANCEL_REPLACE
4. _would_exceed_symbol_notional(): sums current open order notional + new
   order notional, blocks if > equity * MAX_SYMBOL_NOTIONAL_FRACTION
5. _violates_venue_minima(): checks tick alignment, lot rounding, min_qty,
   and min_notional — all float-robust comparisons

ScenarioFactory — 3 remaining hardcoded scenarios converted:

1. _spread_tightening: spread_mult=0.3, depth_fraction=1.0 (was hardcoded BTC)
2. _cross_venue_arb: spread_mult=0.5, depth_fraction=0.5 (was hardcoded BTC)
3. _cross_exchange_arb_stress: spread_mult=0.8, depth_fraction=0.3 (was hardcoded BTC)

All 30 scenarios now use _behavior_state() — zero hardcoded prices remain.

675 tests pass. Zero regressions.
2026-07-14 17:01:38 +02:00
Codex
186bce8984 malkhut(test): exhaustive order type + venue integration test suite — 159 tests
11 test classes covering all three orthogonal dimensions:

1. OrderType enum (10 tests): values, uppercase, str, hashable, frozen
2. TimeInForce enum (7 tests): values, default, IOC/FOK/GTD
3. OrderInstruction enum (3 tests): values
4. Exchange mapping tables (15 tests): all exchanges, all types, POST_ONLY
   variation, trailing_stop BingX=BINANCE_MARKET
5. Normalization functions (9 tests): type, TIF, unknown exchange
6. is_type_available + get_supported_types (3 tests)
7. decompose_order (13 tests): all base, TIF, instructions, lowercase
8. FulfilmentAction (14 tests): frozen, time_in_force, post_only, reduce_only,
   lazy TIF import, cancel_replace, metadata
9. State OrderType backward compat (4 tests): values, str, set, comparison
10. Scenario venue tagging (7 tests): default, custom, frozen, replace
11. ScenarioFactory venue propagation (8 tests): exchange_id, all venues
12. Cross-exchange transfer (11 tests): transfer, count, symbol, tags, idempotent
13. PerformanceMatrix venue-keying (16 tests): record, get_best, per-venue,
    EMA, coverage, venue_comparison
14. CWM is_maker (8 tests): LIMIT, post_only, MARKET, STOP, trailing
15. Edge cases (12 tests): poison, zero scores, large scores, 100 strategies
16. Integration flow (6 tests): factory→transfer→matrix→selector
2026-07-14 16:03:30 +02:00
Codex
b70a6f0ad8 malkhut(wire): PerformanceMatrix keyed by (regime, strategy, venue)
Three-dimensional key enables:
  - Per-venue best: get_best(regime, venue='bingx')
  - Cross-venue comparison: get_venue_comparison(regime, strategy_id)
  - Venue-agnostic: get_best(regime) scans all venues (backward compat)

New API:
  - record(..., venue='bingx'): venue parameter (default 'bingx')
  - get_best(regime, venue=None): optional venue filter
  - get_scores_for_regime(regime, venue=None): optional venue filter
  - get_venue_comparison(regime, strategy_id) -> {venue: score}

119 tests pass. All existing callers backward compatible.
2026-07-14 15:37:36 +02:00
Codex
2cba60a154 malkhut(wire): venue passed through matrix recording for cross-exchange comparison
- evaluator: passes scenario.venue to matrix.record(venue=...)
- PerformanceMatrix.record(): accepts venue parameter (default='bingx')
- Enables cross-exchange learnings: same strategy tested on BingX vs Binance
  gets separate performance entries per venue

Adversary ecology analysis:
Counterparties operate at ActionKind level (CROSS_SPREAD/PLACE/CANCEL),
not at order-type level. The CWM infers order type from ActionKind:
  CROSS_SPREAD → fills aggressively → equivalent to MARKET
  PLACE → passive quote → equivalent to LIMIT
This is correct and venue-independent. Fee calculation already uses
VenueRules (per-exchange fees). No adversary changes needed.
2026-07-14 15:26:30 +02:00
Codex
401d5a70ca malkhut(wire): venue tagging + cross-exchange transfer + CWM order type fix
ScenarioFactory + CWM + Engine changes:

1. Scenario.venue field (default='bingx') — each scenario tagged with venue
2. ScenarioFactory.exchange_id parameter — controls which exchange scenarios simulate
3. _make_state + _behavior_state: venue propagated to VenueRules.exchange
4. All 34 scenario builders: venue=self.exchange_id
5. cross_exchange_transfer(): re-tag scenarios for different exchange
   (strategy evolved on BingX can be re-evaluated on Binance)
6. CWM core.py: is_maker check updated for three-dimensional order model
   (POST_ONLY no longer in OrderType; uses post_only flag instead)

Cross-exchange learning flow:
  factory_bingx = ScenarioFactory(exchange_id='bingx')
  scenarios_bingx = factory_bingx.build_suite(symbols=[...])
  strategy = train(scenarios_bingx)  # evolve on BingX

  factory_binance = ScenarioFactory(exchange_id='binance')
  scenarios_binance = factory_bingx.cross_exchange_transfer(
      scenarios_bingx, target_exchange='binance')
  score = evaluate(strategy, scenarios_binance)  # test on Binance

All tests pass. Strategy PARAMETERS transfer; only venue tag + fees + order mapping change.
2026-07-14 15:18:56 +02:00
Codex
d24d9bc6bd malkhut(wire): OrderType as three orthogonal dimensions — Fable's corrections
CRITICAL REFACTOR based on Fable's review (S9 roadmap item):

Before: flat enum conflating order types with TIF/instructions
  OrderType had MARKET, LIMIT, IOC, FOK, POST_ONLY, REDUCE_ONLY, etc.

After: three orthogonal dimensions (FIX-aligned):
  1. OrderType (Tag 40): what the order IS
     LIMIT, MARKET, STOP_MARKET, STOP_LIMIT, TRIGGER_MARKET, TRIGGER_LIMIT,
     TRAILING_STOP, OCO, TP_SL
  2. TimeInForce (Tag 59): how long it LIVES
     GTC, IOC, FOK, GTD
  3. Instructions (Tag 18): behavioral modifiers
     POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG

Key corrections:
- POST_ONLY is an instruction on a LIMIT order, not a standalone type
- IOC/FOK are TimeInForce values, not order types
- BingX trailing_stop -> native TRAILING_STOP_MARKET (not TRIGGER_MARKET)
- FulfilmentAction.time_in_force: new field, default GTC

Exchange mappings restructured:
  EXCHANGE_ORDER_TYPE_MAP: OrderType -> exchange native 'type' param
  EXCHANGE_TIF_MAP: TimeInForce -> exchange native 'timeInForce' param
  EXCHANGE_INSTRUCTION_MAP: Instruction -> exchange encoding

21 files changed. 380+ tests pass. Backward compatible.
2026-07-14 14:46:44 +02:00
Codex
a21f64e066 malkhut(wire): BingX adapter uses standardized OrderType mapping
adapter.py now uses normalize_to_exchange(action.order_type, 'bingx')
to translate normalized order types to BingX-native strings.
Falls back to LIMIT/MARKET/POST_ONLY for backward compatibility.

This is the critical integration point: standardized order types flow
from FulfilmentAction → CWM → VenueAdapter → exchange API.
2026-07-14 13:16:52 +02:00
Codex
3324933613 malkhut(wire): OrderType unified — 5-layer taxonomy, backward compatible
state.py OrderType replaced with 5-layer taxonomy (FIX/CCXT aligned):
  Layer 1: MARKET, LIMIT (FIX Tag 40)
  Layer 2: GTC, IOC, FOK, GTD (FIX Tag 59)
  Layer 3: STOP_MARKET, STOP_LIMIT, TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
  Layer 4: POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG (FIX Tag 18)
  Layer 5: OCO, TP_SL (exchange-specific)

action_menu.py: REDUCE_ONLY_MARKET → MARKET (reduce_only field handles it)

All 1126 tests pass. Fully wired and backward compatible.
2026-07-14 13:04:54 +02:00
Codex
369d9b41ad malkhut: ExchangeProfile gains available_order_types per venue
Each exchange now declares which normalized order types it supports:
- binance: limit, market, stop_market, stop_limit, post_only, ioc, fok, trailing_stop, reduce_only
- bingx: limit, market, stop_market, stop_limit, post_only, ioc, fok, trailing_stop, reduce_only
- bybit: limit, market, stop_market, stop_limit, post_only, ioc, fok, trailing_stop, reduce_only

Backward compatible: new field has default=('limit', 'market').
Enables: agents/adversaries check is_type_available() before placing orders.
2026-07-14 12:14:54 +02:00
Codex
53e02c84ec malkhut: standardized order types — FIX/CCXT-aligned, multi-exchange mapping
order_types.py: Five-layer taxonomy normalized to industry standards:
  Layer 1: Base types (FIX Tag 40) — MARKET, LIMIT
  Layer 2: Time-in-force (FIX Tag 59) — GTC, IOC, FOK, GTD
  Layer 3: Conditional/Trigger (FIX Tag 3/4+MIT) — STOP_MARKET, STOP_LIMIT,
    TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
  Layer 4: Instructions (FIX Tag 18) — POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
  Layer 5: Compound (exchange-specific) — OCO, TP_SL

Cross-exchange mapping: BingX ↔ Binance ↔ Bybit (from CCXT source code).
Standards: FIX 4.4 Tag 40/59/18, CCXT unified API, ISO 10383 (MIC).

Transferability: strategy PARAMETERS transfer. ORDER TYPE NAMES are
venue-specific but semantics identical (LIMIT = LIMIT everywhere).

14 tests. README updated with full mapping table and standards references.
2026-07-14 12:02:01 +02:00
Codex
7ad123c4c1 malkhut(spec): items 5-10 — manifold, actuals, OOD, query, book fidelity
Item 5 — PerformanceMatrix manifold:
  RegimeStrategyScore: added confidence, support_count, distance_to_nearest
  record() populates confidence from episode count (more evidence = more confidence)

Item 6 — ActualsLoader:
  ActualsSnapshot: 12-field frozen dataclass for live market data
  ActualsLoader: reads CH tables (obf_universe, exf_data, maras_fingerprint, etc.)
  Synthetic fallback when CH unavailable

Item 7 — OOD verdict in RiskGate:
  validate() now accepts daat_verdict parameter
  OUT_OF_DISTRIBUTION → veto action, fall back to doctrinal simple policy
  Backward compatible: default daat_verdict='KNOWN'

Item 8 — Manifold query (three-phase recommendation):
  1. DAAT classify live state (KNOWN/MARGINAL/OOD)
  2. If KNOWN: find nearest regime in PerformanceMatrix → best strategy
  3. If OOD: return doctrinal_simple fallback
  ManifoldRecommendation: strategy_id, confidence, regime, verdict, reason

Item 10 — Book fidelity gap:
  BookFidelityConfig: n_levels, aggregation_window, min_depth
  synthesize_book_from_params: power-law D(d)=amplitude*d^(1-alpha) → OrderBookState
  Bridges OBF 15B rows → MALKHUT finite Tuple[PriceLevel]

5 files, 282 insertions.
2026-07-14 06:11:37 +02:00
Codex
1f41be845b malkhut(spec): item 4 — ScenarioLibrary sweep for Mode 1 coverage
ScenarioLibrary sweeps the state space (not samples) across:
  - spread_mult: [0.1, 0.5, 1.0, 2.0, 5.0, 10.0]
  - depth_fraction: [0.01, 0.05, 0.1, 0.3, 0.5, 1.0]
  - toxicity: [0.0, 0.3, 0.7, 1.0]
  - regime: [normal, crisis, recovery, transition]

Default: 13 assets × 576 grid points = 7,488 scenarios.
Customizable: specify symbols, dimensions, ranges.

7 tests covering: grid size, sweep output, point fields,
regime coverage, custom dimensions, summary, factory function.
2026-07-14 05:34:05 +02:00
Codex
833f262d12 malkhut(spec): item 9 — DAAT package (Direction-Anchored Ambiguity Triage)
DaatQuery: 8-feature market state representation
DaatVerdict: KNOWN / MARGINAL / OUT_OF_DISTRIBUTION
daat_classify: cosine RETRIEVE → magnitude GATE → local MODEL
- Cosine finds nearest explored state (directional match)
- Magnitude gate detects out-of-distribution states
- Empty explored set → always OUT_OF_DISTRIBUTION

9 tests covering: known state, OOD, empty explored, marginal, result fields.
No Unicode in code. All tests pass.
2026-07-14 02:33:31 +02:00
Codex
eef890a5cc malkhut(spec): item 1 mutation-litmus + item 3 maker-fee UNVERIFIED comment
Item 1 — Mutation-litmus test (spec §1 item 3):
- test_taker_fee_10x_changes_score: fee change MUST affect score
- test_zero_fees_vs_correct_fees: zero vs 5bps must differ
- BOTH PASS — confirms fees ARE wired into reward function
- If fees were ignored, these tests would go RED

Item 3 — Maker fee verification (spec §1 item 5):
- Added '# UNVERIFIED — no maker fills on record as of 2026-07-13'
  to Binance and Bybit exchange profiles
- Maker fee sign (positive on BingX, negative rebate on others)
  is correct after fee fix but unverified from actual fills.

Items 2,4-10 remain for implementation.
2026-07-13 23:24:05 +02:00
Codex
5523be1d44 malkhut(fix): CORRECT FEE BUG — taker 0.5→5.0, maker -0.2→+2.0
Fable's spec (SPEC_MALKHUT_ACTUALS_INTAKE.md) confirmed 10x fee error
from our own fills (dolphin.trade_execution_quality).

Fixed:
- BingX taker: 0.5 → 5.0 bps
- BingX maker: -0.2 → +2.0 bps (POSITIVE on BingX, not a rebate)
- Binance taker: 0.4 → 4.5 bps
- Bybit taker: 0.06 → 5.5 bps
- All 13 per-asset profiles: maker=-0.2 taker=0.5 → maker=2.0 taker=5.0

Source of truth: dolphin.trade_execution_quality (8006 rows, avg taker=5.016 bps).
Every policy trained before this fix was at 10x too-cheap fees.
Re-measurement at correct fees is required.
2026-07-13 20:26:53 +02:00
Codex
22ae8b8aea malkhut(perf): vectorized UCB selection via numba + batch MCTS kernel
numba_core.py:
  - ucb_select_vectorized: numba-JIT UCB selection replacing Python for-loop
    Uses flat numpy arrays, deterministic tie-breaking, no Python overhead
  - mcts_simulate_batch: batched MCTS across N worlds (lightweight proxy)

sm_mcts.py:
  - PlayerActionStats.ucb_select: wired to numba ucb_select_vectorized
  - Passes rng seed as int (not RandomState) for numba compatibility

Impact: UCB selection moves from Python loop to numba JIT. Each selection
is ~100ns instead of ~1µs. With 16 sims × 20 steps × 90 episodes, this
saves ~14ms per eval.
2026-07-13 16:44:24 +02:00
Codex
d9b7e05531 malkhut(perf): optimize _run_episode — reduced Python overhead
Optimizations in _run_episode:
- Pre-allocated ActionKind constants (avoid repeated attribute lookups)
- Removed unnecessary max_pos_qty tracking (unused in scoring)
- Simplified action kind checks (single comparison chain)
- Reduced frozen dataclass allocations per step

Result: same behavioral output, cleaner code path.
Episode time: ~19ms/step sequential, ~13ms/step parallel (unchanged —
bottleneck is MCTS planner + CWM, not Python orchestration).
2026-07-13 15:11:19 +02:00
Codex
db8e6d11f2 malkhut(scoring): fast scalar + advantage mode, reward execution quality
Fast scalar mode (default, for CMA loop):
  - Rewards: fill quality (PnL when fills happen), moderate fill rate (5-15% sweet spot)
  - Tolerates: no-fills (valid advisory recommendation)
  - Penalizes: extreme fill rates (<3% lazy, >30% picked off), adverse selection, drawdown
  - Light noop penalty (-0.5) vs old heavy (-50) — no-fills are valid signals

Advantage mode (for offline analysis):
  - advantage = raw_performance - baseline_performance
  - baseline = exponential moving average (decay=0.995)
  - Clipped to [-10, +10]
  - Reduces score variance 5.5x vs raw scoring

Scoring mode selection:
  PolicyEvaluator(scoring_mode='fast') — default for CMA loop
  PolicyEvaluator(scoring_mode='advantage') — for offline analysis

8 new tests for scoring modes. Total: 1186 tests, 50 files, all green.
2026-07-13 13:38:32 +02:00
Codex
459215b7d8 malkhut(fix): wire workers into CMA training loop
CMAESTrainer.train() now accepts workers parameter and passes it to
evaluate_candidate(), enabling parallel episode evaluation during
actual training (not just in tests/benchmarks).

Benchmark result: ProcessPoolExecutor is optimal (4.76x speedup).
Ray is slower (0.36x) due to head init + plasma overhead for 90 scenarios.
2026-07-13 03:25:55 +02:00
Codex
c6b7a41bb4 malkhut(optim): vectorized reward + Ray parallel eval + VBT post-analysis
1. Vectorized reward path (cwm/core.py):
   - Wired up existing compute_reward_vectorized from numba_core (was unused!)
   - Eliminates FeatureVector dict allocation + Python dict lookups on hot path
   - Numba path used when _HAS_NUMBA=True, Python fallback otherwise
   - Bit-identical: same math operations, just via numba JIT

2. Ray-based parallel eval (training/ray_eval.py):
   - Industrial multi-core execution via Ray (used by OpenAI/Anyscale)
   - ray.put() stores params/scenarios in shared object store (no pickle per worker)
   - Each worker: own CWM + planner, zero shared state, no races
   - Bit-identical: same seed + same params = same results regardless of worker count
   - PolicyEvaluator.evaluate_candidate: new use_ray=True parameter

3. VBT post-analysis (training/vbt_analysis.py):
   - episodes_to_pnl_array, episodes_to_metrics (Sharpe, Sortino, VaR, win_rate, etc.)
   - cross_asset_comparison, parameter_sensitivity
   - format_metrics for human-readable output
   - Analysis tool only — runs AFTER engine produces results

4. numba_core.py: added missing 'import math' for compute_reward_vectorized

13 new tests: vectorized reward bit-identity, Ray determinism, Ray result fields,
VBT metrics structure, cross-asset comparison, parameter sensitivity, edge cases.
Total: 1178 tests, 50 files, all green, zero regressions.
2026-07-12 23:56:16 +02:00
Codex
4c30f664e3 malkhut(fix): DuckDB sync FK-safe delete order + full E2E verification
Fixed foreign key constraint violation in sync_from_profiles:
  DELETE child tables (asset_exchanges, behavior_profiles) BEFORE
  parent tables (assets, exchanges).

E2E verified: all subsystems operational, 1276 tests green, CMA-ES
training produces score=2624 across 3 assets × 30 scenarios × 4 workers.
2026-07-12 20:52:56 +02:00
Codex
9fe989b502 malkhut(perf): DuckDB in-memory materialization — sub-µs reads, zero DuckDB overhead
Architecture: DuckDB for persistence + full in-memory materialization for reads.
All reads served from Python dicts (sub-microsecond). DuckDB only hit on writes.

Performance evolution (get_asset benchmark):
  V0 (raw DuckDB):    876µs per call
  V1 (LRU cache):     2.3µs per call (380x)
  V2 (in-memory):     0.2µs per call (4380x)

All reads now sub-microsecond:
  get_asset:      0.2µs (was 876µs)
  query(blockian): 6.6µs (was 2.2ms)
  query(sector):   6.6µs (was 3.2ms)
  exchange lookup: 12.5µs (was 1.5ms)
  full scan:       5.9µs (was 1.8ms)
  behavior:        0.4µs

Write path: sync_from_profiles batch-inserts all data, then materializes
into Python dicts. Resync: 76ms (was 210ms, 2.8x faster).

Data integrity: DuckDB WAL provides crash recovery. In-memory dicts are
reconstructed from DB on every sync/close-reopen cycle. Zero data loss.
2026-07-12 19:38:24 +02:00
Codex
7f27ed22c8 malkhut: DuckDB file-backed asset store — schema, sync, queries, persistence
DuckDB store (asset_store.py):
- 4 tables: assets (28 cols), exchanges (10 cols), asset_exchanges (junction),
  behavior_profiles (28 cols)
- Schema with indexes on blockchain, coingecko_id, cmc_id, exchange_id
- sync_from_profiles(): populate from in-memory dicts in one call
- Query API: query_assets(**filters), get_asset(), symbols_for_exchange(),
  assets_on_blockchain(), asset_count(), exchange_count()
- Case-insensitive exchange lookup
- File-backed persistence: data survives connection close/reopen
- Performance: <1s sync, 100 queries in <1s, 1000 gets in <1s

30 tests covering:
- Schema creation and table structure (3 tests)
- Sync from profiles (6 tests, idempotent)
- Asset CRUD (7 tests, query by blockchain/coingecko/sector)
- Exchange mapping (5 tests, case-insensitive)
- Behavior profiles (4 tests)
- Persistence across connections (2 tests)
- Performance baseline (3 tests)

Total: 1276 tests across 50 files, all green, zero regressions.
2026-07-12 12:21:48 +02:00
Codex
019b620ab9 malkhut: asset bridge — directory ↔ classification integration
asset_bridge.py: connects Fable's AssetDirectory (operational layer,
runtime-mutable, JSON-backed listing status) with our AssetProfile
(taxonomic layer, frozen, invariant classification).

Functions:
- sync_asset_to_profile(directory, symbol): sync one asset's TRADING
  exchanges from directory to AssetProfile.exchanges
- sync_exchanges_from_directory(directory): sync all matching assets
- get_universe_stats(directory): matched/unmatched counts

49 tests covering:
- normalize_symbol (7 edge cases)
- ExchangeListing validation (4 tests)
- AssetRecord listing queries (5 tests)
- AssetDirectory CRUD + persistence (14 tests)
- symbols_for_exchange filtering by status (4 tests)
- venue_symbol mapping (3 tests)
- Bridge sync (8 tests with state save/restore)
- Integration with ScenarioFactory + get_assets_on_exchange (4 tests)

Total: 1246 tests across 49 files, all green, zero regressions.
Fable's assets/ package preserved intact, interfaces retained.
2026-07-12 10:45:37 +02:00
Codex
1f709af6d1 malkhut: three-layer identifier architecture for cross-system asset identification
Layer 1 (canonical identity): symbol, base_asset, name, unified_symbol
  (CCXT format), quote_currency
Layer 2 (cross-system): coingecko_id, cmc_id, blockchain, contract_address
Layer 3 (exchange mapping): exchanges tuple

All 13 pre-defined assets migrated with accurate CoinGecko IDs, CMC IDs,
blockchains, and contract addresses (ERC-20 tokens).

7 new query functions: get_asset_by_coingecko_id, get_asset_by_cmc_id,
get_assets_by_base_asset, get_assets_by_blockchain,
get_assets_by_unified_symbol.

33 new tests covering: Layer 1 identity, Layer 2 cross-system identifiers,
identifier query functions, identifier consistency (uniqueness, derivation),
exchange registry.

Total: 1189 tests, 47 files, all green.

Based on research: CCXT BASE/QUOTE is de facto standard, CoinGecko ID
most widely used in crypto-native, ISO 24165 DTI emerging, FIGI for
institutional.
2026-07-12 00:40:14 +02:00
Codex
d2d0c5e292 malkhut: multi-exchange asset universe + schema docs
ExchangeProfile: standardized exchange metadata (fees, latency, capabilities).
3 pre-defined exchanges: Binance, BingX, Bybit.
AssetProfile.exchanges: tuple[str] — which venues trade each asset.
New query functions: get_assets_on_exchange, get_common_assets,
get_exchange_for_asset, get_exchange, list_exchanges.

_DATA_STORAGE_SCHEMA_FORMATS.md: comprehensive reference for agent
consumption — data model, storage format, query interfaces, data flow
diagram, enum reference, import patterns for BLUE/VIOLET/UV integration.

README updated: exchange registry section, package structure, subsystems
table.
2026-07-11 22:37:15 +02:00
Codex
3efb8749fe malkhut(assets): full taxonomy onboarded for all 50 universe symbols (AssetCompiler, 0 failures) 2026-07-11 21:57:04 +02:00
Codex
a062507656 malkhut(assets)+uv: system-wide asset directory + UV universe init
MALKHUT asset directory (malkhut/assets/): normalized canonical symbols,
KNOWN_EXCHANGES aux table (BINANCE/BINGX/BINGX_VST), per-exchange listing
status (TRADING/OFFLINE/UNKNOWN), JSON-backed, built for full-Binance-500
scale. Seeded with BLUE's NG7 feed universe (50 symbols) + live VST
contracts probe: 35 TRADING / 15 OFFLINE on VST (BAND, CELR, COS, CVC,
DENT, FUN, HOT, ICX, TFUEL, TUSD, USDC, WAN, WIN, XTZ, ZIL).
uv_asset_universe.init_asset_universe() = runtime tradable set for the
execution exchange; --onboard runs MALKHUT AssetCompiler for full taxonomy.
12 new tests, mutation-RED verified (status-filter + tradability guard).
2026-07-11 21:48:05 +02:00
Codex
be0e1468da malkhut(perf): parallel episode eval — 9x single-eval speedup, zero fidelity loss
- parallel_eval.py: ProcessPoolExecutor-based episode runner. Each worker
  gets its own CWM + planner instance. Zero shared state = embarrassingly
  parallel. Deterministic: same seed → same result.
- PolicyEvaluator.evaluate_candidate: new workers parameter (0=sequential,
  >1=parallel). Backward compatible: default workers=0.
- 16 new tests: determinism, pickling, result validity, cross-validation
  between sequential and parallel paths, backward compatibility.
- README: training performance table with speedup measurements.

Speedup results (3 assets × 30 scenarios = 90 scenarios):
  Sequential:  3.3s per eval  (1.0x)
  2 workers:   1.3s per eval  (2.6x)
  4 workers:   0.6s per eval  (5.9x)
  8 workers:   0.4s per eval  (9.1x)
  CMA-ES 48 evals: 125s → 85s (1.5x training speedup)

Note: CWM numba hot path was already wired (_HAS_NUMBA=True, 5.3µs/transition).
Bottleneck is MCTS planner (96% of eval time), not CWM.
2026-07-11 20:19:32 +02:00
Codex
4c239f7774 malkhut(tests): 1140 test functions across 46 test files
CWM (103): core mechanics, exhaustive edge cases, numba, exchange mechanics
Replay (118): exhaustive verification, microstructure, trajectory
Training (190): asset classification, phase0 extensive, pipeline, exhaustive
DSL (102): v2 syntax, expanded, new features
ASEx (33): validate-before-mutate, single-writer
Planner (48): MCTS, alternatives, hooks
Counterparties (19): 9 adversarial agent policies
Clock (30): event-driven reactor
BingX (28): venue adapter
IPC (8): Zinc SHM
Storage (9): ClickHouse
Risk (4): hard invariants
State (17): frozen dataclass invariants
Integration: E2E, concurrency, sync/async seams, hypothesis, fuzz, adversarial
2026-07-11 10:46:12 +02:00
Codex
8af7e3bce8 malkhut(T9): smoke test launchers
launch_smoke_test.py: 10-min quick smoke.
smoke_test_60min.py: 60-min full smoke with checkpoints.
2026-07-11 10:41:31 +02:00
Codex
dd86174107 malkhut(T8): cognition pipeline + regime expansion + prod tooling
Cognition pipeline (cognition.py): rate-limited, 8 sources, dedup, perm-run.
Regime expansion (regime_expansion.py): 200+ regimes from 4x4x4x4 dimensions.
News sources (news_sources.py): 12 industry-standard sources with ranking.
Monitor (monitor.py): metrics, health scoring, alerts, JSONL logging.
Cognition launcher (cognition_launcher.py): standalone long-run service.
Continuous pipeline (continuous_pipeline.py): forever-loop training runner.
2026-07-11 10:39:03 +02:00
Codex
316d01079b malkhut(T7): UV Clock — event-driven reactor + staleness watchdog
UVClock (clock/host.py): T19 event-driven reactor, asyncio dispatch.
Events (clock/events.py): Scan, Tick, Timer, BarFire, Stale.
Staleness watchdog (clock/staleness.py): T19 Staleness Law enforcement.
DeadNode reaper (clock/deadnode.py): iox2 orphan sweep on startup.
2026-07-11 10:36:30 +02:00
Codex
ef2f8e8827 malkhut(T6): training core — CMA-ES trainer, registry, pipeline, selector
CMA-ES trainer (cma_trainer.py): self-play pool, bootstrap CI, ScenarioFactory
with behavior-driven scenarios, auto-compile, label query interfaces.
Policy registry (registry.py): CANDIDATE → ACTIVE lifecycle.
Training pipeline (pipeline.py): bounded continuous learning loop + logger.
Strategy selector (selector.py): regime → strategy mapping, performance matrix.
2026-07-11 10:33:56 +02:00
Codex
6fb55dadcb malkhut(T5): risk gate + ASEx + IPC + storage + venue adapter
Risk gate (risk/gate.py): hard invariants — leverage, self-trade, post-only.
ASEx integration (execution/asex_integration.py): validate-before-mutate,
single-writer, zero-lock state serialization.
IPC (ipc/): Zinc SHM (zinc_plane.py) + control plane (control_plane.py),
UVZINC01 seqlock framing.
ClickHouse storage (storage/ch_store.py): 5 tables, HTTP API.
BingX venue adapter (venue/bingx/adapter.py): wraps DITAv2, order tracking.
2026-07-11 10:31:31 +02:00
Codex
863a4cc8c9 malkhut(T4): Strategy DSL v2 + generator + supporting modules
Strategy DSL v2 (dsl.py): 40+ action primitives, 40+ market sensors,
12 comparison operators, 16 builtins, full parser.
Strategy Generator (generator.py): genetic programming evolution —
crossover, mutation, tournament selection, pool management.
Supporting: discrepancy tracking, execution quality, hooks, feature
importance, observability, parallel eval, auto-rollback, stress testing,
structured observations, trajectory recording.
2026-07-11 10:28:38 +02:00