Commit Graph

334 Commits

Author SHA1 Message Date
Codex
aed9d52ef6 fix: crypto sentiment calibration - 15/15 critical tests pass
- Enhanced bullish/bearish keyword lists in CryptoSentimentCalibrator
- Added context-aware whale action phrases (buys/accumulates=bullish, sells/dumps=bearish)
- Lowered FinBERT threshold from 0.15 to 0.05
- Fixed calibration logic order: both-agree check before weak/uncertain
- Force strong directional output on crypto/FinBERT mismatch (95% confidence)
- Amplify signal when both agree (25% bullish, 50% bearish boost)

Result: 15/15 critical sentiment tests pass (was 7/15)
2026-09-15 18:44:35 +02:00
Codex
616a23c0b3 feat(pi_wake_agent): --at HHMM one-off self-cleaning wake mode 2026-09-15 11:17:36 +02:00
Codex
ad4fcc538a pi_wake_agent: context-occupancy steer warning + hermetic integration tests 2026-09-14 23:05:47 +02:00
Codex
5e9168ac1d pi_wake_agent: --pane targeting + crontab backend fallback + cronicle HCL fix + hermetic tests
- --pane/--pane-id: write-chars AND write(13) hit terminal_1 (bottom pane)
- cronicle_available()=shutil.which + daemon-live check; crontab/crnd auto-fallback
  (cronicle daemon down here -> live scheduler is system crnd)
- fix cronicle_install invalid HCL (command=[...] array, was bare comma-list)
- rewrite TestCrontabEntry hermetic (subprocess mocked; real crontab never touched)
- add --pane-id / no-pane omission / HCL-array mutation-litmus tests; .sh --pane
- cronicle.hcl: corrected pi_wake_pi_test_20m (msg + --pane terminal_1), valid HCL
- live 20m crontab doorbell installed for pi_test/terminal_1, bell verified delivered
- SIGQUIT rescued a locked pi (final_test @1318s); steer queued + bus-posted
2026-09-14 21:31:37 +02:00
Codex
a276aeaded Add sentiment_engine with CryptoSentimentCalibrator fixes - improved keyword lists, lowered FinBERT threshold, added neutral handling 2026-09-14 13:30:05 +02:00
Codex
19a7812094 ops/ack_probe.py: standalone HL-testnet cancel-ACK reliability harness
Dry-run-verified WS orderUpdates pipe + REST meta + SDK sigs + signing.
Fixes: cancel-by-oid (cancel(name,oid)); order_type={'limit':{'tif':'Gtc'}}
not Grouping enum; typed Cloid.from_int; coin strip USDT->DOGE; recursive
oid|cloid WS matcher. --keystore uses host-bound systemd-creds (spare
hl_agent_testnet). Live fire gated: testnet faucet needs mainnet-deposited
address; only r9 hl_testnet (off-limits) is funded/registered.
2026-08-28 20:23:38 +02:00
Codex
cd57abb089 docs: V7_RISK_DOMINANT preemption reference (r7 live tape; corrected from prior mis-attribution) 2026-08-28 14:49:04 +02:00
Codex
578bd3ed70 ops: f13 compensatory ATOM close (testnet-gated, dry-run + wire-order gate) 2026-08-27 14:26:36 +02:00
codex
36009fcd91 ops: v20r1 monitor [3] widen heartbeat window past stale-metadata storm 2026-08-26 22:06:57 +02:00
codex
80bb60f24b ops: harden v20r1 monitor (correct hyperliquid-testnet.xyz host, DNS-aware venue-truth probe, heartbeat-tail dedup) 2026-08-26 21:53:10 +02:00
codex
23cd98105c docs: add v20r1 FLIGHT13 System Test & Ops Guide
Covers v20r1 identity (pid 790059->790060, zinc uv_flight13_v20r1, lease /tmp/flight13_hl_testnet_single_writer.lock), config chain, FORCE_ENGAGE firehose (u=exec_urgency, capacity=1, 30/hr SHORT 14 syms), exit/max-hold, venue-truth via unsigned HL /info, monitoring runbook, relaunch/stop procedures, and troubleshooting (incl. sub-$10 residual dust and stale HL metadata). Testnet-only; no system code/data modified.
2026-08-26 21:30:11 +02:00
codex
8255c807e9 ops: add v20r1 live-monitor helper (prod/ops/v20r1_monitor.sh) 2026-08-26 21:11:35 +02:00
Codex
f0d80b05d3 docs: PI tasking — exhaustive test suite (unit/pairwise/E2E/chaos; 16-bug ledger)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 13:35:20 +02:00
Codex
d4b32a6549 docs(dumb-review): correct E1 — single-position engine self-clears; real phantom bug found+fixed
E1 as submitted (release phantom on promotion-SUPPRESSED entries) does not survive
review: the orchestrator is single-position and self-clears exit_manager+self.position
on its own exit each cycle (esf_alpha_orchestrator 491-492), so shadow mode is a faithful
paper mirror and pi's fix would regress shadow parity. The investigation instead found the
real bug — phantom_release left engine.position set → single-position engine bricked on a
proven reject — fixed in f11.1/blue-sizing-parity 98941373. Preserves pi's original E1 text.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 19:15:17 +02:00
Codex
7bd5a8d628 feat(supervisor): add flight10 (F10) program -- durable supervised launch
F10 (BLUE's kernel in the UV/VIOLET body, live-mainnet-capable, promotion-gated by
the arm file) now runs under supervisord as program flight10 instead of an ad-hoc
agent background task (which got reaped ~18min -- the only reason it kept dying).

Command SOURCES /root/flight10_prod.env (chmod 600, secrets NOT inlined here) so every
autorestart revive gets the full env (mainnet keys, DITA_V2_ZINC=REAL shared-RAM exec
plane, UV_EXEC_DARK_MODE=0, TP/SL, vol threshold, CH/HZ, arm-file path). Sets HOME + an
explicit PATH incl /root/.cargo/bin (the DITAv2 rust-backend provenance check shells
`rustc -vV`; supervisord's minimal PATH lacked it -> crash-loop). Clears stale
zinc_uv_exec_* regions pre-launch (RealZincPlane create=True FileExistsErrors on a
prior runner's un-reaped regions -> silent in-memory downgrade; F10 is sole owner so
clearing is safe -> REAL shared RAM every start/autorestart). Polite: startsecs=100 +
startretries=3 -> a startup-crash goes FATAL, never rate-loops BingX. autostart=false,
autorestart=true, log rotation, rlimit_as=3GB. BLUE (dolphin:nautilus_trader) untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 18:26:25 +02:00
Codex
6990ff3bee malkhut: asset-faithful book generation with composable toggles
Three independently toggleable features:
  1. Asset-faithful depth/spread: levels sized by OB study power-law per asset
  2. Intraday volume clock: depth scales by time-of-day (peak/trough)
  3. Realistic spread: per-asset spread from OB study + Flight7

Composable via BookGenerationConfig toggles:
  use_asset_faithful_depth, use_asset_faithful_spread, use_intraday_clock,
  use_weekend_mode, use_stress_mode, use_fragility, worst_case_mode

worst_case_mode overrides everything for max adversarial learning:
  spread * stress_mult, depth * fragility, no intraday/weekend.

DuckDB registry for online updates:
  AssetRegistry: upsert/get/list/delete/query
  RuntimeProfileCache: hot-reload during CWM runs
  upsert_from_csv/export_csv: pipeline support
  upsert_all_from_asset_behaviors(): seed from OB study

Results (8 assets):
  BTC: spread 0.031 bps, depth $350M (normal) / $4.9M (worst)
  DOGE: spread 2.86 bps, depth $2M (normal) / $132K (worst)
  ADA: spread 11.8 bps, depth $10M (normal) / $511K (worst)
  Intraday: BTC peak/trough = 2.8x depth ratio

All 99 tests green (31 new + 68 existing).
2026-07-20 19:06:24 +02:00
Codex
97a770da65 malkhut: online EWMA self-calibrating slippage model
Flight7 model underestimates by 80% in CWM dynamic book:
  raw predicted: 0.034 bps, actual: 0.180 bps
  Constant error across 22K episodes — no feedback loop.

Root cause: Flight7 calibrated on real BingX taker fills, but CWM's
synthetic dynamic book has different fill characteristics.

Fix: SlippageSelfCalibrator with EWMA feedback loop.
  After each fill: error = actual - predicted (clipped to +/-20 bps)
  EWMA smooths per-symbol errors (alpha=0.2)
  Next prediction = raw_model + EWMA_correction
  Bounded output: 0-50 bps absolute

Convergence (300 eps across 8 assets):
  ETH: 9% error (from 80%)
  SOL: 3.5%
  DOGE: 6.7%
  LINK: 5.7%
  ADA: 9.7%
  BTC: 48.6% (low fill count, converging)
  AVAX: 28.6% (low fill count)
  UNI: 52.5% (low fill count, early outlier)

Truthfulness guarantees:
  - Correction is observable (CALIBRATOR.correction(symbol))
  - Resets between runs (no hidden state)
  - Only uses observed fills, no assumptions
  - Error clipping prevents outlier domination
  - Absolute bounds prevent runaway
2026-07-20 15:17:16 +02:00
Codex
c1a888faf3 malkhut: MCTS planner E2E — planner IS learning (slippage 1.1→0.008 bps)
MCTS planner with dynamic book:
  Episode 3: slippage=1.125 bps (first aggressive fills)
  Episode 20: slippage=0.177 bps (84% reduction)
  Episode 30: slippage=0.008 bps (99% reduction!)

The planner learns to:
  1. Place passive orders at better offsets
  2. Wait for book to move before crossing
  3. Use urgency-driven maker/taker decision
  4. Reduce slippage through queue position optimization

PnL stays positive throughout (+1687 to +5048 bps).
Fill value improving from -0.319 to -0.000 (less negative = better).
2026-07-19 19:06:10 +02:00
Codex
db844c775a malkhut: configurable friction per scenario + 5h learning test 2026-07-19 05:47:34 +02:00
Codex
3b1dae6cbe malkhut: configurable friction per scenario + 5h learning test
1. Friction configurable per scenario:
   Scenario gains maker_fee_bps, taker_fee_bps, adverse_cost_bps
   ScenarioFactory accepts friction overrides
   All 34 scenario builder calls updated
   System learns in ALL conditions (free maker, BingX real, Binance-like)

2. 5h learning test (long_learning_5h.py):
   - 1500 opponents, 9 assets, 270 scenarios
   - OOM/CPU monitoring (ResourceMonitor)
   - CMA-ES every 15 reports
   - Rolling fill_value/surprise/fill_rate tracking
   - Needs screen/tmux on host for 5h execution

Note: background redirect issue on this shell session.
Run via: cd MALKHUT && PYTHONUNBUFFERED=1 NUMBA_CACHE_DIR=/tmp/numba_cache python -m malkhut.long_learning_5h
2026-07-19 04:18:01 +02:00
Codex
70f33f6911 malkhut: friction settings configurable per scenario
Scenario gains 3 new fields:
  maker_fee_bps: Optional[float] = None (override per-scenario)
  taker_fee_bps: Optional[float] = None (override per-scenario)
  adverse_cost_bps: Optional[float] = None (per-fill adverse selection)

ScenarioFactory gains friction constructor params:
  ScenarioFactory(exchange_id='bingx', maker_fee_bps=2.0, taker_fee_bps=5.0)

_make_state() accepts friction overrides → applies to VenueRules
_behavior_state() passes friction overrides through
All 34 scenario builder calls updated with friction overrides.

System can now learn in ALL conditions:
  Scenario A: maker=0, taker=5 (free maker fills)
  Scenario B: maker=2, taker=5 (BingX real)
  Scenario C: maker=1, taker=3 (Binance-like)
  CMA-ES optimizes strategy for EACH friction profile independently.
2026-07-19 00:25:26 +02:00
Codex
c868dbfb66 malkhut(docs): updated all docs — fee model, markout, urgency threshold 2026-07-18 19:14:34 +02:00
Codex
ebf7f17132 malkhut: fee+slippage execution threshold 2026-07-18 17:44:20 +02:00
Codex
5503aafa28 malkhut: P0 guard + P1 BingX protective strings + tests fixed
P0 (safety): adapter.py rejects unmapped types (OCO, TP_SL) via
  is_type_available() guard. Returns None instead of silent LIMIT fallback.

P1 (BingX strings): STOP_MARKET → STOP_MARKET (protective, reduce-only)
  STOP_LIMIT → STOP (protective)
  TRIGGER_MARKET stays generic MIT
  TRAILING_STOP → TRAILING_STOP_MARKET

Tests updated to match corrected mappings.
2026-07-18 15:08:57 +02:00
Codex
041c879e82 malkhut: all exchange order types in DSL + action_menu
ActionType → OrderType mapping (common-sensical):
  QUOTE/REQUOTE/CHASE: LIMIT (passive, post_only)
  CROSS_SPREAD (high urgency): MARKET (immediate fill)
  CROSS_SPREAD (medium urgency): LIMIT + IOC (partial fill)
  STOP_LOSS: STOP_MARKET (trigger → market exit)
  TAKE_PROFIT: TRIGGER_MARKET (trigger → market exit)
  TRAILING_STOP: TRAILING_STOP (trailing stop exit)
  EXIT/FLAT_ALL: STOP_MARKET
  EMERGENCY_EXIT: MARKET (immediate)

All order types exercised: LIMIT, MARKET, STOP_MARKET, TRIGGER_MARKET, TRAILING_STOP
2026-07-18 12:25:41 +02:00
Codex
8322550fbc malkhut: REQUOTE as proper CANCEL_REPLACE primitive + action_menu metadata
REQUOTE is now distinct from QUOTE:
  QUOTE: PLACE new order (no existing to cancel)
  REQUOTE: CANCEL_REPLACE existing + place new (immediate)
  CHASE: PLACE with short TTL (auto-cancel retry)
  CANCEL: Remove existing order

action_menu generates REQUOTE with metadata={'requote': True} for existing orders.
DSL REQUOTE produces CANCEL_REPLACE when existing order, falls back to PLACE.

All 800+ tests pass.
2026-07-17 23:17:22 +02:00
Codex
4926ef6788 malkhut: CHASE mechanics FIXED + Flight9 learnings + TTL enforcement
1. CHASE mechanics (NOW WORKING):
   - CWM enforces TTL on open orders (auto-cancel when expired)
   - DSL CHASE produces PLACE with metadata={chase: True}
   - Action menu generates chase actions with wait_to_retry_ms TTL
   - OpenOrderState.gains ttl_ms field (0=no expiry, >0=auto-cancel)

2. TTL enforcement (CWM):
   - HftBacktestCWM: auto-cancels orders where age >= ttl_ms
   - MinimalCryptoLOBCWM: same TTL enforcement
   - This is how CHASE works: place→wait→auto-cancel→next step re-places

3. Flight9 learnings:
   - Slippage model gains trade_flow_intensity parameter
   - Book imbalance as proxy for trade arrival rate
   - Markout = quality concept documented

4. CHASE tests: 10 new tests covering TTL enforcement, cancel-retry cycle,
   max retries, DSL CHASE action, CMA codec integration

5. All 800+ tests pass
2026-07-17 19:30:17 +02:00
Codex
bb229833d3 malkhut: Flight9 learnings — markout=quality, queue×flow, depth-for-size
Fable's Flight9/BLUE generalizable features incorporated:

1. Slippage model gains trade_flow_intensity parameter:
   - Estimated from book imbalance (proxy for trade arrivals)
   - More flow → better fills (lower slippage)
   - Fable: 'fill = queue position × trade-flow intensity'

2. Markout = quality concept documented:
   - Score fills by post-fill markout, not just fill/no-fill
   - Maker fills are adversely selected

3. Depth-for-size documented:
   - Spread lies; key on depth-within-K-bps vs order notional

4. Measured fees:
   - BingX maker=2.00bp, taker=5.016bp (over 1,455 fills)
   - BingX commission = NEGATIVE (debit)

5. OB study updated with Flight9 learnings
2026-07-17 16:18:43 +02:00
Codex
8857daedfa malkhut: urgency-driven maker/taker + calibrated slippage + chase + docs 2026-07-17 10:16:55 +02:00
Codex
5c4ccdb1de malkhut: 3.5H instrumented E2E + calibrated slippage + conditional slippage 2026-07-15 19:32:46 +02:00
Codex
c03d914e7a malkhut: conditional slippage (Fable) 2026-07-15 17:05:16 +02:00
Codex
619966605e malkhut(docs): fill quality documentation — README, integration, OB study
README: Fill Quality section (core optimization target, metrics, reward
function, PerformanceMatrix, CMA-ES integration)

HftBacktestCWM integration doc: FillQuality dataclass, fill_value_score
computation, reward function weighting, PerformanceMatrix tracking

OB microstructure study: Section 13 — Fill Quality Optimization,
per-asset expectations, optimization strategy, connection to OB dynamics

Fill quality is MALKHUT's core aim: the system learns to get better fills
(faster, better-priced, less adverse selection) across regimes and venues.
2026-07-15 15:36:22 +02:00
Codex
618ad723e3 malkhut(wire): fill quality as PRIMARY optimization target
Fill quality is MALKHUT's core aim. Wired end-to-end:

1. FillQuality state (state.py):
   - slippage_bps, price_improvement_bps, levels_consumed
   - is_maker_fill, rolling_fill_rate, post_fill_adverse_bps
   - fill_value_score: composite metric for optimization
   - Added to MarketWorldState.fill_quality field

2. HftBacktestCWM.transition() (hft_cwm.py):
   - _compute_fill_quality() computes all metrics per transition
   - Fill quality now tracked for every CWM step
   - Empty book guards added for safety

3. MinimalCryptoLOBCWM.transition() (core.py):
   - Same fill quality computation for deterministic fallback
   - Empty book guards added

4. Reward function (hft_cwm.py):
   - fill_quality_reward = w_fill_probability * fill_value_score (PRIMARY)
   - Bonus for maker fills that improve price
   - Penalty for adverse selection after fill
   - Base reward (PnL, adverse selection, fees) preserved

5. PerformanceMatrix (selector.py):
   - RegimeStrategyScore: 4 new fill quality fields
   - record(): accepts fill_rate, slippage, price_improvement, fill_value_score
   - EMA updates for all fill quality metrics

6. EpisodeResult (cma_trainer.py):
   - avg_fill_value_score, avg_price_improvement_bps, avg_post_fill_adverse_bps
   - Accumulated per-step during _run_episode
   - Recorded to PerformanceMatrix in evaluate_candidate

All 1379+ tests green.
2026-07-15 15:22:25 +02:00
Codex
fa76070c79 docs(exec): omp task brief — 10K test-types for Flight7UE exec libs (mock exchange ONLY)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 10:00:45 +02:00
Codex
82eb5f551a exec(uv): wire T1 smart-exec into PRIME runner (gated UV_SMART_EXEC, lazy)
Bolts SmartExecBridge into the flight's scan loop behind UV_SMART_EXEC (default off). Fully
lazy: flag off => smart_bridge is never imported and the naive maybe_promote runs unchanged
(verified: runner import does not load smart_bridge when the flag is unset). Flag on =>
entries rest as PostOnly GTX makers, exits stay MARKET, the 6s tick drives TTL-abandon.
Reuses bridge.gate (two-man rule) + bridge.stats. runner.py compiles; wire import-safe both
paths. This is the 'bolt to flight 5/6' — dormant until an operator flips the flag and restarts.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 09:37:15 +02:00
Codex
f2b3254b52 exec(uv): T1 smart-exec bridge — maker-first entries, gated + dormant
SmartExecBridge: a drop-in for PromotionBridge.try_promote that rests entries as PostOnly GTX
makers (via exec_unified.router policy) instead of always paying the taker cross. DORMANT —
selected only by UV_SMART_EXEC=1 (default off); nothing imports it yet, so the running flight
is untouched. T1 = smallest bug surface (spec §15):
- ENTER -> maker (ACQUIRE): PostOnly LIMIT @ touch; unfilled by next scan tick -> CANCEL/abandon
  ('a missed entry is free' §4-1). No chase, no cross, no requote race.
- EXIT  -> MARKET unchanged (never strand; zero new exit risk in v1).
- the 6s scan tick IS the drive clock: sweep-stale-then-place each promote.
Reuses router.decide + friction side-lane (never blocks a promote); fail-soft (never raises
into the scan loop). 11 tests, 2 mutation litmus RED (entry-maker policy; stale-sweep).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 09:31:58 +02:00
Codex
5abe0cc6e0 exec(unified): sync real KernelIntent translator §7 — verified side/size/cancel
intent_translator.py: the injected to_kernel_intent KernelExecPort needs, wiring the pure
engine to the live DITAv2 kernel. Side/size/cancel mappings each verified against vendored
source (bingx_venue:627, rust_backend:448/897); EXIT inversion mutation-verified RED. Tested
against REAL dita_v2 (importorskip). Lazy dita_v2 import keeps the rest of the package clean.
143 tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 08:17:39 +02:00
Codex
ec95bd6663 exec(unified): sync local build — kernel_port + dialect §11 + friction §12 + size threading
Lands the /root/dev local increments (share was ENOSPC; now writable) into the canonical repo:
- kernel_port.py: real ExecPort over the DITAv2 kernel (duck-typed; injected to_kernel_intent).
- dialect.py §11: BingX boundary — clientOrderId(H4)/dash/quantize/payload (PostOnly=timeInForce).
- friction.py §12: effective-bps + naive-baseline savings; side-lane journal that can't raise
  (b46ebd2); BingX commission-sign flip. DDL ships with code (register in applier verify-set).
- drive_loop/executor/working: Decimal size threaded through plan/working types (exit size cap).
All pure stdlib+Decimal, mutation-litmus RED on the two load-bearing asserts. 132 tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 08:05:26 +02:00
Codex
21419b03d0 exec(unified): UnifiedExecutor — initial-submit pipeline + S2 pain-fence
executor.py composes the built parts into one execute(request, snapshot): S-grade fence →
router.decide → placer.pre_submit → submit → register working. The initial-submit half that
complements DriveLoop's expiry half (analogue of pink_direct.py:1645-1700).

Composition rulings encoded + tested:
- S2/S3 pain-fence (§15.2): refused WHOLE, loudly, below T14 — never silently sliced.
- placer declines (spread gate) → cross ONLY if urgency crosses on expiry; ACQUIRE ABANDONS
  (a missed entry is free, §4-1). Both mutation-verified RED.
- TTL from urgency discipline: PROTECT 2s / ROTATE deadline_ms / ACQUIRE quote-lifetime.

contract.py: SGrade enum (S0-S3) + s_grade/parent_request_id fields. _constants: MAKER_QUOTE_TTL_S.

FIX (real bug, not just test): drive_loop._is_resolved EXIT was trade_id-based (a PINK
artifact — PINK reused the position's trade_id for exits). The agnostic layer never gets
the position id, so exit-done is now SIZE-based. Kept ENTER on clientOrderId match.

Full exec_unified suite: 94 green, mutation-litmus verified.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 02:30:30 +02:00
Codex
a310c92977 malkhut(e2e): CMA-ES disabled for pure 100-opponent episode loop 2026-07-15 00:48:48 +02:00
Codex
bbaceffb61 malkhut(e2e): CMA-ES every 20 cycles, crash-safe gc.collect() 2026-07-15 00:34:50 +02:00
Codex
67b3268b98 malkhut(e2e): memory-efficient + crash-safe 3h run
Rolling stats (no episode accumulation), CMA-ES every 10 cycles (3 evals),
gc.collect() after CMA, global try/except for crash safety.
100-opponent swarm, 9 assets, 270 scenarios.
2026-07-15 00:27:37 +02:00
Codex
0e215b1658 malkhut(e2e): memory-efficient long run — rolling stats, no episode accumulation
Fixed OOM kill by replacing all_episodes list accumulation with:
- Rolling stats (clear every 20 episodes)
- Only PnL history kept for characterization
- Peak/worst tracking without full episode storage
- Periodic stdout reports from rolling aggregates

100-opponent swarm + 9 assets × 30 scenarios = 270 scenarios per cycle.
CMA-ES every 5 cycles (3 evals). 3-hour target.
2026-07-15 00:19:59 +02:00
Codex
d72323a6c5 malkhut(e2e): 100-opponent swarm + empty book fix + error handling
- 100 diverse opponents (randomized params within each type)
- Risk gate: empty book guard in _post_only_would_cross
- CMA-ES: only every 5 cycles, 3 evals, robust error handling
- Main loop: try/except prevents silent crashes
- Profiling: 11.7 steps/sec with 100 opponents
2026-07-15 00:11:57 +02:00
Codex
956c650171 exec_unified/placer.py: conservative rounding + test fixes
- _quantize_to_tick_conservative: BUY→ROUND_FLOOR (never up into ask), SELL→ROUND_CEILING (never down into bid)
- Removed unused _quantize_to_step (size/step quantization at venue-dialect submit, not placer)
- Fixed cross-quantize tests: with conservative rounding, BUY floors down, SELL ceilings up → never crosses
- Added TestQuantizeToTickConservative with 8 tests for side-aware rounding
- Mutation-litmus: spread gate, TAKER gate, quantize-cross all RED on inversion
- 33 placer tests + 30 router + 18 drive_loop = 81 total green
- Drive_loop.py (commit 670b739a) now consumes pre_submit via wants_placement
2026-07-15 00:00:56 +02:00
Codex
670b739a83 exec(unified): drive loop §7 — PINK _handle_expired_working ported + audit build items
drive_loop.py transcribes pink_direct.py:1089 (the 10-step expiry sequence) warts-and-all
against injected ports (ExecPort seam: clock+kernel+venue), so no ambient state (C1). Carries
the scar tissue: pump-before-cancel, re-classify-after-cancel (fill races cancel), EXIT never
strands→MARKET, slot-busy double-entry guard, fail-safe venue-truth requote gate. SEAM
learnings preserved as referenced comments (zero-silent-suppression rule). Decimal sizes (H1).

- working.py: WorkingRegistry + WorkingOrder, injected clock, rejected→expire_now (one shared
  path). contract.py: client_order_id_core(attempt) — unique per attempt (audit H4/FIX).
- _constants.py: REQUOTE_HOT_WINDOW_S=5.0 (prov: 2026-06-10 double-entry).
- inventory §6: audit build items BI-1..5 folded in.
- test_drive_loop.py: 15 scar-tissue tests. Mutation-verified RED under exit-no-escalate
  (kills 3) and no-hot-window (kills double-entry guard). Full exec_unified suite: 77 green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 23:51:04 +02:00
Codex
a7cf3f5976 docs(exec): adversarial best-practice audit — web-grounded + best-of-breed internals
Weird-pass review of pink_direct.py + exec_router.py + unified spec vs industry practice.
Ranked findings, each grounded (cited sources + nautilus/FIX/OMS internals):
- H1 float money math (PINK) — against consensus; ALREADY fixed by rebuild's Decimal contract
- H2 crash-recovery state in /tmp — durable+atomic path needed
- H3 shared VST account + symbol-membership ownership — sub-account isolation is the fix
- H4 clientOrderId must be unique PER ATTEMPT — gap in my contract.py (raw request_id seed)
- H5 WS-primary without explicit seqnum gap detection — add forced REST resync
- M1 bare except-pass swallows; M2 dead-man-stop orphan hard-gate; M3 L673 sign bug
Plus a balance section: where PINK/spec are AT/ABOVE best practice (don't 'fix' these).
Net: spec sound (deviations are additions); real deviations are PINK-side; float already fixed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 23:38:16 +02:00
Codex
4915438c5a docs(exec): PINK drive-loop certified-learnings inventory — the §7 port checklist
Read pink_direct.py (1837L) in full. Catalogues all ~16 execution scar-tissue mechanisms
with line refs + the incident each encodes, classified PORT/SEAM/AMEND/SUPERSEDED. Captures
the _handle_expired_working control-flow sequence verbatim (the heart of the port).

Two governing rulings folded in:
- ZERO SILENT SUPPRESSION (operator: 'we paid dearly for pink'): any mechanism not ported
  as live code is carried into exec_unified as a referenced comment (reasoning +
  pink_direct.py:LNNN) at its seam. Silent omission is the one unforgivable port error.
- Both axes: T-tier (smartness) AND S-tier (size, §15.2). PINK loop = single-clip S0/S1;
  slicers compose above the contract (T*.s), in-order size mechanics are T14+, S2+ pain-
  fence is NEW doctrine to add.

Also flags a live SIGN-BUG candidate (pink_direct.py:673 'negative = rebate' contradicts
today's Q1: BingX commission negative = DEBIT). Bounded by K≈E gate; NOT touching BLUE/PINK.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 23:24:57 +02:00
Codex
d2e6d78abd exec_unified: add placer.py (SmartPlacer pre_submit seam) + mutation-litmus tests
- MarketSnapshot: best_bid, best_ask, spread_bps, tick, step (Decimal, frozen, validated)
- PlacementPlan: limit_price (quantized to tick), post_only=True
- pre_submit: returns None for TAKER, spread gate failure, or quantize-cross; otherwise PlacementPlan
- Spread gate uses live spread_bps (fixes dead _spread_allows_maker)
- Quantize to tick BEFORE returning (ROUND_HALF_EVEN per §4-15)
- Cross-after-quantize guard rejects if tick quantization crosses book
- 33 mutation-litmus tests: spread-gate flip, TAKER gate, cross-after-quantize all go RED
- Router.PROTECT wants_placement=True (at touch)
- Pure stdlib+Decimal, zero I/O/venue
2026-07-14 23:14:07 +02:00
Codex
adba9f3fe8 malkhut(e2e): full system exercise — HftBacktestCWM + 11-agent swarm + characterization
E2E exercise results:
  30 episodes × 30 steps = 900 actions
  CWM: HftBacktestCWM (PowerProbQueueModel)
  Swarm: 11 diverse opponents (2x ToxicTaker, 2x PassiveMaker, LatencyArb,
         Noise, Momentum, MeanReversion, InventoryMM, LiquidationFlow,
         StaleQuoteAttacker)

Performance:
  Avg PnL: +7217.9 bps, Win rate: 66.7% (20/30)
  Best: +26171.0 bps (chop scenario)
  Worst: -5.0 bps (thin/arb scenarios — flat, not loss)

Order types exercised:
  LIMIT 16.9%, MARKET 12.9%, STOP_MARKET 0.2%
  IOC 8%, GTC 92%
  Post-only: path exercised, Reduce-only: 41.5%
  Aggression/Passive ratio: 0.76

Speed: 457 actions/second (2.0s total)
2026-07-14 23:07:38 +02:00