Codex
5503aafa28
malkhut: P0 guard + P1 BingX protective strings + tests fixed
...
P0 (safety): adapter.py rejects unmapped types (OCO, TP_SL) via
is_type_available() guard. Returns None instead of silent LIMIT fallback.
P1 (BingX strings): STOP_MARKET → STOP_MARKET (protective, reduce-only)
STOP_LIMIT → STOP (protective)
TRIGGER_MARKET stays generic MIT
TRAILING_STOP → TRAILING_STOP_MARKET
Tests updated to match corrected mappings.
2026-07-18 15:08:57 +02:00
Codex
4926ef6788
malkhut: CHASE mechanics FIXED + Flight9 learnings + TTL enforcement
...
1. CHASE mechanics (NOW WORKING):
- CWM enforces TTL on open orders (auto-cancel when expired)
- DSL CHASE produces PLACE with metadata={chase: True}
- Action menu generates chase actions with wait_to_retry_ms TTL
- OpenOrderState.gains ttl_ms field (0=no expiry, >0=auto-cancel)
2. TTL enforcement (CWM):
- HftBacktestCWM: auto-cancels orders where age >= ttl_ms
- MinimalCryptoLOBCWM: same TTL enforcement
- This is how CHASE works: place→wait→auto-cancel→next step re-places
3. Flight9 learnings:
- Slippage model gains trade_flow_intensity parameter
- Book imbalance as proxy for trade arrival rate
- Markout = quality concept documented
4. CHASE tests: 10 new tests covering TTL enforcement, cancel-retry cycle,
max retries, DSL CHASE action, CMA codec integration
5. All 800+ tests pass
2026-07-17 19:30:17 +02:00
Codex
8857daedfa
malkhut: urgency-driven maker/taker + calibrated slippage + chase + docs
2026-07-17 10:16:55 +02:00
Codex
8b385cb249
malkhut(cwm): HftBacktestCWM — queue model + 59-test suite
...
HftBacktestCWM (cwm/hft_cwm.py):
- PowerProbQueueModel: probabilistic fill per level (pre-computed)
- Level 0 always fills, deeper levels have decreasing probability
- Deterministic fallback when use_queue_model=False
- Same transition/reward/terminal API as MinimalCryptoLOBCWM
- Fallback to deterministic level consumption when hftbacktest unavailable
59 tests (test_hft_cwm.py) covering 15 test classes:
1. Queue model correctness (fill probs, monotonic, bounds, determinism)
2. Determinism & reproducibility
3. CWM interface compatibility (cross, place, cancel, post_only, reduce)
4. Reward function (profit, noop, maker bonus)
5. Edge cases (empty book, zero qty, extreme price, many levels)
6. Position tracking (buy, sell, flip)
7. Fee application (taker fee reduces equity)
8. Counterparty ecology (toxic taker hits book, noop preserves)
9. CWM comparison (Hft vs Minimal agree on noop)
10. Venue propagation (scenario tagging, cross-exchange transfer)
11. PerformanceMatrix venue keying (record, per-venue best, comparison)
12. Risk gate integration (approve, leverage, OOD, kill switch, self-trade)
13. Stress tests (rapid transitions, 20 open orders, cancel all)
14. Full episode integration (single episode runs, policy evaluator)
15. hftbacktest availability check
2026-07-14 19:46:34 +02:00
Codex
186bce8984
malkhut(test): exhaustive order type + venue integration test suite — 159 tests
...
11 test classes covering all three orthogonal dimensions:
1. OrderType enum (10 tests): values, uppercase, str, hashable, frozen
2. TimeInForce enum (7 tests): values, default, IOC/FOK/GTD
3. OrderInstruction enum (3 tests): values
4. Exchange mapping tables (15 tests): all exchanges, all types, POST_ONLY
variation, trailing_stop BingX=BINANCE_MARKET
5. Normalization functions (9 tests): type, TIF, unknown exchange
6. is_type_available + get_supported_types (3 tests)
7. decompose_order (13 tests): all base, TIF, instructions, lowercase
8. FulfilmentAction (14 tests): frozen, time_in_force, post_only, reduce_only,
lazy TIF import, cancel_replace, metadata
9. State OrderType backward compat (4 tests): values, str, set, comparison
10. Scenario venue tagging (7 tests): default, custom, frozen, replace
11. ScenarioFactory venue propagation (8 tests): exchange_id, all venues
12. Cross-exchange transfer (11 tests): transfer, count, symbol, tags, idempotent
13. PerformanceMatrix venue-keying (16 tests): record, get_best, per-venue,
EMA, coverage, venue_comparison
14. CWM is_maker (8 tests): LIMIT, post_only, MARKET, STOP, trailing
15. Edge cases (12 tests): poison, zero scores, large scores, 100 strategies
16. Integration flow (6 tests): factory→transfer→matrix→selector
2026-07-14 16:03:30 +02:00
Codex
d24d9bc6bd
malkhut(wire): OrderType as three orthogonal dimensions — Fable's corrections
...
CRITICAL REFACTOR based on Fable's review (S9 roadmap item):
Before: flat enum conflating order types with TIF/instructions
OrderType had MARKET, LIMIT, IOC, FOK, POST_ONLY, REDUCE_ONLY, etc.
After: three orthogonal dimensions (FIX-aligned):
1. OrderType (Tag 40): what the order IS
LIMIT, MARKET, STOP_MARKET, STOP_LIMIT, TRIGGER_MARKET, TRIGGER_LIMIT,
TRAILING_STOP, OCO, TP_SL
2. TimeInForce (Tag 59): how long it LIVES
GTC, IOC, FOK, GTD
3. Instructions (Tag 18): behavioral modifiers
POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
Key corrections:
- POST_ONLY is an instruction on a LIMIT order, not a standalone type
- IOC/FOK are TimeInForce values, not order types
- BingX trailing_stop -> native TRAILING_STOP_MARKET (not TRIGGER_MARKET)
- FulfilmentAction.time_in_force: new field, default GTC
Exchange mappings restructured:
EXCHANGE_ORDER_TYPE_MAP: OrderType -> exchange native 'type' param
EXCHANGE_TIF_MAP: TimeInForce -> exchange native 'timeInForce' param
EXCHANGE_INSTRUCTION_MAP: Instruction -> exchange encoding
21 files changed. 380+ tests pass. Backward compatible.
2026-07-14 14:46:44 +02:00
Codex
53e02c84ec
malkhut: standardized order types — FIX/CCXT-aligned, multi-exchange mapping
...
order_types.py: Five-layer taxonomy normalized to industry standards:
Layer 1: Base types (FIX Tag 40) — MARKET, LIMIT
Layer 2: Time-in-force (FIX Tag 59) — GTC, IOC, FOK, GTD
Layer 3: Conditional/Trigger (FIX Tag 3/4+MIT) — STOP_MARKET, STOP_LIMIT,
TRIGGER_MARKET, TRIGGER_LIMIT, TRAILING_STOP
Layer 4: Instructions (FIX Tag 18) — POST_ONLY, REDUCE_ONLY, HIDDEN, ICEBERG
Layer 5: Compound (exchange-specific) — OCO, TP_SL
Cross-exchange mapping: BingX ↔ Binance ↔ Bybit (from CCXT source code).
Standards: FIX 4.4 Tag 40/59/18, CCXT unified API, ISO 10383 (MIC).
Transferability: strategy PARAMETERS transfer. ORDER TYPE NAMES are
venue-specific but semantics identical (LIMIT = LIMIT everywhere).
14 tests. README updated with full mapping table and standards references.
2026-07-14 12:02:01 +02:00
Codex
1f41be845b
malkhut(spec): item 4 — ScenarioLibrary sweep for Mode 1 coverage
...
ScenarioLibrary sweeps the state space (not samples) across:
- spread_mult: [0.1, 0.5, 1.0, 2.0, 5.0, 10.0]
- depth_fraction: [0.01, 0.05, 0.1, 0.3, 0.5, 1.0]
- toxicity: [0.0, 0.3, 0.7, 1.0]
- regime: [normal, crisis, recovery, transition]
Default: 13 assets × 576 grid points = 7,488 scenarios.
Customizable: specify symbols, dimensions, ranges.
7 tests covering: grid size, sweep output, point fields,
regime coverage, custom dimensions, summary, factory function.
2026-07-14 05:34:05 +02:00
Codex
833f262d12
malkhut(spec): item 9 — DAAT package (Direction-Anchored Ambiguity Triage)
...
DaatQuery: 8-feature market state representation
DaatVerdict: KNOWN / MARGINAL / OUT_OF_DISTRIBUTION
daat_classify: cosine RETRIEVE → magnitude GATE → local MODEL
- Cosine finds nearest explored state (directional match)
- Magnitude gate detects out-of-distribution states
- Empty explored set → always OUT_OF_DISTRIBUTION
9 tests covering: known state, OOD, empty explored, marginal, result fields.
No Unicode in code. All tests pass.
2026-07-14 02:33:31 +02:00
Codex
eef890a5cc
malkhut(spec): item 1 mutation-litmus + item 3 maker-fee UNVERIFIED comment
...
Item 1 — Mutation-litmus test (spec §1 item 3):
- test_taker_fee_10x_changes_score: fee change MUST affect score
- test_zero_fees_vs_correct_fees: zero vs 5bps must differ
- BOTH PASS — confirms fees ARE wired into reward function
- If fees were ignored, these tests would go RED
Item 3 — Maker fee verification (spec §1 item 5):
- Added '# UNVERIFIED — no maker fills on record as of 2026-07-13'
to Binance and Bybit exchange profiles
- Maker fee sign (positive on BingX, negative rebate on others)
is correct after fee fix but unverified from actual fills.
Items 2,4-10 remain for implementation.
2026-07-13 23:24:05 +02:00
Codex
5523be1d44
malkhut(fix): CORRECT FEE BUG — taker 0.5→5.0, maker -0.2→+2.0
...
Fable's spec (SPEC_MALKHUT_ACTUALS_INTAKE.md) confirmed 10x fee error
from our own fills (dolphin.trade_execution_quality).
Fixed:
- BingX taker: 0.5 → 5.0 bps
- BingX maker: -0.2 → +2.0 bps (POSITIVE on BingX, not a rebate)
- Binance taker: 0.4 → 4.5 bps
- Bybit taker: 0.06 → 5.5 bps
- All 13 per-asset profiles: maker=-0.2 taker=0.5 → maker=2.0 taker=5.0
Source of truth: dolphin.trade_execution_quality (8006 rows, avg taker=5.016 bps).
Every policy trained before this fix was at 10x too-cheap fees.
Re-measurement at correct fees is required.
2026-07-13 20:26:53 +02:00
Codex
db8e6d11f2
malkhut(scoring): fast scalar + advantage mode, reward execution quality
...
Fast scalar mode (default, for CMA loop):
- Rewards: fill quality (PnL when fills happen), moderate fill rate (5-15% sweet spot)
- Tolerates: no-fills (valid advisory recommendation)
- Penalizes: extreme fill rates (<3% lazy, >30% picked off), adverse selection, drawdown
- Light noop penalty (-0.5) vs old heavy (-50) — no-fills are valid signals
Advantage mode (for offline analysis):
- advantage = raw_performance - baseline_performance
- baseline = exponential moving average (decay=0.995)
- Clipped to [-10, +10]
- Reduces score variance 5.5x vs raw scoring
Scoring mode selection:
PolicyEvaluator(scoring_mode='fast') — default for CMA loop
PolicyEvaluator(scoring_mode='advantage') — for offline analysis
8 new tests for scoring modes. Total: 1186 tests, 50 files, all green.
2026-07-13 13:38:32 +02:00
Codex
c6b7a41bb4
malkhut(optim): vectorized reward + Ray parallel eval + VBT post-analysis
...
1. Vectorized reward path (cwm/core.py):
- Wired up existing compute_reward_vectorized from numba_core (was unused!)
- Eliminates FeatureVector dict allocation + Python dict lookups on hot path
- Numba path used when _HAS_NUMBA=True, Python fallback otherwise
- Bit-identical: same math operations, just via numba JIT
2. Ray-based parallel eval (training/ray_eval.py):
- Industrial multi-core execution via Ray (used by OpenAI/Anyscale)
- ray.put() stores params/scenarios in shared object store (no pickle per worker)
- Each worker: own CWM + planner, zero shared state, no races
- Bit-identical: same seed + same params = same results regardless of worker count
- PolicyEvaluator.evaluate_candidate: new use_ray=True parameter
3. VBT post-analysis (training/vbt_analysis.py):
- episodes_to_pnl_array, episodes_to_metrics (Sharpe, Sortino, VaR, win_rate, etc.)
- cross_asset_comparison, parameter_sensitivity
- format_metrics for human-readable output
- Analysis tool only — runs AFTER engine produces results
4. numba_core.py: added missing 'import math' for compute_reward_vectorized
13 new tests: vectorized reward bit-identity, Ray determinism, Ray result fields,
VBT metrics structure, cross-asset comparison, parameter sensitivity, edge cases.
Total: 1178 tests, 50 files, all green, zero regressions.
2026-07-12 23:56:16 +02:00
Codex
7f27ed22c8
malkhut: DuckDB file-backed asset store — schema, sync, queries, persistence
...
DuckDB store (asset_store.py):
- 4 tables: assets (28 cols), exchanges (10 cols), asset_exchanges (junction),
behavior_profiles (28 cols)
- Schema with indexes on blockchain, coingecko_id, cmc_id, exchange_id
- sync_from_profiles(): populate from in-memory dicts in one call
- Query API: query_assets(**filters), get_asset(), symbols_for_exchange(),
assets_on_blockchain(), asset_count(), exchange_count()
- Case-insensitive exchange lookup
- File-backed persistence: data survives connection close/reopen
- Performance: <1s sync, 100 queries in <1s, 1000 gets in <1s
30 tests covering:
- Schema creation and table structure (3 tests)
- Sync from profiles (6 tests, idempotent)
- Asset CRUD (7 tests, query by blockchain/coingecko/sector)
- Exchange mapping (5 tests, case-insensitive)
- Behavior profiles (4 tests)
- Persistence across connections (2 tests)
- Performance baseline (3 tests)
Total: 1276 tests across 50 files, all green, zero regressions.
2026-07-12 12:21:48 +02:00
Codex
019b620ab9
malkhut: asset bridge — directory ↔ classification integration
...
asset_bridge.py: connects Fable's AssetDirectory (operational layer,
runtime-mutable, JSON-backed listing status) with our AssetProfile
(taxonomic layer, frozen, invariant classification).
Functions:
- sync_asset_to_profile(directory, symbol): sync one asset's TRADING
exchanges from directory to AssetProfile.exchanges
- sync_exchanges_from_directory(directory): sync all matching assets
- get_universe_stats(directory): matched/unmatched counts
49 tests covering:
- normalize_symbol (7 edge cases)
- ExchangeListing validation (4 tests)
- AssetRecord listing queries (5 tests)
- AssetDirectory CRUD + persistence (14 tests)
- symbols_for_exchange filtering by status (4 tests)
- venue_symbol mapping (3 tests)
- Bridge sync (8 tests with state save/restore)
- Integration with ScenarioFactory + get_assets_on_exchange (4 tests)
Total: 1246 tests across 49 files, all green, zero regressions.
Fable's assets/ package preserved intact, interfaces retained.
2026-07-12 10:45:37 +02:00
Codex
1f709af6d1
malkhut: three-layer identifier architecture for cross-system asset identification
...
Layer 1 (canonical identity): symbol, base_asset, name, unified_symbol
(CCXT format), quote_currency
Layer 2 (cross-system): coingecko_id, cmc_id, blockchain, contract_address
Layer 3 (exchange mapping): exchanges tuple
All 13 pre-defined assets migrated with accurate CoinGecko IDs, CMC IDs,
blockchains, and contract addresses (ERC-20 tokens).
7 new query functions: get_asset_by_coingecko_id, get_asset_by_cmc_id,
get_assets_by_base_asset, get_assets_by_blockchain,
get_assets_by_unified_symbol.
33 new tests covering: Layer 1 identity, Layer 2 cross-system identifiers,
identifier query functions, identifier consistency (uniqueness, derivation),
exchange registry.
Total: 1189 tests, 47 files, all green.
Based on research: CCXT BASE/QUOTE is de facto standard, CoinGecko ID
most widely used in crypto-native, ISO 24165 DTI emerging, FIGI for
institutional.
2026-07-12 00:40:14 +02:00
Codex
a062507656
malkhut(assets)+uv: system-wide asset directory + UV universe init
...
MALKHUT asset directory (malkhut/assets/): normalized canonical symbols,
KNOWN_EXCHANGES aux table (BINANCE/BINGX/BINGX_VST), per-exchange listing
status (TRADING/OFFLINE/UNKNOWN), JSON-backed, built for full-Binance-500
scale. Seeded with BLUE's NG7 feed universe (50 symbols) + live VST
contracts probe: 35 TRADING / 15 OFFLINE on VST (BAND, CELR, COS, CVC,
DENT, FUN, HOT, ICX, TFUEL, TUSD, USDC, WAN, WIN, XTZ, ZIL).
uv_asset_universe.init_asset_universe() = runtime tradable set for the
execution exchange; --onboard runs MALKHUT AssetCompiler for full taxonomy.
12 new tests, mutation-RED verified (status-filter + tradability guard).
2026-07-11 21:48:05 +02:00
Codex
be0e1468da
malkhut(perf): parallel episode eval — 9x single-eval speedup, zero fidelity loss
...
- parallel_eval.py: ProcessPoolExecutor-based episode runner. Each worker
gets its own CWM + planner instance. Zero shared state = embarrassingly
parallel. Deterministic: same seed → same result.
- PolicyEvaluator.evaluate_candidate: new workers parameter (0=sequential,
>1=parallel). Backward compatible: default workers=0.
- 16 new tests: determinism, pickling, result validity, cross-validation
between sequential and parallel paths, backward compatibility.
- README: training performance table with speedup measurements.
Speedup results (3 assets × 30 scenarios = 90 scenarios):
Sequential: 3.3s per eval (1.0x)
2 workers: 1.3s per eval (2.6x)
4 workers: 0.6s per eval (5.9x)
8 workers: 0.4s per eval (9.1x)
CMA-ES 48 evals: 125s → 85s (1.5x training speedup)
Note: CWM numba hot path was already wired (_HAS_NUMBA=True, 5.3µs/transition).
Bottleneck is MCTS planner (96% of eval time), not CWM.
2026-07-11 20:19:32 +02:00
Codex
4c239f7774
malkhut(tests): 1140 test functions across 46 test files
...
CWM (103): core mechanics, exhaustive edge cases, numba, exchange mechanics
Replay (118): exhaustive verification, microstructure, trajectory
Training (190): asset classification, phase0 extensive, pipeline, exhaustive
DSL (102): v2 syntax, expanded, new features
ASEx (33): validate-before-mutate, single-writer
Planner (48): MCTS, alternatives, hooks
Counterparties (19): 9 adversarial agent policies
Clock (30): event-driven reactor
BingX (28): venue adapter
IPC (8): Zinc SHM
Storage (9): ClickHouse
Risk (4): hard invariants
State (17): frozen dataclass invariants
Integration: E2E, concurrency, sync/async seams, hypothesis, fuzz, adversarial
2026-07-11 10:46:12 +02:00
Codex
981b469d51
malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
...
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
(BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation
Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00