malkhut(docs + bench): comprehensive update + smoke test script
README updated with: - Vectorized UCB selection (7.7x speedup, 1.13µs/selection) - Batch MCTS kernel (numba-accelerated) - Fast scalar + advantage scoring modes - Updated performance benchmarks (1186 tests, 390 scenarios, 3043 score/min) - Advantage scorer module in package structure smoke_1h.py: standalone training script for extended runs. Total session: 19 commits, 1186 tests, all green. All implementations: parallel eval (7x), vectorized reward (numba), vectorized UCB (7.7x), fast scalar scoring, advantage mode, DuckDB store (sub-µs reads), asset compiler, behavior DSL, multi-exchange support, three-layer identifiers.
This commit is contained in:
@@ -144,6 +144,7 @@ MALKHUT/
|
||||
│ │ ├── parallel_eval.py # ProcessPoolExecutor episode runner
|
||||
│ │ ├── ray_eval.py # Ray-based eval (industrial alternative)
|
||||
│ │ ├── vbt_analysis.py # Post-sim metrics: Sharpe, Sortino, VaR
|
||||
│ │ ├── advantage_scorer.py # Advantage estimation for offline analysis
|
||||
│ │ ├── cognition.py # Rate-limited market regime research
|
||||
│ │ ├── regime_expansion.py # 200+ regimes from dimension combinations
|
||||
│ │ ├── news_sources.py # 12 industry-standard news sources
|
||||
@@ -402,14 +403,14 @@ simple doctrinal tick-exits (C11) ship first via T19 step 3; MALKHUT supersedes
|
||||
|
||||
## DEVELOPMENT STATUS (2026-07-13)
|
||||
|
||||
**1178 test functions. 50 test files. All green. 0 failures. 0 regressions.**
|
||||
**1186 test functions. 50 test files. All green. 0 failures. 0 regressions.**
|
||||
|
||||
### Completed subsystems
|
||||
|
||||
| Subsystem | Module | Tests | Status |
|
||||
|-----------|--------|-------|--------|
|
||||
| **State Model** | `state.py` | 17 | 42 frozen dataclasses, immutable |
|
||||
| **CWM** | `cwm/core.py` + `cwm/numba_core.py` | 103 | Exchange mechanics + numba JIT (5.3µs/transition) + vectorized reward |
|
||||
| **CWM** | `cwm/core.py` + `cwm/numba_core.py` | 103 | Exchange mechanics + numba JIT (5.3µs/transition) + vectorized reward + vectorized UCB + batch MCTS kernel |
|
||||
| **Replay Verification** | `cwm/replay_verify.py` | 65 | Deep comparison, binary search, trajectory recording |
|
||||
| **Planner** | `planner/sm_mcts.py` | 11 | Decoupled UCB/UCT, ≤25ms budget |
|
||||
| **Action Menu** | `planner/action_menu.py` | (in planner) | Compact action space construction |
|
||||
@@ -424,6 +425,7 @@ simple doctrinal tick-exits (C11) ship first via T19 step 3; MALKHUT supersedes
|
||||
| **CMA-ES Training** | `training/cma_trainer.py` | 65 | Behavior-driven, auto-compile, parallel workers, 7x speedup |
|
||||
| **Parallel Eval** | `training/parallel_eval.py` | 16 | ProcessPoolExecutor, 7x CMA-ES speedup |
|
||||
| **Ray Eval** | `training/ray_eval.py` | 5 | Ray-based eval (available, slower for ≤1K scenarios) |
|
||||
| **Scoring Modes** | `cma_trainer.py` + `advantage_scorer.py` | 8 | Fast scalar (CMA loop) + advantage (offline analysis) |
|
||||
| **VBT Analysis** | `training/vbt_analysis.py` | 8 | Post-sim trade metrics: Sharpe, Sortino, VaR, cross-asset |
|
||||
| **Policy Registry** | `training/registry.py` | 14 | CANDIDATE → ACTIVE lifecycle |
|
||||
| **Training Pipeline** | `training/pipeline.py` | 21 | Bounded continuous learning loop |
|
||||
@@ -451,17 +453,18 @@ simple doctrinal tick-exits (C11) ship first via T19 step 3; MALKHUT supersedes
|
||||
|
||||
| Metric | Value |
|
||||
|--------|-------|
|
||||
| CWM transition | 5.3 µs/call (numba JIT) |
|
||||
| CWM throughput | 189K calls/sec |
|
||||
| CWM 100-step episode | 0.64 ms |
|
||||
| CWM reward (numba vectorized) | ~0.3µs (was 2µs with dict) |
|
||||
| CWM throughput | 189K calls/sec (numba JIT) |
|
||||
| CWM per-call latency | 5.3 µs |
|
||||
| CWM reward (numba vectorized) | ~0.3µs |
|
||||
| UCB selection (numba vectorized) | 1.13µs (was 8.7µs, 7.7x speedup) |
|
||||
| Numba fill speedup | 1.8x (batch 100) |
|
||||
| DuckDB asset reads | 0.2µs (in-memory) |
|
||||
| Scenario generation | 390 scenarios in 0.8s |
|
||||
| CMA-ES parallel (8 workers) | 7× speedup, 87% efficiency |
|
||||
| Best CMA-ES score (48 evals) | 2,594 (parallel) vs 1,727 (sequential) |
|
||||
| Score/min (8 workers) | 1,718 (was 164 sequential) |
|
||||
| Peak RAM | 146 MB |
|
||||
| Best CMA-ES score (48 evals) | 8,628 (fast scalar, 100.4 bps PnL) |
|
||||
| Score/min (8 workers) | 3,043 |
|
||||
| Episode throughput (parallel) | 16ms/ep |
|
||||
| Episode throughput (sequential) | 23ms/ep |
|
||||
|
||||
### Bugs found and fixed (22 total)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user