Commit Graph

7 Commits

Author SHA1 Message Date
Codex
401d5a70ca malkhut(wire): venue tagging + cross-exchange transfer + CWM order type fix
ScenarioFactory + CWM + Engine changes:

1. Scenario.venue field (default='bingx') — each scenario tagged with venue
2. ScenarioFactory.exchange_id parameter — controls which exchange scenarios simulate
3. _make_state + _behavior_state: venue propagated to VenueRules.exchange
4. All 34 scenario builders: venue=self.exchange_id
5. cross_exchange_transfer(): re-tag scenarios for different exchange
   (strategy evolved on BingX can be re-evaluated on Binance)
6. CWM core.py: is_maker check updated for three-dimensional order model
   (POST_ONLY no longer in OrderType; uses post_only flag instead)

Cross-exchange learning flow:
  factory_bingx = ScenarioFactory(exchange_id='bingx')
  scenarios_bingx = factory_bingx.build_suite(symbols=[...])
  strategy = train(scenarios_bingx)  # evolve on BingX

  factory_binance = ScenarioFactory(exchange_id='binance')
  scenarios_binance = factory_bingx.cross_exchange_transfer(
      scenarios_bingx, target_exchange='binance')
  score = evaluate(strategy, scenarios_binance)  # test on Binance

All tests pass. Strategy PARAMETERS transfer; only venue tag + fees + order mapping change.
2026-07-14 15:18:56 +02:00
Codex
d9b7e05531 malkhut(perf): optimize _run_episode — reduced Python overhead
Optimizations in _run_episode:
- Pre-allocated ActionKind constants (avoid repeated attribute lookups)
- Removed unnecessary max_pos_qty tracking (unused in scoring)
- Simplified action kind checks (single comparison chain)
- Reduced frozen dataclass allocations per step

Result: same behavioral output, cleaner code path.
Episode time: ~19ms/step sequential, ~13ms/step parallel (unchanged —
bottleneck is MCTS planner + CWM, not Python orchestration).
2026-07-13 15:11:19 +02:00
Codex
db8e6d11f2 malkhut(scoring): fast scalar + advantage mode, reward execution quality
Fast scalar mode (default, for CMA loop):
  - Rewards: fill quality (PnL when fills happen), moderate fill rate (5-15% sweet spot)
  - Tolerates: no-fills (valid advisory recommendation)
  - Penalizes: extreme fill rates (<3% lazy, >30% picked off), adverse selection, drawdown
  - Light noop penalty (-0.5) vs old heavy (-50) — no-fills are valid signals

Advantage mode (for offline analysis):
  - advantage = raw_performance - baseline_performance
  - baseline = exponential moving average (decay=0.995)
  - Clipped to [-10, +10]
  - Reduces score variance 5.5x vs raw scoring

Scoring mode selection:
  PolicyEvaluator(scoring_mode='fast') — default for CMA loop
  PolicyEvaluator(scoring_mode='advantage') — for offline analysis

8 new tests for scoring modes. Total: 1186 tests, 50 files, all green.
2026-07-13 13:38:32 +02:00
Codex
459215b7d8 malkhut(fix): wire workers into CMA training loop
CMAESTrainer.train() now accepts workers parameter and passes it to
evaluate_candidate(), enabling parallel episode evaluation during
actual training (not just in tests/benchmarks).

Benchmark result: ProcessPoolExecutor is optimal (4.76x speedup).
Ray is slower (0.36x) due to head init + plasma overhead for 90 scenarios.
2026-07-13 03:25:55 +02:00
Codex
c6b7a41bb4 malkhut(optim): vectorized reward + Ray parallel eval + VBT post-analysis
1. Vectorized reward path (cwm/core.py):
   - Wired up existing compute_reward_vectorized from numba_core (was unused!)
   - Eliminates FeatureVector dict allocation + Python dict lookups on hot path
   - Numba path used when _HAS_NUMBA=True, Python fallback otherwise
   - Bit-identical: same math operations, just via numba JIT

2. Ray-based parallel eval (training/ray_eval.py):
   - Industrial multi-core execution via Ray (used by OpenAI/Anyscale)
   - ray.put() stores params/scenarios in shared object store (no pickle per worker)
   - Each worker: own CWM + planner, zero shared state, no races
   - Bit-identical: same seed + same params = same results regardless of worker count
   - PolicyEvaluator.evaluate_candidate: new use_ray=True parameter

3. VBT post-analysis (training/vbt_analysis.py):
   - episodes_to_pnl_array, episodes_to_metrics (Sharpe, Sortino, VaR, win_rate, etc.)
   - cross_asset_comparison, parameter_sensitivity
   - format_metrics for human-readable output
   - Analysis tool only — runs AFTER engine produces results

4. numba_core.py: added missing 'import math' for compute_reward_vectorized

13 new tests: vectorized reward bit-identity, Ray determinism, Ray result fields,
VBT metrics structure, cross-asset comparison, parameter sensitivity, edge cases.
Total: 1178 tests, 50 files, all green, zero regressions.
2026-07-12 23:56:16 +02:00
Codex
be0e1468da malkhut(perf): parallel episode eval — 9x single-eval speedup, zero fidelity loss
- parallel_eval.py: ProcessPoolExecutor-based episode runner. Each worker
  gets its own CWM + planner instance. Zero shared state = embarrassingly
  parallel. Deterministic: same seed → same result.
- PolicyEvaluator.evaluate_candidate: new workers parameter (0=sequential,
  >1=parallel). Backward compatible: default workers=0.
- 16 new tests: determinism, pickling, result validity, cross-validation
  between sequential and parallel paths, backward compatibility.
- README: training performance table with speedup measurements.

Speedup results (3 assets × 30 scenarios = 90 scenarios):
  Sequential:  3.3s per eval  (1.0x)
  2 workers:   1.3s per eval  (2.6x)
  4 workers:   0.6s per eval  (5.9x)
  8 workers:   0.4s per eval  (9.1x)
  CMA-ES 48 evals: 125s → 85s (1.5x training speedup)

Note: CWM numba hot path was already wired (_HAS_NUMBA=True, 5.3µs/transition).
Bottleneck is MCTS planner (96% of eval time), not CWM.
2026-07-11 20:19:32 +02:00
Codex
981b469d51 malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
  templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
  rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
  (BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
  BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation

Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00