Codex
db8e6d11f2
malkhut(scoring): fast scalar + advantage mode, reward execution quality
Fast scalar mode (default, for CMA loop):
- Rewards: fill quality (PnL when fills happen), moderate fill rate (5-15% sweet spot)
- Tolerates: no-fills (valid advisory recommendation)
- Penalizes: extreme fill rates (<3% lazy, >30% picked off), adverse selection, drawdown
- Light noop penalty (-0.5) vs old heavy (-50) — no-fills are valid signals
Advantage mode (for offline analysis):
- advantage = raw_performance - baseline_performance
- baseline = exponential moving average (decay=0.995)
- Clipped to [-10, +10]
- Reduces score variance 5.5x vs raw scoring
Scoring mode selection:
PolicyEvaluator(scoring_mode='fast') — default for CMA loop
PolicyEvaluator(scoring_mode='advantage') — for offline analysis
8 new tests for scoring modes. Total: 1186 tests, 50 files, all green.
2026-07-13 13:38:32 +02:00
..
2026-07-11 21:57:04 +02:00
2026-07-11 10:36:30 +02:00
2026-07-12 23:56:16 +02:00
2026-07-11 10:31:31 +02:00
2026-07-11 10:31:31 +02:00
2026-07-11 10:26:01 +02:00
2026-07-11 10:31:31 +02:00
2026-07-12 20:52:56 +02:00
2026-07-13 13:38:32 +02:00
2026-07-13 13:38:32 +02:00
2026-07-11 10:31:31 +02:00
2026-07-11 10:21:27 +02:00
2026-07-11 10:21:27 +02:00
2026-07-11 10:23:44 +02:00
2026-07-11 10:39:03 +02:00
2026-07-11 10:39:03 +02:00
2026-07-11 10:26:01 +02:00
2026-07-11 10:26:01 +02:00
2026-07-11 10:21:27 +02:00
2026-07-11 10:21:27 +02:00
2026-07-11 10:41:31 +02:00
2026-07-11 10:41:31 +02:00
2026-07-11 10:21:27 +02:00