malkhut(optim): vectorized reward + Ray parallel eval + VBT post-analysis
1. Vectorized reward path (cwm/core.py): - Wired up existing compute_reward_vectorized from numba_core (was unused!) - Eliminates FeatureVector dict allocation + Python dict lookups on hot path - Numba path used when _HAS_NUMBA=True, Python fallback otherwise - Bit-identical: same math operations, just via numba JIT 2. Ray-based parallel eval (training/ray_eval.py): - Industrial multi-core execution via Ray (used by OpenAI/Anyscale) - ray.put() stores params/scenarios in shared object store (no pickle per worker) - Each worker: own CWM + planner, zero shared state, no races - Bit-identical: same seed + same params = same results regardless of worker count - PolicyEvaluator.evaluate_candidate: new use_ray=True parameter 3. VBT post-analysis (training/vbt_analysis.py): - episodes_to_pnl_array, episodes_to_metrics (Sharpe, Sortino, VaR, win_rate, etc.) - cross_asset_comparison, parameter_sensitivity - format_metrics for human-readable output - Analysis tool only — runs AFTER engine produces results 4. numba_core.py: added missing 'import math' for compute_reward_vectorized 13 new tests: vectorized reward bit-identity, Ray determinism, Ray result fields, VBT metrics structure, cross-asset comparison, parameter sensitivity, edge cases. Total: 1178 tests, 50 files, all green, zero regressions.
This commit is contained in:
@@ -46,6 +46,7 @@ try:
|
||||
round_tick as _nb_round_tick,
|
||||
round_lot as _nb_round_lot,
|
||||
clip_lots as _nb_clip_lots,
|
||||
compute_reward_vectorized,
|
||||
)
|
||||
_HAS_NUMBA = True
|
||||
except ImportError:
|
||||
@@ -587,6 +588,39 @@ class MinimalCryptoLOBCWM:
|
||||
next_state: MarketWorldState,
|
||||
params: FulfilmentPolicyParams,
|
||||
) -> float:
|
||||
if _HAS_NUMBA:
|
||||
# Fast path: numba-optimized reward computation
|
||||
# Extract values directly — avoids FeatureVector dict allocation
|
||||
b = next_state.book
|
||||
path = next_state.trade_path
|
||||
|
||||
pnl = path.pnl_bps if path else 0.0
|
||||
toxicity = path.orderflow_toxicity if path else 0.0
|
||||
churn = path.queue_churn_score if path else 0.0
|
||||
time_in_loss = path.time_in_loss_s if path else 0.0
|
||||
spread_bps = b.spread_bps if b.bids and b.asks else 0.0
|
||||
|
||||
inv_risk = self._inventory_risk(next_state)
|
||||
tail_risk = self._tail_risk_proxy(next_state)
|
||||
|
||||
is_maker = (action.order_type and
|
||||
action.order_type.value in ("POST_ONLY", "LIMIT"))
|
||||
is_cross = action.kind.value == "CROSS_SPREAD"
|
||||
is_cancel = action.kind.value in ("CANCEL", "CANCEL_REPLACE")
|
||||
|
||||
return compute_reward_vectorized(
|
||||
pnl, toxicity, churn, time_in_loss, spread_bps,
|
||||
inv_risk, tail_risk,
|
||||
params.w_expected_pnl, params.w_adverse_selection,
|
||||
params.w_inventory_risk, params.w_tail_loss, params.w_time_decay,
|
||||
is_maker, prev_state.venue.maker_fee_bps,
|
||||
is_cross, prev_state.venue.taker_fee_bps,
|
||||
is_cancel, params.adverse_toxicity_cancel_threshold,
|
||||
params.queue_churn_cancel_threshold,
|
||||
params.w_queue_priority, params.w_adverse_selection,
|
||||
)
|
||||
|
||||
# Fallback: Python path (no numba)
|
||||
fv = self.feature_extractor.extract(next_state).values
|
||||
|
||||
pnl = fv.get("pnl_bps", 0.0)
|
||||
|
||||
Reference in New Issue
Block a user