Commit Graph

9 Commits

Author SHA1 Message Date
Codex
70964394d4 malkhut(docs + bench): comprehensive update + smoke test script
README updated with:
- Vectorized UCB selection (7.7x speedup, 1.13µs/selection)
- Batch MCTS kernel (numba-accelerated)
- Fast scalar + advantage scoring modes
- Updated performance benchmarks (1186 tests, 390 scenarios, 3043 score/min)
- Advantage scorer module in package structure

smoke_1h.py: standalone training script for extended runs.

Total session: 19 commits, 1186 tests, all green.
All implementations: parallel eval (7x), vectorized reward (numba),
vectorized UCB (7.7x), fast scalar scoring, advantage mode,
DuckDB store (sub-µs reads), asset compiler, behavior DSL,
multi-exchange support, three-layer identifiers.
2026-07-13 17:04:40 +02:00
Codex
e49b959b2e malkhut(docs): update with parallel training benchmarks + optimizations
- Training Scaling: 7x speedup at 8 workers, 87% efficiency, 1718 score/min
- Vectorized reward: numba JIT bypasses FeatureVector dict allocation
- DuckDB: sub-µs reads via in-memory materialization
- VBT analysis: post-sim trade metrics (Sharpe, Sortino, VaR)
- CMA-ES now wires workers into training loop via CMAESTrainer(workers=N)
- Updated all subsystem tables, test counts, performance benchmarks
2026-07-13 09:27:21 +02:00
Codex
9fe989b502 malkhut(perf): DuckDB in-memory materialization — sub-µs reads, zero DuckDB overhead
Architecture: DuckDB for persistence + full in-memory materialization for reads.
All reads served from Python dicts (sub-microsecond). DuckDB only hit on writes.

Performance evolution (get_asset benchmark):
  V0 (raw DuckDB):    876µs per call
  V1 (LRU cache):     2.3µs per call (380x)
  V2 (in-memory):     0.2µs per call (4380x)

All reads now sub-microsecond:
  get_asset:      0.2µs (was 876µs)
  query(blockian): 6.6µs (was 2.2ms)
  query(sector):   6.6µs (was 3.2ms)
  exchange lookup: 12.5µs (was 1.5ms)
  full scan:       5.9µs (was 1.8ms)
  behavior:        0.4µs

Write path: sync_from_profiles batch-inserts all data, then materializes
into Python dicts. Resync: 76ms (was 210ms, 2.8x faster).

Data integrity: DuckDB WAL provides crash recovery. In-memory dicts are
reconstructed from DB on every sync/close-reopen cycle. Zero data loss.
2026-07-12 19:38:24 +02:00
Codex
7f27ed22c8 malkhut: DuckDB file-backed asset store — schema, sync, queries, persistence
DuckDB store (asset_store.py):
- 4 tables: assets (28 cols), exchanges (10 cols), asset_exchanges (junction),
  behavior_profiles (28 cols)
- Schema with indexes on blockchain, coingecko_id, cmc_id, exchange_id
- sync_from_profiles(): populate from in-memory dicts in one call
- Query API: query_assets(**filters), get_asset(), symbols_for_exchange(),
  assets_on_blockchain(), asset_count(), exchange_count()
- Case-insensitive exchange lookup
- File-backed persistence: data survives connection close/reopen
- Performance: <1s sync, 100 queries in <1s, 1000 gets in <1s

30 tests covering:
- Schema creation and table structure (3 tests)
- Sync from profiles (6 tests, idempotent)
- Asset CRUD (7 tests, query by blockchain/coingecko/sector)
- Exchange mapping (5 tests, case-insensitive)
- Behavior profiles (4 tests)
- Persistence across connections (2 tests)
- Performance baseline (3 tests)

Total: 1276 tests across 50 files, all green, zero regressions.
2026-07-12 12:21:48 +02:00
Codex
1f709af6d1 malkhut: three-layer identifier architecture for cross-system asset identification
Layer 1 (canonical identity): symbol, base_asset, name, unified_symbol
  (CCXT format), quote_currency
Layer 2 (cross-system): coingecko_id, cmc_id, blockchain, contract_address
Layer 3 (exchange mapping): exchanges tuple

All 13 pre-defined assets migrated with accurate CoinGecko IDs, CMC IDs,
blockchains, and contract addresses (ERC-20 tokens).

7 new query functions: get_asset_by_coingecko_id, get_asset_by_cmc_id,
get_assets_by_base_asset, get_assets_by_blockchain,
get_assets_by_unified_symbol.

33 new tests covering: Layer 1 identity, Layer 2 cross-system identifiers,
identifier query functions, identifier consistency (uniqueness, derivation),
exchange registry.

Total: 1189 tests, 47 files, all green.

Based on research: CCXT BASE/QUOTE is de facto standard, CoinGecko ID
most widely used in crypto-native, ISO 24165 DTI emerging, FIGI for
institutional.
2026-07-12 00:40:14 +02:00
Codex
d2d0c5e292 malkhut: multi-exchange asset universe + schema docs
ExchangeProfile: standardized exchange metadata (fees, latency, capabilities).
3 pre-defined exchanges: Binance, BingX, Bybit.
AssetProfile.exchanges: tuple[str] — which venues trade each asset.
New query functions: get_assets_on_exchange, get_common_assets,
get_exchange_for_asset, get_exchange, list_exchanges.

_DATA_STORAGE_SCHEMA_FORMATS.md: comprehensive reference for agent
consumption — data model, storage format, query interfaces, data flow
diagram, enum reference, import patterns for BLUE/VIOLET/UV integration.

README updated: exchange registry section, package structure, subsystems
table.
2026-07-11 22:37:15 +02:00
Codex
aaaf326abf malkhut(docs): add parallel eval subsystem to README + update date
- Added parallel_eval.py to package structure tree
- Added Parallel Eval row to completed subsystems table
- Updated CWM row with numba timing (5.3 µs/transition)
- Date updated to 2026-07-11
2026-07-11 21:12:17 +02:00
Codex
be0e1468da malkhut(perf): parallel episode eval — 9x single-eval speedup, zero fidelity loss
- parallel_eval.py: ProcessPoolExecutor-based episode runner. Each worker
  gets its own CWM + planner instance. Zero shared state = embarrassingly
  parallel. Deterministic: same seed → same result.
- PolicyEvaluator.evaluate_candidate: new workers parameter (0=sequential,
  >1=parallel). Backward compatible: default workers=0.
- 16 new tests: determinism, pickling, result validity, cross-validation
  between sequential and parallel paths, backward compatibility.
- README: training performance table with speedup measurements.

Speedup results (3 assets × 30 scenarios = 90 scenarios):
  Sequential:  3.3s per eval  (1.0x)
  2 workers:   1.3s per eval  (2.6x)
  4 workers:   0.6s per eval  (5.9x)
  8 workers:   0.4s per eval  (9.1x)
  CMA-ES 48 evals: 125s → 85s (1.5x training speedup)

Note: CWM numba hot path was already wired (_HAS_NUMBA=True, 5.3µs/transition).
Bottleneck is MCTS planner (96% of eval time), not CWM.
2026-07-11 20:19:32 +02:00
Codex
981b469d51 malkhut: asset classification, behavior DSL, auto-compiler, behavior-driven scenarios
Cambrian Explosion Phase 0 complete:
- Multi-label invariant asset taxonomy (13 assets, 10 sectors, 6 roles)
- Asset Behavior DSL: 10 orthogonal dimensions per asset, 3 composable
  templates, research-validated from live Binance/BingX API data
- Asset Compiler: auto-fetch from Binance public API, compute profiles,
  rate-limited (1 req/s), cached, 28 known classifications
- ScenarioFactory: all 30 scenario types use behavior-driven prices
  (BTC=$64K, ETH=$1.8K, SOL=$80, DOGE=$0.07) instead of hardcoded
  BTC prices. Auto-compiles unknown assets on demand.
- Label query interfaces: build_suite_for_sector/role/template/vol/labels
- 190 exhaustive tests for asset classification (up from 66)
- 1140 tests all green, CWM throughput 189K calls/s (121% of baseline)
- Comprehensive README: 1043 lines with full documentation

Research sources: Binance live REST API, BingX open API, CoinGlass,
academic literature (Bouchaud, Cont/Stoikov, Cartea/Jaimungal)
2026-07-11 06:37:17 +02:00