Files
siloqy/prod/docs/VIOLET_TODO_CRITICAL.md
Codex f8826f613d VIOLET: record PASS4-done + Task 17/18 caveats in the review queue (PR docs)
PASS 4 completed on worktree agent/oa-violet4 (/mnt/vp-oa4). Recorded in the review queue the two
caveats the OA agent reported — also flagged in-code as prominent banner comments + empty marker
commits on that branch:
  - Task 17 exec_intent: explicit reference_price required (ShadowDecision has notional, no price).
  - Task 18 alpha_data_feed: parse/feed-port only, no live NT/Binance in unit code.
For Claude's later PASS-4 review (verify the caveats + non-vacuous tests).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 12:25:38 +02:00

7.7 KiB
Raw Permalink Blame History

VIOLET — CRITICAL TODO / review queue (prominent)

Date: 2026-06-17. Single place for the must-not-forget VIOLET items. Review-later, not now.


🔴 CRITICAL #1 — VIOLET↔BLUE parity is VERY DISAPPOINTING (review + root-cause BEFORE any soak/V4)

Report: prod/VIOLET_dev/reports/violet_parity_20260616_220412.md (+ .json). Window 2026-06-14 20:15 → 2026-06-15 21:00. Produced by PASS-2 Task 4 (parity_report.py).

Headline:

  • Pick-match rate: 0.015 (1.5%) — VIOLET rows 2853, exact pick matches only 43.
  • Same-asset rate: 0.136 (13.6%); no-pick: 2465 / 2853 (86%).
  • BLUE rows 25941 (scan_eval 25278 + trade_events 663).

KEY NUANCE (where the review should START): on the 43 rows that DID align, sizing is near-identical — leverage abs error mean 0.016 / median 0.0 / max 0.135. So the V3.4 sizing math (boost/beta/mc_scale/ob/esof/compose) is NOT the problem; the divergence is in ASSET SELECTION / TIMING / the comparison's ALIGNMENT method. Candidate causes to investigate (do not assume — measure):

  1. Apples-to-oranges population. BLUE scan_eval (25278) is likely per-scan-per-asset evaluations, while VIOLET rows are actuated decisions — the report may be comparing different things. Verify the alignment/join semantics in parity_report.py first.
  2. Selection divergence. VIOLET's VioletAssetSelector (IRP) vs BLUE's live selection over the same scan stream — are they fed the same universe/lookback at the same scan index? The sizing-gap samples (TRX/ATOM/LTC/XLM SHORT with large notional_rel_err) suggest VIOLET fires on assets BLUE sized very differently or didn't pick.
  3. Cadence/actuation. VIOLET actuates at Q=scan; if its scan alignment or dedupe differs, picks land at different scans → counted as no-pick.
  4. The known structural items (OB single-shot before V3.4d; mc_scale; live-factor sourcing) — re-run parity AFTER V3.4d's persistent-OB launcher + the bit-identity fixes to see if the number moves.

Action: full review of parity_report.py + a root-cause pass; fix the alignment OR the selection divergence; re-run. This gates a meaningful DARK soak — a 1.5% pick-match makes the soak uninterpretable. Owner: Claude (me), later.


🟡 #2 — Review the OA-delegated PASS work (NOT yet reviewed)

The parallel agent reports these DONE; none reviewed for correctness / BLUE-compliance yet.

PASS 1 (VIOLET_PART_SPEC_OA_TODO.md):

  • 53bdd90 sizing parity-pin tests
  • 12b768b venue OB provider seam (venue_ob_provider.py)
  • bae9284 base-fraction sizing study (+ archived report)

PASS 2 (VIOLET_PART_SPEC_OA_TODO_PASS2.md):

  • parity_report.py (Task 4 — the report above) + test
  • tradeability.py (Task 6) + test
  • test_violet_v3_decision_latency_gate.py (Task 5) → report violet_v3_decision_latency_2026...
  • test_violet_replay_determinism_gate.py (Task 7) → report violet_replay_determinism_2026...

PASS 3: VIOLET_PART_SPEC_OA_TODO_PASS3.md (issued 2026-06-17) — venue feed / mechanical exits / SL floor / event-restore / slippage / cadence; review when done.

PASS 4: VIOLET_PART_SPEC_OA_TODO_PASS4.md (issued 2026-06-17) — vol gate / V7 exit wrapper / economics ledger / exec-intent / alpha data feed / time exits. DONE on worktree agent/oa-violet4 (/mnt/vp-oa4) — commits 6996a6a/7c8e524/bbebe58/3920665/f328d3a/b39d8a4. Review when done. Caveats to verify in review (flagged in-code as prominent banners + empty marker commits 57ca851/8624dee/2aad0aa/539ea47):

  • Task 17 (exec_intent): to_exec_intent needs an EXPLICIT reference_price arg — ShadowDecision has notional/exposure but NO price field, so qty can't be derived from a decision alone. Integration owner threads the live price (PASS-3 VenueTick) at wiring time.
  • Task 18 (alpha_data_feed): intentionally parse/feed-port ONLY — no live NT node / Binance connection in unit code (NT owns its loop; dummy keys; separate feed process). Live node = the integration owner's job. Confirm tests aren't vacuous (parse-fixture, not a live connection).

PASS 5: VIOLET_PART_SPEC_OA_TODO_PASS5.md (issued 2026-06-17) — MOCK-BINGX execution stack (exec contracts + QuirkProfile seams / order FSM / mock venue / fill reducer / DARK E2E gate). NOTE: BingX "quirks" are SEAM-ONLY here (default OFF) — a later quirk-injection pass + a mandatory real-key boundary smoke are required before V4-live; review when done.

PASS 6: VIOLET_PART_SPEC_OA_TODO_PASS6.md (issued 2026-06-17) — execution INTERNALS (fill-pump/ownership filter, reconcile/zero-wb guard, TTL-requote, orphan handling) + the QUIRK-INJECTION gate that flips PASS-5's QuirkProfile flags ON and proves each handler neutralizes the quirk. Mirrors PINK's production fixes (pink_direct.py). Real-key smoke still MANDATORY before V4-live; review when done.

PASS 7: VIOLET_PART_SPEC_OA_TODO_PASS7.md (issued 2026-06-17) — V5 selection (faithful ARS/IRP ranking + OB Sub-1), multi-asset slot manager, capital allocation, multi-asset flow, and a ranking bit-identity gate vs AlphaAssetSelector. NOTE: this layer is the suspected locus of CRITICAL #1 — PASS 7 PINS the ranking math (narrowing the suspect to timing/join), but the live-aggregate root-cause stays Claude's job. Review when done.

PASS 8: VIOLET_PART_SPEC_OA_TODO_PASS8.md (issued 2026-06-17) — V6 bible CONSUMERS: posture effects engine (5-state entry-gate/flatten/cap), MARAS fingerprint consumer (role TBD from code — no invented modulation), regime read-model, cadence shadow-actuation telemetry, posture/regime parity gate. VIOLET CONSUMES posture/MARAS (MHS owns them); ACB/vol/SL-guard already covered elsewhere. Review when done.

PASS 9: VIOLET_PART_SPEC_OA_TODO_PASS9.md (issued 2026-06-17) — economics/observability completeness: DDL-first exactly-one-row economics sink (PINK duplicate-emission fix), provenance + capital=anchor+Σ-trade-realized, factor divergence v2, soak-readiness aggregator, and the soak-readiness @gate (currently NO_GO by design — gated on CRITICAL #1). Review when done.

OA BACKLOG COMPLETE (PASS 19 issued). PASS 19 carve out the full parallelizable DARK surface of the V0→V6 ladder (~48 independent units, contracts_v3-keyed). What remains is NOT parallelizable and is Claude's: the parity root-cause (CRITICAL #1), review of all passes, integration, E2E, VBT re-cert, the real-key V4 boundary smoke, and live execution.

Action: review each pass for correctness, BLUE-algo compliance, V-TYPES, no-shared-edits, real (non-vacuous) tests. Owner: Claude (me), later.


🟢 #3 — Integration / "sprint" consolidation (my later work)

Once the passes are reviewed, I will: (a) review ALL passes together, (b) integrate them (resolve interfaces, dedupe), (c) test them together, (d) plug into the operational system, (e) E2E test.

Nomenclature note (raised by operator): a "pass" here = a batch of self-contained tasks delegated to one agent — smaller than an Agile sprint (a time-boxed iteration, typ. 1-4 weeks, team-scoped, ending in a shippable increment, with planning/review/retro ceremonies). In this project's existing usage, "Sprint N" already maps to a V-stage bundle (Sprint 1 = V0+V1, Sprint 2 = V2, Sprint 3 = V3) — i.e. an epic/milestone. So the cleanest mapping: V-stage = sprint/epic; "pass" = a sub-sprint work-package / task-bundle within it. Renaming passes to "sprints" would over-claim scope; keep "pass"/"work-package", or call each an "increment". Decide at integration time.


Gating rule

The DARK soak (operator-held) and V4 are NOT meaningfully runnable until CRITICAL #1 is root-caused — a 1.5% pick-match means the shadow is not yet tracking BLUE's decisions.