Files
siloqy/prod/docs/VIOLET_TODO_CRITICAL.md
Codex f8826f613d VIOLET: record PASS4-done + Task 17/18 caveats in the review queue (PR docs)
PASS 4 completed on worktree agent/oa-violet4 (/mnt/vp-oa4). Recorded in the review queue the two
caveats the OA agent reported — also flagged in-code as prominent banner comments + empty marker
commits on that branch:
  - Task 17 exec_intent: explicit reference_price required (ShadowDecision has notional, no price).
  - Task 18 alpha_data_feed: parse/feed-port only, no live NT/Binance in unit code.
For Claude's later PASS-4 review (verify the caveats + non-vacuous tests).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 12:25:38 +02:00

129 lines
7.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# VIOLET — CRITICAL TODO / review queue (prominent)
Date: 2026-06-17. Single place for the must-not-forget VIOLET items. Review-later, not now.
---
## 🔴 CRITICAL #1 — VIOLET↔BLUE parity is VERY DISAPPOINTING (review + root-cause BEFORE any soak/V4)
Report: `prod/VIOLET_dev/reports/violet_parity_20260616_220412.md` (+ `.json`).
Window 2026-06-14 20:15 → 2026-06-15 21:00. Produced by PASS-2 Task 4 (`parity_report.py`).
**Headline:**
- Pick-match rate: **0.015 (1.5%)** — VIOLET rows 2853, exact pick matches only 43.
- Same-asset rate: **0.136 (13.6%)**; no-pick: **2465 / 2853 (86%)**.
- BLUE rows 25941 (scan_eval 25278 + trade_events 663).
**KEY NUANCE (where the review should START):** on the 43 rows that DID align, **sizing is
near-identical** — leverage abs error mean 0.016 / **median 0.0** / max 0.135. So the V3.4
sizing math (boost/beta/mc_scale/ob/esof/compose) is NOT the problem; the divergence is in
**ASSET SELECTION / TIMING / the comparison's ALIGNMENT method**. Candidate causes to
investigate (do not assume — measure):
1. **Apples-to-oranges population.** BLUE `scan_eval` (25278) is likely per-scan-per-asset
*evaluations*, while VIOLET rows are *actuated* decisions — the report may be comparing
different things. Verify the alignment/join semantics in `parity_report.py` first.
2. **Selection divergence.** VIOLET's `VioletAssetSelector` (IRP) vs BLUE's live selection over
the same scan stream — are they fed the same universe/lookback at the same scan index? The
sizing-gap samples (TRX/ATOM/LTC/XLM SHORT with large notional_rel_err) suggest VIOLET fires
on assets BLUE sized very differently or didn't pick.
3. **Cadence/actuation.** VIOLET actuates at Q=scan; if its scan alignment or dedupe differs,
picks land at different scans → counted as no-pick.
4. **The known structural items** (OB single-shot before V3.4d; mc_scale; live-factor sourcing)
— re-run parity AFTER V3.4d's persistent-OB launcher + the bit-identity fixes to see if the
number moves.
**Action:** full review of `parity_report.py` + a root-cause pass; fix the alignment OR the
selection divergence; re-run. This gates a meaningful DARK soak — a 1.5% pick-match makes the
soak uninterpretable. **Owner: Claude (me), later.**
---
## 🟡 #2 — Review the OA-delegated PASS work (NOT yet reviewed)
The parallel agent reports these DONE; none reviewed for correctness / BLUE-compliance yet.
**PASS 1** (`VIOLET_PART_SPEC_OA_TODO.md`):
- `53bdd90` sizing parity-pin tests
- `12b768b` venue OB provider seam (`venue_ob_provider.py`)
- `bae9284` base-fraction sizing study (+ archived report)
**PASS 2** (`VIOLET_PART_SPEC_OA_TODO_PASS2.md`):
- `parity_report.py` (Task 4 — the report above) + test
- `tradeability.py` (Task 6) + test
- `test_violet_v3_decision_latency_gate.py` (Task 5) → report `violet_v3_decision_latency_2026...`
- `test_violet_replay_determinism_gate.py` (Task 7) → report `violet_replay_determinism_2026...`
**PASS 3**: `VIOLET_PART_SPEC_OA_TODO_PASS3.md` (issued 2026-06-17) — venue feed / mechanical
exits / SL floor / event-restore / slippage / cadence; review when done.
**PASS 4**: `VIOLET_PART_SPEC_OA_TODO_PASS4.md` (issued 2026-06-17) — vol gate / V7 exit wrapper /
economics ledger / exec-intent / alpha data feed / time exits. **DONE on worktree
`agent/oa-violet4` (/mnt/vp-oa4)** — commits 6996a6a/7c8e524/bbebe58/3920665/f328d3a/b39d8a4.
Review when done. **Caveats to verify in review (flagged in-code as prominent banners + empty
marker commits 57ca851/8624dee/2aad0aa/539ea47):**
- **Task 17 (exec_intent):** `to_exec_intent` needs an EXPLICIT `reference_price` arg —
`ShadowDecision` has notional/exposure but NO price field, so qty can't be derived from a
decision alone. Integration owner threads the live price (PASS-3 `VenueTick`) at wiring time.
- **Task 18 (alpha_data_feed):** intentionally parse/feed-port ONLY — no live NT node / Binance
connection in unit code (NT owns its loop; dummy keys; separate feed process). Live node = the
integration owner's job. Confirm tests aren't vacuous (parse-fixture, not a live connection).
**PASS 5**: `VIOLET_PART_SPEC_OA_TODO_PASS5.md` (issued 2026-06-17) — MOCK-BINGX execution stack
(exec contracts + QuirkProfile seams / order FSM / mock venue / fill reducer / DARK E2E gate).
NOTE: BingX "quirks" are SEAM-ONLY here (default OFF) — a later quirk-injection pass + a mandatory
real-key boundary smoke are required before V4-live; review when done.
**PASS 6**: `VIOLET_PART_SPEC_OA_TODO_PASS6.md` (issued 2026-06-17) — execution INTERNALS
(fill-pump/ownership filter, reconcile/zero-wb guard, TTL-requote, orphan handling) + the
QUIRK-INJECTION gate that flips PASS-5's QuirkProfile flags ON and proves each handler neutralizes
the quirk. Mirrors PINK's production fixes (pink_direct.py). Real-key smoke still MANDATORY before
V4-live; review when done.
**PASS 7**: `VIOLET_PART_SPEC_OA_TODO_PASS7.md` (issued 2026-06-17) — V5 selection (faithful
ARS/IRP ranking + OB Sub-1), multi-asset slot manager, capital allocation, multi-asset flow, and a
ranking bit-identity gate vs AlphaAssetSelector. NOTE: this layer is the suspected locus of CRITICAL
#1 — PASS 7 PINS the ranking math (narrowing the suspect to timing/join), but the live-aggregate
root-cause stays Claude's job. Review when done.
**PASS 8**: `VIOLET_PART_SPEC_OA_TODO_PASS8.md` (issued 2026-06-17) — V6 bible CONSUMERS: posture
effects engine (5-state entry-gate/flatten/cap), MARAS fingerprint consumer (role TBD from code —
no invented modulation), regime read-model, cadence shadow-actuation telemetry, posture/regime
parity gate. VIOLET CONSUMES posture/MARAS (MHS owns them); ACB/vol/SL-guard already covered
elsewhere. Review when done.
**PASS 9**: `VIOLET_PART_SPEC_OA_TODO_PASS9.md` (issued 2026-06-17) — economics/observability
completeness: DDL-first exactly-one-row economics sink (PINK duplicate-emission fix), provenance +
capital=anchor+Σ-trade-realized, factor divergence v2, soak-readiness aggregator, and the
soak-readiness @gate (currently NO_GO by design — gated on CRITICAL #1). Review when done.
> **OA BACKLOG COMPLETE (PASS 19 issued).** PASS 19 carve out the full parallelizable DARK
> surface of the V0→V6 ladder (~48 independent units, contracts_v3-keyed). What remains is NOT
> parallelizable and is Claude's: the parity root-cause (CRITICAL #1), review of all passes,
> integration, E2E, VBT re-cert, the real-key V4 boundary smoke, and live execution.
**Action:** review each pass for correctness, BLUE-algo compliance, V-TYPES, no-shared-edits,
real (non-vacuous) tests. **Owner: Claude (me), later.**
---
## 🟢 #3 — Integration / "sprint" consolidation (my later work)
Once the passes are reviewed, I will: **(a)** review ALL passes together, **(b)** integrate them
(resolve interfaces, dedupe), **(c)** test them together, **(d)** plug into the operational
system, **(e)** E2E test.
**Nomenclature note (raised by operator):** a "pass" here = a batch of self-contained tasks
delegated to one agent — smaller than an Agile **sprint** (a time-boxed iteration, typ. 1-4
weeks, team-scoped, ending in a shippable increment, with planning/review/retro ceremonies).
In this project's existing usage, **"Sprint N" already maps to a V-stage bundle** (Sprint 1 =
V0+V1, Sprint 2 = V2, Sprint 3 = V3) — i.e. an epic/milestone. So the cleanest mapping:
**V-stage = sprint/epic; "pass" = a sub-sprint work-package / task-bundle within it.** Renaming
passes to "sprints" would over-claim scope; keep "pass"/"work-package", or call each an
"increment". Decide at integration time.
---
## Gating rule
The DARK soak (operator-held) and V4 are NOT meaningfully runnable until CRITICAL #1 is
root-caused — a 1.5% pick-match means the shadow is not yet tracking BLUE's decisions.