From 1b6d280d261f76270bb6e810a1449e9c68084eac Mon Sep 17 00:00:00 2001 From: Codex Date: Tue, 14 Jul 2026 16:55:18 +0200 Subject: [PATCH] =?UTF-8?q?docs(blue-ops):=20verified=20corrections=20from?= =?UTF-8?q?=20live=20incident=20136f5417=20=E2=80=94=20cluster=3Ddolphin,?= =?UTF-8?q?=20SAFETY=20key=3Dlatest=20dict=20(posture=20inside,=20per-bar?= =?UTF-8?q?=20clobber=20=3D>=20write-loop),=20capital=20key=3Dlatest=5Fnau?= =?UTF-8?q?tilus;=20Tier-1=20as=20documented=20was=20a=20no-op?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 4.8 (1M context) --- ...IGHT_TRADE_UPDATING_AND_RESET_PROCEDURE.md | 446 ++++++++++++++++++ 1 file changed, 446 insertions(+) create mode 100644 prod/docs/BLUE_CAPITAL_AND_IN_FLIGHT_TRADE_UPDATING_AND_RESET_PROCEDURE.md diff --git a/prod/docs/BLUE_CAPITAL_AND_IN_FLIGHT_TRADE_UPDATING_AND_RESET_PROCEDURE.md b/prod/docs/BLUE_CAPITAL_AND_IN_FLIGHT_TRADE_UPDATING_AND_RESET_PROCEDURE.md new file mode 100644 index 0000000..e2f222a --- /dev/null +++ b/prod/docs/BLUE_CAPITAL_AND_IN_FLIGHT_TRADE_UPDATING_AND_RESET_PROCEDURE.md @@ -0,0 +1,446 @@ +# BLUE Capital & In-Flight Trade — Updating and Reset Procedure + +**Status:** Verified operational 2026-07-08 · **Scope:** BLUE live mainnet (`nautilus_event_trader.py` PID 1735549+) · **Not for:** PINK, VIOLET, or DITAv2 (each has its own control surfaces) + +--- + +## ⚠ VERIFIED CORRECTIONS 2026-07-14 (Fable, live incident 136f5417) — READ BEFORE ANY TIER + +The procedures below contain THREE errors that made Tier-1/2 inoperable during a live +stuck-trade incident. Verified against the running cluster: + +1. **Cluster name is `dolphin`, NOT `dev`.** The container XML still says `dev` + (Appendix A is stale — the env override IS in effect on the live member; its log + banner prints `[dolphin]`). A bare `HazelcastClient()` (default cluster "dev") + fails auth with `IllegalStateError: Unable to connect to any cluster` — this is + why "hotkeys are dead". Every connect must be: + ```python + hazelcast.HazelcastClient(cluster_name="dolphin", + cluster_members=["127.0.0.1:5701"], + cluster_connect_timeout=30) + ``` +2. **`DOLPHIN_SAFETY` has ONE key, `latest`** — a dict + `{"posture": "APEX", "Rm": ..., "timestamp": ..., "breakdown": {...}}`. + The Tier-1 recipe below (`safety.put("posture", "HIBERNATE")`) writes a key the + engine NEVER reads — it was always a no-op. Correct Tier-1 write: + ```python + m = client.get_map("DOLPHIN_SAFETY").blocking() + snap = m.get("latest"); snap = json.loads(snap) if isinstance(snap, str) else (snap or {}) + snap["posture"] = "HIBERNATE" + m.put("latest", snap) # match the stored type: dict in, dict out + ``` + **AND**: the Rm posture system REWRITES `latest` every bar (~12 s) — a single + external write races the clobber. Loop the write every 1-2 s until the exit is + observed in the trader log (`EXIT: ... HIBERNATE_HALT`), then stop and restore. +3. **Capital snapshot key is `latest_nautilus` (JSON string), not `latest`** in + `DOLPHIN_STATE_BLUE` (boot log agrees: "Capital restored from HZ latest_nautilus"). + `latest` is absent. `DOLPHIN_CONTROL_PLANE` keys verified as documented. + +Also verified 2026-07-14: V7 retract enqueue is EDGE-TRIGGERED (one command per +episode, `nautilus_event_trader` V7 block) — a lost command means V7 re-decides +forever without re-enqueueing. Do not expect V7 to self-recover a stuck retract; +the native lanes (TP / catastrophic floor / MAX_HOLD) run in-process and remain +the real safety net (136f5417 exited via MAX_HOLD at 125 bars, −0.01%). + +--- + +## 0. Principles (read first) + +1. **BLUE's ledger of record is E-anchored** — the BingX wallet balance is truth. Capital fields in HZ/CH are derived mirrors. Interventions update the mirrors; only a restart with `reset_and_seed(live_capital)` re-anchors K=E. +2. **HZ is write-safe** — Hazelcast IMap writes from the terminal are atomic. A bad value is fixable by overwriting. No harm to the running Python process. +3. **Supervisor autorestart is your safety net** — if a restart command kills the process, `autorestart=true` brings it back. Only `supervisorctl stop` holds it down. +4. **Retract has chain-token complexity** — the commit-chain idempotency guard rejects untokenised commands. When retract fails, use HIBERNATE posture or restart instead. + +--- + +## 1. Architecture (who owns what) + +### Live process tree + +``` +PID 1735549 BLUE /mnt/dolphinng5_predict/prod/nautilus_event_trader.py + ↑ supervised by dolphin-supervisord under program:nautilus_trader + → reads DOLPHIN_FEATURES.latest_eigen_scan (HZ) + → writes DOLPHIN_SAFETY, DOLPHIN_PNL_BLUE, DOLPHIN_STATE_BLUE (HZ) + → writes dolphin_uv. (ClickHouse) + → exits 86 when scan path stalls (watchdog; autorestart catches it) + +PID 4025061 UV PRIME shell-launched, NOT supervised + → separate concern (VIOLET/UV control surface, not this doc) +``` + +### Hazelcast map schema (BLUE-relevant only) + +| Map | Key | Value shape | Written by | Read by | +|---|---|---|---|---| +| `DOLPHIN_SAFETY` | `posture` | `"APEX"`/`"STALKER"`/`"TURTLE"`/`"HIBERNATE"` | Risk posture system | Engine `process_bar` | +| `DOLPHIN_CONTROL_PLANE` | `blue_runtime_commands` | `[{action, trade_id, fraction, ...}]` JSON array | Operators, agents | Engine `_drain_runtime_commands` | +| `DOLPHIN_STATE_BLUE` | `latest` | `{capital, posture, trades_executed, ...}` | Engine per-bar | Restore on restart | +| `DOLPHIN_FEATURES` | `latest_eigen_scan` | NG7 scan dict | DolphinNG6 | Engine step_bar | +| `DOLPHIN_PNL_BLUE` | `latest` | `{pnl, capital, ...}` | Engine per-bar | Observability | + +### Supervisor control + +```bash +supervisorctl -c /mnt/dolphinng5_predict/prod/supervisor/dolphin-supervisord.conf + status # all programs + stop nautilus_trader # graceful stop (SIGTERM) + start nautilus_trader # start + restart nautilus_trader # stop + start +``` + +`autorestart=true` means unexpected process death re-launches it. Only `stop` holds it down. + +--- + +## 2. Tier-1: Force-exit via HIBERNATE posture (most reliable) + +Write HIBERNATE to `DOLPHIN_SAFETY`. On the next bar, `_manage_position` catches the posture transition → `_execute_exit("HIBERNATE_HALT")` — full exit machinery fires (TP/SL logic, slippage, CH journal). No chain-token needed. + +### Procedure + +```python +# From any Python with HZ access: +import hazelcast +client = hazelcast.HazelcastClient() +safety = client.get_map("DOLPHIN_SAFETY").blocking() +safety.put("posture", "HIBERNATE") +``` + +Or via the running BLUE process's own HZ bridge — write to the map directly: + +```bash +python3 -c " +import hazelcast +client = hazelcast.HazelcastClient() +safety = client.get_map('DOLPHIN_SAFETY').blocking() +safety.put('posture', 'HIBERNATE') +print('Written HIBERNATE. Engine will exit position on next bar.') +" +``` + +**Verify:** +```bash +python3 -c " +import hazelcast +client = hazelcast.HazelcastClient() +safety = client.get_map('DOLPHIN_SAFETY').blocking() +print(safety.get('posture')) +" +# → 'HIBERNATE' +``` + +### Recovery after exit + +Restore posture to `"APEX"` to allow new entries on the next scan: + +```bash +python3 -c " +import hazelcast +client = hazelcast.HazelcastClient() +safety = client.get_map('DOLPHIN_SAFETY').blocking() +safety.put('posture', 'APEX') +" +``` + +The engine checks posture on EVERY bar (via `_drain_runtime_commands` path), so the +live process picks it up without restart. + +### When HIBERNATE does NOT work (rare) + +The posture transition is checked inside the bar loop. If the bar loop is stuck (stale +scan, HZ outage, process hung), HIBERNATE never fires. Escalate to Tier-2 (retract) or +Tier-3 (restart). + +--- + +## 3. Tier-2: Force-exit via RETRACT control-plane command + +The control plane accepts `blue_runtime_commands` as a JSON array. Each command has +`action`, `trade_id`, `fraction`, and chain-token fields for idempotency. + +### 3.1 Full retract (close entire position) + +Requires the trade's commit-chain metadata (chain_root_trade_id, chain_head_leg_id, +chain_token). These are available from: + - `DOLPHIN_STATE_BLUE["latest"]` → `pending` → `pending.meta.chain_*` + - The CH `dolphin_decisions` or `dolphin_uv.exec_journal` row for the entry + - `dolphin_uv.position_state` where `chain_seq` is the latest leg + +```python +import hazelcast, json +client = hazelcast.HazelcastClient() +ctl = client.get_map("DOLPHIN_CONTROL_PLANE").blocking() + +# Read trade-id and chain-token from CH or from DOLPHIN_STATE_BLUE +state = client.get_map("DOLPHIN_STATE_BLUE").blocking().get("latest") +pending = (state or {}).get("pending", {}) +# pending contains: asset, side, entry_price, quantity, notional, +# chain_root_trade_id, chain_head_leg_id, chain_token, entry_bar + +cmd = { + "action": "RETRACT", + "trade_id": pending.get("trade_id", ""), + "fraction": 1.0, # 1.0 = full close + "chain_root_trade_id": pending.get("chain_root_trade_id", ""), + "chain_head_leg_id": pending.get("chain_head_leg_id", ""), + "chain_token": pending.get("chain_token", ""), + "chain_seq": pending.get("chain_seq", 1), + "reason": "OPERATOR_FORCE_CLOSE", + "command_id": f"retract-{int(time.time())}", +} + +# Queue it (append to existing commands, don't overwrite) +existing = ctl.get("blue_runtime_commands") +queue = json.loads(existing) if existing else [] +queue.append(cmd) +ctl.put("blue_runtime_commands", json.dumps(queue)) +``` + +### 3.2 When retract fails (chain-token mismatch) + +The `_apply_internal_retract` function verifies `chain_root_trade_id`, `chain_head_leg_id`, +`chain_token` against expected values from the in-memory `_chain_state_for_pending`. If +they don't match → `"CHAIN_MISMATCH"` / `"NO_CHAIN_LINK"`. This is the "control paths do +not always work" warning. + +**Known failure modes:** +1. **Stale chain state** — the in-memory `_chain_state_for_pending` may differ from what's + in CH by one leg count or token. Reading chain data from CH state while the engine's + in-memory state is ahead/behind → mismatch. +2. **Partial retract already in flight** — `_processed_retract_set` deduplicates by + `command_id`. Re-issuing the same command_id is a no-op. +3. **No position** — `position is None` → `"NO_POSITION"` (trade already closed). + +**Workaround when retract fails:** fall back to Tier-1 (HIBERNATE) or Tier-3 (restart). +A restart zeroes the in-memory state and re-hydrates from CH, clearing any chain-token +desync. + +### 3.3 Partial retract (reduce position by fraction) + +Same as full but `fraction: 0.5` for 50% reduction. The remainder persists through the +canonical OPEN write gate. + +--- + +## 4. Tier-3: Capital update via SET_CAPITAL control command + +Adjust BLUE's tracked capital without restarting. The engine processes this on the +next bar via `_apply_internal_capital_update`. + +```python +cmd = { + "action": "SET_CAPITAL", + "capital": 85410.0, # new capital value + "reason": "MANUAL_ADJUST_LOSS_RECORDING", + "command_id": f"cap-{int(time.time())}", +} +``` + +Same queue procedure as retract — append to `blue_runtime_commands`. + +**Caveat:** The capital update adjusts the in-memory `eng.capital`. It does NOT anchor +K to the exchange (E). Only `reset_and_seed(live_capital)` on restart does K=E. So a +SET_CAPITAL value will drift from the real wallet if subsequent fills/pnls apply on top +of the manually-set base. + +--- + +## 5. Tier-4: Hard restart (nuclear, use when control paths fail) + +When HIBERNATE doesn't fire and retract chain-tokens are desynced. + +### 5.1 Stop BLUE + +```bash +supervisorctl -c /mnt/dolphinng5_predict/prod/supervisor/dolphin-supervisord.conf stop nautilus_trader +``` + +Wait for the process to die (verify: `ps aux | grep nautilus_event_trader | grep -v grep` +returns empty, or `supervisorctl status` shows `STOPPED`). + +### 5.2 Update capital (optional) + +If the goal is to also adjust capital, write the desired value to HZ before restart. +On restart, the restore path reads `DOLPHIN_STATE_BLUE["latest"]` and picks up the +persisted snapshot. If K=E anchoring is needed (capital drift from the reset_and_seed +path), write the live wallet balance directly: + +```python +import hazelcast, json +client = hazelcast.HazelcastClient() +state = client.get_map("DOLPHIN_STATE_BLUE").blocking() +snap = state.get("latest") or {} +snap["capital"] = 85410.0 # or whatever the corrected value is +state.put("latest", snap) +``` + +Note: this only works if the engine's `_seed_posture_for_restored_position` path picks +up the persisted snapshot faithfully. The more reliable approach is to let the engine +boot normally and apply a `SET_CAPITAL` command after restart (Tier-3), or to modify +the persisted snapshot file. + +### 5.3 Start BLUE + +```bash +supervisorctl -c /mnt/dolphinng5_predict/prod/supervisor/dolphin-supervisord.conf start nautilus_trader +``` + +### 5.4 Verify + +```bash +supervisorctl -c /mnt/dolphinng5_predict/prod/supervisor/dolphin-supervisord.conf status nautilus_trader +# → RUNNING, pid + +python3 -c " +import hazelcast +client = hazelcast.HazelcastClient() +safety = client.get_map('DOLPHIN_SAFETY').blocking() +state = client.get_map('DOLPHIN_STATE_BLUE').blocking() +print('Posture:', safety.get('posture')) +s = state.get('latest') or {} +print('Capital:', s.get('capital')) +print('Trades:', s.get('trades_executed')) +" +``` + +### 5.5 What restart does and doesn't do + +**Does:** +- Calls `reset_and_seed(live_capital)` → zeros stale K accumulators, sets K=E=live BingX balance → `capital_frozen=False` (per PINK Startup fix) +- Restores position state from CH → `_seed_posture_for_restored_position()` hydrates trade_id, asset, entry_price, etc. +- Reconnects HZ listeners (DOLPHIN_FEATURES, DOLPHIN_SAFETY) + +**Does NOT:** +- Auto-exit the position — if a position is open, the engine re-hydrates it and continues managing it (exits fire naturally via TP/max_hold/or by writing HIBERNATE after restart) +- Lose CH data — all decisions, positions, and pnl are in ClickHouse (durable) + +So after a restart with an in-flight position, the position is still there. To force-exit +it post-restart, write HIBERNATE (Tier-1) on the first bar. + +--- + +## 6. Capital basis doctrine (BLUE vs E-anchored) + +BLUE's `eng.capital` is the in-memory capital tracker, updated per-bar with PnL and +drawn-down by entry notional. It is NOT the exchange wallet balance. The two converge +only at restart when `reset_and_seed(live_capital)` sets K=E. + +**When to adjust capital manually:** + +| Scenario | Method | Priority | +|---|---|---| +| Realized loss recorded accurately by engine | No action needed — engine tracks it | — | +| Engine capital drifted from wallet (e.g. fee sign bug, orphan trade) | SET_CAPITAL via control plane | Tier-3 | +| Restart already planned | Update DOLPHIN_STATE_BLUE snapshot before start | Tier-4 variant | +| Need K=E hard anchor | Stop → verify wallet balance → start (reset_and_seed sets K=E) | Tier-4 | + +--- + +## 7. Decision tree + +``` +Position stuck / wrong capital / need to close? +│ +├─ HIBERNATE has been written but position persists? +│ → Bar loop may be stuck. Check HZ liveness (echo "get DOLPHIN_HEARTBEAT"). +│ → If HZ is live: restart. +│ → If HZ is down: restart. +│ +├─ Retract command returned CHAIN_MISMATCH or NO_CHAIN_LINK? +│ → Chain-token desync. Fall back to HIBERNATE (Tier-1) or restart (Tier-4). +│ +├─ Capital needs adjusting mid-flight? +│ → Use SET_CAPITAL via control plane (Tier-3). +│ → If restart is already needed: modify DOLPHIN_STATE_BLUE first (Tier-4). +│ +├─ Pure force-close, no capital change needed? +│ → HIBERNATE (Tier-1) is the safest. One write, one bar, done. +│ +└─ Everything failing, trade must die NOW? + → Stop BLUE (Tier-4) → trade remains open on BingX → write HIBERNATE → start BLUE. + The bar after restart will exit via HIBERNATE_HALT. +``` + +--- + +## 8. Verification & monitoring + +### Read live position state from HZ + +```python +import hazelcast +client = hazelcast.HazelcastClient() +state = client.get_map("DOLPHIN_STATE_BLUE").blocking().get("latest") +pending = (state or {}).get("pending", {}) +print("Asset:", pending.get("asset")) +print("Side:", pending.get("side")) +print("Entry price:", pending.get("entry_price")) +print("Quantity:", pending.get("quantity")) +print("Notional:", pending.get("notional")) +``` + +### Read live position state from CH + +```sql +SELECT * FROM dolphin_uv.position_state +ORDER BY ts DESC LIMIT 1 +FORMAT PrettyCompact; +``` + +### Check BLUE heartbeat + +```python +import hazelcast +client = hazelcast.HazelcastClient() +hb = client.get_map("DOLPHIN_HEARTBEAT").blocking().get("latest") +print(hb) # → {ts, bar_idx, posture, capital, ...} +``` + +--- + +## 9. Related documents + +| Document | Purpose | +|---|---| +| `SYSTEM_BIBLE.md` | Full BLUE system architecture, exit management (§7), posture logic (§11), bar loop (§13) | +| `DITA_V2_OPERATOR_PLAYBOOK.md` | DITAv2 kernel control surface (separate process, not BLUE) | +| `MULTI_EXIT_RETRACTION_SPEC_v1_2026-05-12.md` | Commit-chain retraction protocol deep-dive | +| `AGENT_TERMINAL_DIRECT_INTERVENTION_PROCEDURES.md` | Zellij injection mechanics (inter-agent, not system) | +| `CRITICAL_BLUE_BUGFIXED_DISCARD_RETURNED_TERMINAL_ON_FORCED_EXIT.md` | Bugfix: forced-exit terminal discard (historical) | +| `BLUE_INCIDENT_TODO__EFSM_LONG_RETRACT_FLIPOVER_20260623.md` | Incident: EFSM long retract flipover (historical) | +| supervisor config | `/mnt/dolphinng5_predict/prod/supervisor/dolphin-supervisord.conf` | +| Launcher | `/mnt/dolphinng5_predict/prod/nautilus_event_trader.py` | +| Pink startup/reset | `pink_direct.py::reset_and_seed(live_capital)` — K=E anchoring | + +--- + +--- + +## Appendix A — Hazelcast Cluster Name Mismatch (2026-07-08) + +**Finding during 16:10 load-262 incident:** `nautilus_event_trader.py:110` sets +`HZ_CLUSTER = "dolphin"`, but the Hazelcast Docker container's XML config +(`/opt/hazelcast/config/hazelcast-docker.xml`) has `dev`. +The container env var `HZ_CLUSTERNAME=dolphin` does NOT override the XML value in +Hazelcast 5.x — XML takes precedence. + +**Consequence:** Under normal CPU, the Python client's retry loop reconnects fast enough +that the mismatch is invisible (connecting to cluster `dev` works, connecting to `dolphin` +hangs). But under CPU starvation (load > 260 on 12 cores at 16:10), the retry threads +are preempted long enough that HZ marks the client dead → all map I/O hangs → M2 heartbeat +sensor goes red → control plane interventions fail. + +**Fix options (operator choose):** + +| Fix | Change | Risk | Permanence | +|---|---|---|---| +| A — Container XML | `docker exec` edit XML → `dolphin`, restart container | Brief HZ outage | Survives container restart, not image rebuild | +| B — BLUE source | Edit `nautilus_event_trader.py:110`: `HZ_CLUSTER = "dev"`, restart BLUE | One scan cycle | Survives BLUE restart and container rebuild | + +Trade self-healed at ~16:20 when load subsided and HZ client reconnected. No intervention +needed. The mismatch is a latent SPOF under load — either fix closes it. + +--- + +*Maintained by cmd-PASS8 (PASS1.2) per T2P5-capstone session, 2026-07-08. Update when BLUE architecture changes or new intervention patterns are discovered.*