docs(blue-ops): verified corrections from live incident 136f5417 — cluster=dolphin, SAFETY key=latest dict (posture inside, per-bar clobber => write-loop), capital key=latest_nautilus; Tier-1 as documented was a no-op

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codex
2026-07-14 16:55:18 +02:00
parent 186bce8984
commit 1b6d280d26

View File

@@ -0,0 +1,446 @@
# BLUE Capital & In-Flight Trade — Updating and Reset Procedure
**Status:** Verified operational 2026-07-08 · **Scope:** BLUE live mainnet (`nautilus_event_trader.py` PID 1735549+) · **Not for:** PINK, VIOLET, or DITAv2 (each has its own control surfaces)
---
## ⚠ VERIFIED CORRECTIONS 2026-07-14 (Fable, live incident 136f5417) — READ BEFORE ANY TIER
The procedures below contain THREE errors that made Tier-1/2 inoperable during a live
stuck-trade incident. Verified against the running cluster:
1. **Cluster name is `dolphin`, NOT `dev`.** The container XML still says `dev`
(Appendix A is stale — the env override IS in effect on the live member; its log
banner prints `[dolphin]`). A bare `HazelcastClient()` (default cluster "dev")
fails auth with `IllegalStateError: Unable to connect to any cluster` — this is
why "hotkeys are dead". Every connect must be:
```python
hazelcast.HazelcastClient(cluster_name="dolphin",
cluster_members=["127.0.0.1:5701"],
cluster_connect_timeout=30)
```
2. **`DOLPHIN_SAFETY` has ONE key, `latest`** — a dict
`{"posture": "APEX", "Rm": ..., "timestamp": ..., "breakdown": {...}}`.
The Tier-1 recipe below (`safety.put("posture", "HIBERNATE")`) writes a key the
engine NEVER reads — it was always a no-op. Correct Tier-1 write:
```python
m = client.get_map("DOLPHIN_SAFETY").blocking()
snap = m.get("latest"); snap = json.loads(snap) if isinstance(snap, str) else (snap or {})
snap["posture"] = "HIBERNATE"
m.put("latest", snap) # match the stored type: dict in, dict out
```
**AND**: the Rm posture system REWRITES `latest` every bar (~12 s) — a single
external write races the clobber. Loop the write every 1-2 s until the exit is
observed in the trader log (`EXIT: ... HIBERNATE_HALT`), then stop and restore.
3. **Capital snapshot key is `latest_nautilus` (JSON string), not `latest`** in
`DOLPHIN_STATE_BLUE` (boot log agrees: "Capital restored from HZ latest_nautilus").
`latest` is absent. `DOLPHIN_CONTROL_PLANE` keys verified as documented.
Also verified 2026-07-14: V7 retract enqueue is EDGE-TRIGGERED (one command per
episode, `nautilus_event_trader` V7 block) — a lost command means V7 re-decides
forever without re-enqueueing. Do not expect V7 to self-recover a stuck retract;
the native lanes (TP / catastrophic floor / MAX_HOLD) run in-process and remain
the real safety net (136f5417 exited via MAX_HOLD at 125 bars, −0.01%).
---
## 0. Principles (read first)
1. **BLUE's ledger of record is E-anchored** — the BingX wallet balance is truth. Capital fields in HZ/CH are derived mirrors. Interventions update the mirrors; only a restart with `reset_and_seed(live_capital)` re-anchors K=E.
2. **HZ is write-safe** — Hazelcast IMap writes from the terminal are atomic. A bad value is fixable by overwriting. No harm to the running Python process.
3. **Supervisor autorestart is your safety net** — if a restart command kills the process, `autorestart=true` brings it back. Only `supervisorctl stop` holds it down.
4. **Retract has chain-token complexity** — the commit-chain idempotency guard rejects untokenised commands. When retract fails, use HIBERNATE posture or restart instead.
---
## 1. Architecture (who owns what)
### Live process tree
```
PID 1735549 BLUE /mnt/dolphinng5_predict/prod/nautilus_event_trader.py
↑ supervised by dolphin-supervisord under program:nautilus_trader
→ reads DOLPHIN_FEATURES.latest_eigen_scan (HZ)
→ writes DOLPHIN_SAFETY, DOLPHIN_PNL_BLUE, DOLPHIN_STATE_BLUE (HZ)
→ writes dolphin_uv. (ClickHouse)
→ exits 86 when scan path stalls (watchdog; autorestart catches it)
PID 4025061 UV PRIME shell-launched, NOT supervised
→ separate concern (VIOLET/UV control surface, not this doc)
```
### Hazelcast map schema (BLUE-relevant only)
| Map | Key | Value shape | Written by | Read by |
|---|---|---|---|---|
| `DOLPHIN_SAFETY` | `posture` | `"APEX"`/`"STALKER"`/`"TURTLE"`/`"HIBERNATE"` | Risk posture system | Engine `process_bar` |
| `DOLPHIN_CONTROL_PLANE` | `blue_runtime_commands` | `[{action, trade_id, fraction, ...}]` JSON array | Operators, agents | Engine `_drain_runtime_commands` |
| `DOLPHIN_STATE_BLUE` | `latest` | `{capital, posture, trades_executed, ...}` | Engine per-bar | Restore on restart |
| `DOLPHIN_FEATURES` | `latest_eigen_scan` | NG7 scan dict | DolphinNG6 | Engine step_bar |
| `DOLPHIN_PNL_BLUE` | `latest` | `{pnl, capital, ...}` | Engine per-bar | Observability |
### Supervisor control
```bash
supervisorctl -c /mnt/dolphinng5_predict/prod/supervisor/dolphin-supervisord.conf
status # all programs
stop nautilus_trader # graceful stop (SIGTERM)
start nautilus_trader # start
restart nautilus_trader # stop + start
```
`autorestart=true` means unexpected process death re-launches it. Only `stop` holds it down.
---
## 2. Tier-1: Force-exit via HIBERNATE posture (most reliable)
Write HIBERNATE to `DOLPHIN_SAFETY`. On the next bar, `_manage_position` catches the posture transition → `_execute_exit("HIBERNATE_HALT")` — full exit machinery fires (TP/SL logic, slippage, CH journal). No chain-token needed.
### Procedure
```python
# From any Python with HZ access:
import hazelcast
client = hazelcast.HazelcastClient()
safety = client.get_map("DOLPHIN_SAFETY").blocking()
safety.put("posture", "HIBERNATE")
```
Or via the running BLUE process's own HZ bridge — write to the map directly:
```bash
python3 -c "
import hazelcast
client = hazelcast.HazelcastClient()
safety = client.get_map('DOLPHIN_SAFETY').blocking()
safety.put('posture', 'HIBERNATE')
print('Written HIBERNATE. Engine will exit position on next bar.')
"
```
**Verify:**
```bash
python3 -c "
import hazelcast
client = hazelcast.HazelcastClient()
safety = client.get_map('DOLPHIN_SAFETY').blocking()
print(safety.get('posture'))
"
# → 'HIBERNATE'
```
### Recovery after exit
Restore posture to `"APEX"` to allow new entries on the next scan:
```bash
python3 -c "
import hazelcast
client = hazelcast.HazelcastClient()
safety = client.get_map('DOLPHIN_SAFETY').blocking()
safety.put('posture', 'APEX')
"
```
The engine checks posture on EVERY bar (via `_drain_runtime_commands` path), so the
live process picks it up without restart.
### When HIBERNATE does NOT work (rare)
The posture transition is checked inside the bar loop. If the bar loop is stuck (stale
scan, HZ outage, process hung), HIBERNATE never fires. Escalate to Tier-2 (retract) or
Tier-3 (restart).
---
## 3. Tier-2: Force-exit via RETRACT control-plane command
The control plane accepts `blue_runtime_commands` as a JSON array. Each command has
`action`, `trade_id`, `fraction`, and chain-token fields for idempotency.
### 3.1 Full retract (close entire position)
Requires the trade's commit-chain metadata (chain_root_trade_id, chain_head_leg_id,
chain_token). These are available from:
- `DOLPHIN_STATE_BLUE["latest"]` → `pending` → `pending.meta.chain_*`
- The CH `dolphin_decisions` or `dolphin_uv.exec_journal` row for the entry
- `dolphin_uv.position_state` where `chain_seq` is the latest leg
```python
import hazelcast, json
client = hazelcast.HazelcastClient()
ctl = client.get_map("DOLPHIN_CONTROL_PLANE").blocking()
# Read trade-id and chain-token from CH or from DOLPHIN_STATE_BLUE
state = client.get_map("DOLPHIN_STATE_BLUE").blocking().get("latest")
pending = (state or {}).get("pending", {})
# pending contains: asset, side, entry_price, quantity, notional,
# chain_root_trade_id, chain_head_leg_id, chain_token, entry_bar
cmd = {
"action": "RETRACT",
"trade_id": pending.get("trade_id", ""),
"fraction": 1.0, # 1.0 = full close
"chain_root_trade_id": pending.get("chain_root_trade_id", ""),
"chain_head_leg_id": pending.get("chain_head_leg_id", ""),
"chain_token": pending.get("chain_token", ""),
"chain_seq": pending.get("chain_seq", 1),
"reason": "OPERATOR_FORCE_CLOSE",
"command_id": f"retract-{int(time.time())}",
}
# Queue it (append to existing commands, don't overwrite)
existing = ctl.get("blue_runtime_commands")
queue = json.loads(existing) if existing else []
queue.append(cmd)
ctl.put("blue_runtime_commands", json.dumps(queue))
```
### 3.2 When retract fails (chain-token mismatch)
The `_apply_internal_retract` function verifies `chain_root_trade_id`, `chain_head_leg_id`,
`chain_token` against expected values from the in-memory `_chain_state_for_pending`. If
they don't match → `"CHAIN_MISMATCH"` / `"NO_CHAIN_LINK"`. This is the "control paths do
not always work" warning.
**Known failure modes:**
1. **Stale chain state** — the in-memory `_chain_state_for_pending` may differ from what's
in CH by one leg count or token. Reading chain data from CH state while the engine's
in-memory state is ahead/behind → mismatch.
2. **Partial retract already in flight** — `_processed_retract_set` deduplicates by
`command_id`. Re-issuing the same command_id is a no-op.
3. **No position** — `position is None` → `"NO_POSITION"` (trade already closed).
**Workaround when retract fails:** fall back to Tier-1 (HIBERNATE) or Tier-3 (restart).
A restart zeroes the in-memory state and re-hydrates from CH, clearing any chain-token
desync.
### 3.3 Partial retract (reduce position by fraction)
Same as full but `fraction: 0.5` for 50% reduction. The remainder persists through the
canonical OPEN write gate.
---
## 4. Tier-3: Capital update via SET_CAPITAL control command
Adjust BLUE's tracked capital without restarting. The engine processes this on the
next bar via `_apply_internal_capital_update`.
```python
cmd = {
"action": "SET_CAPITAL",
"capital": 85410.0, # new capital value
"reason": "MANUAL_ADJUST_LOSS_RECORDING",
"command_id": f"cap-{int(time.time())}",
}
```
Same queue procedure as retract — append to `blue_runtime_commands`.
**Caveat:** The capital update adjusts the in-memory `eng.capital`. It does NOT anchor
K to the exchange (E). Only `reset_and_seed(live_capital)` on restart does K=E. So a
SET_CAPITAL value will drift from the real wallet if subsequent fills/pnls apply on top
of the manually-set base.
---
## 5. Tier-4: Hard restart (nuclear, use when control paths fail)
When HIBERNATE doesn't fire and retract chain-tokens are desynced.
### 5.1 Stop BLUE
```bash
supervisorctl -c /mnt/dolphinng5_predict/prod/supervisor/dolphin-supervisord.conf stop nautilus_trader
```
Wait for the process to die (verify: `ps aux | grep nautilus_event_trader | grep -v grep`
returns empty, or `supervisorctl status` shows `STOPPED`).
### 5.2 Update capital (optional)
If the goal is to also adjust capital, write the desired value to HZ before restart.
On restart, the restore path reads `DOLPHIN_STATE_BLUE["latest"]` and picks up the
persisted snapshot. If K=E anchoring is needed (capital drift from the reset_and_seed
path), write the live wallet balance directly:
```python
import hazelcast, json
client = hazelcast.HazelcastClient()
state = client.get_map("DOLPHIN_STATE_BLUE").blocking()
snap = state.get("latest") or {}
snap["capital"] = 85410.0 # or whatever the corrected value is
state.put("latest", snap)
```
Note: this only works if the engine's `_seed_posture_for_restored_position` path picks
up the persisted snapshot faithfully. The more reliable approach is to let the engine
boot normally and apply a `SET_CAPITAL` command after restart (Tier-3), or to modify
the persisted snapshot file.
### 5.3 Start BLUE
```bash
supervisorctl -c /mnt/dolphinng5_predict/prod/supervisor/dolphin-supervisord.conf start nautilus_trader
```
### 5.4 Verify
```bash
supervisorctl -c /mnt/dolphinng5_predict/prod/supervisor/dolphin-supervisord.conf status nautilus_trader
# → RUNNING, pid <new_pid>
python3 -c "
import hazelcast
client = hazelcast.HazelcastClient()
safety = client.get_map('DOLPHIN_SAFETY').blocking()
state = client.get_map('DOLPHIN_STATE_BLUE').blocking()
print('Posture:', safety.get('posture'))
s = state.get('latest') or {}
print('Capital:', s.get('capital'))
print('Trades:', s.get('trades_executed'))
"
```
### 5.5 What restart does and doesn't do
**Does:**
- Calls `reset_and_seed(live_capital)` → zeros stale K accumulators, sets K=E=live BingX balance → `capital_frozen=False` (per PINK Startup fix)
- Restores position state from CH → `_seed_posture_for_restored_position()` hydrates trade_id, asset, entry_price, etc.
- Reconnects HZ listeners (DOLPHIN_FEATURES, DOLPHIN_SAFETY)
**Does NOT:**
- Auto-exit the position — if a position is open, the engine re-hydrates it and continues managing it (exits fire naturally via TP/max_hold/or by writing HIBERNATE after restart)
- Lose CH data — all decisions, positions, and pnl are in ClickHouse (durable)
So after a restart with an in-flight position, the position is still there. To force-exit
it post-restart, write HIBERNATE (Tier-1) on the first bar.
---
## 6. Capital basis doctrine (BLUE vs E-anchored)
BLUE's `eng.capital` is the in-memory capital tracker, updated per-bar with PnL and
drawn-down by entry notional. It is NOT the exchange wallet balance. The two converge
only at restart when `reset_and_seed(live_capital)` sets K=E.
**When to adjust capital manually:**
| Scenario | Method | Priority |
|---|---|---|
| Realized loss recorded accurately by engine | No action needed — engine tracks it | — |
| Engine capital drifted from wallet (e.g. fee sign bug, orphan trade) | SET_CAPITAL via control plane | Tier-3 |
| Restart already planned | Update DOLPHIN_STATE_BLUE snapshot before start | Tier-4 variant |
| Need K=E hard anchor | Stop → verify wallet balance → start (reset_and_seed sets K=E) | Tier-4 |
---
## 7. Decision tree
```
Position stuck / wrong capital / need to close?
│
├─ HIBERNATE has been written but position persists?
│ → Bar loop may be stuck. Check HZ liveness (echo "get DOLPHIN_HEARTBEAT").
│ → If HZ is live: restart.
│ → If HZ is down: restart.
│
├─ Retract command returned CHAIN_MISMATCH or NO_CHAIN_LINK?
│ → Chain-token desync. Fall back to HIBERNATE (Tier-1) or restart (Tier-4).
│
├─ Capital needs adjusting mid-flight?
│ → Use SET_CAPITAL via control plane (Tier-3).
│ → If restart is already needed: modify DOLPHIN_STATE_BLUE first (Tier-4).
│
├─ Pure force-close, no capital change needed?
│ → HIBERNATE (Tier-1) is the safest. One write, one bar, done.
│
└─ Everything failing, trade must die NOW?
→ Stop BLUE (Tier-4) → trade remains open on BingX → write HIBERNATE → start BLUE.
The bar after restart will exit via HIBERNATE_HALT.
```
---
## 8. Verification & monitoring
### Read live position state from HZ
```python
import hazelcast
client = hazelcast.HazelcastClient()
state = client.get_map("DOLPHIN_STATE_BLUE").blocking().get("latest")
pending = (state or {}).get("pending", {})
print("Asset:", pending.get("asset"))
print("Side:", pending.get("side"))
print("Entry price:", pending.get("entry_price"))
print("Quantity:", pending.get("quantity"))
print("Notional:", pending.get("notional"))
```
### Read live position state from CH
```sql
SELECT * FROM dolphin_uv.position_state
ORDER BY ts DESC LIMIT 1
FORMAT PrettyCompact;
```
### Check BLUE heartbeat
```python
import hazelcast
client = hazelcast.HazelcastClient()
hb = client.get_map("DOLPHIN_HEARTBEAT").blocking().get("latest")
print(hb) # → {ts, bar_idx, posture, capital, ...}
```
---
## 9. Related documents
| Document | Purpose |
|---|---|
| `SYSTEM_BIBLE.md` | Full BLUE system architecture, exit management (§7), posture logic (§11), bar loop (§13) |
| `DITA_V2_OPERATOR_PLAYBOOK.md` | DITAv2 kernel control surface (separate process, not BLUE) |
| `MULTI_EXIT_RETRACTION_SPEC_v1_2026-05-12.md` | Commit-chain retraction protocol deep-dive |
| `AGENT_TERMINAL_DIRECT_INTERVENTION_PROCEDURES.md` | Zellij injection mechanics (inter-agent, not system) |
| `CRITICAL_BLUE_BUGFIXED_DISCARD_RETURNED_TERMINAL_ON_FORCED_EXIT.md` | Bugfix: forced-exit terminal discard (historical) |
| `BLUE_INCIDENT_TODO__EFSM_LONG_RETRACT_FLIPOVER_20260623.md` | Incident: EFSM long retract flipover (historical) |
| supervisor config | `/mnt/dolphinng5_predict/prod/supervisor/dolphin-supervisord.conf` |
| Launcher | `/mnt/dolphinng5_predict/prod/nautilus_event_trader.py` |
| Pink startup/reset | `pink_direct.py::reset_and_seed(live_capital)` — K=E anchoring |
---
---
## Appendix A — Hazelcast Cluster Name Mismatch (2026-07-08)
**Finding during 16:10 load-262 incident:** `nautilus_event_trader.py:110` sets
`HZ_CLUSTER = "dolphin"`, but the Hazelcast Docker container's XML config
(`/opt/hazelcast/config/hazelcast-docker.xml`) has `<cluster-name>dev</cluster-name>`.
The container env var `HZ_CLUSTERNAME=dolphin` does NOT override the XML value in
Hazelcast 5.x — XML takes precedence.
**Consequence:** Under normal CPU, the Python client's retry loop reconnects fast enough
that the mismatch is invisible (connecting to cluster `dev` works, connecting to `dolphin`
hangs). But under CPU starvation (load > 260 on 12 cores at 16:10), the retry threads
are preempted long enough that HZ marks the client dead → all map I/O hangs → M2 heartbeat
sensor goes red → control plane interventions fail.
**Fix options (operator choose):**
| Fix | Change | Risk | Permanence |
|---|---|---|---|
| A — Container XML | `docker exec` edit XML → `dolphin`, restart container | Brief HZ outage | Survives container restart, not image rebuild |
| B — BLUE source | Edit `nautilus_event_trader.py:110`: `HZ_CLUSTER = "dev"`, restart BLUE | One scan cycle | Survives BLUE restart and container rebuild |
Trade self-healed at ~16:20 when load subsided and HZ client reconnected. No intervention
needed. The mismatch is a latent SPOF under load — either fix closes it.
---
*Maintained by cmd-PASS8 (PASS1.2) per T2P5-capstone session, 2026-07-08. Update when BLUE architecture changes or new intervention patterns are discovered.*