dita_v2: unknown != flat (VenuePostAckError) + lossless telemetry lane

BUG CLASS (doc: BUGCLASS_INDETERMINATE_OUTCOME_20260713.md): an operation with an external
side effect has THREE outcomes — NOT_ATTEMPTED, ATTEMPTED_REFUSED, ATTEMPTED_INDETERMINATE
— and rollback is sound only for the first two. Collapsing the third into 'failed' is what
orphaned 6 live SHORTs: a post-ack TypeError reached rust_backend's 'except Exception ->
synthetic REJECTED -> FSM rollback', which asserted 'no order exists' about an order that
was already filled. Telemetry never had a veto; it hijacked the failure channel.

FIXES (no new seams, no re-architecture):
- venue.py: VenuePostAckError — typed channel meaning THE EFFECT EXISTS. Carries receipt.
- bingx_venue submit/submit_async: point-of-no-return fence. Post-ack bookkeeping failures
  raise VenuePostAckError instead of a bare exception.
- rust_backend (BOTH submit paths): catch VenuePostAckError FIRST -> no synthetic REJECT,
  no rollback. Slot stays working; E-feed FULL_FILL / reconcile settles the truth.

LOSSLESS TELEMETRY (HJ: 'DITAv2 exists precisely because seams dropped 40% of inputs'):
drop-oldest is data loss and is GONE. Exec path appends O(1) to an unbounded queue and
returns. A SEPARATE spiller thread (which never touches the plane, so a wedged plane cannot
starve it) parks the backlog above HWM into a durable append-only spool; the publisher
replays the spool when the plane recovers.
Proven: wedged-forever plane + 200k records -> 0 dropped, 195903 durable on disk, 4096 in
memory, 7.4 us/call on the exec path. Lossless AND memory-bounded. Healthy plane: 2000/2000.

STILL BROKEN, flagged to codex: the pre-ack branch rolls back on TIMEOUT — but a timeout is
the definition of INDETERMINATE (the order may have filled). Same bug class, older, live.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Codex
2026-07-13 10:31:39 +02:00
parent b13b54e769
commit 3d75e4b59d
4 changed files with 405 additions and 40 deletions

View File

@@ -11,10 +11,16 @@ import asyncio
import concurrent.futures
import inspect
import itertools
import atexit
import json
import os
import re
import tempfile
import threading
import time
from collections import deque
from datetime import datetime, timezone
from pathlib import Path
from dataclasses import replace
from typing import Any, Iterable, List, Optional
@@ -38,6 +44,7 @@ from .contracts import (
from .utils import json_safe
from .utils import safe_float
from .venue import VenueAdapter
from .venue import VenuePostAckError
def _row_text(row: dict[str, Any], *keys: str, default: str = "") -> str:
@@ -283,42 +290,177 @@ class BingxVenueAdapter(VenueAdapter):
except Exception:
pass
# ── telemetry lane ───────────────────────────────────────────────────────
# Telemetry is NOT on the execution priority ladder (P0 kill/emergency,
# P1 SL/ADVSL, P2 exits, P3 ENTER, P4 maintenance) — it ranks below all of
# it. It therefore must never sit ON the exec path, never block it, and
# never raise into it.
# ── telemetry lane — NON-BLOCKING **AND** LOSSLESS ───────────────────────
# Two standing principles, both non-negotiable, and they do NOT conflict:
#
# The exec path does exactly one thing: append a tuple of primitives to a
# bounded ring and return. deque(maxlen=N).append is O(1), never blocks, and
# drops the OLDEST record when full — telemetry loss is always preferable to
# exec-path backpressure. Snapshot construction and the plane write happen on
# the drain lane, off the critical section entirely.
_TEL_RING_MAX = 4096
# 1. NOTHING below the execution ladder may block or raise into the exec
# path. Telemetry is not even ON the ladder (P0 kill .. P4 maintenance);
# it ranks below all of it.
# 2. NO DATA LOSS. DITAv2 exists precisely because sync/async seams were
# silently dropping ~40% of inputs. A drop-oldest ring is that same sin
# wearing a bounded-memory hat, and it discards forensic evidence exactly
# when things go wrong — i.e. exactly when it is needed (2026-07-13).
#
# Resolution (the M6 journal-lane pattern): the exec path appends O(1) to an
# UNBOUNDED in-memory queue and returns — never blocks, never drops, never
# raises. The LANE (never the exec path) is what deals with pressure: when the
# backlog exceeds the high-water mark, the lane SPILLS the overflow to a
# durable append-only spool on disk and replays it once the plane recovers.
# Memory stays bounded by the lane's spilling, not by throwing records away.
#
# Records are dropped: NEVER. Not when the plane wedges, not when it explodes,
# not on shutdown. They go to disk instead.
_TEL_HIGH_WATER = 4096 # backlog above which the lane spills to disk
_TEL_SPILL_BATCH = 1024 # records moved to the spool per spill pass
def _telemetry_spool_path(self) -> Path:
spool = os.environ.get("UV_TELEMETRY_SPOOL")
base = Path(spool) if spool else Path(tempfile.gettempdir()) / "uv_telemetry_spool"
base.mkdir(parents=True, exist_ok=True)
return base / "venue_telemetry.jsonl"
def _start_telemetry_lane(self) -> None:
if getattr(self, "_tel_thread", None) is not None:
return
self._tel_ring: Any = deque(maxlen=self._TEL_RING_MAX)
self._tel_q: Any = deque() # UNBOUNDED: the exec path never drops
self._tel_wake = threading.Event()
self._tel_dropped = 0
self._tel_spilled = 0 # records parked on disk (recoverable)
self._tel_replayed = 0
self._tel_dropped = 0 # MUST stay 0 — asserted by tests
self._tel_spool_lock = threading.Lock()
thread = threading.Thread(
target=self._drain_telemetry, name="bingx-telemetry-lane", daemon=True
)
self._tel_thread = thread
thread.start()
# SEPARATE SPILL DUTY. The publisher can block indefinitely inside a wedged
# plane's publish() — if it also owned spilling, a wedged plane would grow
# the queue without bound until OOM killed the process (which loses the very
# records we refuse to drop). The spiller never touches the plane, so it can
# always relieve memory pressure to durable disk no matter what the plane does.
spiller = threading.Thread(
target=self._spill_telemetry, name="bingx-telemetry-spill", daemon=True
)
self._tel_spill_thread = spiller
spiller.start()
atexit.register(self._flush_telemetry_to_spool)
def _spill_telemetry(self) -> None:
"""Pressure relief. Never touches the plane — cannot be wedged by it."""
while True:
time.sleep(0.2)
q = getattr(self, "_tel_q", None)
if q is None or len(q) <= self._TEL_HIGH_WATER:
continue
try:
with self._tel_spool_lock:
with self._telemetry_spool_path().open("a", encoding="utf-8") as fh:
while len(q) > self._TEL_HIGH_WATER:
try:
fields = q.popleft()
except IndexError:
break
fh.write(json.dumps(json_safe(fields), default=str) + "\n")
self._tel_spilled += 1
fh.flush()
except Exception as exc:
import logging as _log
_log.getLogger(__name__).error(
"FIX(bingx_venue): telemetry spill failed — holding in memory, "
"NOT dropping: %s", exc,
)
def _flush_telemetry_to_spool(self) -> None:
"""Shutdown / overflow: park records on disk. Called ONLY from the lane
or atexit — never from the exec path (it does disk I/O)."""
q = getattr(self, "_tel_q", None)
if not q:
return
try:
with self._tel_spool_lock:
with self._telemetry_spool_path().open("a", encoding="utf-8") as fh:
while True:
try:
fields = q.popleft()
except IndexError:
break
fh.write(json.dumps(json_safe(fields), default=str) + "\n")
self._tel_spilled += 1
fh.flush()
os.fsync(fh.fileno())
except Exception as exc:
import logging as _log
_log.getLogger(__name__).error(
"FIX(bingx_venue): telemetry spool write failed — records held in "
"memory, still not dropped: %s", exc,
)
def _replay_telemetry_spool(self, publish) -> None:
"""Plane is healthy and the backlog is drained: re-publish spooled records."""
path = self._telemetry_spool_path()
try:
with self._tel_spool_lock:
if not path.exists() or path.stat().st_size == 0:
return
lines = path.read_text(encoding="utf-8").splitlines()
path.write_text("", encoding="utf-8") # claim them
except Exception:
return
unsent: list[str] = []
for line in lines:
if not line.strip():
continue
try:
fields = json.loads(line)
fields["timestamp"] = datetime.now(timezone.utc)
publish(VenueTelemetrySnapshot(**fields))
self._tel_replayed += 1
except Exception:
unsent.append(line) # failed -> back to disk
if unsent:
try:
with self._tel_spool_lock:
with path.open("a", encoding="utf-8") as fh:
fh.write("\n".join(unsent) + "\n")
except Exception:
pass
def _drain_telemetry(self) -> None:
"""Low-priority lane: build snapshots + publish. Never touches the exec path."""
"""Low-priority lane: spill under pressure, publish, replay. Never drops."""
while True:
self._tel_wake.wait(timeout=0.5)
self._tel_wake.clear()
ring = getattr(self, "_tel_ring", None)
if ring is None:
q = getattr(self, "_tel_q", None)
if q is None:
continue
# Pressure valve: the LANE spills to durable disk — it never discards.
# This is what keeps memory bounded, in place of drop-oldest.
if len(q) > self._TEL_HIGH_WATER:
try:
with self._tel_spool_lock:
with self._telemetry_spool_path().open("a", encoding="utf-8") as fh:
for _ in range(self._TEL_SPILL_BATCH):
try:
fields = q.popleft()
except IndexError:
break
fh.write(json.dumps(json_safe(fields), default=str) + "\n")
self._tel_spilled += 1
fh.flush()
except Exception as exc:
import logging as _log
_log.getLogger(__name__).error(
"FIX(bingx_venue): telemetry spill failed — holding in memory, "
"NOT dropping: %s", exc,
)
while True:
try:
fields = ring.popleft()
fields = q.popleft()
except IndexError:
break
try:
@@ -327,15 +469,27 @@ class BingxVenueAdapter(VenueAdapter):
getattr(plane, "publish_venue", None) if plane is not None else None
)
if publish is None:
continue
q.appendleft(fields) # no plane yet: keep it, never lose it
break
publish(VenueTelemetrySnapshot(**fields))
except Exception as exc: # lane-local: never escapes to exec
import logging as _log
_log.getLogger(__name__).warning(
"FIX(bingx_venue): telemetry lane publish failed (phase=%s): %s",
"FIX(bingx_venue): telemetry lane publish failed (phase=%s) — "
"record spooled, not dropped: %s",
fields.get("phase", "?"), exc,
)
self._tel_q.appendleft(fields)
self._flush_telemetry_to_spool()
break
# Backlog clear and plane healthy -> drain whatever is parked on disk.
if not q:
plane = self._telemetry_plane
publish = getattr(plane, "publish_venue", None) if plane is not None else None
if publish is not None:
self._replay_telemetry_spool(publish)
def set_telemetry_plane(self, zinc_plane: Any | None) -> None:
self._telemetry_plane = zinc_plane
@@ -359,19 +513,18 @@ class BingxVenueAdapter(VenueAdapter):
client_order_id: str = "",
details: dict[str, Any] | None = None,
) -> None:
# EXEC PATH. Non-blocking and total: pack primitives, append to the ring,
# return. No plane write, no snapshot construction, no I/O, no lock, no
# raise. Telemetry ranks below every rung of the execution ladder and gets
# dropped (oldest-first) before it is ever allowed to cost the order path
# a microsecond.
# EXEC PATH. Non-blocking, total, and LOSSLESS: pack primitives, append to
# an UNBOUNDED queue, return. No plane write, no snapshot construction, no
# disk I/O, no lock, no raise — and no drop, ever. Pressure is handled by
# the lane (spill to durable spool), never by discarding records here.
#
# Everything below used to run INLINE, immediately after the venue had
# already accepted the order — so a raise here reached the caller's submit
# guard, which synthesised a REJECTED and rolled the FSM back while the
# venue kept the position (6 orphans, 2026-07-13).
try:
ring = getattr(self, "_tel_ring", None)
if ring is None or self._telemetry_plane is None:
q = getattr(self, "_tel_q", None)
if q is None or self._telemetry_plane is None:
return
slot_id = 0
trade_id = ""
@@ -391,12 +544,11 @@ class BingxVenueAdapter(VenueAdapter):
trade_id = str(order.internal_trade_id or trade_id)
asset = str(order.metadata.get("asset") or asset)
side = order.side or side
if len(ring) == ring.maxlen:
# Full: deque drops the oldest on append. Count it — silent
# telemetry loss must still be observable.
self._tel_dropped = getattr(self, "_tel_dropped", 0) + 1
# No drop check: there is nothing to drop. The queue is unbounded and
# the LANE relieves pressure by spilling to a durable spool. _tel_dropped
# exists only so a test can assert it stays 0 forever.
# Timestamp is stamped HERE (event time), not on the drain lane.
ring.append(
q.append(
{
"phase": phase,
"status": status,
@@ -699,9 +851,17 @@ class BingxVenueAdapter(VenueAdapter):
)
legacy = self._legacy_intent(intent, kernel=getattr(self, "_kernel_ref", None))
submitted = replace(intent, target_size=float(legacy.target_size))
# ── PRE-ACK ── safe to roll back if this raises.
receipt = self._call_backend("submit_intent", legacy)
events = self._events_from_submit(submitted, receipt, None, None)
ack_row = dict(getattr(receipt, "raw_ack", {}) or {})
# ── POINT OF NO RETURN ── order is LIVE; see submit_async().
try:
events = self._events_from_submit(submitted, receipt, None, None)
ack_row = dict(getattr(receipt, "raw_ack", {}) or {})
except Exception as exc:
raise VenuePostAckError(
f"post-ack failure after venue accepted the order: {exc}",
receipt=receipt,
) from exc
# Post-ack: the order is LIVE at the venue. Telemetry is observability,
# never a veto — an exception escaping here reaches the caller's submit
# guard, which synthesises a REJECTED event and rolls the FSM back while
@@ -757,9 +917,20 @@ class BingxVenueAdapter(VenueAdapter):
)
legacy = self._legacy_intent(intent, kernel=getattr(self, "_kernel_ref", None))
submitted = replace(intent, target_size=float(legacy.target_size))
# ── PRE-ACK ── an exception here means the venue never took the order.
# Safe for the caller to roll the FSM back.
receipt = await self.backend.submit_intent(legacy)
events = self._events_from_submit(submitted, receipt, None, None)
ack_row = dict(getattr(receipt, "raw_ack", {}) or {})
# ── POINT OF NO RETURN ── the order is LIVE at the venue from here on.
# Nothing below may reach the caller as a bare exception: that channel
# means "no order exists", and the caller rolls back to flat on it.
try:
events = self._events_from_submit(submitted, receipt, None, None)
ack_row = dict(getattr(receipt, "raw_ack", {}) or {})
except Exception as exc:
raise VenuePostAckError(
f"post-ack failure after venue accepted the order: {exc}",
receipt=receipt,
) from exc
# Post-ack: the order is LIVE at the venue. Telemetry is observability,
# never a veto — an exception escaping here reaches the caller's submit
# guard, which synthesises a REJECTED event and rolls the FSM back while