Skip to content

Coalesce Mongo reconnect recovery - #368

Merged
witbrock merged 1 commit into
mainfrom
codex/jvnautosci-2631-reconnect-singleflight
Aug 12, 2026
Merged

witbrock merged 1 commit into
mainfrom
codex/jvnautosci-2631-reconnect-singleflight

Conversation

@witbrock

Copy link
Copy Markdown
Member

What changed

  • Single-flight physical Mongo client construction per invalidation generation.
  • Fence and close stale unpublished clients without closing published clients still held by in-flight work.
  • Make degraded-route selection and invalidation atomic.
  • Coalesce explicit reconnect attempts and retry backoff while preserving each caller's retry budget.
  • Add deterministic concurrency coverage for success, failure, retry, invalidation, route rotation, cancellation, and health-preflight races.

Why

A transient Mongo or network failure could cause many concurrent requests and workers to independently invalidate the client, probe routes, and construct new clients. That multiplied a single outage into repeated Atlas handshakes, long request tails, and lease risk.

Impact

Concurrent callers now share one physical recovery attempt at a time. A later independent call can still retry after a failed wave, and stronger degraded-route intent receives one bounded atomic follow-up. Existing direct, DNS fallback, SSH tunnel, writable-primary, replica-set, and invalidate-without-close behaviour is preserved.

Validation

  • 106 focused Mongo, reconnect, fallback, index-guard, startup, observability, diagnostics, URI-redaction, and admin-health tests passed.
  • 56 neighbouring durable-workflow and execution recovery tests passed.
  • Both new single-flight suites passed 10 consecutive local runs; independent stress review ran them 100 consecutive times.
  • Black, Ruff on the new tests, Python compilation, and git diff --check passed.
  • Two independent concurrency/compatibility reviews found no remaining stop-ship race.

Residual limitation

A completion retained briefly for a lagging reconnect caller is not tagged with the Mongo generation. An unrelated direct invalidation in that tiny interval could yield a stale success receipt; the subsequent canonical operation detects it, and a fresh reconnect call cannot reuse that completion.

Jira: JVNAUTOSCI-2631

@witbrock
witbrock marked this pull request as ready for review August 12, 2026 08:29

assert len(results) == 12
assert {result["client"] for result in results} == {client}
assert connect_count == 1
time.sleep(_exponential_backoff(base_ms, attempt_num, max_ms))
_execute_reconnect_attempt(reconnect_attempt)
error = None
except BaseException as exc: # noqa: BLE001 - release waiting callers
build_error: BaseException | None = None
try:
candidate, clear_preference_on_failure = _build_real_connection_candidate()
except BaseException as exc: # noqa: BLE001 - finalise cancellation waves
result = call()
with result_lock:
results.append(result)
except BaseException as exc: # noqa: BLE001 - asserted by caller
try:
start.wait(timeout=2)
results.append(operation())
except BaseException as exc: # noqa: BLE001 - asserted by the test thread
@witbrock
witbrock merged commit 78eeebc into main Aug 12, 2026
4 checks passed
@witbrock
witbrock deleted the codex/jvnautosci-2631-reconnect-singleflight branch August 12, 2026 08:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant