Conversation
…ent loop When a reconnecting daemon's state-log ack had been pruned, the daemon-status handler awaited `replay_persisted_params_for_daemon` inline. That replay sends every persisted param one round-trip at a time, each bounded only by `TCP_READ_TIMEOUT` (30s). A slow or half-dead daemon could therefore stall the whole coordinator for N x 30s per affected dataflow: no CLI requests, no heartbeats, and no other daemons' events. Spawn the replay as the other two replay call sites already do. It reports back through a new internal `ParamFallbackReplayFinished` event, which advances the daemon's ack to the `state_log_sequence` captured *before* the replay started (never moving it backwards), and only if every param was replayed and the daemon is still on the connection the replay was sent on. A per-daemon in-flight mark, keyed by connection, keeps a later status report on the same connection from starting a duplicate replay; a reconnect starts a fresh one. Fixes #3684 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016ZCEkLBowwvtAMhopshZwt
|
Merging to
After your PR is submitted to the merge queue, this comment will be automatically updated with its status. If the PR fails, failure details will also be posted here |
Port of #3682: cargo-deny fails the Audit job on every PR because yoke-derive 0.8.3 was yanked. No-op once #3682 lands on main. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016ZCEkLBowwvtAMhopshZwt
|
🤖 This is a fully automated review by Claude Code. No human has checked it. I found one issue. Running the fallback replay off the event loop opens the #3683 race on this path. Example: a daemon reconnects with a pruned ack. The fallback replay starts with This path is race-free on I checked the tests by reading them. I didn't run them with the fix reverted, because they call functions that don't exist on Generated by Claude Code |
|
Automated review by Claude (fully automated, not a human review) — reviewed at 2763e88. The race with The pruned-log fallback replays the same params that the same status report has already sent to this daemon. This means #3684 could likely be fixed with much less code. Drop the second replay in the pruned branch and only advance the ack, either from the outcome of the ready-barrier replay or on the next successful I found this by reading the code ( How the two PRs interact. Whichever lands second needs the Generated by Claude Code |
Fixes #3684
Problem
Suppose a reconnecting daemon's state-log ack has been pruned. The daemon-status handler then called
handle_pruned_state_catchup_fallback(...).awaitinline, on the coordinator's sequential event loop. That awaitsreplay_persisted_params_for_daemon, which sends every persisted param onesend_and_receiveat a time, each bounded only byTCP_READ_TIMEOUT(30s). A slow or half-dead daemon could stall the whole coordinator for N × 30s per affected dataflow. During that time it handled no CLI requests, no heartbeats, and no other daemons' events. The other two replay call sites alreadytokio::spawnthe replay.Fix
start_pruned_state_catchup_fallback(replaceshandle_pruned_state_catchup_fallback) is now synchronous. It spawns the replay and returns immediately.Event::ParamFallbackReplayFinished. It travels on a coordinator-internal channel merged into the event stream (Coordinator::internal_events).finish_pruned_state_catchup_fallbackhandles that event:state_log_sequencecaptured before the replay started, as the issue requires. Entries appended while the replay runs are not skipped. The ack never moves backwards past aStateCatchUpAckthat arrived meanwhile.RunningDataflow::fallback_replay_in_flightmaps a daemon to the connection ID its replay is in flight on. A later status report on the same connection does not start a duplicate. A reconnect does start a fresh replay, so a replay stuck on a dead connection cannot block the new one.Note: status reports are sent once per daemon connection. That is why the result comes back as an event rather than being polled at the next status report.
Related: #3683 / #3685 also touches
replay_persisted_params_for_daemon. The two PRs are independent but will need a trivial merge (one extra argument at the call site) depending on merge order.Also carries the
yoke-derive0.8.3 → 0.8.4Cargo.lockbump from #3682, because the yanked crate fails the Audit job on every PR. It becomes a no-op once #3682 lands.Validation
Class: C (coordinator)
cargo test -p dora-coordinator: ✅ 172 + 23 + 5 passedbinaries/coordinator/src/tests.rs:fallback_replay_does_not_wait_for_an_unresponsive_daemon: the daemon never replies. The call returns, the replay is in flight, the ack is unchanged, and a second status report does not duplicate the replay.fallback_replay_reports_back_and_acks_the_sequence_it_started_at:state_log_sequencegoes from 10 to 12 while the replay runs, and the ack lands at 10.fallback_replay_result_from_a_replaced_connection_is_ignoredfailed_fallback_replay_leaves_the_ack_and_never_moves_it_backfallback_replay_keeps_ack_unchanged_when_daemon_is_disconnectedandfallback_replay_respects_backoff_windowwere adapted to the new signature.cargo clippy -p dora-coordinator --all-targets -- -D warnings: ✅,cargo fmt --check: ✅make qa-fast: everything passes locally exceptaudit/typos(not installed here) andbreaking-changes(shallow clone, no tags). CI covers all three.Eventgains a variant.dora-coordinatoris not in the semver-checked crate list.cargo test --test fault-tolerance-e2e -- --test-threads=1): ✅ 16 passed on 2763e88🤖 Generated with Claude Code
https://claude.ai/code/session_016ZCEkLBowwvtAMhopshZwt