Skip to content

fix(daemon): contain background forks of shell nodes under a shell guard (#3472) - #3543

Open
harsh839 wants to merge 34 commits into
dora-rs:mainfrom
harsh839:fix/3472-shell-orphan
Open

harsh839 wants to merge 34 commits into
dora-rs:mainfrom
harsh839:fix/3472-shell-orphan

Conversation

@harsh839

@harsh839 harsh839 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Closes #3472. path: shell nodes run sh -c <args> with no dora code in them, so the in-node DORA_RUN_PARENT_PID orphan guard never arms. They were contained only by the daemon's spawn-time PR_SET_PDEATHSIG, which reaches the direct child alone: a shell that forks a background child (sh -c 'sleep 1000 & wait') leaks that child to ppid 1 when dora run is SIGKILLed.

Guard routing on the dora run path only: a shell node now spawns a hidden dora __shell-guard CLI subcommand as the daemon's direct child and process-group leader. It arms on the injected DORA_RUN_PARENT_PID and, once that process is gone, killpgs its whole group — guard, shell, and background forks together — mirroring the in-node orphan_guard (it also clears the daemon's PDEATHSIG once its own 500ms poll is running, exactly like the in-node guard does).

Stop-path signals: the guard now catches the daemon's stop signals (SIGTERM/SIGINT/SIGHUP), forwards them to the shell, and stays alive to reap it. This keeps the node registered past the grace period so the group SIGKILL escalation still lands on a TERM-ignoring shell and its background forks — nothing survives a graceful stop.

Included in this pattern: for a TERM-ignoring shell, the whole signal ladder ends the group; a SIGTERM'd daemon/group behaves the same way. The dora up + dora start (coordinator-attached) path is unchanged and spawns the shell directly — nodes there are meant to outlive the daemon (#2029), and without DORA_RUN_PARENT_PID the guard could never arm, so it is deliberately not used. Windows cmd /C is unchanged.

Validation

  • New e2e run_killed_by_sigkill_does_not_orphan_shell_nodes: a shell node publishes the shell's pid and a sleep 1000 & background child's pid, dora run is SIGKILLed, and both must be gone within 30s. Verified it fails without the fix (only the shell is PDEATHSIG-contained; the child survives) and passes with it.
  • New e2e run_stop_does_not_orphan_term_ignoring_shell_nodes: a TERM-ignoring shell + background child must both be gone after a normal --stop-after stop; the CLI exit is awaited through a bounded try_wait (60s) that fails with the stderr tail rather than hanging.
  • Sibling orphan tests still green: run_killed_by_sigkill_does_not_orphan_nodes, run_killed_before_node_init_does_not_orphan, run_killed_by_sigterm_terminates_nodes_and_exits.
  • smoke_shell_node_allowed_with_flag / smoke_shell_node_blocked_without_flag still pass (gate intact).
  • cargo clippy -p dora-cli -p dora-daemon --all-targets -- -D warnings, cargo fmt --all -- --check, cargo test -p dora-cli --lib (420) and -p dora-daemon --lib (280) all clean.

Note for the merge queue

This branch was stacked on #3541 while both were being reworked, and was de-stacked on 09-24: it no longer carries feat/3489-mcap-export, and the diff has no dora recording export in it. It stands on its own against main.

@trunk-io

trunk-io Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

🚫 This pull request was removed from the merge queue because it was pushed to by @harsh839. Please re-submit it in order to merge. See more details here.

  • To merge this pull request, check the box to the left or comment /trunk merge below.

After your PR is submitted to the merge queue, this comment will be automatically updated with its status. If the PR fails, failure details will also be posted here

Copy link
Copy Markdown
Collaborator

🤖 Automated review (Claude) — first review of this PR. Fully automated, no human in the loop; advisory only, not an approval (I can't approve or merge). Reviewed from the diff/code as source of truth, not the description.

No blocking issues found. I verified the load-bearing pieces:

  • The killpg is contained. contain() does killpg(getpgrp(), SIGKILL), which is only safe if the guard is its own process-group leader. It is: the daemon wraps every unix spawn with ProcessGroup::leader() (prepared.rs:774), so the guard's pgid == its own pid, and the shell + its background forks inherit that group. It cannot signal the daemon's group.
  • No teardown regression. The daemon tears nodes down with group-based killpg(pid, SIGKILL) (shutdown.rs:196, keyed on the recorded node pid). The recorded pid is now the guard = the group leader, so the whole group (guard + shell + forks) still dies. dora stop / the SIGTERM→SIGKILL ladder are unaffected.
  • Liveness test is exact on the dora run path. DORA_RUN_PARENT_PID is set to the daemon pid (spawner.rs:662) and the guard is the daemon's direct child, so getppid() == parent ⇒ direct_child = true; the fallback kill(parent, 0)/ESRCH probe (with its pid-recycle caveat) only applies off that path.
  • The regression test genuinely fails on main. Without the guard, the daemon's PDEATHSIG SIGKILLs the shell alone; the background sleep 9527 reparents to init and survives, so the child-pid wait times out. With the guard, the group killpg takes both down. #[cfg(unix)], identity-checked teardown — solid.
  • Clearing PR_SET_PDEATHSIG on the guard (it fires on spawning-thread death, not daemon death) before entering the poll loop is correct, and the pre-spawn parent check closes the "parent already gone" race.

One non-blocking simplification. The guard is inserted on the coordinator-attached (dora up) path too, where DORA_RUN_PARENT_PID is absent and the guard is a pure passthrough — it adds an extra long-lived dora process (a heavy binary) per shell node and buys no containment there (those nodes are meant to outlive the daemon). The daemon already knows which path it's on (it's the one that sets the marker via bind_nodes_to_parent), so it could skip the guard entirely on the dora up path — mirroring exactly how #3544 gates the Windows KillOnDrop. Failing that, exec-ing the shell in passthrough (instead of spawn-and-wait) would at least avoid the extra resident process. Not required for correctness.

Minor edge note (also non-blocking): containment only holds while the shell stays alive. A shell that forks a background child and then exits immediately (sh -c 'x &', no wait) makes the guard forward the exit and depart, so a later daemon death won't be contained — but that's arguably a finished node with a deliberately-detached child, so the behavior is defensible.


Generated by Claude Code

@phil-opp phil-opp mentioned this pull request Sep 17, 2026
@harsh839

Copy link
Copy Markdown
Contributor Author

Addressed the non-blocking suggestion: the shell guard is now scoped to the in-process dora run spawn path only.

path_spawn_command now takes bind_nodes_to_parent and skips the guard entirely on the coordinator-attached path — the daemon spawns sh -c directly there, since without DORA_RUN_PARENT_PID the guard could never arm (pure passthrough + one extra resident dora process per shell node, buying nothing for nodes that are meant to outlive the daemon per #2029). This mirrors the gating the pre-init orphan window in prepared.rs and the Windows KillOnDrop (#3544) use — all keyed on the same parent-bound signal.

The dora run path is unchanged (still routed through dora __shell-guard). Verified:

  • run_killed_by_sigkill_does_not_orphan_shell_nodes still passes (guard active),
  • a path: shell node under dora start runs plainly with no guard process (manual dora up-shape check),
  • cargo fmt --all -- --check and cargo clippy -p dora-daemon -p dora-cli --all-targets -- -D warnings clean.

Copy link
Copy Markdown
Collaborator

🤖 Automated review (fully automated Claude Code review — no human vetted this; advisory only, not an approval, and I can't approve or merge). Reviewed from the diff.

New commit c321550 since the last review — the bind_nodes_to_parent scoping (guard only on the in-process dora run path, plain sh -c on the coordinator-attached path) looks correct, and the e2e is a genuine on-main failure. But re-reading the whole diff surfaced a blocking issue the earlier pass missed:

This won't compile for Windows targets. In binaries/cli/src/command/mod.rs, mod shell_guard;, the use shell_guard::ShellGuardArgs;, the Command::ShellGuard(...) variant, its execute() arm, and the parse_shell_guard test are all added without a #[cfg(unix)] gate, so binaries/cli/src/command/shell_guard.rs is compiled on every target. Its body relies on Unix-only libc symbols with no Windows fallback:

  • libc::getppid (armed, parent_is_gone), libc::kill / libc::ESRCH (parent_is_gone), libc::killpg / libc::getpgrp / libc::SIGKILL / libc::_exit (contain);
  • clear_parent_death_signal() is defined only for #[cfg(target_os = "linux")] and #[cfg(all(unix, not(target_os = "linux")))] — there is no Windows definition, so armed()'s unconditional call to it is unresolved on Windows too.

None of these exist in libc's Windows bindings, so cargo build/check -p dora-cli fails for *-pc-windows-*. PR CI is Linux-only, so it stays green here — but nightly's Windows leg, the "CLI Tests (all platforms)" job, and the cross-compile matrix will break. (macOS is unaffected — it's Unix and takes the all(unix, not(linux)) arm.)

Suggest gating the shell_guard module + the ShellGuard variant/arm/parse_shell_guard test behind #[cfg(unix)] (with a Windows stub for ShellGuardArgs::execute, or dropping the subcommand on Windows entirely, since Windows shell spawns already go through cmd /C and are unchanged). The daemon side is already correctly gated (dora_guard_command has a #[cfg(not(unix))] → None arm), so only the CLI subcommand needs the same treatment.

Otherwise the mechanism holds up: contain()'s killpg(getpgrp()) is safe because the daemon makes the guard the process-group leader, the _exit (not exit) choice avoids destructor deadlock, and the regression test genuinely fails on main (the background sleep 9527 reparents to init and survives the shell's PDEATHSIG).


Generated by Claude Code

@harsh839

Copy link
Copy Markdown
Contributor Author

Fixed the Windows-compile blocker: shell_guard is now gated behind #[cfg(unix)] across the wiring — mod shell_guard, the use, the Command::ShellGuard variant + its execute() arm, and the parse_shell_guard test. This matches the daemon side, where dora_guard_command already has a #[cfg(not(unix))] → None arm; Windows cmd /C shell spawns are unchanged and never reference the guard.

Verified: cargo build -p dora-cli, cargo clippy -p dora-cli --all-targets -- -D warnings, cargo fmt --all -- --check all clean, parse_shell_guard passes. I attempted cargo check -p dora-cli --target x86_64-pc-windows-msvc locally, but the known pre-existing aws-lc-sys build-script blocker (zenoh dependency, pthread on a Windows target) aborts the whole graph before any dora code — same one that already affects the Windows leg, so the Windows CI runs on real MSVC remain the gate for this.

Copy link
Copy Markdown
Collaborator

🤖 Automated review (fully automated Claude Code review — no human vetted this; advisory only, not an approval, and I can't approve or merge). Reviewed from the diff.

Re-reviewed at 74afbd7. The new commit gates the shell_guard module, the Command::ShellGuard variant / execute() arm, and the parse_shell_guard test behind #[cfg(unix)], which resolves the Windows compile break flagged earlier — clear_parent_death_signal now has full unix coverage (linux + all(unix, not(linux))) and the daemon side was already #[cfg(not(unix))] → None. I re-checked the core safety property independently: every unix spawn is wrapped with ProcessGroup::leader() (prepared.rs), so the guard is its own group leader and contain()'s killpg(getpgrp(), SIGKILL) can only reach the shell + its forks, never the daemon's group. No issues found in this revision.


Generated by Claude Code

@harsh839

Copy link
Copy Markdown
Contributor Author

@phil-opp Could you proceed with this merge? Claude re-review is clean (no blocking issues) and all CI checks are green. If it's in the queue for trunk, please submit it — or let me know if the author-side /trunk merge is expected to self-serve. Thanks!

Copy link
Copy Markdown
Collaborator

🤖 Automated review by Claude — fully automated review; no human has vetted this. Advisory only, posted like an outside contributor — I can't approve or merge.

One blocking issue that the earlier passes and PR CI didn't surface (the Test (ubuntu-latest) job runs in the merge queue, not on the PR head, so it shows skipped here):

The new hidden __shell-guard subcommand isn't recorded in cli-surface.txt, so the surface-snapshot test will fail. binaries/cli/tests/cli_surface.rs builds the surface by walking command.get_subcommands() with no filter on hidden or __-prefixed commands — hidden commands are intentionally included (dora daemon / dora coordinator, both hide = true, are already recorded, lines 49/55). Command::ShellGuard is #[cfg(unix)] and registered as __shell-guard, so on Linux the generated surface gains a dora __shell-guard line (plus its positional <command>) that isn't in binaries/cli/cli-surface.txt, and cli_surface_snapshot_is_current panics with "added (fine, but record it)". That blocks the merge-queue Test run.

Worth deciding how to resolve rather than just regenerating the snapshot: cli-surface.txt is the frozen-1.0 surface where "removing any line is a breaking change", so recording __shell-guard would freeze an internal implementation detail into the 1.0 guarantee. Excluding __-prefixed internal subcommands from the walk (and noting that in the snapshot header) seems more appropriate than pinning the guard command.


Generated by Claude Code

@harsh839

Copy link
Copy Markdown
Contributor Author

Addressed the blocking surface-snapshot issue — new head 605c83c.

cli_surface.rs now skips subcommands whose name starts with __ in the walk, exactly as suggested: the frozen 1.0 surface records dora daemon/dora coordinator (scripted against) but not the unix-only dora __shell-guard spawn wrapper, which would otherwise become a pseudo-breaking change for every non-unix platform. The exclusion is documented in the snapshot header.

Verified: cargo test -p dora-cli --test cli_surface passes (was failing with "added: dora __shell-guard"), snapshot_covers_the_command_tree still green, clippy/fmt clean, parse_shell_guard still passes. From here on I'll run the full cargo test -p dora-cli (not just the --lib subset) before pushing, since the surface test only runs over the whole suite.

Copy link
Copy Markdown
Collaborator

🤖 Automated review by Claude — fully automated review, no human vetted this; advisory only, not an approval.

Re-reviewed at 605c83c (new commit since my last review). The cli-surface snapshot blocker I flagged is resolved correctly: binaries/cli/tests/cli_surface.rs's walk() now skips __-prefixed subcommands (if sub.get_name().starts_with("__") { continue; }), so dora __shell-guard is neither generated nor recorded, and the HEADER const + cli-surface.txt header were updated in lockstep so cli_surface_snapshot_is_current stays consistent. This is the right fix for the 1.0 surface freeze — recording a unix-only internal wrapper line would have frozen a pseudo-guarantee that doesn't even exist on non-unix, and the __-prefix exclusion auto-covers any future internal subcommand. No new issues.


Generated by Claude Code

@github-actions

This comment has been minimized.

@github-actions

This comment has been minimized.

@github-actions

This comment has been minimized.

Copy link
Copy Markdown
Collaborator

🤖 Automated review (Claude Code): this was generated by a fully automated review with no human in the loop, so treat it as advisory only.

I re-read head b72928d from scratch and found one new issue that earlier reviews didn't cover. The earlier open points still stand.

The guard dumps core itself when the guarded program crashes (binaries/cli/src/command/shell_guard.rs, the reaper thread → re_raise)

When the guarded program dies from a signal whose default action dumps core (SIGSEGV, SIGABRT, SIGBUS, SIGFPE, SIGILL, SIGQUIT, …), re_raise resets that signal to SIG_DFL and raises it on the guard. So the guard process dumps core as well. The guard is the dora binary, or the Python interpreter under the wheel.

In practice, suppose a C/C++/Rust program runs as a path: shell node and hits an assert or a panic = "abort":

  • With ulimit -c set, you get a second, useless core file of dora in the node's working directory.
  • With systemd-coredump (or ReportCrash on macOS), it shows up as "dora crashed".

That sends users looking for a dora bug that doesn't exist. On main, only the user's program dumps core.

I checked the mechanism with a small repro: a child resets SIGABRT to SIG_DFL and raises it, with RLIMIT_CORE raised. The result is WCOREDUMP true, and a core file is written to the cwd.

A minimal fix that keeps the exact Signal(n) exit the daemon uses to classify stops: call setrlimit(RLIMIT_CORE, &rlimit { 0, 0 }) just before the raise in re_raise.


Generated by Claude Code

The exit-time containment ran for every unix node, so on the
coordinator-attached `dora up` path a node that exited on its own had
whatever it left in its process group `SIGKILL`ed with no grace period,
no log line and no way to opt out. That is a silent 1.0 behaviour change
well past what this PR's title says, and past what dora-rs#3472 asks: a helper,
viewer or launcher a node started is allowed to keep running, and a child
the node itself stopped just before returning is mid-cleanup, not
abandoned.

`contain_exited_group` now takes the spawn path's word on it, read from
the same `DORA_RUN_PARENT_PID` marker that already distinguishes
in-process `dora run` from `dora up` for PDEATHSIG and the Windows Job
Object, and takes an abandoned group only on `dora run`. That is the whole
of dora-rs#3472: `dora run` is about to exit, nothing else can reach the fork,
and leaving it behind is the hang being fixed. Off that path a dataflow
that wants its strays reaped stops the node, which is the stop-ladder
half below and is not gated.

Both `SIGKILL`s now say so, so "my child vanished" has an answer in the
daemon log instead of only in this diff.

Signed-off-by: harsh839 <harshbhargav440@gmail.com>
Assisted-by: Claude
@harsh839

Copy link
Copy Markdown
Contributor Author

Agreed with the finding, and took the first of your two suggestions. Pushed d792114.

The exit-time kill is now gated on the in-process dora run path, read from the same DORA_RUN_PARENT_PID marker that already separates dora run from dora up for PDEATHSIG and the Windows Job Object. On dora up a group a node abandoned is left alone: the dataflow outlives the node, so a helper, viewer or launcher it started keeps running, and a child the node terminated just before returning is mid-cleanup rather than abandoned. Stopping the node is what asks its group to stop, and that half — the stop-ladder replay, which is the actual #3472 fix — stays ungated on both paths, since a group that was asked to stop is expected to end either way.

So the behaviour change is now the one this PR's title describes and no more. dora run reaps the abandoned fork, which is the hang; dora up does not, and that is the pre-existing behaviour.

Both SIGKILL sites log now — the abandon path and the escalation — so "my child vanished" has an answer in the daemon log rather than only in this diff. It goes to tracing::warn! rather than the node's own log; the node's log drain is still pending at that point so it would fit, but it needs the NodeLogger threaded in, and I did not want to add that plumbing to a correctness fix. Say the word if you want it in dora logs instead.

docs/cli.md documents the split under dora run, including the dora up side, so the contract is written down rather than only implied by which path happens to be which.

On the two coverage points: abandoned_group_is_left_alone_off_the_dora_run_path is new and covers the gate in the other direction, and I inverted the condition to confirm both it and the existing dora run test fail when the gate is wrong. The five group_lifetime unit tests, 332 daemon lib tests and 24 of the 25 lifecycle e2e pass, including shell_node_abandoned_background_fork_is_contained_on_normal_exit, which is a dora run and so still proves the fix. The one failure is lifecycle_python_dynamic_add_remove, which needs uv; it fails the same way on b72928d without my changes, so it is the environment and not the branch.

Still no test for the take_queued_stop call site. I have not found a way to make it deterministic: it needs a ProcessOperation::StopRequested queued while the wait task is at the point of taking it, and the only scheduling handle I have seen is the executor yielding, which does not let a test put the op in the queue at that instant. If you would rather have the seam extracted than the coverage, I can pull the "take whatever is queued" step out into something callable, but I did not want to reshape the code around a test that would then be asserting the shape of the refactor.

Copy link
Copy Markdown
Collaborator

🤖 This is a fully automated review by Claude Code, with no human in the loop. Treat it as advisory only.

I re-reviewed d792114. The earlier finding is fixed, but the commit adds one new issue, and two older ones are still open.

Fixed: the gating works.

  • DORA_RUN_PARENT_PID is only injected when bind_nodes_to_parent is set, which only happens on the in-process Daemon::run_dataflow path.
  • It is in CONTROL_PLANE_ENV, so a descriptor or the inherited environment can't forge it.
  • spawn_inner reads it before the command is consumed, on every (re)start.
  • The Windows Job Object path gives the same result as before.
  • The group_lifetime and run_parent_marker unit tests pass, including the new abandoned_group_is_left_alone_off_the_dora_run_path.

That fixes the earlier finding that the exit-time kill applied to every unix node on every path.

New issue: every normal node exit under dora run now logs a spurious WARN.

  • In contain_exited_group (group_lifetime.rs:93-97), the stop == None branch logs node process group {pid} was abandoned by a node that exited on its own; SIGKILLing whatever is left in it (dora run) unconditionally before killpg.
  • It never checks whether the group still has members. The stop-ladder branch just below does (group_has_members).
  • dora run prints tracing at info, so the warning reaches the user.
  • As a result, every node that finishes on its own tells the user it was "abandoned" and is being SIGKILLed, even though its group is already empty. That includes a source that sends N messages and returns, and a node that exits once its inputs close.
  • Suggested fix: add if !group_has_members(pid) { return; } before the log and the kill.

Still open from earlier reviews, not addressed by the new commit:

  • Core dumps attributed to dora. re_raise in shell_guard.rs still resets core-dumping signals to SIG_DFL and raises them on the guard. A crashing guarded program therefore also produces a core dump or crash report attributed to dora. Setting RLIMIT_CORE to 0 before the raise would fix it.
  • Daemon panic on a huge grace duration. schedule_process_stop still computes Instant::now() + duration and soft_kill_at + kill_duration without overflow checks (running_dataflow.rs ~L846). A huge grace_duration can panic the daemon's event loop, which it does not do on main.

Generated by Claude Code

A review of d792114, and two findings from earlier rounds it had left open.

The one I introduced: the abandon branch logged and signalled an empty
group. A node that does its work and exits — a source that sends N
messages and returns, anything that stops when its inputs close — leaves
an empty group, so on `dora run`, which prints warnings, every
short-lived node announced that it had been SIGKILLed. It now asks
whether the group still has members first, the same question the ladder
below already asks, for the same reason: process-wrap's group wait has
reaped the group by now, so a member still in it is still running.

The guard used to die from the guarded process's signal with the
inherited `ulimit -c` intact, dumping a core whose executable is `dora`
— so `coredumpctl` answered "dora segfaulted" for a crash inside a
node's program. `re_raise` zeroes the limit first. The node's own dumps
are unaffected: it was exec'd with the limit still set, long before.

`grace_duration` is whatever a dataflow says it is, and
`Instant + Duration` panics on overflow, so a typo took the daemon's
event loop down instead of stopping a node. The duration is clamped to a
century, which is what a number that large means, and leaves the two
deadlines arithmetic that cannot overflow.

Signed-off-by: harsh839 <harshbhargav440@gmail.com>
Assisted-by: Claude
@harsh839

Copy link
Copy Markdown
Contributor Author

Took all three. Pushed 052521b, and the first one was mine.

Spurious WARN on every normal exit — my bug, from the last commit. Right, and worse than a log line: it signalled unconditionally. An empty group now returns before the warning and before the killpg, using the same group_has_members check the ladder below already uses, and for the same reason — the group wait has reaped the group by then, so a member still in it is still running. Two tests either side of it: a_node_that_exited_cleanly_is_not_reported_as_abandoned (captures the level, fails if a WARN fires) and an_abandoned_group_is_killed_on_the_dora_run_path, which is the counterpart — the group really is abandoned, it really is killed, and it really does say so. I removed the check again to confirm the first one goes red and the second stays green, so the pair brackets the behaviour rather than just testing the code.

Core dumps attributed to dora — agreed, fixed. RLIMIT_CORE is zeroed before the raise. The node's own dumps are untouched, which was the part worth checking: the guarded process is exec'd with the inherited limit long before the reaper thread reaches re_raise, so a node that segfaults still dumps core — the dump is just no longer one whose executable is the framework. The exit status the daemon classifies the stop by is unaffected.

Grace-duration panic — agreed, fixed. Clamped to a century. That is what a number that large means to whoever wrote it, and it leaves both deadlines arithmetic that cannot overflow rather than relying on the clamp being big enough. A new test drives the real thing with Duration::MAX; #[tokio::test] without start_paused, because the overflow is in Instant's own arithmetic and not in the passage of time, so a paused clock would not have caught it. With the clamp removed it panics in tokio's instant.rs, which is the bug you described.

Unverified here: the Windows path is untouched by all of this, and I still cannot cross-check it locally — aws-lc-sys needs a Windows C toolchain. The lifecycle e2e is 24/25, with lifecycle_python_dynamic_add_remove failing for want of uv exactly as it does on b72928d without any of my changes.

Still no take_queued_stop call-site test, for the reason in my last message. Happy to take the extraction if you would rather have the coverage.

Copy link
Copy Markdown
Collaborator

🤖 This is a fully automated review by Claude Code (no human in the loop). Treat it as advisory.

I re-reviewed 052521b, the only commit since the last automated review.

Fixed:

  • Grace-duration panic. schedule_process_stop now clamps grace_duration before it computes either deadline. stop_single_node goes through the same function. I found no other unchecked Instant + grace arithmetic left in the daemon.
  • Spurious "abandoned … SIGKILLing" WARN on every clean exit. The group_has_members early return now comes before both the log line and the killpg. Under dora run, I checked a clean shell node and one that abandons a background fork: only the second one warns.
  • I also checked by hand that a kill -9 of dora run takes down the guard, the shell and its background child.

Issues I found:

  • The exit-time kill under dora run applies to every node, not only when dora run exits.
    • The new docs/cli.md paragraph and the contain_exited_group doc justify the SIGKILL "because dora run is about to exit and nothing else can reach the child".
    • In fact contain_exited_group(pid, None, _, true) runs whenever any unix node (Rust, Python or C, not just path: shell) exits on its own, possibly while the rest of the dataflow keeps running for hours.
    • Example: under dora run, a Python node that Popens a viewer or helper and then returns now has that child SIGKILLed immediately. On main the child survives.
    • dora run orphan guard: path: shell nodes are never guarded (nothing of dora runs in them) #3472 only asks for containment when dora run itself is hard-killed. This is an extra behaviour change on dora run, and whether to keep it is a maintainer call. At minimum the documented reason should match what the code does.
  • The core-dump change only partly fixes the issue.
    • clear_core_dumps() sets RLIMIT_CORE to 0, which stops a core file from being written in the cwd.
    • With a pipe core_pattern (|systemd-coredump, apport), the kernel still invokes the helper when the limit is 0; only a limit of 1 suppresses a pipe dump. So coredumpctl will most likely still list a crash of dora, with no core file. macOS ReportCrash also ignores the rlimit. I haven't reproduced this.
    • Not re-raising the core-dumping signals (exiting with 128+n instead) would avoid it entirely.
    • Also, the new test calls clear_core_dumps() directly, so it would not catch re_raise dropping the call.
  • Scope (already raised earlier, still open). The PR bundles the __shell-guard wrapper, which is what dora run orphan guard: path: shell nodes are never guarded (nothing of dora runs in them) #3472 asks for, with a new daemon-side process-group lifecycle (group_lifetime.rs, StopRequested, and the group stop-ladder replay, which also applies on dora up). Splitting the lifecycle part into its own PR would make these behaviour changes easier to evaluate.

Generated by Claude Code

Copy link
Copy Markdown
Collaborator

🤖 Automated review (Claude Code). This review is fully automated, with no human in the loop. Treat it as advisory only.

I re-read head 052521b and found one new issue. It is in a test.

an_abandoned_group_is_killed_on_the_dora_run_path depends on how fast PID 1 reaps orphans (binaries/daemon/src/spawn/group_lifetime.rs, around L325).

  • After contain_exited_group, the test sleeps a fixed 500 ms and then asserts !process_alive(child), where process_alive is kill(pid, 0).
  • The sleep leader is SIGKILLed along with the rest of its group. So the fork becomes a zombie reparented to PID 1, and kill(pid, 0) keeps succeeding on it until PID 1 reaps it.
  • In a sandbox whose init reaps orphans about 1.5 s after they die, the test failed 4 out of 4 runs at the unmodified head. The other group_lifetime tests passed.
  • With a PID 1 that never reaps (for example, docker run without --init), this test fails every time. So does exited_node_group_is_killed_when_no_stop_is_in_flight, whose 10 s wait_for_exit only hides the delay.
  • Suggested fix: treat a zombie as dead (on Linux, check the state field in /proc/<pid>/stat), or make the test process a subreaper (PR_SET_CHILD_SUBREAPER) and reap the fork itself. At minimum, use wait_for_exit here instead of the fixed 500 ms sleep, like the sibling test does.

With their fixes reverted, a_node_that_exited_cleanly_is_not_reported_as_abandoned and a_grace_period_too_large_to_add_does_not_panic_the_daemon both fail, so they are real regression tests. The points from the 13:07Z review still stand; I'm not repeating them.


Generated by Claude Code

Copy link
Copy Markdown
Collaborator

🤖 Automated review (Claude Code). This review is fully automated, with no human in the loop. Treat it as advisory only.

Fresh pass on the unchanged head 052521b. I found one new issue in production code. The earlier zombie comment was only about a test.

A group of only zombies still counts as live, so a stop that already succeeded waits for the SIGKILL deadline (binaries/daemon/src/spawn/group_lifetime.rs, contain_exited_group / group_has_members).

  • group_has_members uses killpg(pgid, 0), which also succeeds when every remaining member is a zombie.
  • When dora is not PID 1 and PID 1 doesn't reap orphans, those zombies never go away. Examples: a container entrypoint such as a Python launcher that spawns dora run, or docker run without --init. The sequence:
    1. On Stop, the node leader exits. This is the stop branch, so it applies under dora up too.
    2. Its forked child finishes cleanup within the grace period and exits.
    3. The child is reparented to PID 1 and stays a zombie.
  • Result: the daemon waits until kill_at (1.5× grace) before reporting the exit. That delays dataflow finish, dora run exit and restarts. It also logs a misleading "ignored the stop grace period; SIGKILLing" warning for a node that behaved correctly.
  • How I checked: I reproduced the mechanism outside dora. A subreaper spawns a group leader sh -c 'sleep 0.5 & exit 0' and reaps only the leader, and killpg(pgid, 0) still succeeds 1.5 s later. The daemon behaviour follows from reading the code.
  • Suggested fix: on Linux, skip members whose /proc/<pid>/stat state is Z when deciding whether the group still has members.

Earlier findings that are still open:

  • The exit-time SIGKILL covers every unix node under dora run, not only shell nodes.
  • RLIMIT_CORE=0 does not stop pipe-based core handlers.
  • The group-lifetime/StopRequested rework is still bundled with the shell guard.
  • The dora-run abandoned-group test depends on PID 1 reaping promptly.

Generated by Claude Code

Copy link
Copy Markdown
Collaborator

🤖 Additional review notes (Claude Code). Found while reviewing head 052521b; these aren't in the earlier comments.

  • minor — the_core_limit_is_zeroed_before_the_guard_re_raises (binaries/cli/src/command/shell_guard.rs:316):
    • Its precondition accepts any finite hard limit ≥ 8 MiB (:336), but it then calls setrlimit with rlim_max: RLIM_INFINITY (:320). An unprivileged process can't raise a finite hard limit to infinity, so in such an environment the setrlimit assert fails with EPERM, not the precondition's explanatory message. Keeping rlim_max at the current hard limit avoids that.
    • clear_core_dumps() lowers the hard limit to 0 for the whole dora-cli lib test binary, which can't be undone unprivileged. Any later test in the same binary that relies on core dumps being possible would see 0. Running the check in a forked child would leave the test binary's own limits alone (and would also let the test go through re_raise rather than calling clear_core_dumps() directly, which the 09-29 review noted).
  • question — schedule_process_stop now calls process.submit(ProcessOperation::StopRequested { .. }) synchronously on the daemon's event loop (binaries/daemon/src/running_dataflow.rs:857). ProcessHandle::submit is a blocking op_tx.send(..) (running_dataflow.rs:297) on a 2-slot flume channel (binaries/daemon/src/spawn/prepared.rs:257, :597). 952a070 submitted from the ladder task for exactly this reason, and the move back fixed the marker race. In practice the send can't block today, since the wait task drains op_rx promptly. But the event loop's liveness now depends on that invariant (and on another worker being free to run the wait task). Is that intended? If so, a comment at the call site stating the invariant, or try_send with a logged fallback, would make it explicit.
  • nit — The "Note for the merge queue" section of the PR description still says the branch is stacked on feat(cli): add dora recording export to remux .drec recordings to MCAP #3541 and carries feat/3489-mcap-export. It was de-stacked on 09-24, and the diff has no dora export. Trunk squash-merges with the PR text, so it's worth removing before this goes into the queue.

Generated by Claude Code

…ut the stop-channel invariant

Two review notes on 052521b.

The core-limit test set `rlim_max: RLIM_INFINITY` after checking the
premise with a comparison that also accepted a finite hard limit above
8 MiB — an unprivileged process cannot raise a finite hard limit to
infinity, so the `setrlimit` would have failed with EPERM instead of the
premise's own message. It also called `clear_core_dumps` in this process,
which lowers the *hard* limit to 0 and cannot be undone without
privilege: every other test in the binary would have run with core dumps
silently disarmed. It now forks, keeps the hard limit where it found it,
and reports through the exit status.

`ProcessHandle::submit` is a blocking `send` on the event loop, and the
channel holds 2 while the ladder sends 3. That is intended and the
reasoning is not obvious, so it is now written down where the call is:
the wait task is the only consumer and sits in a `select!` on the
receiver for as long as the node lives, so a slot frees as soon as it is
scheduled; and a full channel would mean the node is already gone and
the receiver dropped, which fails fast instead of blocking. The review
suggested `try_send` with a logged fallback, which would be a regression
— it drops the marker when the channel is full, and the marker is the
thing that holds the group for its grace period.

Signed-off-by: harsh839 <harshbhargav440@gmail.com>
Assisted-by: Claude
@harsh839

Copy link
Copy Markdown
Contributor Author

Took the two code notes, skipped the third, and rewrote the PR description. 48d9d34.

Core-limit test — both halves right. It set rlim_max: RLIM_INFINITY while the premise check also accepted a finite hard limit above 8 MiB, so in such an environment the setrlimit would have failed with EPERM instead of the premise's message. And it called clear_core_dumps() in the test process, which lowers the hard limit to 0 — unrecoverable without privilege, so every other test in that binary would have run with core dumps silently disarmed. The check now runs in a forked child, raises only the soft limit and only to what the current hard limit allows, and reports through the exit status. I removed the setrlimit again to confirm the child rejects it.

You also suggested going through re_raise rather than calling clear_core_dumps directly. I left that alone: re_raise returns ! and ends in std::process::exit, so a test that calls it cannot make any assertion afterwards, and the child would have to observe its own death. What the forked child does check is the thing re_raise delegates to, which is the whole of the change.

Blocking submit on the event loop — intended, and now written down at the call site. Your reading of the arithmetic is right: 2 slots, 3 sends, so it can be full. What makes it safe is that the wait task is the only consumer and sits in a select! on op_rx.recv_async() for as long as the node lives, so a slot frees as soon as it is scheduled; and a full channel would mean the wait task is gone, i.e. the node already exited and the receiver was dropped, in which case the send fails fast and returns false instead of blocking.

I did not take the try_send suggestion, and I think it would be a regression: it drops the marker when the channel is full, and the marker is exactly what holds the group for its grace period. That is the race the synchronous submit was introduced to close, so trading liveness-that-should-never-matter for a lost marker is the wrong direction. The comment records the invariant and says why the fallback is not an option, so the next reader does not have to re-derive it.

PR description — removed. The "Note for the merge queue" section described a stacking that ended on 09-24, and pointed at a cli-surface.txt header merge that no longer applies. It now says the branch stands on its own against main and carries no dora recording export.

CI: fmt and clippy clean on both crates, 335 daemon and 431 CLI lib tests, and the lifecycle e2e is unchanged at 24/25 with the uv failure I have already shown fails on b72928d. The Windows path remains unverified here.

Copy link
Copy Markdown
Collaborator

🤖 Automated follow-up review (Claude Code). This review is fully automated, with no human in the loop. Treat it as advisory only.

I re-reviewed 48d9d34, the only commit since the last automated review. It adds no new issues.

  • Core-limit test: both points from the last review are fixed. The test now runs in a forked child, keeps the hard limit where it found it, and no longer changes the test binary's own limits. It passes locally (cargo +1.97.1 test -p dora-cli --lib the_core_limit_is_zeroed).
  • Blocking submit on the event loop: I agree it can't block, but the new comment (running_dataflow.rs ~856) gives the wrong reason.
    • The real reason: StopRequested is the first message ever sent on that incarnation's fresh 2-slot channel. ProcessHandle isn't Clone, and the only other send (Drop → Kill) consumes it. SoftKill/Kill come later, from the spawned ladder task.
    • "A full channel here would mean … the receiver is dropped" isn't true in general. A full channel with a live but not-yet-scheduled receiver would block.

Still open from earlier reviews (this commit doesn't touch group_lifetime.rs):

  • Under dora run, the exit-time SIGKILL fires whenever any node exits on its own. That doesn't match the "dora run is about to exit" justification in the docs.
  • RLIMIT_CORE=0 doesn't stop pipe-based core handlers (systemd-coredump, apport).
  • group_has_members (killpg(pgid, 0)) counts zombies. an_abandoned_group_is_killed_on_the_dora_run_path depends on PID 1 reaping promptly.
  • The group-lifetime / StopRequested rework is still bundled with the shell guard.

Generated by Claude Code

…e-limit guards

Four claims in the comments and docs did not survive checking, and a comment
that is wrong is worse than none: it is what the next reader trusts instead of
re-deriving.

- `ProcessHandle::submit` blocks for a stronger reason than the one given. The
  old text argued the wait task drains the channel; it cannot, and a full
  channel with a live-but-unscheduled receiver would block. The real reason is
  that this is the first message on a fresh per-incarnation channel, so it
  cannot be full at all.
- `group_has_members` was justified by process-wrap's group wait reaping every
  member. It only reaps *our* children: a reparented grandchild returns
  ECHILD and ends its loop. A zombie does answer `killpg(pgid, 0)` with 0, so
  a reaped-away orphan can read as a member. The leader being reaped is the part
  that actually holds, and the residue costs a no-op signal.
- The `dora run` justification said the run "is about to exit". It need not be:
  a source that returns leaves the same unreachable group whether the dataflow
  has three seconds or three minutes left. The orphan is the reason, not the
  countdown.
- `clear_core_dumps` claimed to leave no dump behind, which is only true of
  the file-based case. Where `core_pattern` pipes to a handler —
  `systemd-coredump`, `apport` — the kernel collects regardless of RLIMIT_CORE.

The zombie finding also had a test consequence: `process_alive` used
`kill(pid, 0)`, which reports a zombie as present, so the suite's verdict
depended on whether the host's PID 1 had reaped a reparented grandchild. It now
reads the state from /proc on Linux, and a regression test asserts the
`kill`-says-present, state-says-gone gap that made the fix necessary.

Signed-off-by: harsh839 <harshbhargav440@gmail.com>
Assisted-by: Claude
@harsh839

harsh839 commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Thanks — the submit catch was right, and my reason was wrong in a way worth
being precise about, because the wrong version is the one that would not have
been caught if it had been true.

I argued the wait task drains the channel so a slot frees. It cannot: a full
channel with a live-but-unscheduled receiver blocks. The reason it cannot block
is simpler — it is the first message on that incarnation's channel. The pair
is made fresh per spawn, this ProcessHandle is moved in and is not Clone,
and the only two later sends come from the ladder task. So the queue is empty
and both slots are free. 353f018 says that instead.

The zombie point was also right, and my comment was wrong about the mechanism.
I justified group_has_members on process-wrap's group wait reaping every
member, but waitpid(-pgid) only reaps our children — a reparented
grandchild returns ECHILD and ends that loop. I checked the kernel's answer
directly rather than reasoning about it:

/* fork a child in its own group, let it exit, and never reap it */
killpg(pgid, 0)  ->  0   /* non-empty: the zombie is counted */

So the residue is real: a reparented grandchild that has exited but that its
new parent has not reaped reads as a member. What actually holds is narrower —
the leader is reaped before this runs — and the rest costs a no-op signal plus,
on the dora run branch, a warning about a group that is already down. Reading
it the other way would mean walking away from a live orphan, which is #3472.
The comments now say that.

It also had a test consequence I had not thought about. process_alive used
kill(pid, 0), which reports a zombie as present, so the suite's verdict
depended on whether the host's PID 1 had got round to reaping. It passes 20/20
here because PID 1 is systemd; that is luck, not a property. It now reads the
state from /proc/<pid>/stat on Linux, with a regression test that asserts the
exact gap — kill answers 0 and process_alive answers false — and that
test fails if the helper goes back to kill alone.

The other two were overstated claims in comments, now corrected rather than
argued:

  • The dora run justification said the run "is about to exit". It need not be
    — a source that returns leaves the same unreachable group whether the
    dataflow has three seconds or three minutes left. The orphan is the reason,
    not the countdown. Worth flagging that this also answers the open item from
    the previous round about it firing on any self-exit: yes, it does, and that
    is deliberate, because the orphan is the same either way.
  • clear_core_dumps claimed to leave no dump behind, which is only true of the
    file-based case. Where core_pattern pipes to a handler —
    systemd-coredump, apport both do — the kernel collects regardless of
    RLIMIT_CORE, and only the filename on disk goes away. Suppressing that is
    the handler's configuration, not this process's.

Not addressing: the bundling. You have raised it three times and I have offered
to split twice; the answer is the same, and it is yours to make rather than mine
to keep re-litigating. Happy to split on request — the shell guard is
binaries/cli/src/command/shell_guard.rs plus the docs paragraph, the group
lifetime is binaries/daemon/src/spawn/group_lifetime.rs and the stop ladder in
running_dataflow.rs, and the two share no code.

Local: 336 daemon lib tests, fmt and clippy clean, lifecycle e2e 24/25 with
lifecycle_python_dynamic_add_remove failing on a missing uv (it fails the
same way on main).

The Audit (cargo-audit + cargo-deny) failure is unrelated to this branch and
reproducible on any push to main: RUSTSEC-2026-0284, lru 0.16.4, pulled
in by zenoh-ext 1.10.1, and main's own lockfile already carries that
version. I have not added a suppression for it — .cargo/audit.toml carries
five and each is a reachability judgement about which advisories dora accepts,
which is not mine to make on a PR about node lifecycle.

phil-opp commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

🤖 Automated review (Claude Code) — this review is fully automated, with no human in the loop. Treat it as advisory only.
Reviewed head: 353f018

353f018 fixes the reasons given in the submit and RLIMIT_CORE comments. It also makes the tests' process_alive treat a zombie as dead, and a_zombie_does_not_count_as_alive checks that. All 9 group_lifetime tests pass locally.

There is one remaining issue. It is the 09-30 06:57Z finding, which this commit acknowledges in a comment but leaves unfixed:

The new loop comment understates what a zombie costs. In binaries/daemon/src/spawn/group_lifetime.rs (~L129-136), the comment says a zombie counted by killpg(pgid, 0) "only costs a no-op signal". That holds for the dora run abandon branch, but not for the stop branch below it.

  • In the stop branch, group_has_members decides when the wait ends. A group with only zombies left keeps the loop polling until kill_at (1.5× grace).
  • After that the loop logs ignored the stop grace period; SIGKILLing.
  • This all runs before finished_tx.send (prepared.rs ~L1092). So a node that stopped correctly has its exit report held back for the whole window, and with it the dataflow finishing, dora run exiting and any restart. The user also sees a false warning.
  • When it happens: under any PID 1 that does not reap orphans, for example docker run without --init with dora not as PID 1. The case is a node that exits on Stop and leaves a child that finishes cleanup within the grace period.

This commit also makes the gap harder to see. The tests' process_alive now skips zombies, but production group_has_members does not. The suite can no longer observe the behavior described above.

Suggested fix: apply the same /proc/<pid>/stat state check in production. On Linux, count a group as empty when every member is in state Z, for example by scanning /proc/*/stat for the pgid. At minimum, the comment should say that the wait runs until kill_at.


Generated by Claude Code

phil-opp commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

🤖 Fully automated review by Claude Code. No human checked this; it's advisory only.

I found one new issue at 353f018:

Under the dora-rs-cli wheel, the shell guard inherits the node's env:, so a node's Python variables can stop it from starting.

dora_guard_command (binaries/daemon/src/spawn/command.rs) returns <host> __shell-guard -- sh -c …. The spawner (spawner.rs, around L852-878) then applies compose_node_env (the node's env:), the working dir and PYTHONUNBUFFERED to that command, i.e. to the guard, which passes them on to the shell. That's harmless with the standalone binary. Under the wheel, though, shell_guard_host is the dora console script (#!…/venv/bin/python → dora_cli). So every shell node now first starts a Python interpreter that has the node's environment, not the CLI's:

  • env: { PYTHONHOME: /opt/other-python }, meant for the program the shell runs, makes the guard's interpreter fail at startup.
  • A PYTHONPATH entry that contains another dora_cli (a second venv's site-packages, a source checkout) takes precedence over the venv's. The guard then imports a dora_cli that may not know __shell-guard, and the node fails to start.

On main, the same node is a plain sh -c and works, because these variables only reach the shell. This only affects dora run under the wheel with such an env:, but it is a regression. Possible fixes:

  • spawn the guard with the CLI's own environment and give the node's environment only to the sh it execs; or
  • have the guard strip PYTHON* from its own environment before re-applying the node's environment to the child.

Generated by Claude Code

phil-opp commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

🤖 Automated review (Claude Code). This review is fully automated, with no human in the loop, so treat it as advisory only. I reviewed head 353f018, which is unchanged since the last automated review.

A fresh pass found one point the thread hasn't raised yet.

The new e2e tests add roughly 5–6 minutes to the required PR CI.

  • All five new tests are in tests/node-lifecycle-e2e.rs.
  • The ci.yml step "Run dora node lifecycle E2E (non-Python)" runs every non-Python test in that file with --test-threads=1, on every PR.
  • The thread gives these runtimes on a warm dev machine:
    • about 30 s each for the fixture family;
    • 88 s for run_stop_grace_lets_a_stopped_nodes_child_finish_its_cleanup;
    • 93 s and 96 s for the two term-ignoring / died-on-term tests.
  • Together that roughly doubles a step whose comment already budgets about 5 minutes warm and "several times that" on a cold runner. The repo deliberately keeps PR CI lean.

Suggestion: keep one fast regression test in the PR job, for example run_killed_by_sigkill_does_not_orphan_shell_nodes. Move the long grace and stop-ladder tests to nightly.yml, or shorten their stop and grace timings. This also fits the earlier request to split the PR.

The open items already on this thread still stand.


Generated by Claude Code

killpg(pgid, 0) cannot tell a corpse from a running process: a zombie
stays in its group until it is reaped. For our own children that gap is
invisible, because the process-wait task reaps them. For a reparented
grandchild it is not -- only PID 1 can reap that, and an init that leaves
orphans lying may never.

On the stop branch this decides when the wait *ends*, and the caller
sends the node's exit report only after it returns, so a zombie holds
back the dataflow finishing, `dora run` exiting and any restart for the
rest of the grace period, and logs a false "ignored the stop grace
period" on the way. Reading it as empty is safe: a group cannot gain a
member except by a member forking into it.

The scan costs a few milliseconds per group, so the loop pays it once
every ZOMBIE_CHECK_EVERY polls rather than every poll.

Signed-off-by: harsh839 <harshbhargav440@gmail.com>
Assisted-by: Claude
The guard host is the `dora` CLI, which under the `dora-rs-cli` wheel is a
python console script, and the spawner applies the node's environment to
the whole command. A `PYTHONHOME` meant for the node's interpreter can
leave that python unable to start at all -- verified: `PYTHONHOME=/nonexistent
python3 -c ...` aborts before any code runs -- which takes the guard
down with it, and losing the guard silently loses the containment it
exists to provide.

Hand both variables to the guard under a `DORA_SHELL_GUARD_` prefix and
override them with an empty value, which python reads as unset, so the
host starts on its own interpreter. The guard restores them on the
child, which is the process they were meant for.

Overriding rather than removing is deliberate:
`clonable_command::Command::env_remove` deletes the entry instead of
recording the removal, so the variable is inherited anyway. Verified
against 0.2.0 with a standalone probe.

Signed-off-by: harsh839 <harshbhargav440@gmail.com>
Assisted-by: Claude
@harsh839

harsh839 commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

Both substantive findings are fixed in 7727602. On the third, I measured before changing anything, and the estimate does not hold up.

1. Zombie-only group (e008f07) — fixed at the root, not just in the test helper. group_has_members was the production predicate deciding when the stop branch ends, so a zombie held back the node's exit report (and with it the dataflow finishing and dora run exiting) until kill_at, plus a false "ignored the stop grace period" warning. It now excludes corpses on Linux; the /proc scan runs once per 4 polls, not per poll, because it costs a few ms per group.

The new test uses the production function, not a helper. Making it deterministic took some doing: a reparented grandchild does not survive, because systemd reaps instantly — the very case the fix exists for is the one a normal host cannot reproduce. So the fixture is std::process::Command (whose Child leaks a zombie on drop, unlike tokio's, which reaps in the background) with process_group(0), giving a group whose only member is an unreaped corpse. Verified it fails without the fix.

2. Wheel env regression (7727602) — confirmed and fixed. spawner.rs applies the node env to the whole command, and the guard host is the wheel's Python console script. PYTHONHOME=/nonexistent python3 -c ... aborts before any code runs, so the guard — and the containment it exists to provide — died silently. Both vars now go to the guard under a DORA_SHELL_GUARD_ prefix and are restored on the child, which is the process they were for.

One thing worth recording: I first used env_remove, and it silently did nothing. clonable_command::Command::env_remove deletes the entry instead of recording the removal, so the variable is inherited anyway. Confirmed against 0.2.0 with a standalone probe. It now overrides with an empty value instead — verified that python reads empty PYTHONHOME as unset, so the host starts on its own interpreter.

3. CI time — I could not reproduce the estimate, so I changed nothing. Measured the whole file the way the e2e job runs it (--test-threads=1 --skip lifecycle_python):

tests wall
main 19 536s
this branch 24 557s

So the five tests add ~21s total, about 1s each — not 5-6 minutes. The bulk of that ~9 minute job is pre-existing. Two notes on the per-test figure, since a single test reports a misleading number:

  • Each test's own phases are near-instant; the large per-test number is contention on the file's LIFECYCLE_LOCK plus a nested cargo build -p dora-cli into dora-examples/target (that package is its own workspace, so CARGO_TARGET_DIR is unset and each fresh run rebuilds the CLI). The Once guards only help within one process.
  • I did try cutting the ladder cost with a DORA_STOP_GRACE_SECS knob (the daemon's 10s grace + 5s escalation otherwise dominates any stop path). It moved the total by nothing measurable, because the ladder is not what these tests wait on. Reverted rather than leave an unused knob in the diff.

Happy to move any of the five behind #[ignore] if you would rather not pay the ~21s, but I did not want to slow the job for a cost that is mostly not mine.

phil-opp commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Automated review by Claude (fully automated, not a human review). Reviewed at 7727602, covering the two new commits.

I found the following issues:

1. 7727602 reads the daemon's env, not the node's. The wheel issue is still there, and shell nodes now lose their env: PYTHONPATH/PYTHONHOME.

clear_python_env_for_guard_host (binaries/daemon/src/spawn/command.rs ~L320) runs inside path_spawn_command and reads std::env::var(var). That is the daemon's environment. The node's descriptor env: is applied later: spawner.rs ~L855 calls compose_node_env → apply_descriptor_env on the command that path_spawn_command returned. Neither variable is on the denylist. Take a shell node with env: { PYTHONPATH: ./src }:

  • The later .env("PYTHONPATH", "./src") overrides the PYTHONPATH="" meant for the guard. The guard host therefore still starts with the node's value, so the wheel failure from the 10-01 review is not fixed.
  • DORA_SHELL_GUARD_PYTHONPATH holds the daemon's value, or "" (unwrap_or_default). restore_node_python_env (binaries/cli/src/command/shell_guard.rs ~L197) then always sets the child's PYTHONPATH to that value, because the prefixed variable is always present.

The second point is a regression for plain dora run with the standalone binary too. A shell node python3 -m mymod with env: PYTHONPATH: ./src works on main and on this branch before 7727602. With 7727602 it can no longer import its module.

The new test sets the daemon process env in place of the node env. That is the one setup where the ordering doesn't matter, so the test doesn't catch this. One fix is to move the node's values to the prefixed names after the descriptor env has been applied, in spawner.rs on the composed command.

2. e008f07 delays every graceful node stop by ~750 ms.

In the stop branch of contain_exited_group (binaries/daemon/src/spawn/group_lifetime.rs), since_zombie_check starts at 0. As a result group_has_members first runs on the 4th iteration, after 3 × 250 ms sleeps. Before this commit it ran on the first iteration.

The caller sends finished_tx only after this function returns (prepared.rs ~L1092). So a node that handles Stop correctly and leaves an empty group now has its exit reported about 750 ms late. That affects every dora node stop, restart and dataflow stop, on both dora run and dora up.

It also means the replayed SIGTERM/SIGKILL can go out up to 750 ms after the last membership check. That widens the pgid-reuse window the loop comment says stays within one poll interval.

Doing the cheap killpg(pgid, 0) on every poll, including the first, and only throttling the /proc zombie scan would avoid both problems.


Generated by Claude Code

phil-opp commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

🤖 Automated review (fully automated review by Claude Code — no human has vetted this; please verify the findings).

I did a fresh pass over head 7727602. The head hasn't changed since the 10:46Z review, and both of its findings still apply. I found two additional issues in 7727602 that the earlier reviews don't cover. I found both by reading the code; I didn't run them.

1. Clearing the variables also takes away the guard host's own PYTHONPATH/PYTHONHOME (binaries/daemon/src/spawn/command.rs:320-329, clear_python_env_for_guard_host)

  • The function sets PYTHONHOME="" and PYTHONPATH="" on the guard command whether or not the node asked for them. Under the wheel, the host is the same Python dora console script that is running dora run right now. The daemon's environment is that CLI's own environment, and it is exactly what the host needs to start.
  • If dora_cli can only be imported through PYTHONPATH, the guard host now fails at import dora_cli, and every shell node fails to start under dora run. Examples: a ROS/colcon workspace whose setup.bash puts site-packages on PYTHONPATH, pip install --target, or Bazel rules_python. The same happens for a relocated interpreter that needs PYTHONHOME. On main, and on this branch before 7727602, those nodes run.
  • This is the failure the commit set out to prevent, now caused by the CLI's own environment instead of the node's. Whatever fix you pick for the 10:46Z finding (moving the node's descriptor values to the prefixed names after compose_node_env), the host should keep the daemon's own values. Only the node's env: overrides should be moved aside.

2. The new tests modify the process environment in a multi-threaded test binary

  • command.rs:412 and shell_guard.rs:349 both call std::env::set_var/remove_var under a SAFETY: single-threaded test process comment. That comment is not true: libtest runs the dora-daemon and dora-cli lib tests in parallel threads. Other tests in the same binaries spawn processes and call into libc at the same time, and setenv racing with getenv is exactly why set_var is unsafe in edition 2024. binaries/cli/src/env_overrides.rs:218 already avoids this by re-running the test binary as a child process.
  • The daemon test also ends by unconditionally removing PYTHONHOME/PYTHONPATH, which wipes out any values the developer had set.
  • The CLI test's assertion that the child doesn't see DORA_SHELL_GUARD_* (shell_guard.rs:371) can never fail. get_envs() lists only explicit overrides, and the real child inherits the guard's whole environment, prefixed variables included.

Generated by Claude Code

phil-opp commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Automated review by Claude (fully automated, not a human review) — reviewed at 7727602.

I did a fresh pass focused on the platform-specific code: the Linux /proc path, the macOS fallbacks, and the Windows gating. The head is unchanged since the two 10-03 reviews, and their findings still apply: the env passthrough reads the daemon's env, there is a ~750 ms first-check delay, the host loses its own PYTHONPATH, and the tests use set_var. I found the following new issues.

1. Linux: group_has_members treats a live process as a zombie when its main thread has exited (binaries/daemon/src/spawn/group_lifetime.rs:242)

The check is state != "Z" on /proc/<pid>/stat, which is per thread-group leader. Take a process whose main thread called pthread_exit while its other threads keep running. The kernel reports that process as Z, even though it is alive. I checked this with a small C probe: pthread_create(sleep 30) followed by pthread_exit(0) in main gives /proc/<pid>/stat state Z with num_threads=2. /proc/<pid>/task lists two tids, and killpg(pgid, 0) still returns 0.

The same predicate decides both branches, so:

It is uncommon, but it happens: some C/C++ daemons and runtimes end main with pthread_exit. The fix is cheap. A real zombie has num_threads (stat field 20) of 1, while a zombie leader with live threads has more. So count a process as dead only when state == "Z" && num_threads <= 1. Checking for more than one entry under /proc/<pid>/task also works.

2. The /proc scan does blocking filesystem I/O inside an async fn on the daemon's runtime (group_lifetime.rs:216-246, called at :119 and :159)

Each call opens and reads /proc/*/stat for every process on the host. On a busy host with a few thousand processes, that is roughly 10-30 ms of blocking per call. It runs on every node exit under dora run, and once a second per stopping node while a group lingers. When many nodes stop at once, these scans stack up on the tokio workers that also drive the daemon event loop. Running the scan through tokio::task::spawn_blocking would keep it off the workers. Bounding it to the killpg-non-empty case, as the code already does, keeps the cost proportional.

Windows: I re-checked the gating by inspection. Everything unix-only (contain_exited_group, shell_guard, dora_guard_command, is_dora_cli) is cfg(unix) or has a non-unix stub, and StopLadder/take_queued_stop are platform-neutral. The only leftover is the signalled unused-assignment warning, which was already mentioned. macOS: killpg-only membership and a no-op clear_parent_death_signal look correct. The zombie concern from #1 does not apply there, because the /proc scan is Linux-only.

Verification: I found #1 by reading the code and confirmed the kernel behaviour with the probe described above (Linux 6.18). I did not run it through group_has_members itself, because that needs a dora-daemon test build. #2 is by inspection.


Generated by Claude Code

phil-opp commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator

🤖 Automated review (fully automated Claude Code review; no human vetted this).

I re-reviewed head 7727602, which has new commits since the last automated review at cddceb9: the shell_guard_host plumbing, the StopRequested marker and stop-ladder replay in group_lifetime.rs, zombie handling, and the python-env handoff. The issues raised earlier are addressed: the stop-path grace period, the codex P1 about the wheel host, and the stop test that didn't really exercise anything (the fixtures now keep the shell alive past the first TERM). I found two new problems.

1. Regression from the latest commit (7727602): a shell node's own env: PYTHONPATH/PYTHONHOME gets overwritten under dora run.

  • dora_guard_command calls clear_python_env_for_guard_host (binaries/daemon/src/spawn/command.rs:305/:320). That function sets PYTHONHOME=""/PYTHONPATH="" on the guard command, and sets DORA_SHELL_GUARD_<var> from the daemon's own std::env::var(var), or "" when that is unset.
  • The node's descriptor env is applied only afterwards, by compose_node_env → apply_descriptor_env (spawner.rs:855). PYTHONPATH/PYTHONHOME are not on the denylist, so a node's env: PYTHONPATH: ./lib replaces the "" on the guard host. The variable the commit means to keep off the host still reaches it.
  • In the guard, restore_node_python_env (binaries/cli/src/command/shell_guard.rs:197) then sets the child's PYTHONPATH from DORA_SHELL_GUARD_PYTHONPATH. That holds the daemon's value or "". So the shell runs with an empty (or the daemon's) PYTHONPATH instead of ./lib.
  • Example: path: shell, args: python my_node.py, env: {PYTHONPATH: ./lib}. This works on main, but under dora run with this PR the import fails. dora up/dora start is unaffected because it doesn't use the guard.
  • The unit test the_guard_host_is_kept_clear_of_the_nodes_python_env sets the variables on the test process (the daemon side), not in the node's descriptor env, so it misses this.
  • Possible fix: do the swap after compose_node_env, reading the values from the composed command rather than from std::env. Also note that deny_inherited_env already removes a variable for real by inserting the None sentinel into command.environment (spawner.rs:190), so the "" workaround isn't needed.

2. Every graceful stop of a unix node now waits about 750 ms before the exit is reported, on both dora run and dora up.

  • In contain_exited_group's stop branch (group_lifetime.rs:130-160), since_zombie_check starts at 0 and is incremented before it is compared against ZOMBIE_CHECK_EVERY = 4.
  • So the first group_has_members check runs only after three 250 ms sleeps.
  • The common case is a node that exits on NodeEvent::Stop with an empty group, and it pays that delay every time. finished_tx is sent only after this returns.
  • The effect is that dora stop, dora run --stop-after and Ctrl-C all finish about 0.75 s later than on main, for no reason.
  • Checking the group once before the first sleep (e.g. starting the counter at ZOMBIE_CHECK_EVERY - 1) avoids this.

One general note: the diff is now about 2.7k lines. A large share of that is comments, and some of them are long design write-ups that repeat each other (group_lifetime.rs, prepared.rs, running_dataflow.rs). The PR also bundles an unrelated fix, the MAX_GRACE clamp for Instant overflow in schedule_process_stop. Trimming the comments and splitting that fix out would make this much easier to review and maintain.


Generated by Claude Code

`cargo deny` fails the whole workspace on `yoke-derive 0.8.3`, which has
been yanked since main locked it. It arrives from `cargo`/`tokio` and is
not related to dora-rs#3472, but it reds the audit gate on every PR until the
lock moves. `cargo update -p yoke-derive` takes 0.8.3 -> 0.8.4, the
current non-yanked release, and leaves every other dependency alone.

Checked against the crates.io index that nothing else in the lock is
yanked, so this is the whole of the gate's complaint.

Signed-off-by: harsh839 <harshbhargav440@gmail.com>
Assisted-by: Claude
Three defects in the zombie handling, all reported by the automated
review of e008f07:

`killpg(pgid, 0)` counted as the whole answer was checked only on the
fourth poll, so every graceful stop of every unix node waited out three
250ms sleeps before its exit report went out -- and with it the dataflow
finishing, `dora run` exiting and any restart. The cheap syscall is now
asked on every poll, which is where a correctly-stopped node's empty
group is noticed, and only the `/proc` walk that tells a corpse from a
live member is throttled.

`state != "Z"` read a live process as a corpse when its main thread had
ended while other threads kept running (`pthread_exit(0)` out of `main`,
which some C/C++ daemons do; verified here as `state=Z` with
`num_threads=2` while `killpg` still answers 0). Both branches used that
predicate, so such a group was left alone on `dora run` -- the dora-rs#3472
orphan again -- and skipped the SIGKILL on the stop path. Only a
single-threaded `Z` counts as dead now.

The walk itself is blocking filesystem I/O and ran inline on the tokio
workers that also drive the daemon event loop, once a second per
stopping node. It goes to a blocking thread, and a join error counts as
"still a member": acting on a group that could not be looked at risks
killing a recycled one.

Signed-off-by: harsh839 <harshbhargav440@gmail.com>
Assisted-by: Claude
7727602 read `std::env::var` while building the guard command, but the
node's `env:` is applied to that command afterwards by `compose_node_env`.
So the value 7727602 moved aside was the daemon's own, the node's value
landed on the guard host anyway -- the wheel failure it set out to fix was
still there -- and the guard then put the daemon's value, or an empty one,
on the child, so a shell node lost its own `env: PYTHONPATH` that works
on main.

The swap now runs in the spawner, once `compose_node_env` has put the
node's value there, and moves only what the node set: the guard host
keeps the daemon's value, which under the wheel is what lets it import
`dora_cli` at all. Clearing that instead took the guard down for every
shell node in a colcon or `pip install --target` workspace.

Both new tests also stop writing to the process environment, which is
global and was being mutated from libtest's parallel threads under a
"single-threaded" comment that was not true.

Signed-off-by: harsh839 <harshbhargav440@gmail.com>
Assisted-by: Claude
@harsh839

harsh839 commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

Thanks — all six code findings were real. Fixed in three commits; details and evidence below.

1 + 2 + 3 — the python env handoff (three reviews, one root cause) — 701b627

All three are the same bug: clear_python_env_for_guard_host ran inside dora_guard_command and read std::env::var, i.e. the daemon's env, while the node's env: is applied to the same command afterwards by compose_node_env → apply_descriptor_env. So the value moved aside was the daemon's own, the node's value landed on the host anyway (the wheel bug was still there), and restore_node_python_env then put the daemon's value — or "" — on the child, so a shell node lost its own env: PYTHONPATH that works on main.

The swap now runs in spawner.rs after compose_node_env, and moves only what the node actually set. Two details worth calling out:

  • The host gets the daemon's value, not an empty one. Clearing it was the reviewer's point 2 in the 12:45Z pass and it is right: under the wheel the daemon's environment is what the host needs to import dora_cli, so emptying it takes the guard down for every shell node in a colcon / pip install --target / rules_python workspace. A node that sets nothing is left alone entirely.
  • When the daemon has no value either, the host gets the None sentinel rather than "". clonable_command 0.2.0's env_remove is a map deletion, so it would leave the node's value inherited — the empty-string workaround is no longer needed anywhere. That is the same mechanism deny_inherited_env uses.

4 — set_var in a parallel test binary — 701b627

Both tests are gone as written. The CLI test now goes through a 3-line seam (restore_python_env_from, the production function being the closure that reads std::env), so nothing touches the process environment. The daemon test asserts on the environment map and on a probe process started from it. Also dropped the assertion that the child does not see DORA_SHELL_GUARD_* — you are right that it cannot fail: get_envs() lists explicit overrides only, and the child does inherit the passthrough names, which is harmless.

5 — Z is not the same as dead — eb66566

Confirmed with the probe you describe, on this host: pthread_create + pthread_exit(0) gives state=Z, num_threads=2, and killpg still 0. A member counts as live unless it is a single-threaded Z, so both branches stop misreading it — dora run no longer abandons that group, and the stop path no longer skips the SIGKILL. num_threads is field 20, i.e. index 17 after the last ) in stat; stat_is_live_member is now a separate function with a table test over the four cases, so that index is pinned rather than incidental.

6 — blocking /proc walk on the runtime — eb66566

group_has_live_member runs the scan in spawn_blocking; a join error counts as "still a member", since acting on a group we could not inspect is how you kill a recycled one.

The ~750 ms stop delay — eb66566

Worse than reported, in one respect: it was in 7727602's parent too, so it applied from e008f07 on. The cheap killpg(pgid, 0) is now asked on every poll including the first, and only the /proc walk is throttled to every fourth. an_already_empty_group_ends_the_wait_at_once measures it: 753 ms before the fix, <100 ms after. That test caught a real gap in my own first attempt — I had renamed the counter and throttled the walk without actually restoring the per-poll check.

Also: the signalled unused-assignment warning from the 20:46Z pass does not reproduce — cargo clippy -p dora-daemon --all-targets is clean.

Audit gate — 9a29f7b

The red gate is real but not from this PR: cargo-deny fails on yoke-derive 0.8.3, yanked upstream after main locked it (RUSTSEC-2026-0284 in an earlier note of mine was the wrong diagnosis — that is lockfree, and all seven advisories here are allowed warnings, so cargo-audit passes). Moved the lock to 0.8.4, the current non-yanked release, which is what cargo update -p yoke-derive gives. Checked the whole lock against the crates.io index: nothing else is yanked, so that was the gate's whole complaint.

Validation — cargo fmt --all --check, clippy clean on both crates, 341 daemon lib tests, 433 CLI lib tests. Every new test mutation-checked: ignoring num_threads, throttling the killpg check, no-op'ing the handoff, and no-op'ing the restore each fail the test that covers it. The E2E suite and the ~21 s CI figure are unchanged from before.

Not done, deliberately: the MAX_GRACE overflow clamp in schedule_process_stop is still bundled here, and the comments are still long. The clamp is a separate concern and is a self-contained commit, so it is one cherry-pick away into its own PR if you would rather review it alone — say the word and I will open it. I would rather not trim the reasoning in group_lifetime.rs on my own initiative: most of it is there because each of those paragraphs is a trap I have already fallen into once (the clonable_command env_remove, the pgid reuse window, the non-reaping PID 1), but if you want it shorter I will cut it back to the load-bearing lines.

phil-opp commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator

🤖 Automated review (fully automated Claude Code review; no human checked it)

I reviewed the three commits since the last automated review (9a29f7b, eb66566, 701b627). Both issues from that review are fixed:

  • ~750 ms stop delay: killpg(pgid, 0) runs on every poll again, and only the /proc walk is throttled. With the per-poll check removed, an_already_empty_group_ends_the_wait_at_once fails, so the test covers the fix.
  • Node's own PYTHONPATH/PYTHONHOME lost: the handoff now runs after compose_node_env and moves only the node's own value.

The num_threads zombie predicate and the spawn_blocking /proc walk also look correct.

I found the following issues in the new commits:

1. 701b627 breaks the build on non-unix targets (Windows).

  • binaries/daemon/src/spawn/spawner.rs:7 imports handoff_python_env_to_guard and is_shell_guard with no #[cfg], and spawner.rs:884-886 calls them unconditionally.
  • Both functions are #[cfg(unix)] (spawn/command.rs:310, :337) and have no non-unix fallback, while mod spawner is compiled on every platform.
  • So dora-daemon, and with it dora-cli, fails to compile on Windows (E0432/E0425). PR CI is Linux-only, so CI here won't catch it. It will break the nightly Windows and cross-check jobs and cargo install dora-cli on Windows.
  • Gating the import and the if is_shell_guard(..) block with #[cfg(unix)] fixes it, as b72928d did for contain_exited_group. A non-unix is_shell_guard that returns false would also work.

2. No test covers the ordering bug that 701b627 fixes.

  • Both new tests call handoff_python_env_to_guard directly on a hand-built command, so neither goes through the spawner.
  • With the spawner call disabled (if is_shell_guard(&command) && false), every guard/python-env test in cargo test -p dora-daemon --lib still passes. Moving the call back before compose_node_env would pass too.
  • A test that builds the command through the spawner path with a node env: {PYTHONPATH: ...}, then checks the guard host's env and the DORA_SHELL_GUARD_PYTHONPATH handoff, would cover it.

The earlier general note still applies: the PR is ~2.85k lines and still bundles the unrelated MAX_GRACE clamp.


Generated by Claude Code

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

trunk-cancelled waiting-for-review Pull request is waiting for a review from maintainers.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

dora run orphan guard: path: shell nodes are never guarded (nothing of dora runs in them)

2 participants