Files
msd-core/.changeset/sunny-wasps-cheer.md
Tom Boucher 5e2055ab90 fix(#3660): reap a bounded check's orphaned worker after its own timeout kill (#4615)
* test(#3660): prove a bounded node-test check orphans its worker on timeout

Regression test only, no fix yet: `execFileSync`'s timeout kills the direct
`node --test` runner but never the per-file worker it forks by default since
Node 22 (`--test-isolation=process`). The worker is reparented to PID 1 and
can busy-loop forever while the bounded-check verdict still reports a clean
fail-closed timeout.

Adds three tests driven through the real, uninjected defaultRunCheck path:
a hanging subject's worker must not survive the call, a control proving the
liveness probe can actually distinguish alive-vs-dead, and a non-hanging
failure proving the reap-gating logic added by the next commit doesn't
change the ordinary-failure return shape.

Expected RED on this commit (src/prohibition-enforcement.cts is unchanged).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3660): reap a bounded check's descendant worker after its own timeout kill

execFileSync's timeout only signals the direct child (the node --test runner);
since Node 22, node --test forks a per-file WORKER by default
(--test-isolation=process), so a hung subject's worker survives the bound,
gets reparented to PID 1, and busy-loops forever while the verdict still
reports a clean fail-closed timeout.

Adds execFileSyncReaping (wraps execFileSync, detached:true on POSIX) and
reapDescendants(pid): POSIX process.kill(-pid, 'SIGKILL') against the
process group, Windows an absolute-path taskkill /PID <pid> /T /F (never a
bare PATH-resolved name -- PR #3681 review minor-9). The reap fires ONLY
when this call's own timeout killed the child (the thrown error carries a
signal) -- an ordinary non-zero-exit failure has signal:null and is left
alone, which is the fix for PR #3681's Blocker-3 (that attempt reaped on
every throw, risking a PGID-reuse collateral kill on a ordinary red run).

All four execFileSync(process.execPath, ...) call sites now route through
execFileSyncReaping: runNodeTestWithSubject, defaultRunCheck's node-test and
lint-rule arms, defaultProveFailFirst's lint-rule arm (its node-test arm
reuses runNodeTestWithSubject).

RED proven on 0bd741fbc2b2ad0792fdf2361de14b68b2e3aea3 (test-only commit,
gsd-test outcome:failed, exactly the new orphan-detection test failing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3660): address code-review nits on the reap doc comments

- Clarify execFileSyncReaping's gate covers a maxBuffer-triggered kill too,
  not just a timeout -- both set .signal, both are "this call's own bound".
- Note reapDescendants' POSIX catch swallows any errno, not only ESRCH.

No behavior change (tsc --noEmit clean, no-op for the compiler).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* debug(#3660): fix hardcoded Windows path + add temp diagnostics for CI reap failure

Real defect #1 (fixed for good): taskkillPath() had a 'C:\Windows' literal
fallback, tripping tests/hardcoded-paths.test.cjs's repo-wide scanner. Now
returns null when neither SystemRoot nor windir is set, and the caller skips
the Windows reap rather than guessing a path.

Real defect #2 (under investigation): the prior GREEN gsd-test run showed the
#3660 orphan-detection test STILL failing on linux-node24 even with the fix
applied -- the worker survived. Isolated diagnostic scripts against the exact
same execFileSync({detached:true})+process.kill(-pid) mechanism, including
one using a REAL node --test worker, both confirm the mechanism works
correctly on macOS (group-kill reaches the worker). This commit adds
TEMPORARY stderr instrumentation (GSD-DEBUG-3660 tags) around the reap
attempt to get direct evidence from the actual Linux CI environment before
guessing further. Will be removed once the root cause is confirmed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3660): gate the reap on err.code === 'ETIMEDOUT', not err.signal

Root cause of the prior GREEN run's failure, confirmed with real evidence
from linux-node24 CI: execFileSync's thrown error on a genuine timeout-kill
does NOT reliably set `.signal` -- on that environment it came back
`signal: null, code: 'ETIMEDOUT', status: 7`, so the reap gate never fired.
A separate macOS/Node run of the identical scenario showed `signal: 'SIGTERM'`
for the same case -- neither field alone is safe across platforms/versions,
but `code === 'ETIMEDOUT'` was present and correct in both. Verified via
temporary stderr instrumentation on a real gsd-test run (now removed) before
landing this, rather than guessing from the macOS-only result.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* debug(#3660): round-2 instrumentation -- ETIMEDOUT gate fix alone didn't work

The err.code === 'ETIMEDOUT' gate fix (previous commit) did not resolve the
failure -- same test still red on real Linux CI. Adding probes around the
actual process.kill(-pid, 'SIGKILL') call itself to see whether it throws,
and whether the group is observably alive/dead before and after, since the
gate may now be firing correctly but the kill may not be reaching the
worker's process group on this environment. Temporary, will be removed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3660): the fix was already correct -- the TEST's liveness probe was not

Root cause of the two prior red rounds, confirmed via process-group probes on
real Linux CI: process.kill(-pid, 'SIGKILL') succeeds (no throw) every time
the ETIMEDOUT gate fires -- the worker genuinely IS killed. But
process.kill(pid, 0) cannot tell a truly-running process from an
already-killed ZOMBIE stuck unreaped: this bench's container has no init
process collecting arbitrary orphans, so a killed worker (reparented to PID 1
on death) sits as a zombie forever, still answering kill(pid,0) with "exists"
even though it is fully dead and burning zero CPU -- which is the actual harm
#3660 is about.

Test now reads /proc/<pid>/stat's process-state field on Linux and treats 'Z'
(zombie) as dead, falling back to the plain kill(pid,0) probe elsewhere (no
/proc on macOS/Windows).

Also strips the round-2 GSD-DEBUG-3660b instrumentation now that its evidence
has been used and the real root cause is fixed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3660): merge duplicate doc comment, fix stale err.signal reference

Leftover artifacts from the multi-round debugging: taskkillPath had two
stacked doc comments (an edit only replaced the function body, not the
original comment above it); a test comment still said "err.signal" after
the gate was changed to err.code === 'ETIMEDOUT'. Comment-only, no behavior
change (tsc --noEmit no-op).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3660): close a vacuous-test gap and harden isAlive's error handling

Code review finding (major): none of the three #3660 regression tests ever
asserted isAlive(pid) === true for a genuinely running process -- the real
code path is fully synchronous, so there's no natural window to observe
"alive" before "dead" inside those tests. A probe that always returned false
would have passed all three vacuously. Added a standalone test proving
isAlive(process.pid) reports true, using this test's own unambiguously-alive
process, running before the three existing tests.

Also hardened isAlive's /proc read-failure handling (minor finding): only
ENOENT (process genuinely gone) now means "dead"; any other read error
(EACCES, EIO, ...) reports "alive" (inconclusive) rather than risking a
false "dead" that would silently mask a real regression.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3660): add changeset fragment

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3660): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3660): bound the taskkill spawnSync with a timeout

CI caught it: local/require-subprocess-timeout (DEFECT.UNBOUNDED-SUBPROCESS)
flagged the new spawnSync(taskkill, ...) call in reapDescendants' Windows
branch for having no timeout. 5s bound -- a local OS command, not a network
call; reapDescendants already treats any failure (including a hypothetical
hang) identically via its existing try/catch, so the bound costs nothing and
just prevents a stuck taskkill from blocking the caller forever.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 17:55:05 -04:00

6 lines
557 B
Markdown

---
type: Fixed
pr: 4615
---
**A hung bounded test check no longer leaks a permanent CPU-pegging orphan process.** `node --test`'s per-file worker subprocess (the process default since Node 22) survived a timed-out check's own kill signal, which only reached the direct runner -- the worker was reparented to PID 1 and could busy-loop forever, consuming a full core, with no visible indication anything was wrong. The bounded check now reaps the whole process tree (POSIX process-group SIGKILL, Windows `taskkill /T /F`) when its own timeout fires. (#3660)