The `full test (windows-latest, 22)` job intermittently got CANCELLED at its
20m wall-clock cap with no failed test step — a false-negative gate (recurrence
of #869). Root cause: a unit test leaves an open event-loop handle, so the
chunk's `node --test` child hangs ~150s on Windows after its last test prints;
two such stalls push the already-~13m job past 20m.
Fix (defense in depth):
- run-tests.cjs: pass --test-force-exit (Node >=22; engines requires >=22.0.0)
so the runner exits once all tests finish regardless of lingering handles —
the durable backstop. Account for the flag in the argv-length ceiling.
- run-tests.cjs: add a per-chunk execFileSync timeout (default 600000ms, env
RUN_TESTS_CHUNK_TIMEOUT_MS) that fails loudly with a diagnostic naming the
chunk's files, so a hung chunk can never silently eat the job budget.
- perf-316 test: terminate both Worker threads on all paths (afterEach +
finally) so they cannot outlive the test.
- locking-bugs test: kill spawned children in a finally that wraps the whole
spawn -> waitFor -> barrier-release -> Promise.all sequence, so a barrier
timeout no longer leaks live child processes.
- Refresh the stale synckit comment (synckit/SDK bridge was removed).
Regression tests in run-tests-harness: a hung chunk hits the per-chunk timeout
and fails with a clear message; force-exit lets a chunk with a leaked handle
exit cleanly.
Closes#1051
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>