The `full test (windows-latest, 22)` job intermittently got CANCELLED at its
20m wall-clock cap with no failed test step — a false-negative gate (recurrence
of #869). Root cause: a unit test leaves an open event-loop handle, so the
chunk's `node --test` child hangs ~150s on Windows after its last test prints;
two such stalls push the already-~13m job past 20m.
Fix (defense in depth):
- run-tests.cjs: pass --test-force-exit (Node >=22; engines requires >=22.0.0)
so the runner exits once all tests finish regardless of lingering handles —
the durable backstop. Account for the flag in the argv-length ceiling.
- run-tests.cjs: add a per-chunk execFileSync timeout (default 600000ms, env
RUN_TESTS_CHUNK_TIMEOUT_MS) that fails loudly with a diagnostic naming the
chunk's files, so a hung chunk can never silently eat the job budget.
- perf-316 test: terminate both Worker threads on all paths (afterEach +
finally) so they cannot outlive the test.
- locking-bugs test: kill spawned children in a finally that wraps the whole
spawn -> waitFor -> barrier-release -> Promise.all sequence, so a barrier
timeout no longer leaks live child processes.
- Refresh the stale synckit comment (synckit/SDK bridge was removed).
Regression tests in run-tests-harness: a hung chunk hits the per-chunk timeout
and fails with a clear message; force-exit lets a chunk with a leaked handle
exit cleanly.
Closes#1051
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
extractCurrentMilestone reads STATE.md via planningDir(cwd), which is
workstream-aware (honours GSD_PROJECT/GSD_WORKSTREAM). The fixtures write
STATE.md to the plain <tmp>/.planning/STATE.md, so a developer shell inside a
GSD workstream (GSD_WORKSTREAM exported) redirected the read to a non-existent
workstream subdir -> version=null -> closed milestone sections leaked into the
slice and assertions failed. Clean CI/Docker env never hit it. Not a Node-26
regex bug; reproduces identically on any Node with GSD_WORKSTREAM set.
- scripts/run-tests.cjs: strip GSD_PROJECT/GSD_WORKSTREAM before spawning test
children so the local runner env matches clean CI/Docker.
- tests/roadmap-phase-fallback.test.cjs: file-level beforeEach/afterEach
save/delete/restore of both vars; new regression test pinning workstream-aware
STATE.md resolution.
- tests/run-tests-harness.test.cjs: guard asserting the runner strips both vars
(so removing the deletion fails clean CI).
Closes#872
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Windows CreateProcess caps lpCommandLine at 32,767 chars. The original
`execFileSync(node, ['--test', ...546 paths])` exceeded that on every
Windows runner and exited within ~70ms with no test output. Linux/macOS
allow ~2 MB ARG_MAX so the same call worked there.
`scripts/run-tests.cjs` now splits selected files into chunks that keep
each spawn's argv under 28,000 chars (operator-overridable via
RUN_TESTS_MAX_CMDLINE_CHARS), runs them sequentially, and reports the
first non-zero exit. Cross-platform regression test forces chunking with
a low ceiling and asserts the `run-tests: chunk N/M …` stderr marker.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
scripts/run-tests.cjs gains `--suite <name>` filtering using a filename
suffix convention (`*.security.test.cjs`, `*.integration.test.cjs`, …).
Files with no marker are `unit` (the default fast lane); files with a
marker land in the matching suite. No `--suite` flag preserves the prior
behavior of running every test (backcompat for `npm test` and
`npm run test:coverage`).
New package scripts wire the suites to stable entrypoints:
test:unit, test:integration, test:install, test:security, test:slow,
test:coverage:unit, test:coverage:all. Unknown suite → exit 2 with the
list of valid suites; empty suite → exit 0 with a stderr notice so empty
lanes (e.g. `security` before adversarial tests land) don't gate CI.
CI matrix grows from `ubuntu × {22,24}` + a single macOS lane to
`{ubuntu, macos, windows} × {22, 24, 26}`. `fail-fast: false` so one
lane failure doesn't cancel siblings. Node 26 is `continue-on-error`
until actions/setup-node stabilises that image. PR CI runs unit +
integration + security on every cell; `install` and `slow` only on
`main` push. A dedicated `coverage` job runs `test:coverage:unit` on
ubuntu/Node 24 and uploads the report.
Grouping policy lives in docs/TESTING-SUITES.md with a pointer from
CONTRIBUTING.md. New harness test covers arg parsing, filter selection,
empty-suite behavior, and failure propagation.
Closes#3597.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>