* test(#3271): guard against a folded suite appearing twice in one host Adds local/no-duplicate-fold-marker, an AST rule that reports the second and every subsequent `folded:<name>` marker in a host file, plus RuleTester cases and a tree-wide regression assertion. Failing-first on purpose: the rule is registered at error and the 25 duplicated regions are still present, so eslint and the new tree-wide test are RED. The deletions land in the next commit. The marker key is the whitespace-delimited token after `folded:` — not the issue's `[a-z0-9-]*` slice, which truncates at `.` and false-positives on tests/model-resolver.test.cjs where feat-443-effort-fast-mode.integration and feat-443-effort-fast-mode are two distinct folded suites. Refs #3271 * fix(#3271): delete 25 duplicated folded suites from three install hosts Three consolidated install suites each carried a verbatim second copy of a contiguous run of #1969 B1 folded blocks. Byte-identical, constant offset, and green — each duplicated block registered and ran twice on every lane. tests/install.test.cjs 5981-9937 (3957 lines, 18 blocks) tests/install-minimal-hooks.test.cjs 2734-4015 (1282 lines, 5 blocks) tests/install-write-confinement.test.cjs 1754-2321 ( 568 lines, 2 blocks) Introduced by6d072435d(#1975 re-applying #1970's hunks on a tree that already had them, 2026-07-03) — one stale-base re-application, three files, one commit. Verified by marker-count bisect: 1 at4f779eda4and0cc7a1a42, 2 from6d072435donward. The later copy is deleted in each case, so every file returns to what its authoring batch produced and blame on the surviving lines stays accurate. local/no-duplicate-fold-marker, red on the previous commit, is now green. tests/model-resolver.test.cjs is untouched: the issue lists it, but its two blocks are folded from two different files and are not identical. It is a false positive of the issue's own grep, whose `[a-z0-9-]*` key truncates at `.`. Fixes #3271 * test(#3271): property-test marker identity and pin the alias non-goal Three review findings, all fixed inline: 1. foldMarkerOf is a parser and carried no fast-check property test. Raised independently by the /code-review standards axis and the isolated adversarial pass; the file already establishes the fc.property-driving-ruleTester idiom for a sibling rule. Added, two arms over markers generated from [a-z0-9-._]: the same marker twice always reports exactly once against firstLine 1, and two distinct markers never collide. The alphabet includes `.` on purpose — an implementation keyed on the issue's [a-z0-9-]* slice passes arm 1 and fails arm 2, which is exactly the model-resolver false positive. 2. meta.docs.category was the novel value 'Test hygiene'; all 16 sibling local rules use 'Best Practices', 'Portability' or 'Reliability'. Now 'Best Practices'. 3. A call through a further alias (const d = __foldDescribe) was unreported and undocumented — accidental rather than deliberate. It is now the fourth entry in the rule's documented non-goals, with the reason, and pinned by a valid RuleTester case so it cannot drift silently. Refs #3271 * test(#3271): name the step and elapsed time when a baseline build fails buildBaselineAtRef runs four bounded steps and, when one exceeded its bound, threw a bare "spawnSync ETIMEDOUT" naming neither the step nor how long anything took. Diagnosing one real failure took four separate experiments to recover information the throw already had. Each step is now timed, and any throw carries the breakdown: which step failed, its elapsed time, the timings of every step that completed before it, all three bounds, and the tail of the child's captured stdout/stderr. The failure message is deliberately the carrier. On the remote runner the captured output field comes back empty in failures.json while error and stack survive verbatim, so the message is the only channel that reaches a reader of a remote verdict. Refs #3271 * fix(#3271): size the baseline generator bound for the machine it runs on Instrumentation from a real remote-runner failure gave the breakdown: git-worktree-add=15.1s npm-run-build-lib=19.8s gen-emitted-baseline=FAILED@300.1s Steps 1 and 2 are comfortable. Only the generator exceeds its bound, and it is not hung — it needs more than 300s there. Measured ladder for that step: ~22s idle in a container, ~39s end-to-end in a clean container, ~142s with 8 CPU burners on 8 cores, and >300s under the real suite. Its cost is 19 sequential installer spawns, and spawn latency is exactly where a container degrades worst (3.9x slower than host, against 1.1x for file IO) — which is why a CPU-only load test did not reproduce it and why four earlier hypotheses (container slowness, network, shallow clone, CPU contention) all measured clean. The 300s bound was sized on an idle machine for a step that never runs on one. Under the remote runner the on-disk baseline cache is structurally absent — CI restores it via actions/cache keyed on github.event.pull_request.base.sha, a key that exists only inside GitHub Actions — so this slow path runs on every remote verification. The result: this gate has passed 0 times in 754 runs, failing 80 times and never once executing successfully. Raised to the 600000ms ceiling that local/no-unbounded-spawn treats as the largest meaningful bound; the other two bounds are untouched. This makes the gate RUN, which is the point: the alternative considered and rejected was degrading the timeout to a skip, and that was measured to turn the suite green with the gate silently not running at all. The real remedy is making the cache reachable from the remote runner so the in-job build returns to being the rare fallback ADR-2719 §5 describes. That is a gsd-test-runner change, not one this repo can make. Refs #3271 * fix(#3271): tolerate an overlay source that vanishes mid-walk Observed on the remote runner, three runs across three different branches: ENOENT: no such file or directory, link '/work/hooks/dist/gsd-config-reload.js' -> '/tmp/gsd-2930-overlay-6nOZay/hooks/dist/gsd-config-reload.js' buildOverlayRepo enumerates names with readdirSync and then acts on each one, so statSync, copyFileSync and linkSync all sit in a TOCTOU window. hooks/dist is regenerated by an ATOMIC REPLACE (scripts/build-hooks.js unlinks and renames), so any concurrently running test that rebuilds hooks retires a just-listed name mid-walk and the overlay dies on it. linkOrCopyFile already tolerated EXDEV and EPERM; ENOENT went straight through. On ENOENT the source is now re-examined ONCE rather than slept on. An atomic rename is a single syscall, so by the time the failure surfaces the successor is either already in place (the retry succeeds) or the path has genuinely left the tree, in which case there is nothing to mirror and the leaf is skipped. No sleep and no spin: a timing-based wait here would be the very flake being fixed. Every other errno still propagates untouched, so a real permission or IO fault stays a hard failure. Five tests hold the boundary: gone-for-good skips without retrying, mid-replace retries exactly once and places the file, EACCES still throws, a real linkSync ENOENT is injected by monkeypatching fs and restoring it in a finally (never a mode-bit trick, which root bypasses), and isMissingPath accepts only ENOENT. Refs #3271 * fix(#3271): order the timeout ladder inward-out and lock it Two review blockers, both real. The generator bound had been raised to 600000ms — exactly the whole-chunk timeout in scripts/run-tests.cjs:973. A step bound equal to the chunk ceiling loses the race: the chunk is killed first and the failure arrives as an opaque "no failed step" kill, so the per-step diagnostic added a commit earlier was built and then made unreachable in the same change. Separately the #2767 test declared a per-test timeout of 300000ms, BELOW the inner bound it was meant to permit, so it could still die at the exact 300s ceiling this was supposed to lift — via node:test's timeout rather than spawnSync's. Its sibling declared 900000ms, above the chunk ceiling, which is the same opaque-kill hazard from the other direction. The three bounds only produce a useful failure if they fire inward-out, so they now do: step 360s, per-test 480s, chunk 600s. 360s is ~3x the passing observation (91.6s / 115.8s) and 20% above the censored 300.1s timeout, while leaving 240s of chunk headroom for every other file sharing it. Four tests lock the ordering, including a drift guard on the exported values — without it, editing a call site's literal timeout would leave the ordering assertions passing while the real ladder inverted. Also from review: - err.gsdBaselineStep and err.gsdBaselineTimings were written and never read anywhere in the tree; only the rewritten message is consumed. Removed rather than kept as speculative surface. - buildOverlayRepo discarded placeVanishableLeaf's boolean at both call sites, so a vanished leaf left the overlay with no accounting at all. It now collects the skipped paths and warns once. Not thrown: a source that left the tree really is not part of the snapshot, and throwing would reintroduce the crash the tolerance removes — but silence would let a dropped leaf resurface later as an unrelated missing-file assertion. - The instrumentation commit shipped no test. One now drives a real failure and asserts the message names the step, its elapsed time, and the bounds. Refs #3271 * chore(#3271): backfill the changeset PR number * fix(#3271): bound a hook fan-out as its own class, not as a bare probe CI failure on PR #3285, job full test (windows-latest, 22, shard 2/3) — every other lane green, including windows-latest node 24 across all three shards: not ok 1 - blocks push when any to-be-pushed commit matches local blocked regex error: bash .githooks\pre-push failed — outcome=timed_out exitCode=null stderr= duration_ms: 15040.2168 A bound, not a hang: the test supplies stdin via input:, so the hook is not blocked reading its ref list, and the duration lands exactly on the 15000ms bound. The site used PROBE_TIMEOUT_MS, which tests/helpers/timeouts.cjs documents as "a single short CLI query or node -e probe against a temp fixture". This is not that. It spawns bash running .githooks/pre-push, and the hook then invokes a MOCK git that is itself a bash script, so one runHook is roughly four Git Bash spawns. On Windows each is Defender-scanned and the first hook test in a file pays cold start on top. That module's own docstring warns against precisely this: a call site that differs from its class must not be forced onto a shared value that does not describe it. HOOK_FANOUT_TIMEOUT_MS is that missing class — 60000ms, 4x the bound that failed and half INSTALL_TIMEOUT_MS, which is the right order: a hook fan-out is much lighter than a full installer run and far heavier than reading back a version string. Two tests lock the ordering against both neighbours, including one asserting real margin over the censored 15040ms observation, since a bound that merely matched what was measured would be the same defect again. Scoped deliberately: the other ~360 runHook sites keep their current bounds. This adds the norm and applies it where a real failure demonstrated the need, rather than sweeping a value across sites with no evidence for any of them. Refs #3271 --------- Co-authored-by: sim <sim@local>
826 lines
31 KiB
JavaScript
826 lines
31 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* Tests for tests/helpers/process-seam.cjs (the spawnSync-based subprocess
|
|
* seam) and its runGsdTools adapter in tests/helpers.cjs.
|
|
*
|
|
* Contract: .gsd/phase/test-3055-process-seam-module/40-design.md
|
|
* Matrix: .gsd/phase/test-3055-process-seam-module/50-test-matrix.md
|
|
*
|
|
* Rows 1-24 exercise the seam directly. Rows 25-35 exercise the
|
|
* `runGsdTools` adapter's contract-parity guarantee for its 136 callers.
|
|
*
|
|
* Assertions are on typed fields only (outcome/exitCode/timedOut/signal/
|
|
* killed/code) or on structured JSON a fixture prints to stdout — never on
|
|
* raw stdout/stderr text via .includes()/assert.match(), per CONTRIBUTING
|
|
* "Prohibited: Raw Text Matching on Test Outputs".
|
|
*/
|
|
|
|
const { test, describe, beforeEach, afterEach } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
|
|
const { createTempDir, cleanup, runGsdTools, TOOLS_PATH } = require('./helpers.cjs');
|
|
const processSeam = require('./helpers/process-seam.cjs');
|
|
const { runNode, runGit, runHook, OUTCOME, toSeamResult } = processSeam;
|
|
|
|
// ---- fixture sources -------------------------------------------------
|
|
|
|
// argv: [exitCode?, stdoutPayload?, stderrPayload?]
|
|
const FIXTURE_EXIT = [
|
|
"const code = Number(process.argv[2] || '0');",
|
|
'const stdoutPayload = process.argv[3];',
|
|
'const stderrPayload = process.argv[4];',
|
|
"if (stdoutPayload) console.log(stdoutPayload);",
|
|
"if (stderrPayload) console.error(stderrPayload);",
|
|
'process.exitCode = code;',
|
|
].join('\n');
|
|
|
|
// argv: [sleepMs, stdoutMarker?, stderrMarker?] — writes markers via a
|
|
// synchronous fd write (never buffered console.log) so partial output
|
|
// survives a kill even when the child never reaches a clean exit.
|
|
const FIXTURE_SLEEPER = [
|
|
"const fs = require('fs');",
|
|
"const sleepMs = Number(process.argv[2] || '0');",
|
|
'const stdoutMarker = process.argv[3];',
|
|
'const stderrMarker = process.argv[4];',
|
|
"if (stdoutMarker) fs.writeSync(1, stdoutMarker + '\\n');",
|
|
"if (stderrMarker) fs.writeSync(2, stderrMarker + '\\n');",
|
|
'setTimeout(() => {}, sleepMs);',
|
|
].join('\n');
|
|
|
|
// Echoes received argv back as JSON — proves argv arrives literal/unmodified.
|
|
const FIXTURE_ECHO_ARGV = [
|
|
'console.log(JSON.stringify({ argv: process.argv.slice(2) }));',
|
|
].join('\n');
|
|
|
|
// Echoes stdin back as JSON.
|
|
const FIXTURE_ECHO_STDIN = [
|
|
"let data = '';",
|
|
"process.stdin.on('data', (chunk) => { data += chunk; });",
|
|
"process.stdin.on('end', () => {",
|
|
' console.log(JSON.stringify({ received: data, hadData: data.length > 0 }));',
|
|
'});',
|
|
].join('\n');
|
|
|
|
// Kills its own process with SIGKILL — simulates an external kill (OOM
|
|
// killer, `process.kill(pid, 'SIGKILL')` from outside) that spawnSync does
|
|
// NOT populate `result.error` for on this runtime.
|
|
const FIXTURE_SUICIDE = [
|
|
"process.kill(process.pid, 'SIGKILL');",
|
|
'setTimeout(() => {}, 5000);',
|
|
].join('\n');
|
|
|
|
// argv: [exitCode?] — writes well past the 1MB default maxBuffer, then
|
|
// attempts a clean exit with exitCode (which the overflow kill preempts).
|
|
const FIXTURE_OVERFLOW = [
|
|
"const exitCode = Number(process.argv[2] || '0');",
|
|
'process.exitCode = exitCode;',
|
|
'for (let i = 0; i < 300; i += 1) {',
|
|
" process.stdout.write('x'.repeat(10000));",
|
|
'}',
|
|
].join('\n');
|
|
|
|
const BUFFER_OVERFLOW_CODES = ['ENOBUFS', 'ERR_CHILD_PROCESS_STDIO_MAXBUFFER'];
|
|
|
|
function writeFixture(dir, name, source) {
|
|
const fixturePath = path.join(dir, name);
|
|
fs.writeFileSync(fixturePath, source);
|
|
return fixturePath;
|
|
}
|
|
|
|
/**
|
|
* Monkeypatch processSeam.runNode for the duration of `block`, restoring the
|
|
* original in `finally`. Standalone helper (no test-context access), per
|
|
* CONTRIBUTING's try/finally carve-out and the CLAUDE.md IO-fault-injection
|
|
* pattern (save original, override, restore in finally).
|
|
*/
|
|
function withMockedRunNode(impl, block) {
|
|
const original = processSeam.runNode;
|
|
let callCount = 0;
|
|
processSeam.runNode = (...args) => {
|
|
callCount += 1;
|
|
return impl(callCount, ...args);
|
|
};
|
|
try {
|
|
block(() => callCount);
|
|
} finally {
|
|
processSeam.runNode = original;
|
|
}
|
|
}
|
|
|
|
describe('process-seam', () => {
|
|
let tmpDir;
|
|
|
|
beforeEach(() => {
|
|
tmpDir = createTempDir('process-seam-');
|
|
});
|
|
|
|
afterEach(() => {
|
|
cleanup(tmpDir);
|
|
});
|
|
|
|
test('exit 0 reports EXITED with the child stdout', () => {
|
|
const fixture = writeFixture(tmpDir, 'exit.cjs', FIXTURE_EXIT);
|
|
const result = runNode([fixture, '0', JSON.stringify({ ok: true })]);
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.equal(result.exitCode, 0);
|
|
assert.equal(result.timedOut, false);
|
|
assert.equal(result.signal, null);
|
|
assert.equal(result.killed, false);
|
|
assert.equal(result.code, null);
|
|
assert.equal(typeof result.stdout, 'string');
|
|
assert.deepStrictEqual(JSON.parse(result.stdout.trim()), { ok: true });
|
|
});
|
|
|
|
test('non-zero exit is EXITED, not a failure-to-run', () => {
|
|
const fixture = writeFixture(tmpDir, 'exit.cjs', FIXTURE_EXIT);
|
|
const result = runNode([fixture, '5', '', 'boom']);
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.equal(result.exitCode, 5);
|
|
assert.equal(result.timedOut, false);
|
|
assert.ok(result.stderr.length > 0);
|
|
});
|
|
|
|
test('empty stderr with non-zero exit keeps exitCode', () => {
|
|
const fixture = writeFixture(tmpDir, 'exit.cjs', FIXTURE_EXIT);
|
|
const result = runNode([fixture, '9']);
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.equal(result.exitCode, 9);
|
|
assert.equal(result.stderr, '');
|
|
assert.notEqual(result.code, undefined);
|
|
assert.equal(result.code, null);
|
|
});
|
|
|
|
test('a child that overruns is TIMED_OUT as data, not a throw', () => {
|
|
const fixture = writeFixture(tmpDir, 'sleeper.cjs', FIXTURE_SLEEPER);
|
|
const result = runNode([fixture, '5000'], { timeoutMs: 300 });
|
|
assert.equal(result.outcome, OUTCOME.TIMED_OUT);
|
|
assert.equal(result.timedOut, true);
|
|
assert.equal(result.killed, true);
|
|
assert.equal(result.exitCode, null);
|
|
});
|
|
|
|
test('timedOut does not depend on signal presence (Windows)', () => {
|
|
const fixture = writeFixture(tmpDir, 'sleeper.cjs', FIXTURE_SLEEPER);
|
|
const result = runNode([fixture, '5000'], { timeoutMs: 300 });
|
|
// The assertion below is intentionally the whole point of this test: it
|
|
// proves timedOut alone, without ever branching on result.signal. See
|
|
// 40-design.md row 6 — signal is null on Windows and must not be
|
|
// load-bearing for this flag.
|
|
assert.equal(result.timedOut, true);
|
|
});
|
|
|
|
test('a timeout still returns string stdout/stderr (partial content is platform-dependent)', () => {
|
|
const fixture = writeFixture(tmpDir, 'sleeper.cjs', FIXTURE_SLEEPER);
|
|
const marker = JSON.stringify({ partial: true });
|
|
const result = runNode([fixture, '5000', marker], { timeoutMs: 300 });
|
|
assert.equal(result.outcome, OUTCOME.TIMED_OUT);
|
|
assert.equal(result.timedOut, true);
|
|
assert.equal(typeof result.stdout, 'string');
|
|
assert.equal(typeof result.stderr, 'string');
|
|
// spawnSync preserves partial child output on a timeout on darwin, but discards
|
|
// it on Linux (verified on node 22 and 24). The seam passes through whatever
|
|
// spawnSync gives it, so the cross-platform contract is only that these are
|
|
// strings — the partial content itself is asserted where it is actually available.
|
|
if (process.platform === 'darwin') {
|
|
assert.ok(result.stdout.length > 0, 'darwin preserves partial stdout on timeout');
|
|
assert.deepStrictEqual(JSON.parse(result.stdout.trim()), { partial: true });
|
|
}
|
|
});
|
|
|
|
test('maxBuffer overflow is not misreported as exit 1', () => {
|
|
const fixture = writeFixture(tmpDir, 'overflow.cjs', FIXTURE_OVERFLOW);
|
|
const result = runNode([fixture, '0']);
|
|
assert.equal(result.outcome, OUTCOME.BUFFER_OVERFLOW);
|
|
assert.notEqual(result.exitCode, 1);
|
|
assert.equal(result.exitCode, null);
|
|
assert.ok(BUFFER_OVERFLOW_CODES.includes(result.code));
|
|
});
|
|
|
|
test('a missing binary is SPAWN_FAILED, not a timeout', () => {
|
|
// git exists on the test host; a nonexistent cwd makes the OS-level
|
|
// spawn itself fail with ENOENT (uv_spawn), the same failure class as a
|
|
// missing binary, without requiring the seam to expose a raw command
|
|
// parameter callers could point at an arbitrary executable name.
|
|
const result = runGit(['status'], { cwd: path.join(tmpDir, 'does-not-exist') });
|
|
assert.equal(result.outcome, OUTCOME.SPAWN_FAILED);
|
|
assert.equal(result.code, 'ENOENT');
|
|
assert.equal(result.exitCode, null);
|
|
assert.equal(result.timedOut, false);
|
|
});
|
|
|
|
// Regression for the Windows CI failure on PR #3066: a `status === null`
|
|
// catch-all previously misclassified this as TIMED_OUT (the process never
|
|
// even started), which drove the adapter into a pointless retry. An
|
|
// oversized argv errors at the OS spawn boundary before the child exists
|
|
// at all — E2BIG on Linux/macOS, ENAMETOOLONG on Windows — and must
|
|
// classify as SPAWN_FAILED regardless of which errno the platform uses.
|
|
test('an oversized argv is SPAWN_FAILED, not TIMED_OUT, cross-platform', () => {
|
|
const fixture = writeFixture(tmpDir, 'echo-argv.cjs', FIXTURE_ECHO_ARGV);
|
|
const oversizedArg = 'x'.repeat(4 * 1024 * 1024);
|
|
const result = runNode([fixture, oversizedArg]);
|
|
assert.equal(result.outcome, OUTCOME.SPAWN_FAILED);
|
|
assert.equal(result.timedOut, false);
|
|
assert.equal(typeof result.code, 'string');
|
|
assert.ok(result.code.length > 0);
|
|
});
|
|
|
|
test('omitting timeoutMs still bounds the call', () => {
|
|
const fixture = writeFixture(tmpDir, 'exit.cjs', FIXTURE_EXIT);
|
|
const withDefault = runNode([fixture, '0']);
|
|
const withExplicitDefault = runNode([fixture, '0'], { timeoutMs: 60000 });
|
|
// Omitting timeoutMs must resolve to the same bounded code path as
|
|
// explicitly passing the documented default — never a distinct
|
|
// "unbounded" branch.
|
|
assert.deepStrictEqual(
|
|
{ outcome: withDefault.outcome, exitCode: withDefault.exitCode, timedOut: withDefault.timedOut },
|
|
{ outcome: withExplicitDefault.outcome, exitCode: withExplicitDefault.exitCode, timedOut: withExplicitDefault.timedOut }
|
|
);
|
|
});
|
|
|
|
test('child finishing just under the bound is EXITED', () => {
|
|
const fixture = writeFixture(tmpDir, 'sleeper.cjs', FIXTURE_SLEEPER);
|
|
const result = runNode([fixture, '50'], { timeoutMs: 5000 });
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.equal(result.timedOut, false);
|
|
});
|
|
|
|
test('at-the-bound child yields one deterministic outcome', () => {
|
|
const fixture = writeFixture(tmpDir, 'sleeper.cjs', FIXTURE_SLEEPER);
|
|
const result = runNode([fixture, '300'], { timeoutMs: 300 });
|
|
// Either outcome is acceptable at the exact bound (OS/scheduler
|
|
// jitter decides which side of the race wins) — what must never happen
|
|
// is an outcome outside the pair, or fields inconsistent with whichever
|
|
// branch fired. This is the non-flaky formulation of "at the limit".
|
|
assert.ok([OUTCOME.EXITED, OUTCOME.TIMED_OUT].includes(result.outcome));
|
|
if (result.outcome === OUTCOME.EXITED) {
|
|
assert.equal(result.timedOut, false);
|
|
assert.notEqual(result.exitCode, null);
|
|
} else {
|
|
assert.equal(result.timedOut, true);
|
|
assert.equal(result.exitCode, null);
|
|
}
|
|
});
|
|
|
|
test('toSeamResult classifies a raced status+ETIMEDOUT as EXITED, not TIMED_OUT', () => {
|
|
// Synthetic reproduction of the exact-bound race: spawnSync's timer
|
|
// fired (error.code === 'ETIMEDOUT') just as the child finished on its
|
|
// own (status: 0). Evidence (a real status) must outrank the attached
|
|
// error — the old discrimination order checked error.code first and
|
|
// reported TIMED_OUT with exitCode: 0, an incoherent shape.
|
|
const result = toSeamResult({
|
|
status: 0,
|
|
error: { code: 'ETIMEDOUT' },
|
|
signal: null,
|
|
stdout: '',
|
|
stderr: '',
|
|
});
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.equal(result.timedOut, false);
|
|
assert.equal(result.exitCode, 0);
|
|
});
|
|
|
|
test('child overrunning the bound is TIMED_OUT', () => {
|
|
const fixture = writeFixture(tmpDir, 'sleeper.cjs', FIXTURE_SLEEPER);
|
|
const result = runNode([fixture, '5000'], { timeoutMs: 200 });
|
|
assert.equal(result.outcome, OUTCOME.TIMED_OUT);
|
|
});
|
|
|
|
test('invalid timeoutMs is rejected, not coerced to unbounded', () => {
|
|
const fixture = writeFixture(tmpDir, 'exit.cjs', FIXTURE_EXIT);
|
|
for (const invalid of [0, -5, NaN, 'abc', Infinity]) {
|
|
assert.throws(() => runNode([fixture, '0'], { timeoutMs: invalid }), TypeError);
|
|
}
|
|
});
|
|
|
|
test('input is delivered to the child on stdin', () => {
|
|
const fixture = writeFixture(tmpDir, 'echo-stdin.cjs', FIXTURE_ECHO_STDIN);
|
|
const result = runNode([fixture], { input: 'hello-stdin' });
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.deepStrictEqual(JSON.parse(result.stdout.trim()), {
|
|
received: 'hello-stdin',
|
|
hadData: true,
|
|
});
|
|
});
|
|
|
|
test('omitted input does not close stdin as empty string', () => {
|
|
const fixture = writeFixture(tmpDir, 'echo-stdin.cjs', FIXTURE_ECHO_STDIN);
|
|
const result = runNode([fixture]);
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.deepStrictEqual(JSON.parse(result.stdout.trim()), {
|
|
received: '',
|
|
hadData: false,
|
|
});
|
|
});
|
|
|
|
test('caller cannot override encoding into Buffers', () => {
|
|
const fixture = writeFixture(tmpDir, 'exit.cjs', FIXTURE_EXIT);
|
|
const result = runNode([fixture, '0', 'marker'], { encoding: 'buffer' });
|
|
assert.equal(typeof result.stdout, 'string');
|
|
assert.equal(Buffer.isBuffer(result.stdout), false);
|
|
});
|
|
|
|
test('argv metacharacters are not interpreted by a shell', () => {
|
|
const fixture = writeFixture(tmpDir, 'echo-argv.cjs', FIXTURE_ECHO_ARGV);
|
|
const hostileArgv = [';', '&&', '$(ls)', '`ls`', '| cat', '> /tmp/x'];
|
|
const result = runNode([fixture, ...hostileArgv]);
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.deepStrictEqual(JSON.parse(result.stdout.trim()).argv, hostileArgv);
|
|
});
|
|
|
|
test('flag-shaped argv values are not re-parsed', () => {
|
|
const fixture = writeFixture(tmpDir, 'echo-argv.cjs', FIXTURE_ECHO_ARGV);
|
|
const flagLikeArgv = ['--weird', '--timeoutMs=1', '-x'];
|
|
const result = runNode([fixture, ...flagLikeArgv]);
|
|
assert.deepStrictEqual(JSON.parse(result.stdout.trim()).argv, flagLikeArgv);
|
|
});
|
|
|
|
test('long and unicode argv survive the seam', () => {
|
|
const fixture = writeFixture(tmpDir, 'echo-argv.cjs', FIXTURE_ECHO_ARGV);
|
|
const longArgv = ['x'.repeat(5000), '日本語テスト', '🚀emoji🚀', 'café'];
|
|
const result = runNode([fixture, ...longArgv]);
|
|
assert.deepStrictEqual(JSON.parse(result.stdout.trim()).argv, longArgv);
|
|
});
|
|
|
|
test('a custom killSignal is reported, not normalized', () => {
|
|
const fixture = writeFixture(tmpDir, 'sleeper.cjs', FIXTURE_SLEEPER);
|
|
const result = runNode([fixture, '5000'], { timeoutMs: 200, killSignal: 'SIGINT' });
|
|
assert.equal(result.outcome, OUTCOME.TIMED_OUT);
|
|
if (process.platform !== 'win32') {
|
|
assert.equal(result.signal, 'SIGINT');
|
|
}
|
|
});
|
|
|
|
test('timeout with stderr reports one outcome, keeps both fields', () => {
|
|
const fixture = writeFixture(tmpDir, 'sleeper.cjs', FIXTURE_SLEEPER);
|
|
const result = runNode([fixture, '5000', '', 'err-marker'], { timeoutMs: 300 });
|
|
assert.equal(result.outcome, OUTCOME.TIMED_OUT);
|
|
assert.equal(typeof result.stdout, 'string');
|
|
assert.equal(typeof result.stderr, 'string');
|
|
// spawnSync preserves partial child output on a timeout on darwin, but discards
|
|
// it on Linux (verified on node 22 and 24). The seam passes through whatever
|
|
// spawnSync gives it, so the cross-platform contract is only that these are
|
|
// strings — the partial content itself is asserted where it is actually available.
|
|
if (process.platform === 'darwin') {
|
|
assert.ok(result.stderr.length > 0, 'darwin preserves partial stderr on timeout');
|
|
}
|
|
});
|
|
|
|
test('overflow is not masked by an exit code', () => {
|
|
const fixture = writeFixture(tmpDir, 'overflow.cjs', FIXTURE_OVERFLOW);
|
|
const result = runNode([fixture, '7']);
|
|
assert.equal(result.outcome, OUTCOME.BUFFER_OVERFLOW);
|
|
assert.equal(result.exitCode, null);
|
|
});
|
|
|
|
test('consecutive timeouts do not share state', () => {
|
|
const fixture = writeFixture(tmpDir, 'sleeper.cjs', FIXTURE_SLEEPER);
|
|
const first = runNode([fixture, '5000'], { timeoutMs: 200 });
|
|
const second = runNode([fixture, '5000'], { timeoutMs: 200 });
|
|
assert.equal(first.outcome, OUTCOME.TIMED_OUT);
|
|
assert.equal(second.outcome, OUTCOME.TIMED_OUT);
|
|
});
|
|
|
|
test('OUTCOME enum keys are locked', () => {
|
|
assert.equal(Object.isFrozen(OUTCOME), true);
|
|
assert.deepStrictEqual(Object.keys(OUTCOME).sort(), [
|
|
'BUFFER_OVERFLOW',
|
|
'EXITED',
|
|
'KILLED',
|
|
'SPAWN_FAILED',
|
|
'TIMED_OUT',
|
|
]);
|
|
assert.deepStrictEqual(OUTCOME, {
|
|
EXITED: 'exited',
|
|
KILLED: 'killed',
|
|
TIMED_OUT: 'timed_out',
|
|
BUFFER_OVERFLOW: 'buffer_overflow',
|
|
SPAWN_FAILED: 'spawn_failed',
|
|
});
|
|
assert.throws(() => {
|
|
OUTCOME.EXITED = 'nope';
|
|
}, TypeError);
|
|
});
|
|
|
|
test('a child killed by an external signal is KILLED, not EXITED', (t) => {
|
|
if (process.platform === 'win32') {
|
|
t.skip('signal semantics differ on win32 — see design doc row on Windows signals');
|
|
return;
|
|
}
|
|
const fixture = writeFixture(tmpDir, 'suicide.cjs', FIXTURE_SUICIDE);
|
|
const result = runNode([fixture], { timeoutMs: 5000 });
|
|
assert.equal(result.outcome, OUTCOME.KILLED);
|
|
assert.equal(result.killed, true);
|
|
assert.equal(result.timedOut, false);
|
|
assert.equal(result.signal, 'SIGKILL');
|
|
assert.equal(result.exitCode, null);
|
|
});
|
|
|
|
// Smoke coverage for runHook's export (matches the invocation shape used by
|
|
// tests/read-guard.test.cjs:34 and tests/workflow-guard.test.cjs:28 —
|
|
// process.execPath + [HOOK_PATH, ...args] with a JSON stdin payload).
|
|
test('runHook invokes a real hook via process.execPath', () => {
|
|
const hookPath = path.join(__dirname, '..', 'hooks', 'gsd-read-guard.js');
|
|
const payload = JSON.stringify({ tool_name: 'Read', tool_input: {} });
|
|
const env = {
|
|
...process.env,
|
|
CLAUDE_SESSION_ID: '',
|
|
CLAUDECODE: '',
|
|
CLAUDE_CODE_ENTRYPOINT: '',
|
|
CLAUDE_CODE_SSE_PORT: '',
|
|
CLAUDE_PROJECT_DIR: '',
|
|
};
|
|
const result = runHook(hookPath, [], { input: payload, env, timeoutMs: 5000 });
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.equal(typeof result.stdout, 'string');
|
|
});
|
|
});
|
|
|
|
describe('runHook interpreter option', () => {
|
|
let tmpDir;
|
|
|
|
beforeEach(() => {
|
|
tmpDir = createTempDir('process-seam-interpreter-');
|
|
});
|
|
|
|
afterEach(() => {
|
|
cleanup(tmpDir);
|
|
});
|
|
|
|
test('omitting interpreter still spawns via process.execPath', () => {
|
|
const hookPath = writeFixture(tmpDir, 'hook.cjs', FIXTURE_ECHO_ARGV);
|
|
const result = runHook(hookPath, ['a', 'b']);
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.deepStrictEqual(JSON.parse(result.stdout.trim()).argv, ['a', 'b']);
|
|
});
|
|
|
|
// bash availability is checked, never assumed — a Windows host or a
|
|
// node-only container may not have bash on PATH.
|
|
function isBashAvailable() {
|
|
if (process.platform === 'win32') return false;
|
|
const probeResult = runHook('-c', ['exit 0'], { interpreter: 'bash' });
|
|
return probeResult.outcome !== OUTCOME.SPAWN_FAILED;
|
|
}
|
|
|
|
const bashAvailable = isBashAvailable();
|
|
|
|
test('interpreter: bash runs a bash script and reports EXITED', (t) => {
|
|
if (!bashAvailable) {
|
|
t.skip('bash is not available on this host');
|
|
return;
|
|
}
|
|
const scriptPath = path.join(tmpDir, 'hook.sh');
|
|
fs.writeFileSync(
|
|
scriptPath,
|
|
[
|
|
'#!/usr/bin/env bash',
|
|
'echo \'{"ok":true}\'',
|
|
'exit 0',
|
|
].join('\n')
|
|
);
|
|
const result = runHook(scriptPath, [], { interpreter: 'bash' });
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.equal(result.exitCode, 0);
|
|
assert.deepStrictEqual(JSON.parse(result.stdout.trim()), { ok: true });
|
|
});
|
|
|
|
test('interpreter is not forwarded into spawnSync options', () => {
|
|
const hookPath = writeFixture(tmpDir, 'hook.cjs', FIXTURE_EXIT);
|
|
// If `interpreter` leaked into spawnOptions, spawnSync would receive an
|
|
// unexpected string-valued option alongside a valid timeoutMs; the
|
|
// seam's contract-validation for timeoutMs must still pass through
|
|
// untouched and the call must complete without throwing.
|
|
const result = runHook(hookPath, ['0'], { interpreter: process.execPath, timeoutMs: 5000 });
|
|
assert.equal(result.outcome, OUTCOME.EXITED);
|
|
assert.equal(result.exitCode, 0);
|
|
});
|
|
|
|
test('interpreter: bash on a script past timeoutMs is TIMED_OUT', (t) => {
|
|
if (!bashAvailable) {
|
|
t.skip('bash is not available on this host');
|
|
return;
|
|
}
|
|
const scriptPath = path.join(tmpDir, 'sleeper.sh');
|
|
fs.writeFileSync(
|
|
scriptPath,
|
|
[
|
|
'#!/usr/bin/env bash',
|
|
'sleep 5',
|
|
].join('\n')
|
|
);
|
|
const result = runHook(scriptPath, [], { interpreter: 'bash', timeoutMs: 300 });
|
|
assert.equal(result.outcome, OUTCOME.TIMED_OUT);
|
|
assert.equal(result.timedOut, true);
|
|
assert.equal(result.exitCode, null);
|
|
});
|
|
});
|
|
|
|
describe('runGsdTools adapter (process-seam parity)', () => {
|
|
let tmpDir;
|
|
|
|
beforeEach(() => {
|
|
tmpDir = createTempDir('process-seam-adapter-');
|
|
});
|
|
|
|
afterEach(() => {
|
|
cleanup(tmpDir);
|
|
});
|
|
|
|
test('adapter returns the legacy success shape', () => {
|
|
const result = runGsdTools(['--help'], tmpDir);
|
|
assert.equal(result.success, true);
|
|
assert.equal(result.exitCode, 0);
|
|
assert.equal(typeof result.output, 'string');
|
|
});
|
|
|
|
test('adapter returns the legacy failure shape', () => {
|
|
const result = runGsdTools(['this-is-not-a-real-command'], tmpDir);
|
|
assert.equal(result.success, false);
|
|
assert.equal(typeof result.error, 'string');
|
|
assert.ok(result.error.length > 0);
|
|
assert.equal(typeof result.exitCode, 'number');
|
|
assert.notEqual(result.exitCode, 0);
|
|
});
|
|
|
|
test('adapter reproduces the empty-stderr diagnostic verbatim', () => {
|
|
withMockedRunNode(
|
|
() => ({
|
|
outcome: OUTCOME.EXITED,
|
|
exitCode: 9,
|
|
stdout: '',
|
|
stderr: '',
|
|
timedOut: false,
|
|
signal: null,
|
|
killed: false,
|
|
code: null,
|
|
}),
|
|
() => {
|
|
const result = runGsdTools(['a', 'b'], tmpDir);
|
|
const expected = `Command failed: ${process.execPath} ${TOOLS_PATH} a b [stderr: (empty) exit:9]`;
|
|
assert.equal(result.success, false);
|
|
assert.equal(result.error, expected);
|
|
assert.equal(result.exitCode, 9);
|
|
}
|
|
);
|
|
});
|
|
|
|
test('adapter still retries once on kill', () => {
|
|
withMockedRunNode(
|
|
(callCount) => (callCount === 1
|
|
? {
|
|
outcome: OUTCOME.TIMED_OUT,
|
|
exitCode: null,
|
|
stdout: '',
|
|
stderr: '',
|
|
timedOut: true,
|
|
signal: 'SIGTERM',
|
|
killed: true,
|
|
code: 'ETIMEDOUT',
|
|
}
|
|
: {
|
|
outcome: OUTCOME.EXITED,
|
|
exitCode: 0,
|
|
stdout: 'ok\n',
|
|
stderr: '',
|
|
timedOut: false,
|
|
signal: null,
|
|
killed: false,
|
|
code: null,
|
|
}),
|
|
(getCallCount) => {
|
|
const result = runGsdTools(['x'], tmpDir);
|
|
assert.equal(result.success, true);
|
|
assert.equal(result.output, 'ok');
|
|
assert.equal(getCallCount(), 2);
|
|
}
|
|
);
|
|
});
|
|
|
|
test('adapter retries once on KILLED, mirroring TIMED_OUT', () => {
|
|
withMockedRunNode(
|
|
(callCount) => (callCount === 1
|
|
? {
|
|
outcome: OUTCOME.KILLED,
|
|
exitCode: null,
|
|
stdout: '',
|
|
stderr: '',
|
|
timedOut: false,
|
|
signal: 'SIGKILL',
|
|
killed: true,
|
|
code: null,
|
|
}
|
|
: {
|
|
outcome: OUTCOME.EXITED,
|
|
exitCode: 0,
|
|
stdout: 'ok\n',
|
|
stderr: '',
|
|
timedOut: false,
|
|
signal: null,
|
|
killed: false,
|
|
code: null,
|
|
}),
|
|
(getCallCount) => {
|
|
const result = runGsdTools(['x'], tmpDir);
|
|
assert.equal(result.success, true);
|
|
assert.equal(result.output, 'ok');
|
|
assert.equal(getCallCount(), 2);
|
|
}
|
|
);
|
|
});
|
|
|
|
test('adapter throws resource-starvation when KILLED persists after retry', () => {
|
|
withMockedRunNode(
|
|
() => ({
|
|
outcome: OUTCOME.KILLED,
|
|
exitCode: null,
|
|
stdout: 'partial-out',
|
|
stderr: 'partial-err',
|
|
timedOut: false,
|
|
signal: 'SIGKILL',
|
|
killed: true,
|
|
code: null,
|
|
}),
|
|
(getCallCount) => {
|
|
const expected =
|
|
`[runGsdTools: resource-starvation / subprocess-kill after retry] ` +
|
|
`gsd-tools was killed before completion ` +
|
|
`(signal=SIGKILL, code=null, killed=true). ` +
|
|
`This indicates host OOM or scheduler contention, not a product bug. ` +
|
|
`stdout=partial-out stderr=partial-err`;
|
|
assert.throws(
|
|
() => runGsdTools(['x'], tmpDir),
|
|
(err) => err instanceof Error && err.message === expected
|
|
);
|
|
assert.equal(getCallCount(), 2);
|
|
}
|
|
);
|
|
});
|
|
|
|
test('adapter still throws resource-starvation after retry', () => {
|
|
withMockedRunNode(
|
|
() => ({
|
|
outcome: OUTCOME.TIMED_OUT,
|
|
exitCode: null,
|
|
stdout: 'partial-out',
|
|
stderr: 'partial-err',
|
|
timedOut: true,
|
|
signal: 'SIGTERM',
|
|
killed: true,
|
|
code: 'ETIMEDOUT',
|
|
}),
|
|
() => {
|
|
const expected =
|
|
`[runGsdTools: resource-starvation / subprocess-kill after retry] ` +
|
|
`gsd-tools was killed before completion ` +
|
|
`(signal=SIGTERM, code=ETIMEDOUT, killed=true). ` +
|
|
`This indicates host OOM or scheduler contention, not a product bug. ` +
|
|
`stdout=partial-out stderr=partial-err`;
|
|
assert.throws(
|
|
() => runGsdTools(['x'], tmpDir),
|
|
(err) => err instanceof Error && err.message === expected
|
|
);
|
|
}
|
|
);
|
|
});
|
|
|
|
test('adapter retries exactly once, not twice', () => {
|
|
withMockedRunNode(
|
|
() => ({
|
|
outcome: OUTCOME.TIMED_OUT,
|
|
exitCode: null,
|
|
stdout: '',
|
|
stderr: '',
|
|
timedOut: true,
|
|
signal: 'SIGTERM',
|
|
killed: true,
|
|
code: 'ETIMEDOUT',
|
|
}),
|
|
(getCallCount) => {
|
|
assert.throws(() => runGsdTools(['x'], tmpDir));
|
|
assert.equal(getCallCount(), 2);
|
|
}
|
|
);
|
|
});
|
|
|
|
test('adapter still accepts a shell-style string', () => {
|
|
const result = runGsdTools('--help', tmpDir);
|
|
assert.equal(result.success, true);
|
|
assert.equal(result.exitCode, 0);
|
|
});
|
|
|
|
test('adapter still accepts an argv array', () => {
|
|
const result = runGsdTools(['--help'], tmpDir);
|
|
assert.equal(result.success, true);
|
|
assert.equal(result.exitCode, 0);
|
|
});
|
|
|
|
test('adapter preserves env override precedence', () => {
|
|
withMockedRunNode(
|
|
(_callCount, _args, options) => {
|
|
// Capture the merged env the adapter built, then return a fast
|
|
// EXITED result — this test is about the merge, not gsd-tools.
|
|
withMockedRunNode.capturedEnv = options.env;
|
|
return {
|
|
outcome: OUTCOME.EXITED,
|
|
exitCode: 0,
|
|
stdout: '',
|
|
stderr: '',
|
|
timedOut: false,
|
|
signal: null,
|
|
killed: false,
|
|
code: null,
|
|
};
|
|
},
|
|
() => {
|
|
runGsdTools(['x'], tmpDir, { GSD_SESSION_KEY: 'override-value' });
|
|
// TEST_ENV_BASE defaults GSD_SESSION_KEY to '' — the caller's env
|
|
// argument must win over it.
|
|
assert.equal(withMockedRunNode.capturedEnv.GSD_SESSION_KEY, 'override-value');
|
|
}
|
|
);
|
|
});
|
|
|
|
// Legacy-shape contract: the SEAM reports exitCode: null for
|
|
// BUFFER_OVERFLOW (asserted above, at the seam level), but a real caller
|
|
// (tests/context-predicates-query.test.cjs) asserts
|
|
// `typeof r.exitCode === 'number'` on the ADAPTER's legacy shape, matching
|
|
// the pre-seam execFileSync helper's `err.status ?? 1`. This is the
|
|
// adapter-level contract, deliberately distinct from the seam-level one.
|
|
test('adapter does not retry a buffer overflow, and coerces exitCode to a number', () => {
|
|
withMockedRunNode(
|
|
() => ({
|
|
outcome: OUTCOME.BUFFER_OVERFLOW,
|
|
exitCode: null,
|
|
stdout: 'x'.repeat(20),
|
|
stderr: '',
|
|
timedOut: false,
|
|
signal: 'SIGTERM',
|
|
killed: true,
|
|
code: 'ENOBUFS',
|
|
}),
|
|
(getCallCount) => {
|
|
const result = runGsdTools(['x'], tmpDir);
|
|
assert.equal(getCallCount(), 1);
|
|
assert.equal(result.success, false);
|
|
assert.equal(typeof result.exitCode, 'number');
|
|
assert.equal(result.exitCode, 1);
|
|
}
|
|
);
|
|
});
|
|
|
|
// Same legacy-shape contract as the buffer-overflow test above, for
|
|
// SPAWN_FAILED (the Windows CI regression case: an oversized argv).
|
|
test('adapter does not retry a spawn failure, and coerces exitCode to a number', () => {
|
|
withMockedRunNode(
|
|
() => ({
|
|
outcome: OUTCOME.SPAWN_FAILED,
|
|
exitCode: null,
|
|
stdout: '',
|
|
stderr: '',
|
|
timedOut: false,
|
|
signal: null,
|
|
killed: false,
|
|
code: 'ENOENT',
|
|
}),
|
|
(getCallCount) => {
|
|
const result = runGsdTools(['x'], tmpDir);
|
|
assert.equal(getCallCount(), 1);
|
|
assert.equal(result.success, false);
|
|
assert.equal(typeof result.exitCode, 'number');
|
|
assert.equal(result.exitCode, 1);
|
|
}
|
|
);
|
|
});
|
|
});
|
|
|
|
describe('#3271: hook fan-out timeout class', () => {
|
|
const {
|
|
PROBE_TIMEOUT_MS: PROBE,
|
|
HOOK_FANOUT_TIMEOUT_MS: HOOK_FANOUT,
|
|
INSTALL_TIMEOUT_MS: INSTALL,
|
|
} = require('./helpers/timeouts.cjs');
|
|
|
|
test('a hook fan-out is bounded above a bare probe and below a full install', () => {
|
|
// The ordering IS the claim: a hook that shells out several times is heavier
|
|
// than reading back a version string and lighter than running bin/install.js.
|
|
// CI recorded a Windows timeout at exactly the probe bound (PR #3285,
|
|
// windows-latest node 22 shard 2/3) while every other lane passed the same
|
|
// commit — the bound was sized for the wrong class.
|
|
assert.ok(PROBE < HOOK_FANOUT, `probe ${PROBE}ms must be under hook fan-out ${HOOK_FANOUT}ms`);
|
|
assert.ok(HOOK_FANOUT < INSTALL, `hook fan-out ${HOOK_FANOUT}ms must be under install ${INSTALL}ms`);
|
|
});
|
|
|
|
test('the fan-out bound clears the duration that actually timed out', () => {
|
|
// Observed: 15040ms, censored at the 15000ms probe bound, so the real need is
|
|
// unknown and above it. A bound that merely matched the observation would be
|
|
// the same defect again.
|
|
const OBSERVED_TIMEOUT_MS = 15040;
|
|
assert.ok(
|
|
HOOK_FANOUT >= OBSERVED_TIMEOUT_MS * 3,
|
|
`hook fan-out ${HOOK_FANOUT}ms must clear the censored ${OBSERVED_TIMEOUT_MS}ms observation with real margin`,
|
|
);
|
|
});
|
|
});
|