* refactor(#3180): one owner for completion ratio, a prompt-layer drift guard, and a written behavior contract The 2026-08-08 coverage audit on #3180 found the epic's copy counts were a lower bound for the third consecutive time, and that two derivation families had never been named at all. ADR-3180 gains Decision 7 — a normative behavior contract that says what the right answer IS for each derivation, not merely who owns it. A reviewer with no written rule can only ask "does this look like the others", which is how a fifth copy passes review. Decision 4 gains (d) scan surface is every authored surface and an owner FILE is never exempt, only its named functions; and (e) a surface that cannot be consolidated today ships ratcheted, never unguarded. Completion ratio: `clampPercent` sat exported and unused beside six hand-inlined copies of its own body across five modules. All six now route through it; `clampPercentFromFraction` is added for the one caller that already held a fraction. Every migration is behaviour-identical — clampPercent's first line IS the `total > 0 ? … : 0` ternary each copy carried. Guarded by lint-completion-ratio-drift.cjs, which reports zero re-derivations with no file-level exemption. Prompt layer: workflow markdown re-derives live-plan counting in raw shell (#1762), invisible to every `src/`-scoped guard. lint-planning-prompt-drift.cjs scans it with a shrink-only baseline of the 7 sites that exist today — new sites fail, and a baseline entry that stops firing fails too, so an acknowledgment can never outlive the thing it describes. lint-milestone-window-drift.cjs stops exempting its owner file wholesale; only the four named canonical functions are exempt now. The blanket exemption was pointed at the one file most likely to grow the next copy, and it had. Refs #3180 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3180): link Phases 6-8 sub-issues (#3216, #3217, #3218) from ADR-3180 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3180): address orthogonal review — consumer-output identity tests, count-keyed ratchet, property coverage Five findings from the two orthogonal review passes, all fixed. Decision 4(c) breach: the completion-ratio identity test asserted at the OWNER, which is exactly the bypass that decision exists to close — a consumer can call clampPercent and then post-process locally, leaving both the lint and an owner-level test green. It now drives `roadmap analyze`, `query progress` and `stats` and asserts on their own output, over a fixture containing a `status: superseded` plan so a consumer that re-counted raw files would report 60 where the owner reports 75. Decision 4(e) breach: ratchet entries named the epic (#3180) rather than the issue that removes them. They name Phase 8 (#3218) now. The ratchet keyed on (file, text) alone, so plan-phase.md's two byte-identical sites were one indistinguishable key and migrating either would have left the guard green with the other alive. Entries carry an occurrence count; fewer than acknowledged fails as a partial migration, more fails as a new copy. Adds the missing MAX_REGEX_LITERAL_LEN boundary coverage the sibling guard's test already had, and the fast-check property tests CONTRIBUTING requires for clamp/budget-limit functions. Refs #3180 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: stop wrapping a nested double-spawn in a 15s wall-clock budget (bug #641 probes) `tests/ci-test-scope.test.cjs`'s `bug #641` block spawned `run-tests.cjs` under PROBE_TIMEOUT_MS=15000; that child then spawned a nested `node --test`. A fixed wall-clock budget around a double spawn, running inside a container that is concurrently executing the full ~31k-test suite, fails by construction under load. Confirmed against three full matrix runs. Every failure was shaped `null !== 0` — the child was KILLED, never an assertion about the thing under test. One captured probe had already printed the correct resolution (`suite="all" files=2: a.test.cjs b.test.cjs`) and was killed anyway. It reproduces on `next` alone: 5 failures on linux-node22, 0 on linux-node24. The victim subset varies by run and by lane. What these tests are actually about is suite-token RESOLUTION — `unit` as a bare token in --files/--files-from. Executing the seeded trivial files is incidental and is the entire timeout surface, so the assertions move in-process against the same functions `main()` calls, in the same order. `parseArgs`, `selectExplicitFiles`, `selectFiles` and `walkTestFiles` are exported for that; no behavior, signature or logic changed. No coverage lost: `tests/run-tests-harness.test.cjs` already spawns the harness for real and asserts exit codes end to end, on a 120s budget. Pre-existing on `next`, fixed here rather than deferred. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: delete the three elapsed-time assertions CLAUDE.md forbids asserting on wall-clock time. Three assertions did, and all three are load-sensitive: on a saturated bench each can fail while the code under test is correct. In every case the load-bearing assertion sits on the line above and the timing line adds no discrimination. run-with-timeout: the stated worry — "was this 124 the cap firing or the 30s harness backstop?" — is already answered by the assertion above it. A backstop kills by signal, which surfaces as status null, never 124. Observed directly this session: three matrix runs produced exactly that null shape from killed children. normalize-test-command and context-predicates: both bounded a ReDoS check. A threshold only ever separates "fast" from "slightly slow", which is bench load, not correctness — catastrophic backtracking on 800 KB of input does not take 251ms, it does not finish at all. A real regression therefore shows up as the suite being killed on that test, which is louder and more reliable than a number. The structural assertions (returned unchanged; cleanly rejected) are what actually carry those tests, and they stay. The sweep now reports zero elapsed-time assertions in tests/. The remaining Date.now() uses are unique-path suffixes, barrier deadlines, fixture timestamps and fake mtimes — none of them assertions. Pre-existing on `next`, fixed here rather than deferred. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3180): backfill changeset PR number (#3223) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3180): key the prompt-drift ratchet on POSIX paths so it works on Windows The baseline keys on (file, trimmed text). `file` came from scanTree's `path.relative()`, which uses NATIVE separators, while the committed baseline stores POSIX. On Windows every violation was therefore unmatched — reported as FRESH — and every baseline entry matched nothing — reported as STALE. The guard failed 100% of the time there, on both CI shards: ✖ scanRepo(repoRoot) matches the baseline exactly: zero fresh AND zero stale + { file: 'gsd-core\\workflows\\execute-plan.md', ... } The remote runner this repo gates on is Linux-only and cannot see this class at all; the GitHub Actions Windows lane is what caught it. Normalization is unconditional — never gated on process.platform. A platform-conditional normalizer makes the POSIX path the special case and leaves the Windows branch unexercised on every other OS, which is the same blind spot in a different place. It is applied at one seam inside findPromptDrift, which builds `file` on every returned violation, so the baseline key, the --update writer, the stderr report and the tests all consume one normalized value. The regression tests drive a Windows-shaped relPath directly and run on every OS rather than skipping off-Windows — a test that only runs on the platform where the bug lives is why this escaped. They include a sanity check that un-normalized input does NOT match, so the assertion cannot pass vacuously. Audited the three sibling guards: none keys against a committed cross-platform baseline, and their exemption keys are path.join-built, so producer and consumer share the native convention. Left correct code alone rather than making them look alike. scripts/lib/drift-scan.cjs is untouched — normalizing there would break those three on Windows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
346 lines
17 KiB
JavaScript
346 lines
17 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* #2351 — `gsd_run run-with-timeout`: a portable, coreutils-independent
|
|
* wall-clock cap for a spawned command, plus the parity guard that keeps
|
|
* hardcoded GNU `timeout` from reappearing in workflow/agent/reference markdown.
|
|
*
|
|
* Root cause it fixes: workflow gates hardcoded `timeout <n> <cmd>`. `timeout`
|
|
* is GNU coreutils; stock macOS ships neither it nor `gtimeout`, so the call
|
|
* exited 127 ("command not found") and a passing build/test was misreported as
|
|
* a FAILURE. The verb replaces every such call with a Node-based cap that keeps
|
|
* GNU `timeout`'s exit-code contract (124 on timeout) on every platform.
|
|
*
|
|
* These are behavioral tests driven through the real CLI entrypoint
|
|
* (spawnSync of gsd-tools.cjs), never source-text assertions. The parity block
|
|
* consumes the lint module's typed findings, not grepped text.
|
|
*/
|
|
|
|
const { describe, test } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const { spawnSync } = require('node:child_process');
|
|
const os = require('node:os');
|
|
const path = require('node:path');
|
|
const fs = require('node:fs');
|
|
const { setTimeout: sleep } = require('node:timers/promises');
|
|
const { createTempDir, cleanup } = require('./helpers.cjs');
|
|
|
|
const ROOT = path.join(__dirname, '..');
|
|
const GSD_TOOLS = path.join(ROOT, 'gsd-core', 'bin', 'gsd-tools.cjs');
|
|
const NODE = process.execPath;
|
|
|
|
// run-with-timeout intercepts before gsd-tools' cwd/workstream resolution, so it
|
|
// needs no project fixture — run from a neutral temp dir to prove independence.
|
|
function runVerb(args, opts = {}) {
|
|
return spawnSync(NODE, [GSD_TOOLS, 'run-with-timeout', ...args], {
|
|
cwd: os.tmpdir(),
|
|
encoding: 'utf8',
|
|
timeout: 30000, // test-harness backstop; the verb's own cap is what we assert
|
|
...opts,
|
|
});
|
|
}
|
|
|
|
// A guaranteed-hanging child, cross-platform (no reliance on `sleep`).
|
|
const HANG = [NODE, '-e', 'setTimeout(() => {}, 60000)'];
|
|
// A guaranteed-fast child.
|
|
const OK = [NODE, '-e', 'process.exit(0)'];
|
|
|
|
describe('#2351 run-with-timeout — exit-code contract', () => {
|
|
test('passes a fast zero-exit command through as exit 0', () => {
|
|
const r = runVerb(['5', '--', ...OK]);
|
|
assert.equal(r.status, 0);
|
|
});
|
|
|
|
test('passes a non-zero exit code through unchanged', () => {
|
|
const r = runVerb(['5', '--', NODE, '-e', 'process.exit(7)']);
|
|
assert.equal(r.status, 7);
|
|
});
|
|
|
|
test('exits 124 when the wall-clock budget is exceeded (matches GNU timeout)', () => {
|
|
const r = runVerb(['1', '--', ...HANG]);
|
|
// exit 124 is itself the discriminator: a harness backstop kill surfaces
|
|
// as status === null (signal), never as 124 — so no elapsed-time
|
|
// assertion is needed or allowed here.
|
|
assert.equal(r.status, 124, 'a timed-out command must exit 124');
|
|
});
|
|
|
|
test('exits 127 when the command is not found (matches GNU timeout)', () => {
|
|
const r = runVerb(['5', '--', 'this-command-does-not-exist-2351']);
|
|
assert.equal(r.status, 127);
|
|
});
|
|
|
|
test('runs without a timer when <seconds> is 0 — a slow child is NOT killed', () => {
|
|
// Proves 0 = no timer (not "timer fired at 0ms"): a child that outlives any
|
|
// mis-armed timer must still exit 0. A wrongly-armed 0ms timer would give 124.
|
|
const r = runVerb(['0', '--', NODE, '-e', 'setTimeout(() => process.exit(0), 1500)']);
|
|
assert.equal(r.status, 0, '<seconds> 0 must run untimed');
|
|
});
|
|
|
|
test('a budget past the 32-bit setTimeout ceiling does not spuriously time out', () => {
|
|
// secs*1000 > 2**31-1 → Node clamps setTimeout to 1ms → an immediate false 124
|
|
// unless the delay is capped. The fast child must still exit 0, with no warning.
|
|
const r = runVerb(['3000000', '--', ...OK]);
|
|
assert.equal(r.status, 0, 'oversized budget must not fire an immediate timeout');
|
|
assert.doesNotMatch(r.stderr || '', /TimeoutOverflowWarning/, 'delay must be clamped');
|
|
});
|
|
});
|
|
|
|
describe('#2351 run-with-timeout — argument handling (negative matrix)', () => {
|
|
test('missing <seconds> is a usage error (exit 2), not a crash', () => {
|
|
const r = runVerb([]);
|
|
assert.equal(r.status, 2);
|
|
assert.match(r.stderr, /missing <seconds>/);
|
|
assert.doesNotMatch(r.stderr, /at Object|at Module|\.cjs:\d+/, 'no stack trace in usage error');
|
|
});
|
|
|
|
test('non-numeric <seconds> is a usage error (exit 2)', () => {
|
|
const r = runVerb(['not-a-number', '--', ...OK]);
|
|
assert.equal(r.status, 2);
|
|
assert.match(r.stderr, /invalid <seconds>/);
|
|
});
|
|
|
|
test('blank / whitespace <seconds> is a usage error — never a silent unbounded run', () => {
|
|
for (const blank of ['', ' ']) {
|
|
const r = runVerb([blank, '--', ...OK]);
|
|
assert.equal(r.status, 2, `blank seconds ${JSON.stringify(blank)} must error, not disable the timer`);
|
|
assert.match(r.stderr, /invalid <seconds>/);
|
|
}
|
|
});
|
|
|
|
test('missing <command> is a usage error (exit 2)', () => {
|
|
const r = runVerb(['5']);
|
|
assert.equal(r.status, 2);
|
|
assert.match(r.stderr, /missing <command>/);
|
|
});
|
|
|
|
test('a trailing `s` on the duration is accepted (GNU-style unit)', () => {
|
|
const r = runVerb(['5s', '--', ...OK]);
|
|
assert.equal(r.status, 0);
|
|
});
|
|
|
|
test('the `--` separator is optional', () => {
|
|
const r = runVerb(['5', ...OK]);
|
|
assert.equal(r.status, 0);
|
|
});
|
|
|
|
test("the wrapped command's argv is opaque — gsd-tools flags are NOT consumed", (t) => {
|
|
// --raw / --cwd / --pick are gsd-tools' own global flags. They must reach the
|
|
// wrapped command verbatim, not be stripped by the dispatcher. Use a script
|
|
// FILE, not `node -e` — node parses leading --flags after -e as its OWN options
|
|
// ("bad option", exit 9); after a script path it treats them as argv.
|
|
const dir = createTempDir('rwt-argv');
|
|
t.after(() => cleanup(dir));
|
|
const script = path.join(dir, 'argcheck.js');
|
|
fs.writeFileSync(script,
|
|
'process.exit(process.argv.slice(2).join(",") === "--raw,--cwd,x,--pick,y" ? 0 : 3);');
|
|
const r = runVerb(['5', '--', NODE, script, '--raw', '--cwd', 'x', '--pick', 'y']);
|
|
assert.equal(r.status, 0, 'wrapped --raw/--cwd/--pick must be passed through untouched');
|
|
});
|
|
|
|
test('the `query` meta-prefix form is accepted', () => {
|
|
const r = spawnSync(NODE, [GSD_TOOLS, 'query', 'run-with-timeout', '5', '--', ...OK], {
|
|
cwd: os.tmpdir(), encoding: 'utf8', timeout: 30000,
|
|
});
|
|
assert.equal(r.status, 0);
|
|
});
|
|
});
|
|
|
|
describe('#2351 run-with-timeout — kill semantics (POSIX process groups)', () => {
|
|
const posix = process.platform !== 'win32';
|
|
|
|
test('a timed-out command whose descendant traps SIGTERM is still reaped — no orphan/hang (C1)', { skip: !posix }, async (t) => {
|
|
// The HIGH-severity regression: the direct child exits on SIGTERM fast, but a
|
|
// descendant that IGNORES SIGTERM survives holding the inherited stdio. The
|
|
// whole process group must be SIGKILL-reaped, or a captured/piped gate hangs
|
|
// on the orphan. A node parent spawns the trapping child in the SAME process
|
|
// group (spawn WITHOUT `detached` → the child inherits the parent's pgid,
|
|
// deterministically, on every platform). We deliberately avoid `bash … &`:
|
|
// bash job-control can move a backgrounded job into its own process group,
|
|
// which no group-kill (nor GNU `timeout`) can reach. The parent exits fast on
|
|
// SIGTERM; the child ignores it and must still be reaped.
|
|
//
|
|
// Liveness is detected by a HEARTBEAT the child rewrites every 100ms — NOT
|
|
// `kill(pid, 0)`: a SIGKILL'd orphan lingers as a zombie until reaped, and a
|
|
// container's PID 1 reaps slowly, so `kill(pid,0)` reads a dead child as
|
|
// "alive". A reaped child stops ticking; a genuine orphan keeps ticking.
|
|
//
|
|
// The child writes its FIRST heartbeat synchronously at startup, before
|
|
// arming the interval, and the kill window is 3s rather than 1s. Both are
|
|
// load-independence requirements, not cosmetics: with the first write
|
|
// deferred to the interval's initial 100ms tick inside a 1s window, a loaded
|
|
// CI container can group-kill before that tick ever lands, and the existence
|
|
// check below then fails for a reason that has nothing to do with reaping —
|
|
// the behavior under test is the FREEZE assertion further down, which is
|
|
// unaffected by writing one extra sample at t=0.
|
|
const dir = createTempDir('rwt-c1');
|
|
t.after(() => cleanup(dir));
|
|
const parentFile = path.join(dir, 'parent.js');
|
|
const hbFile = path.join(dir, 'heartbeat');
|
|
fs.writeFileSync(parentFile, [
|
|
'const cp = require("child_process");',
|
|
"const childCode = 'const fs=require(\"fs\");const hb=process.argv[1];const tick=()=>fs.writeFileSync(hb,String(Date.now()));process.on(\"SIGTERM\",()=>{});tick();setInterval(tick,100);';",
|
|
'cp.spawn(process.execPath, ["-e", childCode, process.argv[2]], { stdio: "ignore" });',
|
|
'process.on("SIGTERM", () => process.exit(0));',
|
|
'setInterval(() => {}, 1000);',
|
|
].join('\n'));
|
|
const r = runVerb(['3', '--', NODE, parentFile, hbFile], { timeout: 20000 });
|
|
assert.equal(r.status, 124, 'must report a timeout (124), not hang');
|
|
assert.ok(fs.existsSync(hbFile), 'child heartbeat should exist');
|
|
await sleep(300); // let any in-flight write settle after the SIGKILL
|
|
const first = fs.readFileSync(hbFile, 'utf8');
|
|
await sleep(600); // >> the 100ms heartbeat interval
|
|
const second = fs.readFileSync(hbFile, 'utf8');
|
|
assert.equal(second, first, 'descendant must be reaped (heartbeat frozen), not orphaned and still ticking');
|
|
});
|
|
|
|
test('a command killed by a signal exits 128+signum (bash convention)', { skip: !posix }, () => {
|
|
const r = runVerb(['10', '--', 'bash', '-c', 'kill -TERM $$']);
|
|
assert.equal(r.status, 143, 'self-SIGTERM (15) → 128+15 = 143');
|
|
});
|
|
});
|
|
|
|
describe('#2667 run-with-timeout — Windows .cmd/.bat/.exe spawn mediation (CVE-2024-27980)', () => {
|
|
// Node's CVE-2024-27980 hardening throws EINVAL when child_process.spawn is
|
|
// given a .cmd/.bat without a shell. run-with-timeout now mediates .cmd/.bat
|
|
// on win32 via an explicit `cmd.exe /d /s /c <cmd> ...args` argv ARRAY (not
|
|
// shell:true — that space-joins unescaped args per DEP0190), while leaving
|
|
// every `bash`/argv-array caller unchanged (the recorded no-shell-for-argv-
|
|
// array contract). .exe is INTENTIONALLY excluded — real PEs spawn fine
|
|
// directly and mediating them breaks the timeout reap + risks arg mis-parse.
|
|
const isWin = process.platform === 'win32';
|
|
|
|
test('win32 RED: a .cmd shim runs (exit 0, non-empty stdout) — pre-fix this threw EINVAL → exit 125 / empty stdout', { skip: !isWin ? 'win32-only' : false }, () => {
|
|
const dir = createTempDir('rwt-2667-cmd');
|
|
try {
|
|
// A .cmd shim that echoes JSON to stdout (mimics fallow.cmd audit --format json).
|
|
const shim = path.join(dir, 'fake.cmd');
|
|
fs.writeFileSync(shim, '@echo {"verdict":"clean"}\r\n', 'utf8');
|
|
const r = runVerb(['10', '--', shim]);
|
|
assert.equal(r.status, 0, `expected the .cmd shim to run (exit 0); got ${r.status}. stderr: ${r.stderr}`);
|
|
assert.ok((r.stdout || '').includes('clean'), `expected non-empty JSON stdout from the .cmd shim; got: ${r.stdout}`);
|
|
} finally {
|
|
cleanup(dir);
|
|
}
|
|
});
|
|
|
|
test('win32: a .bat shim is also mediated (exit 0, non-empty stdout)', { skip: !isWin ? 'win32-only' : false }, () => {
|
|
const dir = createTempDir('rwt-2667-bat');
|
|
try {
|
|
const shim = path.join(dir, 'fake.bat');
|
|
fs.writeFileSync(shim, '@echo {"verdict":"clean"}\r\n', 'utf8');
|
|
const r = runVerb(['10', '--', shim]);
|
|
assert.equal(r.status, 0, `expected the .bat shim to run (exit 0); got ${r.status}. stderr: ${r.stderr}`);
|
|
assert.ok((r.stdout || '').length > 0, 'expected non-empty stdout from the .bat shim');
|
|
} finally {
|
|
cleanup(dir);
|
|
}
|
|
});
|
|
|
|
test('win32 negative-space: a .exe (node.exe) is spawned DIRECTLY, not mediated — no cmd.exe wrap', { skip: !isWin ? 'win32-only' : false }, () => {
|
|
// .exe is intentionally excluded from the gate: real PE executables spawn
|
|
// fine directly, and wrapping them in cmd.exe /c breaks the timeout cap's
|
|
// process-group reap AND risks cmd.exe mis-parsing args (e.g. -e "code()").
|
|
// node.exe -e "process.exit(0)" must exit 0 directly.
|
|
const r = runVerb(['10', '--', process.execPath, '-e', 'process.exit(0)']);
|
|
assert.equal(r.status, 0, `expected node.exe to run directly (exit 0); got ${r.status}. stderr: ${r.stderr}`);
|
|
});
|
|
|
|
test('POSIX negative-space: a bash -c caller is unchanged (no shell:true added) — argv stays array-only', { skip: isWin ? 'posix-only' : false }, () => {
|
|
// The fix's gate (win32 && .cmd/.bat) skips `bash` on POSIX: behavior
|
|
// must be identical to before. `bash -c 'echo ok'` exits 0 with stdout "ok".
|
|
const r = runVerb(['10', '--', 'bash', '-c', 'echo ok']);
|
|
assert.equal(r.status, 0, `expected bash caller to still work (exit 0); got ${r.status}`);
|
|
assert.equal((r.stdout || '').trim(), 'ok', 'expected stdout "ok" from the unchanged bash caller');
|
|
});
|
|
});
|
|
|
|
describe('#2351 run-with-timeout — coreutils independence (the regression)', () => {
|
|
// The whole point: no dependency on GNU `timeout`/`gtimeout`. Prove it by
|
|
// scrubbing PATH so neither could be found, and driving the child by absolute
|
|
// path. Before #2351 the gates called `timeout …` directly and exited 127 here.
|
|
const scrubbedEnv = { ...process.env, PATH: '' };
|
|
|
|
test('a real zero-exit command passes even with an empty PATH (no coreutils)', () => {
|
|
const r = runVerb(['5', '--', ...OK], { env: scrubbedEnv });
|
|
assert.equal(r.status, 0, 'must pass (exit 0), not 127, when coreutils is absent');
|
|
});
|
|
|
|
test('a genuine timeout is still detected (exit 124) with an empty PATH', () => {
|
|
const r = runVerb(['1', '--', ...HANG], { env: scrubbedEnv });
|
|
assert.equal(r.status, 124);
|
|
});
|
|
});
|
|
|
|
describe('#2351 parity guard — no hardcoded timeout in workflow/agent/reference/command md', () => {
|
|
const { findRawTimeoutInvocations, scan, DEFAULT_ROOTS } = require('../scripts/lint-portable-timeout.cjs');
|
|
|
|
// Decoy fixtures sourced from the issue report (an author independent of the
|
|
// detector), per the fixture-provenance rule (#2371): the exact bug forms the
|
|
// guard must catch.
|
|
const BUG_FORMS = [
|
|
'timeout 300 bash -c "$BUILD_CMD" 2>&1',
|
|
'timeout "$TEST_GATE_TIMEOUT" bash -c "$TEST_CMD" 2>&1',
|
|
'timeout "$TEST_GATE_TIMEOUT" bash -c "$AUDIT_TEST_CMD" 2>&1 | tail -20',
|
|
'echo "$TASK_PROMPT" | timeout "${CROSS_AI_TIMEOUT}s" ${CROSS_AI_CMD} > out 2>err',
|
|
'timeout 120 "$FALLOW_BIN" audit --format json --quiet',
|
|
'REVIEW_OUTPUT=$(echo "$X" | timeout 120 ${REVIEW_CMD} 2>/tmp/e.log)',
|
|
'timeout 30s bash "$probe"',
|
|
'gtimeout 60 bash -c "x"',
|
|
// GNU long options / no-space short opt / arithmetic budget (review #2351 C3)
|
|
'timeout --kill-after=5 30 bash -c x',
|
|
'timeout --foreground 30 bash -c x',
|
|
'timeout --signal=KILL 30 bash -c x',
|
|
'timeout -k5 30 bash -c x',
|
|
'timeout $((60*5)) bash -c x',
|
|
'cmd && timeout 30 bash x',
|
|
];
|
|
|
|
// Forms that must NEVER be flagged: the approved verb, portable capability
|
|
// probes, config keys, the agy flag, prose, and variable names.
|
|
const CLEAN_FORMS = [
|
|
'gsd_run run-with-timeout 300 -- bash -c "$BUILD_CMD"',
|
|
'echo "$X" | gsd_run run-with-timeout "${CROSS_AI_TIMEOUT}" -- ${CROSS_AI_CMD}',
|
|
'_AGY_KILLER="$(command -v timeout 2>/dev/null || command -v gtimeout 2>/dev/null || true)"',
|
|
'"$_AGY_KILLER" 600 agy --print-timeout 540s "$@" -p "$PROMPT"',
|
|
'TEST_GATE_TIMEOUT=$(gsd_run query config-get workflow.test_gate_timeout || echo "600")',
|
|
'echo "⚠ test gate timed out after ${TEST_GATE_TIMEOUT}s"',
|
|
'# bound the build with a 5-minute timeout',
|
|
'const TIMEOUT = 300;',
|
|
// prose mentions of "timeout <n>" mid-sentence must not be flagged (review #2351 C4)
|
|
'increase the timeout 30 seconds if the runner is slow',
|
|
'# we replaced timeout 300 with the run-with-timeout verb',
|
|
'sometimeout 30 is not the timeout command',
|
|
];
|
|
|
|
for (const form of BUG_FORMS) {
|
|
test(`flags a bare timeout invocation: ${form.slice(0, 42)}…`, () => {
|
|
const findings = findRawTimeoutInvocations(form);
|
|
assert.equal(findings.length, 1, `should flag: ${form}`);
|
|
assert.equal(findings[0].line, 1);
|
|
});
|
|
}
|
|
|
|
for (const form of CLEAN_FORMS) {
|
|
test(`does not flag a portable/unrelated form: ${form.slice(0, 42)}…`, () => {
|
|
assert.deepEqual(findRawTimeoutInvocations(form), [], `should NOT flag: ${form}`);
|
|
});
|
|
}
|
|
|
|
test('multi-line input reports the correct line numbers', () => {
|
|
const text = ['clean line', 'gsd_run run-with-timeout 5 -- true', 'timeout 30 bash x'].join('\n');
|
|
const findings = findRawTimeoutInvocations(text);
|
|
assert.equal(findings.length, 1);
|
|
assert.equal(findings[0].line, 3);
|
|
});
|
|
|
|
test('every shipped workflow/agent/reference/command surface is clean (regression)', () => {
|
|
// Would have returned 10 offenders before the #2351 conversions landed.
|
|
const offenders = scan(DEFAULT_ROOTS);
|
|
assert.deepEqual(
|
|
offenders,
|
|
[],
|
|
`hardcoded timeout still present:\n${offenders.map((o) => `${o.file}:${o.line} ${o.snippet}`).join('\n')}`,
|
|
);
|
|
});
|
|
});
|