Files
msd-core/tests/fix-2590-workflow-script-contract.test.cjs
Tom Boucher 9faacc0c15 test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist

Migrates the final 170 unbounded sync spawn sites across 49 files, then
removes the allowlist entirely. local/no-unbounded-spawn now runs with no
exemption surface across tests/**: there is no file to add a name to.

drift-detection's throw-native git() helper routes to gitOrThrow -- bare
runGit would have taken 16 call sites quiet on failure. commands.test.cjs
has two independently-scoped runGsdTools/runCli helpers, one already bounded
and one not; they are kept distinct rather than unified, the same trap as the
two same-named git() helpers in Wave 1.

runNpm's bound was erasable. Its options spread callerOptions after the
defaults, so an explicit timeout:undefined silently dropped the 180000ms
bound -- the rule flagged it and was right; it was not a false positive. Fixed
by destructuring with a default, with a test that fails when the default is
removed.

Two sites stay on a raw spawn with an explicit timeout because the seam
cannot express them: one needs shell:true for npm.cmd on Windows, one
redirects stdout to a real fd. Both are the rule's own documented second
option, not an escape from it.

Closure verified rather than asserted: the derivation scan reports 0 unbounded
spawn helpers and 0 unbounded direct git call sites, and a temporary file
carrying an unbounded spawn still errors with the allowlist gone.

Closes #3064.

* test(#3148): close a hole in the guard's own eslint-disable ban

The ban listed only the top level of tests/, so it was blind to 37 .cjs
files under tests/helpers, qa, observability, fixtures and dispatch. With the
allowlist deleted this test is the sole remaining way to detect someone
silencing the rule inline, so the gap was load-bearing: a nested file could
carry an unbounded spawn plus an eslint-disable and pass everything.

Proven before and after. A probe planted under tests/helpers with both was
invisible to the guard and clean under eslint; after making the listing
recursive the guard fails on it. The scanned set goes from 771 files to 808.

Pre-existing since the guard shipped, but this wave is what promoted it to
sole defense, so it is fixed here rather than filed.

Also converts the last hand-rolled throw check to throwIfFailed and the last
re-derived legacy shape to compose toLegacyResult, which makes the epic's
none-remain claim true rather than nearly true. toLegacyResult itself is not
widened -- eight callers depend on its shape and one consumer does not
justify changing a shared contract.

* fix(#3148): correct seam incoherence at the bound and a slow review-lane error path

Two real failures from the remote runner, both fixed at the cause.

The seam could return outcome TIMED_OUT together with exitCode 0. At the
exact bound spawnSync reports ETIMEDOUT while the child has already exited
with a real status, and toSeamResult classified on the error code while
passing status straight through -- an incoherent pair its own boundary test
was written to catch, and did. A status that is not null is direct evidence
the child exited on its own, so it now decides the outcome before the
error-code branches run. process-seam.cjs was deliberately untouched by every
earlier wave; this is a defect in the module itself, kept surgical, with a
unit test that fails against the old logic.

review-lane with an unknown subcommand fell through to its usage error only
after loading the capability registry and building a per-lane plan, which
spawns one child process per lane -- up to twelve. The error path took
~1288ms instead of ~119ms, and under bench load it outran a caller's spawn
timeout and was killed before writing anything, which is the empty stdout and
stderr CI saw. It now fails fast before any of that work begins.

This is the epic's first production change. It is user-facing, so it carries
a changeset rather than a no-changelog label.

* test(#3148): replace a real-race timeout test with a deterministic one

E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a
warm container git finishes first, spawnSync returns status 0 with no error
at all, the seam correctly classifies EXITED, and gitOrThrow correctly does
not throw -- so the test failed on both lanes. A probe confirms a genuine
timeout always carries status null, so this was never the seam misbehaving.

Raising the bound would only lengthen the odds, which is the same defect with
better luck. The test now drives gitOrThrow against a stubbed runGit that
returns a synthetic TIMED_OUT result, so it asserts exactly what it always
meant to -- that a timeout propagates as a throw -- with no timing
dependence. Five consecutive runs are identical where the old one varied.

I wrote this test in Wave 0; it is a real-race test by construction and
CLAUDE.md says to replace those rather than re-run them.

* chore(#3148): backfill changeset PR number 3192

---------

Co-authored-by: sim <sim@local>
2026-08-07 21:03:50 -04:00

281 lines
12 KiB
JavaScript

/**
* #2590 — every emitted Workflow script was rejected by the Workflow tool.
*
* `emitWorkflowScript` generated four constructs the tool does not accept. The
* first was fatal on its own, so the backend could never dispatch a wave:
*
* 1. no `export const meta = {…}` first statement -> whole script rejected
* 2. `resumeFromRunId("<id>")` -> "resumeFromRunId is not defined"
* (it is a Workflow TOOL INPUT parameter, not a script function)
* 3. `budget(<n>)` -> "budget is not a function"
* (`budget` is a read-only object { total, spent(), remaining() })
* 4. `parallel(agent(…), agent(…))` -> "parallel() expects an array of functions"
*
* Plus two secondary defects that kept the emitted script from ever being
* REACHED, which is why this shipped undetected:
*
* 5. nothing resolved the Agent SDK version, so gate 5 returned
* `agent_sdk_version_unknown` on every automated run
* 6. the runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging
* from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any
* invocation without --runtime reported `runtime_not_claude`
*
* The script assertions parse the emitted text as a real ES module rather than
* pattern-matching it, so a syntactically invalid script fails outright.
*/
'use strict';
process.env.GSD_TEST_MODE = '1';
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { createTempDir, cleanup } = require('./helpers.cjs');
const { runNode } = require('./helpers/process-seam.cjs');
const { throwIfFailed } = require('./helpers/git-fixture.cjs');
const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs');
const core = require('../gsd-core/bin/lib/claude-orchestration.cjs');
const TOOLS = path.join(__dirname, '..', 'gsd-core', 'bin', 'gsd-tools.cjs');
function emit(overrides) {
const input = Object.assign({
phaseDir: '.planning/phases/01',
runId: 'execute-1',
waves: [{ id: 'wave-1', plans: [{ id: '01-01', brief: 'noop', files_modified: ['a.ts'] }] }],
}, overrides || {});
const r = core.emitWorkflowScript(input);
assert.ok(r.ok, `emit failed: ${JSON.stringify(r)}`);
return r;
}
/** First non-comment, non-blank line — the script's first actual statement. */
function firstStatement(script) {
return script.split('\n').map((l) => l.trim())
.find((l) => l.length > 0 && !l.startsWith('//')) || '';
}
describe('#2590: emitted Workflow scripts satisfy the Workflow tool contract', () => {
test('the emitted script is syntactically valid as an ES module', () => {
// `export const meta` + top-level `await` only parse in module context —
// which is exactly the context the Workflow tool runs the script in.
const { script } = emit({
waves: [
{ id: 'w1', plans: [
{ id: 'a', brief: 'one', files_modified: ['a.ts'] },
{ id: 'b', brief: 'two', files_modified: ['b.ts'] },
] },
{ id: 'w2', plans: [{ id: 'c', brief: 'three', files_modified: ['c.ts'] }] },
],
});
// .mjs so node parses it in module context, inside a helper temp dir so
// cleanup() carries the Windows-EBUSY retry budget.
const dir = createTempDir('gsd-2590-parse-');
const f = path.join(dir, 'emitted.mjs');
fs.writeFileSync(f, script);
try {
const result = runNode(['--check', f], { timeoutMs: PROBE_TIMEOUT_MS });
throwIfFailed(result, `node --check ${f} (emitted script must parse)`);
} finally {
cleanup(dir);
}
});
test('1. `export const meta` is the first statement', () => {
const { script } = emit();
assert.match(
firstStatement(script),
/^export const meta = \{/,
'the Workflow tool rejects any script whose first statement is not the meta block',
);
});
test('meta.phases titles match the emitted phase() calls exactly', () => {
// The tool matches phase titles by exact string; a mismatch silently splits
// progress into an unnamed group.
const { script } = emit({
waves: [
{ id: 'alpha', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] },
{ id: 'beta', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] },
],
});
const metaTitles = [...script.matchAll(/\{ title: "([^"]+)"/g)].map((m) => m[1]);
const phaseTitles = [...script.matchAll(/^phase\("([^"]+)"\)/gm)].map((m) => m[1]);
assert.deepEqual(metaTitles, ['Wave alpha', 'Wave beta']);
assert.deepEqual(phaseTitles, metaTitles, 'phase() titles must match meta.phases exactly');
});
test('duplicate wave ids are rejected (phase titles must map 1:1)', () => {
// Two waves sharing an id emit two identical `phase("Wave x")` calls and two
// identical meta.phases entries; the tool matches titles by exact string, so
// the second wave's agents would be attributed to the first's progress group.
const r = core.emitWorkflowScript({
phaseDir: '.planning/phases/01',
runId: 'execute-1',
waves: [
{ id: 'dup', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] },
{ id: 'dup', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] },
],
});
assert.equal(r.ok, false, 'duplicate wave ids must be rejected, not silently merged');
assert.match(String(r.reason), /duplicate wave id/);
});
test('distinct wave ids are still accepted (the boundary either side)', () => {
const r = core.emitWorkflowScript({
phaseDir: '.planning/phases/01',
runId: 'execute-1',
waves: [
{ id: 'w1', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] },
{ id: 'w2', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] },
],
});
assert.equal(r.ok, true, `distinct wave ids must pass: ${JSON.stringify(r)}`);
});
test('2. resumeFromRunId is never CALLED (it is a tool input, not a function)', () => {
const { script, summary } = emit({ runId: 'execute-7' });
assert.ok(
!/^\s*resumeFromRunId\s*\(/m.test(script),
'calling resumeFromRunId() throws "resumeFromRunId is not defined"',
);
// The run id must still reach the caller, which passes it as the tool input.
assert.equal(summary.resumeRunId, 'execute-7');
});
test('3. budget is never CALLED, at and around the boundary', () => {
// budgetTokens is floored at > 0; check 0 (rejected), 1 (accepted), and a
// large value — none may produce a budget(...) call.
for (const tokens of [0, 1, 500000]) {
const { script, summary } = emit({ budgetTokens: tokens });
assert.ok(
!/^\s*budget\s*\(/m.test(script),
`budgetTokens=${tokens}: calling budget() throws "budget is not a function"`,
);
assert.equal(summary.budgetTokens, tokens > 0 ? tokens : null);
}
});
test('4. parallel() receives an array of thunks, not agent() results', () => {
const { script } = emit({
waves: [{ id: 'w', plans: [
{ id: 'a', brief: 'one', files_modified: ['a.ts'] },
{ id: 'b', brief: 'two', files_modified: ['b.ts'] },
] }],
});
assert.ok(/parallel\(\[/.test(script), 'parallel() expects an array of functions');
assert.ok(
!/parallel\(\s*agent\(/.test(script),
'passing agent() results directly both throws and starts every agent eagerly',
);
// Each agent must be wrapped in a thunk so parallel() can bound concurrency.
const agents = [...script.matchAll(/agent\("/g)].length;
const thunks = [...script.matchAll(/\(\) => agent\("/g)].length;
assert.equal(thunks, agents, 'every agent() must be wrapped in a () => thunk');
});
test('single-plan stages also emit an array (regression: the 1-plan branch)', () => {
// The pre-fix code had a SEPARATE single-plan branch that emitted
// `parallel(\n agent(...)\n)` — valid-looking but the same defect.
const { script } = emit();
assert.ok(/parallel\(\[/.test(script));
assert.equal([...script.matchAll(/\(\) => agent\("/g)].length, 1);
});
test('per-plan worktree isolation still mirrors use_worktree', () => {
const { script } = emit({
waves: [{ id: 'w', plans: [
{ id: 'a', brief: 'iso', files_modified: ['a.ts'] },
{ id: 'b', brief: 'noiso', files_modified: ['b.ts'], use_worktree: false },
] }],
});
assert.match(script, /agent\("iso", \{ agentType: "gsd-executor", isolation: "worktree" \}\)/);
assert.match(script, /agent\("noiso", \{ agentType: "gsd-executor" \}\)/);
});
});
describe('#2590: the backend is reachable without hand-passed flags', () => {
function repro() {
const dir = createTempDir('gsd-2590-repro-');
fs.mkdirSync(path.join(dir, '.planning'), { recursive: true });
fs.writeFileSync(
path.join(dir, '.planning', 'config.json'),
JSON.stringify({ claude_orchestration: { enabled: true } }),
);
fs.writeFileSync(
path.join(dir, 'waves.json'),
JSON.stringify({ waves: [{ id: 'wave-1', plans: [{ id: '01-01', brief: 'noop', files_modified: ['a.ts'] }] }] }),
);
return dir;
}
function resolve(dir, extraArgs) {
const result = runNode([
TOOLS, 'claude-orchestration', 'resolve-wave-dispatch',
'--waves', 'waves.json', '--run-id', 'execute-1',
'--phase-dir', '.planning/phases/01', '--raw',
...(extraArgs || []),
], { cwd: dir, timeoutMs: PROBE_TIMEOUT_MS });
throwIfFailed(result, 'gsd-tools claude-orchestration resolve-wave-dispatch');
return JSON.parse(result.stdout);
}
test('5+6. no --runtime and no --agent-sdk-version still reaches the version gate', () => {
const dir = repro();
try {
const r = resolve(dir);
// Pre-fix this was `agent_sdk_version_unknown` (nothing resolved a
// version) or `runtime_not_claude` (the divergent fallback). Either is a
// regression; the version gate must now be reached and answer truthfully.
assert.notEqual(r.reason, 'agent_sdk_version_unknown',
'the router must resolve the installed SDK version itself');
assert.notEqual(r.reason, 'runtime_not_claude',
'runtime must fall back to the canonical config.runtime > claude chain');
} finally {
cleanup(dir);
}
});
test('an SDK version above the floor activates the workflow backend end to end', () => {
const dir = repro();
try {
const r = resolve(dir, ['--agent-sdk-version', '0.3.149']);
assert.equal(r.backend, 'workflow', `expected workflow backend, got ${JSON.stringify(r)}`);
assert.ok(typeof r.script === 'string' && r.script.length > 0);
assert.match(firstStatement(r.script), /^export const meta = \{/);
} finally {
cleanup(dir);
}
});
test('an explicit --agent-sdk-version still wins over the installed one', () => {
const dir = repro();
try {
// A deliberately ancient pin must be honored (and decline), proving the
// flag is not ignored now that a fallback exists.
const r = resolve(dir, ['--agent-sdk-version', '0.0.1']);
assert.equal(r.backend, 'inline');
assert.equal(r.reason, 'agent_sdk_version_below_floor');
} finally {
cleanup(dir);
}
});
test('GSD_AGENT_SDK_VERSION is honored between the flag and the installed version', () => {
const dir = repro();
try {
const result = runNode([
TOOLS, 'claude-orchestration', 'resolve-wave-dispatch',
'--waves', 'waves.json', '--run-id', 'execute-1',
'--phase-dir', '.planning/phases/01', '--raw',
], { cwd: dir, env: { ...process.env, GSD_AGENT_SDK_VERSION: '0.3.149' }, timeoutMs: PROBE_TIMEOUT_MS });
throwIfFailed(result, 'gsd-tools claude-orchestration resolve-wave-dispatch (GSD_AGENT_SDK_VERSION)');
assert.equal(JSON.parse(result.stdout).backend, 'workflow');
} finally {
cleanup(dir);
}
});
});