* test(#3148): bound the long tail and delete the allowlist Migrates the final 170 unbounded sync spawn sites across 49 files, then removes the allowlist entirely. local/no-unbounded-spawn now runs with no exemption surface across tests/**: there is no file to add a name to. drift-detection's throw-native git() helper routes to gitOrThrow -- bare runGit would have taken 16 call sites quiet on failure. commands.test.cjs has two independently-scoped runGsdTools/runCli helpers, one already bounded and one not; they are kept distinct rather than unified, the same trap as the two same-named git() helpers in Wave 1. runNpm's bound was erasable. Its options spread callerOptions after the defaults, so an explicit timeout:undefined silently dropped the 180000ms bound -- the rule flagged it and was right; it was not a false positive. Fixed by destructuring with a default, with a test that fails when the default is removed. Two sites stay on a raw spawn with an explicit timeout because the seam cannot express them: one needs shell:true for npm.cmd on Windows, one redirects stdout to a real fd. Both are the rule's own documented second option, not an escape from it. Closure verified rather than asserted: the derivation scan reports 0 unbounded spawn helpers and 0 unbounded direct git call sites, and a temporary file carrying an unbounded spawn still errors with the allowlist gone. Closes #3064. * test(#3148): close a hole in the guard's own eslint-disable ban The ban listed only the top level of tests/, so it was blind to 37 .cjs files under tests/helpers, qa, observability, fixtures and dispatch. With the allowlist deleted this test is the sole remaining way to detect someone silencing the rule inline, so the gap was load-bearing: a nested file could carry an unbounded spawn plus an eslint-disable and pass everything. Proven before and after. A probe planted under tests/helpers with both was invisible to the guard and clean under eslint; after making the listing recursive the guard fails on it. The scanned set goes from 771 files to 808. Pre-existing since the guard shipped, but this wave is what promoted it to sole defense, so it is fixed here rather than filed. Also converts the last hand-rolled throw check to throwIfFailed and the last re-derived legacy shape to compose toLegacyResult, which makes the epic's none-remain claim true rather than nearly true. toLegacyResult itself is not widened -- eight callers depend on its shape and one consumer does not justify changing a shared contract. * fix(#3148): correct seam incoherence at the bound and a slow review-lane error path Two real failures from the remote runner, both fixed at the cause. The seam could return outcome TIMED_OUT together with exitCode 0. At the exact bound spawnSync reports ETIMEDOUT while the child has already exited with a real status, and toSeamResult classified on the error code while passing status straight through -- an incoherent pair its own boundary test was written to catch, and did. A status that is not null is direct evidence the child exited on its own, so it now decides the outcome before the error-code branches run. process-seam.cjs was deliberately untouched by every earlier wave; this is a defect in the module itself, kept surgical, with a unit test that fails against the old logic. review-lane with an unknown subcommand fell through to its usage error only after loading the capability registry and building a per-lane plan, which spawns one child process per lane -- up to twelve. The error path took ~1288ms instead of ~119ms, and under bench load it outran a caller's spawn timeout and was killed before writing anything, which is the empty stdout and stderr CI saw. It now fails fast before any of that work begins. This is the epic's first production change. It is user-facing, so it carries a changeset rather than a no-changelog label. * test(#3148): replace a real-race timeout test with a deterministic one E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a warm container git finishes first, spawnSync returns status 0 with no error at all, the seam correctly classifies EXITED, and gitOrThrow correctly does not throw -- so the test failed on both lanes. A probe confirms a genuine timeout always carries status null, so this was never the seam misbehaving. Raising the bound would only lengthen the odds, which is the same defect with better luck. The test now drives gitOrThrow against a stubbed runGit that returns a synthetic TIMED_OUT result, so it asserts exactly what it always meant to -- that a timeout propagates as a throw -- with no timing dependence. Five consecutive runs are identical where the old one varied. I wrote this test in Wave 0; it is a real-race test by construction and CLAUDE.md says to replace those rather than re-run them. * chore(#3148): backfill changeset PR number 3192 --------- Co-authored-by: sim <sim@local>
289 lines
14 KiB
JavaScript
289 lines
14 KiB
JavaScript
// allow-test-rule: source-text-is-the-product
|
|
// The autonomous command and workflow markdown are runtime-loaded contracts.
|
|
// Checking their text verifies the shipped slash-command behavior.
|
|
|
|
'use strict';
|
|
|
|
const { describe, test } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
const { runNode } = require('./helpers/process-seam.cjs');
|
|
const { throwIfFailed } = require('./helpers/git-fixture.cjs');
|
|
const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs');
|
|
|
|
const REPO_ROOT = path.join(__dirname, '..');
|
|
const COMMAND_PATH = path.join(REPO_ROOT, 'commands', 'gsd', 'autonomous.md');
|
|
const WORKFLOW_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'autonomous.md');
|
|
const COMMANDS_DOC_PATH = path.join(REPO_ROOT, 'docs', 'COMMANDS.md');
|
|
const HOW_TO_PATH = path.join(REPO_ROOT, 'docs', 'how-to', 'run-phases-autonomously.md');
|
|
const TOOLS = path.join(REPO_ROOT, 'gsd-core', 'bin', 'gsd-tools.cjs');
|
|
// #2994: fragmentization moved the five converge-gated regions out of the host
|
|
// autonomous.md into dedicated step files (state:plan-strategy-converge) —
|
|
// see docs/reference/workflow-fragments.md. Tests that assert on this moved
|
|
// content read the step file directly rather than the host.
|
|
const STEP_FAIL_FAST_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'autonomous', 'steps', 'converge-fail-fast.md');
|
|
const STEP_DISPATCH_BG_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'autonomous', 'steps', 'converge-dispatch-bg.md');
|
|
const STEP_DISPATCH_INLINE_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'autonomous', 'steps', 'converge-dispatch-inline.md');
|
|
const STEP_LOOP_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'autonomous', 'steps', 'converge-loop.md');
|
|
|
|
function read(filePath) {
|
|
return fs.readFileSync(filePath, 'utf8');
|
|
}
|
|
|
|
describe('autonomous --converge flag (#711)', () => {
|
|
test('command advertises --converge and documents --cross-ai as alias', () => {
|
|
const command = read(COMMAND_PATH);
|
|
|
|
assert.match(
|
|
command,
|
|
/^argument-hint:.*--converge/m,
|
|
'autonomous command should advertise --converge in argument-hint',
|
|
);
|
|
assert.match(command, /--cross-ai/, 'autonomous command should document --cross-ai alias');
|
|
assert.match(
|
|
command,
|
|
/workflow\.plan_review_convergence=true/,
|
|
'autonomous command should mention the existing convergence feature gate',
|
|
);
|
|
});
|
|
|
|
test('workflow parses converge aliases into a plan strategy', () => {
|
|
const workflow = read(WORKFLOW_PATH);
|
|
|
|
assert.match(workflow, /PLAN_STRATEGY="local"/, 'workflow should default to local planning');
|
|
assert.match(workflow, /PLAN_STRATEGY="converge"/, 'workflow should opt into converge planning');
|
|
assert.match(workflow, /converge\|cross-ai/, 'workflow should accept --converge and --cross-ai');
|
|
});
|
|
|
|
test('workflow fails fast when convergence is requested but disabled', () => {
|
|
// #2994: this check lives in the converge-fail-fast step file now
|
|
// (state:plan-strategy-converge) — the host only carries the gated
|
|
// conditional-read stub.
|
|
const workflow = read(WORKFLOW_PATH);
|
|
const step = read(STEP_FAIL_FAST_PATH);
|
|
|
|
assert.match(
|
|
workflow,
|
|
/gsd:section id="converge-fail-fast" when="state:plan-strategy-converge"/,
|
|
'workflow should gate the fail-fast check behind state:plan-strategy-converge',
|
|
);
|
|
assert.match(
|
|
step,
|
|
/config-get workflow\.plan_review_convergence/,
|
|
'converge-fail-fast step should check workflow.plan_review_convergence before planning',
|
|
);
|
|
assert.match(
|
|
step,
|
|
/gsd config-set workflow\.plan_review_convergence true/,
|
|
'converge-fail-fast step should print the enable command instead of silently downgrading',
|
|
);
|
|
});
|
|
|
|
test('workflow routes planning through plan-review-convergence when enabled', () => {
|
|
// #2994: the converge dispatch/loop bodies live in dedicated step files
|
|
// now (state:plan-strategy-converge) — only the local-planning fallback
|
|
// remains inline in the host.
|
|
const workflow = read(WORKFLOW_PATH);
|
|
const dispatchInline = read(STEP_DISPATCH_INLINE_PATH);
|
|
const loop = read(STEP_LOOP_PATH);
|
|
const dispatchBg = read(STEP_DISPATCH_BG_PATH);
|
|
|
|
assert.match(
|
|
dispatchInline,
|
|
/Skill\(skill="gsd-plan-review-convergence", args="\$\{PHASE_NUM\} \$\{CONVERGENCE_ARGS\}"\)/,
|
|
'inline converge dispatch step should call gsd-plan-review-convergence',
|
|
);
|
|
assert.match(
|
|
loop,
|
|
/Skill\(skill="gsd-plan-review-convergence", args="\$\{PHASE_NUM\} \$\{CONVERGENCE_ARGS\}"\)/,
|
|
'default converge loop step should call gsd-plan-review-convergence',
|
|
);
|
|
assert.match(
|
|
dispatchBg,
|
|
/Run plan convergence for phase \$\{PHASE_NUM\}: Skill\(skill=\\"gsd-plan-review-convergence\\"/,
|
|
'interactive converge mode should dispatch plan convergence in the background agent',
|
|
);
|
|
assert.match(
|
|
workflow,
|
|
/Skill\(skill="gsd-plan-phase", args="\$\{PHASE_NUM\}"\)/,
|
|
'local planning path should remain available for default autonomous runs',
|
|
);
|
|
});
|
|
|
|
test('workflow forwards reviewer flags and max cycles to convergence', () => {
|
|
const workflow = read(WORKFLOW_PATH);
|
|
// Non-lane convergence controls remain hand-written literals in the workflow.
|
|
const convergenceControls = ['--all', '--text'];
|
|
// Reviewer lane flags that were formerly hand-enumerated in the workflow text.
|
|
// They must now be DERIVED at runtime via `gsd_run review-lane flags`, not listed.
|
|
const formerlyHardcodedLaneFlags = [
|
|
'--codex',
|
|
'--gemini',
|
|
'--claude',
|
|
'--opencode',
|
|
'--ollama',
|
|
'--lm-studio',
|
|
'--llama-cpp',
|
|
];
|
|
// The literal-absence guard below excludes '--claude': the runtime-launcher
|
|
// preamble legitimately contains an unrelated "npx ... --claude --local"
|
|
// install-runtime flag, so a substring match on '--claude' would false-positive
|
|
// against that literal, not against a re-added reviewer-flag list.
|
|
const antiParityLaneFlags = formerlyHardcodedLaneFlags.filter((flag) => flag !== '--claude');
|
|
|
|
assert.match(workflow, /CONVERGENCE_ARGS/, 'workflow should build convergence pass-through args');
|
|
assert.match(
|
|
workflow,
|
|
/gsd_run review-lane flags/,
|
|
'workflow should derive reviewer flags from the review-lane roster instead of hand-listing them',
|
|
);
|
|
for (const flag of convergenceControls) {
|
|
assert.ok(workflow.includes(flag), `workflow should pass through ${flag}`);
|
|
}
|
|
assert.match(workflow, /--max-cycles/, 'workflow should pass through --max-cycles N');
|
|
|
|
// Anti-parity guard (deliberately inverted polarity): the whole point of the
|
|
// review-lane-flags derivation is that reviewer lane flags are declared ONCE
|
|
// (in the review-lane roster) and never hand-listed again in workflow prose.
|
|
// If a future edit re-adds a hardcoded reviewer-flag list here, that is the
|
|
// regression this test exists to catch — so this assertion must FAIL when
|
|
// any of these flags reappear as literals in the workflow text.
|
|
for (const flag of antiParityLaneFlags) {
|
|
assert.ok(
|
|
!workflow.includes(flag),
|
|
`workflow should NOT hand-enumerate reviewer lane flag ${flag}; it must be derived via review-lane flags`,
|
|
);
|
|
}
|
|
|
|
// Behavioral coverage: prove the roster the workflow derives from actually
|
|
// yields the flags this test used to hardcode, so the derivation is not vacuous.
|
|
const laneFlagsResult = runNode([TOOLS, 'review-lane', 'flags'], { timeoutMs: PROBE_TIMEOUT_MS });
|
|
throwIfFailed(laneFlagsResult, `node ${TOOLS} review-lane flags`);
|
|
const laneFlags = laneFlagsResult.stdout.split('\n').filter(Boolean);
|
|
for (const flag of formerlyHardcodedLaneFlags) {
|
|
assert.ok(laneFlags.includes(flag), `review-lane flags should include ${flag}`);
|
|
}
|
|
});
|
|
|
|
test('docs show autonomous convergence usage', () => {
|
|
const commandsDoc = read(COMMANDS_DOC_PATH);
|
|
const howTo = read(HOW_TO_PATH);
|
|
|
|
assert.match(commandsDoc, /--converge/, 'COMMANDS.md should document --converge');
|
|
assert.match(commandsDoc, /--cross-ai/, 'COMMANDS.md should document --cross-ai alias');
|
|
assert.match(howTo, /\/gsd-autonomous --only 4 --converge/, 'how-to should show single-phase converge usage');
|
|
});
|
|
});
|
|
|
|
describe('autonomous verification deferral contract', () => {
|
|
test('workflow records explicit deferred states instead of silently advancing (#1525)', () => {
|
|
const workflow = read(WORKFLOW_PATH);
|
|
|
|
assert.match(workflow, /verification_deferred_human/);
|
|
assert.match(workflow, /verification_deferred_gaps/);
|
|
assert.match(workflow, /Deferred Verification/);
|
|
assert.match(workflow, /gsd:verify-work \$\{PHASE_NUM\}/);
|
|
assert.match(workflow, /gsd:plan-phase \$\{PHASE_NUM\} --gaps/);
|
|
assert.match(
|
|
workflow,
|
|
/\| \$\{PHASE_NUM\} \| verification_deferred_human \| \/gsd:verify-work \$\{PHASE_NUM\} \|/,
|
|
'human deferral must persist the exact deferred STATE row',
|
|
);
|
|
assert.match(
|
|
workflow,
|
|
/\| \$\{PHASE_NUM\} \| verification_deferred_gaps \| \/gsd:plan-phase \$\{PHASE_NUM\} --gaps \|/,
|
|
'gap deferral must persist the exact deferred STATE row',
|
|
);
|
|
assert.doesNotMatch(
|
|
workflow,
|
|
/Human validation deferred` and proceed to iterate step/,
|
|
'human-needed deferral must not silently proceed to the next phase',
|
|
);
|
|
assert.doesNotMatch(
|
|
workflow,
|
|
/Gaps deferred` and proceed to iterate step/,
|
|
'gap deferral must not silently proceed to the next phase',
|
|
);
|
|
assert.match(
|
|
workflow,
|
|
/Skip deferred phases on autonomous re-entry/,
|
|
'reruns must explicitly skip deferred verification phases',
|
|
);
|
|
assert.match(
|
|
workflow,
|
|
/Deferred Verification \(Skipped on Re-entry\)/,
|
|
'workflow should surface skipped deferred phases and their resume commands',
|
|
);
|
|
});
|
|
|
|
test('workflow runs normal transition post-processing after passed verification (#1526)', () => {
|
|
const workflow = read(WORKFLOW_PATH);
|
|
const passedIdx = workflow.indexOf('**If `passed`:**');
|
|
const transitionIdx = workflow.indexOf('transition.md', passedIdx);
|
|
const iterateIdx = workflow.indexOf('Proceed to iterate step', passedIdx);
|
|
|
|
assert.ok(transitionIdx > passedIdx, 'passed verification must invoke transition.md');
|
|
assert.ok(
|
|
transitionIdx < iterateIdx,
|
|
'normal transition post-processing must run before autonomous iterates',
|
|
);
|
|
});
|
|
|
|
test('workflow reads canonical verification status before human-needed promotion (#1522)', () => {
|
|
const workflow = read(WORKFLOW_PATH);
|
|
const waitIdx = workflow.indexOf('After execute, read canonical verification');
|
|
const humanNeededIdx = workflow.indexOf('**If `human_needed`:**', waitIdx);
|
|
const promoteIdx = workflow.indexOf('set VERIFICATION frontmatter `status: passed`', humanNeededIdx);
|
|
const section = workflow.slice(waitIdx, humanNeededIdx);
|
|
|
|
assert.ok(waitIdx !== -1, 'workflow must document the post-execution verification read');
|
|
assert.ok(humanNeededIdx > waitIdx, 'human_needed branch must follow verification status read');
|
|
assert.ok(promoteIdx > humanNeededIdx, 'human_needed branch must contain the promotion action');
|
|
// #2589: the verification read uses the native --pick flag (no jq dependency).
|
|
// String-based check (not a regex literal) so the assertion stays robust to
|
|
// shell metacharacters in the snippet and parses cleanly under espree.
|
|
assert.ok(
|
|
section.includes('VERIFY_STATUS=$(gsd_run query verification.status "${PHASE_DIR}" --pick status 2>/dev/null || true)'),
|
|
'autonomous must route human validation through canonical verification.status via the native --pick flag',
|
|
);
|
|
assert.doesNotMatch(
|
|
section,
|
|
/grep "\^status:"/,
|
|
'autonomous must not route stale human_needed reports from raw frontmatter',
|
|
);
|
|
});
|
|
|
|
test('workflow discovers incomplete phases from canonical verification projection (#1522)', () => {
|
|
const workflow = read(WORKFLOW_PATH);
|
|
const discoverStart = workflow.indexOf('<step name="discover_phases">');
|
|
const discoverEnd = workflow.indexOf('</step>', discoverStart);
|
|
const iterateStart = workflow.indexOf('<step name="iterate">');
|
|
const iterateEnd = workflow.indexOf('</step>', iterateStart);
|
|
const discoverStep = workflow.slice(discoverStart, discoverEnd);
|
|
const iterateStep = workflow.slice(iterateStart, iterateEnd);
|
|
|
|
assert.match(discoverStep, /INIT_MANAGER=\$\(gsd_run query init\.manager\)/);
|
|
assert.ok(
|
|
discoverStep.includes('if [[ "$INIT_MANAGER" == @file:* ]]; then INIT_MANAGER=$(cat "${INIT_MANAGER#@file:}"); fi'),
|
|
'autonomous discovery must dereference large init.manager payloads before parsing',
|
|
);
|
|
assert.match(discoverStep, /phase_complete !== true/);
|
|
assert.match(discoverStep, /verification_status !== "passed"/);
|
|
assert.match(discoverStep, /STATE_CONTENT=\$\(cat \.planning\/STATE\.md 2>\/dev\/null \|\| true\)/);
|
|
assert.match(discoverStep, /drop any phase whose number appears in the deferred-phase map/);
|
|
assert.doesNotMatch(discoverStep, /ROADMAP=\$\(gsd_run query roadmap\.analyze\)/);
|
|
assert.doesNotMatch(discoverStep, /disk_status !== "complete"/);
|
|
|
|
assert.match(iterateStep, /INIT_MANAGER=\$\(gsd_run query init\.manager\)/);
|
|
assert.ok(
|
|
iterateStep.includes('if [[ "$INIT_MANAGER" == @file:* ]]; then INIT_MANAGER=$(cat "${INIT_MANAGER#@file:}"); fi'),
|
|
'autonomous iteration must dereference large init.manager payloads before parsing',
|
|
);
|
|
assert.match(iterateStep, /phase_complete !== true/);
|
|
assert.match(iterateStep, /verification_status !== "passed"/);
|
|
assert.match(iterateStep, /STATE_CONTENT=\$\(cat \.planning\/STATE\.md 2>\/dev\/null \|\| true\)/);
|
|
assert.match(iterateStep, /drop deferred phases from the autonomous queue/);
|
|
});
|
|
});
|