Folds 15 tests/fix-*.test.cjs regression files (191 test() blocks) into their module's main suite, per the wave decomposition of #3315 (H3 of epic #3053). 187 blocks land in 8 existing suites (4 exact-duplicate cases dropped, documented inline); 4 blocks move via git mv into 2 new suite files with no prior coverage to merge into. Zero production behavior change. Also tightens two H1 (#3313) ratchets that the fold's own file-count reduction moved past their grace window, per the ratchets' documented dual failure mode (a stale/too-loose baseline fails exactly like a novel violation): - lint-test-file-count.allowlist.json: removes the stale "audit" entry (folding fix-2766 into tests/uat.test.cjs drops that module back to its 2-file cap). - lint-allow-test-rule-refs.ceiling.json: lowers maxFiles 314 -> 309, the real post-fold high-water mark (gsd-test's own repo-baseline test caught this — CI, not a human, found it). Two orthogonal review passes (Standards+Spec code-review, isolated security-review) found and this commit fixes two issues before push: a genuinely-distinct #2287 test case (file-absent vs. file-present- resolved) that a prior fold pass had wrongly dropped as a duplicate — restored verbatim into tests/uat.test.cjs; and a missing same-line allow-test-rule citation on the #2196 block in tests/debug-session-management.test.cjs, added for consistency with its sibling #2257 block. lint-removed-but-needed also caught two stale doc references to the now-folded-away fix-2285-claude-orchestration-wiring.test.cjs filename (docs/adr/1143-claude-orchestration-capability.md, gsd-core/references/execute-phase-response-language.md) — updated both to point at tests/claude-orchestration.test.cjs, its new home. Co-authored-by: sim <sim@local>
This commit is contained in:
@@ -9,9 +9,9 @@
|
||||
|
||||
## Why this is still `Proposed` (audited 2026-07-17)
|
||||
|
||||
Confirmed shipped, on-tree: the capability is real and registered, not vaporware. `capabilities/claude-orchestration/capability.json` exists with detection + emission (`detectWorkflowBackend` / `emitWorkflowScript`) in `src/claude-orchestration.cts` (compiled to `gsd-core/bin/lib/claude-orchestration.cjs`), federated config (`claude_orchestration.enabled` / `execution_backend` / `min_agent_sdk_version`), and 1,552 lines of tests across `tests/claude-orchestration.test.cjs`, `tests/claude-orchestration-command-router.test.cjs`, and `tests/fix-2285-claude-orchestration-wiring.test.cjs`. The previously-fatal wiring bug, #2285 ("claude-orchestration capability (#1143) registered as active but never wired into execute-phase orchestrator prompt"), is closed COMPLETED (2026-07-15) — one day before this audit — and the owning feature issue #1143 is also closed COMPLETED.
|
||||
Confirmed shipped, on-tree: the capability is real and registered, not vaporware. `capabilities/claude-orchestration/capability.json` exists with detection + emission (`detectWorkflowBackend` / `emitWorkflowScript`) in `src/claude-orchestration.cts` (compiled to `gsd-core/bin/lib/claude-orchestration.cjs`), federated config (`claude_orchestration.enabled` / `execution_backend` / `min_agent_sdk_version`), and 1,552 lines of tests across `tests/claude-orchestration.test.cjs` (which now includes the #2285 wiring-fix regression coverage, folded in per #3334) and `tests/claude-orchestration-command-router.test.cjs`. The previously-fatal wiring bug, #2285 ("claude-orchestration capability (#1143) registered as active but never wired into execute-phase orchestrator prompt"), is closed COMPLETED (2026-07-15) — one day before this audit — and the owning feature issue #1143 is also closed COMPLETED.
|
||||
|
||||
**The blocker.** The ADR sets its own bar for ratification in its own Amendment (above): "flipping to Accepted follows maintainer sign-off on the E2E behaviour once exercised on Claude Code with the Workflow tool present." No such exercise is recorded anywhere in issues, PRs, or tests. Every test in the three files above operates at the contract or CLI-subprocess layer — asserting the *shape* of an emitted script or the return value of `resolve-wave-dispatch` — none constructs or executes an actual Workflow-tool run (`grep -rn "Workflow(" tests/claude-orchestration*.test.cjs tests/fix-2285-*.test.cjs` returns no hits). Two further gaps sit inside the ADR's own Decision section: (1) Decision §1's claimed net effect — "wave parallelism, the plan-checker, and the verifier are restored" — is narrower than what shipped: `capability.json`'s own description says "the plan-checker and verifier remain inline until separately wired — this capability delivers the parallel-execution backend, not those gates"; (2) Decision §3's fold-in of the `gsd-ultraplan-phase` skill into the capability's `skills[]` has not happened — `capability.json` still shows `"skills": []`, and no follow-up issue for the migration the Amendment promises exists (searched via `gh issue list --search`, no result).
|
||||
**The blocker.** The ADR sets its own bar for ratification in its own Amendment (above): "flipping to Accepted follows maintainer sign-off on the E2E behaviour once exercised on Claude Code with the Workflow tool present." No such exercise is recorded anywhere in issues, PRs, or tests. Every test in the two files above operates at the contract or CLI-subprocess layer — asserting the *shape* of an emitted script or the return value of `resolve-wave-dispatch` — none constructs or executes an actual Workflow-tool run (`grep -rn "Workflow(" tests/claude-orchestration*.test.cjs` returns no hits). Two further gaps sit inside the ADR's own Decision section: (1) Decision §1's claimed net effect — "wave parallelism, the plan-checker, and the verifier are restored" — is narrower than what shipped: `capability.json`'s own description says "the plan-checker and verifier remain inline until separately wired — this capability delivers the parallel-execution backend, not those gates"; (2) Decision §3's fold-in of the `gsd-ultraplan-phase` skill into the capability's `skills[]` has not happened — `capability.json` still shows `"skills": []`, and no follow-up issue for the migration the Amendment promises exists (searched via `gh issue list --search`, no result).
|
||||
|
||||
**Unblock condition.** Ratify once: (a) a real Claude Code session with the Workflow tool present and `claude_orchestration.enabled=true` drives an `execute-phase` wave through the Workflow backend, and the result is recorded (issue comment, PR, or a test that actually builds/executes a `Workflow` script rather than asserting emitted-script shape) — that is the maintainer sign-off the ADR itself asks for; and (b) the Decision section's "plan-checker and verifier restored" language is reconciled with the shipped scope (either corrected to match, or backed by a tracked issue for the deferred wiring `capability.json` already discloses). The `skills[]` migration (item 3) is lower priority since it is openly disclosed as deferred rather than silently dropped, but should carry a tracked issue number before ratification so it doesn't quietly vanish.
|
||||
|
||||
|
||||
@@ -4,4 +4,4 @@
|
||||
|
||||
The literal report templates embedded in this workflow (`## Execution Plan`, `## Phase {X}: {Name} Execution Complete`, `## ⚠ Phase {X}: {Name} — Gaps Found`, etc.) are a structural source, not literal output to copy verbatim — render their prose translated into `{response_language}` while keeping headings' structural markers, table columns, IDs, commands, and file paths unchanged.
|
||||
|
||||
This directive was extracted from `workflows/execute-phase.md` to keep that file under the frozen pre-phase-6 byte ceiling (ADR-857 Phase 6 capstone, `tests/fix-2285-claude-orchestration-wiring.test.cjs`). The `@-reference` is eager, so the runtime still loads this content alongside the workflow — the extraction is purely a file-size discipline, not a lazy-load optimization.
|
||||
This directive was extracted from `workflows/execute-phase.md` to keep that file under the frozen pre-phase-6 byte ceiling (ADR-857 Phase 6 capstone, `tests/claude-orchestration.test.cjs`). The `@-reference` is eager, so the runtime still loads this content alongside the workflow — the extraction is purely a file-size discipline, not a lazy-load optimization.
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
{
|
||||
"maxFiles": 314,
|
||||
"maxFiles": 309,
|
||||
"grace": 3
|
||||
}
|
||||
|
||||
@@ -146,14 +146,6 @@
|
||||
"prompt-budget.test.cjs"
|
||||
],
|
||||
"issue": "2929"
|
||||
},
|
||||
"audit": {
|
||||
"files": [
|
||||
"audit-command-cutover.test.cjs",
|
||||
"audit-fix-command.test.cjs",
|
||||
"fix-2766-audit-uat-archived-and-table-shapes.test.cjs"
|
||||
],
|
||||
"issue": "2766"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -16,9 +16,14 @@ const path = require('node:path');
|
||||
|
||||
const fc = require('fast-check');
|
||||
|
||||
// #2590: gsd-tools subprocess tests in this file spawn the real CLI; keep it in
|
||||
// test mode for the whole process.
|
||||
process.env.GSD_TEST_MODE = '1';
|
||||
|
||||
const {
|
||||
detectWorkflowBackend,
|
||||
emitWorkflowScript,
|
||||
resolveWaveDispatch,
|
||||
WORKFLOW_TOOL_FLOOR_VERSION,
|
||||
BACKEND_VALUES,
|
||||
compareSemver,
|
||||
@@ -34,9 +39,15 @@ const {
|
||||
stripGeneratedComment,
|
||||
} = require('../scripts/gen-capability-registry.cjs');
|
||||
|
||||
const { runGsdTools, createTempDir, cleanup } = require('./helpers.cjs');
|
||||
const { runNode } = require('./helpers/process-seam.cjs');
|
||||
const { throwIfFailed } = require('./helpers/git-fixture.cjs');
|
||||
const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs');
|
||||
|
||||
const ROOT = path.resolve(__dirname, '..');
|
||||
const CAP_PATH = path.join(ROOT, 'capabilities', 'claude-orchestration', 'capability.json');
|
||||
const REGISTRY_PATH = path.join(ROOT, 'gsd-core', 'bin', 'lib', 'capability-registry.cjs');
|
||||
const WORKFLOW_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'execute-phase.md');
|
||||
|
||||
// ─── Fixtures ─────────────────────────────────────────────────────────────────
|
||||
|
||||
@@ -1019,3 +1030,871 @@ describe('#2686 — Workflow backend model threading', () => {
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── #2285 — resolveWaveDispatch fail-closed gate ladder (folded from
|
||||
// fix-2285-claude-orchestration-wiring.test.cjs) ──────────────────────────────
|
||||
//
|
||||
// #2285 — the `claude-orchestration` capability (Workflow backend, #1143) was
|
||||
// registered `active` but fully INERT: `detectWorkflowBackend`/`emitWorkflowScript`
|
||||
// had zero callers outside their own CLI router, and execute-phase.md declared
|
||||
// `execute:wave:pre` as a hook point in its frontmatter but never rendered it —
|
||||
// the wave loop only ever dispatched `execute:pre`, `execute:wave:post`, and
|
||||
// `execute:post`. `claude_orchestration.enabled:true` therefore had no effect on
|
||||
// a real execute-phase run.
|
||||
//
|
||||
// Fix (Approach B):
|
||||
// 1. execute-phase.md now renders `execute:wave:pre` immediately before each
|
||||
// wave's agents are dispatched (step 2.75, before step 3's Agent() loop).
|
||||
// 2. The claude-orchestration contribution moved from `execute:wave:post`
|
||||
// (fires too late — after the wave already dispatched inline) to
|
||||
// `execute:wave:pre` (fires before dispatch, where a backend selector
|
||||
// actually has to run to matter).
|
||||
// 3. `resolveWaveDispatch` in src/claude-orchestration.cts composes
|
||||
// `detectWorkflowBackend` + `emitWorkflowScript` into ONE decision seam,
|
||||
// giving both functions a real caller outside their CLI router and outside
|
||||
// tests. It is also exposed via `gsd-tools claude-orchestration
|
||||
// resolve-wave-dispatch`.
|
||||
//
|
||||
// These drive the real seam (no source-grep on implementation files) and
|
||||
// assert the fail-closed contract: disabled or any gate miss => inline,
|
||||
// byte-identical to today's dispatch shape.
|
||||
|
||||
/** A host-integration descriptor whose dispatch axis signals Workflow-tool capability (#2285 fixture shape). */
|
||||
const CAPABLE_HOST_2285 = { dispatch: { nested: true, background: true } };
|
||||
/** A descriptor that fails the nested/background dispatch gate. */
|
||||
const INCAPABLE_HOST = { dispatch: { nested: false, background: true } };
|
||||
|
||||
const ABOVE_FLOOR_SDK = '0.3.150';
|
||||
const AT_FLOOR_SDK = WORKFLOW_TOOL_FLOOR_VERSION; // '0.3.149'
|
||||
const BELOW_FLOOR_SDK = '0.3.148';
|
||||
|
||||
function enabledConfig(overrides = {}) {
|
||||
return {
|
||||
'claude_orchestration.enabled': true,
|
||||
'claude_orchestration.execution_backend': 'auto',
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
function singleWave() {
|
||||
return {
|
||||
phaseDir: '.planning/phases/01-foo',
|
||||
runId: 'run-2285-1',
|
||||
waves: [
|
||||
{
|
||||
id: 'w1',
|
||||
plans: [
|
||||
{ id: 'p1', brief: 'Implement the foo module', files_modified: ['src/foo.cts'] },
|
||||
],
|
||||
},
|
||||
],
|
||||
};
|
||||
}
|
||||
|
||||
function baseInput(overrides = {}) {
|
||||
return {
|
||||
runtimeId: 'claude',
|
||||
hostIntegration: CAPABLE_HOST_2285,
|
||||
agentSdkVersion: ABOVE_FLOOR_SDK,
|
||||
config: enabledConfig(),
|
||||
...singleWave(),
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
// ─── Section A: happy path — every gate satisfied → workflow backend ────────
|
||||
|
||||
describe('A. resolveWaveDispatch — enabled + all gates satisfied → workflow backend with emitted script', () => {
|
||||
test('[happy] enabled, claude runtime, capable host, SDK above floor, auto backend → backend:"workflow"', () => {
|
||||
const result = resolveWaveDispatch(baseInput());
|
||||
assert.strictEqual(result.backend, 'workflow');
|
||||
assert.strictEqual(result.reason, 'workflow_backend_active');
|
||||
assert.ok(typeof result.script === 'string' && result.script.length > 0, 'script must be a non-empty string');
|
||||
// #2590: resumeFromRunId is a Workflow TOOL INPUT, not a script function —
|
||||
// calling it threw "resumeFromRunId is not defined". The run id must still
|
||||
// reach the caller (which passes it as that input), but never as a call.
|
||||
assert.ok(!/^\s*resumeFromRunId\s*\(/m.test(result.script), 'must not CALL resumeFromRunId');
|
||||
assert.strictEqual(result.summary.resumeRunId, 'run-2285-1');
|
||||
assert.match(result.script, /agentType: "gsd-executor", isolation: "worktree"/);
|
||||
assert.ok(result.summary && result.summary.plans === 1, 'summary.plans must reflect the manifest');
|
||||
});
|
||||
|
||||
test('[happy] execution_backend explicitly "workflow" (not just "auto") also activates', () => {
|
||||
const result = resolveWaveDispatch(baseInput({
|
||||
config: enabledConfig({ 'claude_orchestration.execution_backend': 'workflow' }),
|
||||
}));
|
||||
assert.strictEqual(result.backend, 'workflow');
|
||||
});
|
||||
|
||||
test('[bva] SDK version boundary: floor-1 → inline, floor exact → workflow, floor+1 → workflow', () => {
|
||||
const below = resolveWaveDispatch(baseInput({ agentSdkVersion: BELOW_FLOOR_SDK }));
|
||||
assert.strictEqual(below.backend, 'inline', 'below floor must be inline');
|
||||
assert.strictEqual(below.reason, 'agent_sdk_version_below_floor');
|
||||
|
||||
const at = resolveWaveDispatch(baseInput({ agentSdkVersion: AT_FLOOR_SDK }));
|
||||
assert.strictEqual(at.backend, 'workflow', 'exactly at floor must activate workflow');
|
||||
|
||||
const above = resolveWaveDispatch(baseInput({ agentSdkVersion: ABOVE_FLOOR_SDK }));
|
||||
assert.strictEqual(above.backend, 'workflow', 'above floor must activate workflow');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section B: fail-closed contract — disabled / each gate individually failing → inline ──
|
||||
|
||||
describe('B. resolveWaveDispatch — fail-closed contract: disabled or any gate miss → inline, matches detectWorkflowBackend 1:1', () => {
|
||||
const GATE_MISS_CASES = [
|
||||
{
|
||||
label: 'capability disabled',
|
||||
overrides: { config: {} },
|
||||
expectedReason: 'capability_disabled',
|
||||
},
|
||||
{
|
||||
label: 'capability explicitly disabled',
|
||||
overrides: { config: enabledConfig({ 'claude_orchestration.enabled': false }) },
|
||||
expectedReason: 'capability_disabled',
|
||||
},
|
||||
{
|
||||
label: 'runtime is not claude',
|
||||
overrides: { runtimeId: 'codex' },
|
||||
expectedReason: 'runtime_not_claude',
|
||||
},
|
||||
{
|
||||
label: 'execution_backend explicitly "inline"',
|
||||
overrides: { config: enabledConfig({ 'claude_orchestration.execution_backend': 'inline' }) },
|
||||
expectedReason: 'backend_inline',
|
||||
},
|
||||
{
|
||||
label: 'host descriptor incapable (nested:false)',
|
||||
overrides: { hostIntegration: INCAPABLE_HOST },
|
||||
expectedReason: 'workflow_tool_unavailable',
|
||||
},
|
||||
{
|
||||
label: 'host descriptor missing entirely',
|
||||
overrides: { hostIntegration: null },
|
||||
expectedReason: 'workflow_tool_unavailable',
|
||||
},
|
||||
{
|
||||
label: 'agent SDK version missing',
|
||||
overrides: { agentSdkVersion: undefined },
|
||||
expectedReason: 'agent_sdk_version_unknown',
|
||||
},
|
||||
{
|
||||
label: 'agent SDK version malformed (not semver)',
|
||||
overrides: { agentSdkVersion: 'not-a-version' },
|
||||
expectedReason: 'agent_sdk_version_unknown',
|
||||
},
|
||||
{
|
||||
label: 'agent SDK version below floor',
|
||||
overrides: { agentSdkVersion: BELOW_FLOOR_SDK },
|
||||
expectedReason: 'agent_sdk_version_below_floor',
|
||||
},
|
||||
];
|
||||
|
||||
for (const { label, overrides, expectedReason } of GATE_MISS_CASES) {
|
||||
test(`[negative] ${label} → backend:"inline", reason:"${expectedReason}"`, () => {
|
||||
const input = baseInput(overrides);
|
||||
const result = resolveWaveDispatch(input);
|
||||
|
||||
assert.strictEqual(result.backend, 'inline', `${label}: must resolve to inline`);
|
||||
assert.strictEqual(result.reason, expectedReason, `${label}: reason mismatch`);
|
||||
|
||||
// Fail-closed CONTRACT: today's (byte-identical) inline dispatch carries no
|
||||
// script/summary. Verify the shape never leaks emitter fields on a gate miss.
|
||||
assert.deepStrictEqual(
|
||||
Object.keys(result).sort(),
|
||||
['backend', 'reason'],
|
||||
`${label}: inline result must be exactly {backend, reason}, got keys: ${Object.keys(result).join(',')}`,
|
||||
);
|
||||
|
||||
// Parity: resolveWaveDispatch must not reimplement the gate ladder — its
|
||||
// reason for a detect-side miss must be IDENTICAL to calling
|
||||
// detectWorkflowBackend directly with the same gate-relevant fields.
|
||||
const direct = detectWorkflowBackend({
|
||||
runtimeId: input.runtimeId,
|
||||
hostIntegration: input.hostIntegration,
|
||||
config: input.config,
|
||||
agentSdkVersion: input.agentSdkVersion,
|
||||
});
|
||||
assert.strictEqual(direct.backend, 'inline', `${label}: detectWorkflowBackend parity check must also be inline`);
|
||||
assert.strictEqual(result.reason, direct.reason, `${label}: resolveWaveDispatch must surface detectWorkflowBackend's own reason verbatim`);
|
||||
});
|
||||
}
|
||||
|
||||
test('[negative] null/undefined/non-object input → inline, reason:"invalid_input" (never throws)', () => {
|
||||
assert.deepStrictEqual(resolveWaveDispatch(null), { backend: 'inline', reason: 'invalid_input' });
|
||||
assert.deepStrictEqual(resolveWaveDispatch(undefined), { backend: 'inline', reason: 'invalid_input' });
|
||||
assert.deepStrictEqual(resolveWaveDispatch('not-an-object'), { backend: 'inline', reason: 'invalid_input' });
|
||||
});
|
||||
|
||||
test('[happy] a valid, dispatch-ready waves manifest never flips a gate-missed decision to workflow', () => {
|
||||
// Prove the gate ladder short-circuits BEFORE emitWorkflowScript ever runs:
|
||||
// even with a perfectly valid wave manifest, a disabled capability stays inline.
|
||||
const result = resolveWaveDispatch(baseInput({ config: {}, ...singleWave() }));
|
||||
assert.strictEqual(result.backend, 'inline');
|
||||
assert.strictEqual(result.reason, 'capability_disabled');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section C: composition correctness — detect + emit have a real, non-CLI, non-test caller ──
|
||||
|
||||
describe('C. resolveWaveDispatch composes detectWorkflowBackend + emitWorkflowScript (the seam itself)', () => {
|
||||
test('[happy] resolveWaveDispatch is exported as a function from the core module', () => {
|
||||
assert.strictEqual(typeof resolveWaveDispatch, 'function');
|
||||
});
|
||||
|
||||
test('[happy] on a workflow-hit, the emitted script/summary are IDENTICAL to calling emitWorkflowScript directly with the same wave data', () => {
|
||||
const input = baseInput();
|
||||
const composed = resolveWaveDispatch(input);
|
||||
assert.strictEqual(composed.backend, 'workflow');
|
||||
|
||||
const directEmit = emitWorkflowScript({
|
||||
phaseDir: input.phaseDir,
|
||||
waves: input.waves,
|
||||
runId: input.runId,
|
||||
});
|
||||
assert.strictEqual(directEmit.ok, true);
|
||||
assert.strictEqual(composed.script, directEmit.script, 'resolveWaveDispatch must not re-implement emission — script must match emitWorkflowScript byte-for-byte');
|
||||
assert.deepStrictEqual(composed.summary, directEmit.summary);
|
||||
});
|
||||
|
||||
test('[negative] detect-hit but a malformed wave manifest (emit failure) → inline, carrying emitWorkflowScript\'s own failure reason', () => {
|
||||
const input = baseInput({ waves: [] }); // emitWorkflowScript rejects empty waves
|
||||
const result = resolveWaveDispatch(input);
|
||||
assert.strictEqual(result.backend, 'inline');
|
||||
|
||||
const directEmit = emitWorkflowScript({ phaseDir: input.phaseDir, waves: input.waves, runId: input.runId });
|
||||
assert.strictEqual(directEmit.ok, false);
|
||||
assert.strictEqual(result.reason, 'emit_failed: ' + directEmit.reason, 'the emit failure reason must be surfaced verbatim, prefixed');
|
||||
|
||||
// Still byte-identical inline shape — no partial/broken script ever leaks.
|
||||
assert.deepStrictEqual(Object.keys(result).sort(), ['backend', 'reason']);
|
||||
});
|
||||
|
||||
test('[happy] the CLI subcommand `claude-orchestration resolve-wave-dispatch` is ALSO a caller and matches the pure function output', (t) => {
|
||||
const tmp = createTempDir('fix-2285-');
|
||||
t.after(() => cleanup(tmp));
|
||||
fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(
|
||||
path.join(tmp, '.planning', 'config.json'),
|
||||
JSON.stringify({ claude_orchestration: { enabled: true, execution_backend: 'auto' } }),
|
||||
);
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, JSON.stringify({ waves: singleWave().waves }));
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', wavesPath,
|
||||
'--run-id', 'run-2285-1',
|
||||
'--phase-dir', '.planning/phases/01-foo',
|
||||
'--runtime', 'claude',
|
||||
'--agent-sdk-version', ABOVE_FLOOR_SDK,
|
||||
// #2686: the router now DEFAULTS the executor model from project config,
|
||||
// so pin it on both sides — otherwise this compares a config-resolved CLI
|
||||
// run against a pure call that was given no model, and the equality this
|
||||
// test exists to prove would be testing the default instead of the seam.
|
||||
'--executor-model', 'sonnet',
|
||||
'--raw',
|
||||
], tmp);
|
||||
assert.strictEqual(res.success, true, 'CLI command must succeed; stderr: ' + (res.error || ''));
|
||||
const parsed = JSON.parse(res.output);
|
||||
|
||||
const direct = resolveWaveDispatch(baseInput({ executorModel: 'sonnet' }));
|
||||
assert.strictEqual(parsed.backend, direct.backend);
|
||||
assert.strictEqual(parsed.script, direct.script);
|
||||
assert.deepStrictEqual(parsed.summary, direct.summary);
|
||||
});
|
||||
|
||||
test('[negative] CLI subcommand fails closed to inline exactly like the pure function when disabled', (t) => {
|
||||
const tmp = createTempDir('fix-2285-off-');
|
||||
t.after(() => cleanup(tmp));
|
||||
fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(path.join(tmp, '.planning', 'config.json'), '{}');
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, JSON.stringify({ waves: singleWave().waves }));
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', wavesPath, '--run-id', 'run-x', '--raw',
|
||||
], tmp);
|
||||
assert.strictEqual(res.success, true, 'CLI command must succeed (fail-closed, not error); stderr: ' + (res.error || ''));
|
||||
const parsed = JSON.parse(res.output);
|
||||
assert.strictEqual(parsed.backend, 'inline');
|
||||
assert.strictEqual(parsed.reason, 'capability_disabled');
|
||||
assert.deepStrictEqual(Object.keys(parsed).sort(), ['backend', 'reason']);
|
||||
});
|
||||
|
||||
test('property: for ANY input, resolveWaveDispatch never throws, backend is always "inline"|"workflow", and "inline" results carry exactly {backend, reason}', () => {
|
||||
fc.assert(fc.property(
|
||||
fc.record({
|
||||
runtimeId: fc.oneof(fc.constant('claude'), fc.constant('codex'), fc.constant(undefined), fc.string()),
|
||||
agentSdkVersion: fc.oneof(fc.constant(ABOVE_FLOOR_SDK), fc.constant(BELOW_FLOOR_SDK), fc.constant(undefined), fc.string()),
|
||||
enabled: fc.boolean(),
|
||||
capableHost: fc.boolean(),
|
||||
backendPref: fc.constantFrom('auto', 'workflow', 'inline'),
|
||||
}),
|
||||
({ runtimeId, agentSdkVersion, enabled, capableHost, backendPref }) => {
|
||||
const input = {
|
||||
runtimeId,
|
||||
hostIntegration: capableHost ? CAPABLE_HOST_2285 : INCAPABLE_HOST,
|
||||
agentSdkVersion,
|
||||
config: {
|
||||
'claude_orchestration.enabled': enabled,
|
||||
'claude_orchestration.execution_backend': backendPref,
|
||||
},
|
||||
...singleWave(),
|
||||
};
|
||||
const result = resolveWaveDispatch(input);
|
||||
assert.ok(result.backend === 'inline' || result.backend === 'workflow');
|
||||
if (result.backend === 'inline') {
|
||||
assert.deepStrictEqual(Object.keys(result).sort(), ['backend', 'reason']);
|
||||
} else {
|
||||
assert.ok(typeof result.script === 'string' && result.script.length > 0);
|
||||
}
|
||||
},
|
||||
));
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section D: capability declaration now targets execute:wave:pre ─────────
|
||||
|
||||
describe('D. capability.json declares the contribution at execute:wave:pre (#2285)', () => {
|
||||
test('[happy] contribution point is execute:wave:pre, not execute:wave:post', () => {
|
||||
const cap = JSON.parse(fs.readFileSync(CAP_PATH, 'utf8'));
|
||||
const wavePreContrib = cap.contributions.find((c) => c.point === 'execute:wave:pre');
|
||||
assert.ok(wavePreContrib, 'capability.json must declare a contribution at execute:wave:pre');
|
||||
assert.strictEqual(wavePreContrib.into, 'executor');
|
||||
assert.strictEqual(wavePreContrib.when, 'claude_orchestration.enabled');
|
||||
assert.strictEqual(wavePreContrib.onError, 'skip');
|
||||
assert.strictEqual(wavePreContrib.fragment.path, 'fragments/execute-wave-pre.md');
|
||||
|
||||
const wavePostContrib = cap.contributions.find((c) => c.point === 'execute:wave:post');
|
||||
assert.strictEqual(wavePostContrib, undefined, 'the capability must no longer contribute at execute:wave:post');
|
||||
});
|
||||
|
||||
test('[happy] the declared fragment file exists on disk', () => {
|
||||
const fragPath = path.join(ROOT, 'capabilities', 'claude-orchestration', 'fragments', 'execute-wave-pre.md');
|
||||
assert.ok(fs.existsSync(fragPath), 'fragments/execute-wave-pre.md must exist');
|
||||
const content = fs.readFileSync(fragPath, 'utf8');
|
||||
assert.match(content, /execute:wave:pre/);
|
||||
assert.match(content, /resolve-wave-dispatch/);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section E: source-contract guard — execute-phase.md renders execute:wave:pre BEFORE dispatch ──
|
||||
|
||||
// allow-test-rule: source-text-is-the-product, see #2285 — reads gsd-core/workflows/execute-phase.md
|
||||
// prose to verify the render-hooks call site + ordering. The workflow markdown IS the runtime
|
||||
// contract executed by the orchestrator; there is no behavioral seam to drive this assertion
|
||||
// through other than the rendered prose itself.
|
||||
|
||||
describe('E. execute-phase.md actually renders execute:wave:pre (the dead hook is now live)', () => {
|
||||
test('[happy] execute-phase.md invokes `loop render-hooks execute:wave:pre`', () => {
|
||||
const doc = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
||||
assert.ok(
|
||||
/loop render-hooks execute:wave:pre/.test(doc),
|
||||
'execute-phase.md must dispatch execute:wave:pre hooks (was declared in frontmatter but never rendered — #2285)',
|
||||
);
|
||||
});
|
||||
|
||||
test('[happy] the execute:wave:pre render-hooks call site appears BEFORE the wave\'s Agent() dispatch (pre-wave, not post)', () => {
|
||||
const doc = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
||||
const preHooksIdx = doc.indexOf('loop render-hooks execute:wave:pre');
|
||||
// Anchor on the actual per-wave dispatch call (step 3), not the generic
|
||||
// `subagent_type="gsd-executor"` mention in <runtime_compatibility> near the
|
||||
// top of the file — that mention predates the wave loop entirely and would
|
||||
// give a false "before" reading.
|
||||
const agentDispatchIdx = doc.indexOf('description="Execute plan {plan_number}');
|
||||
assert.ok(preHooksIdx !== -1, 'execute:wave:pre render-hooks call site must exist');
|
||||
assert.ok(agentDispatchIdx !== -1, 'the gsd-executor Agent() dispatch call site (step 3) must exist');
|
||||
assert.ok(
|
||||
preHooksIdx < agentDispatchIdx,
|
||||
`execute:wave:pre render-hooks (idx ${preHooksIdx}) must appear BEFORE the wave's Agent() dispatch (idx ${agentDispatchIdx}) — it is a pre-wave hook`,
|
||||
);
|
||||
});
|
||||
|
||||
test('[happy] the frontmatter still declares all four execute:* points (regression guard)', () => {
|
||||
const doc = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
||||
const frontmatterMatch = doc.match(/points:\s*(.+)/);
|
||||
assert.ok(frontmatterMatch, 'frontmatter must declare a points: line');
|
||||
for (const point of ['execute:pre', 'execute:wave:pre', 'execute:wave:post', 'execute:post']) {
|
||||
assert.ok(frontmatterMatch[1].includes(point), `frontmatter points: line must include ${point}`);
|
||||
}
|
||||
});
|
||||
|
||||
test('[happy] execute:wave:post is still rendered too (regression guard — did not accidentally remove the post-wave gate dispatch)', () => {
|
||||
const doc = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
||||
assert.ok(
|
||||
/loop render-hooks execute:wave:post/.test(doc),
|
||||
'execute-phase.md must still dispatch execute:wave:post hooks (drift/ui gates unaffected by #2285)',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section F: orthogonal-review finding 1 — submodule plans never forced into worktree isolation ──
|
||||
//
|
||||
// #2772 / #2285 finding 1: emitWorkflowScript previously hardcoded
|
||||
// `isolation: "worktree"` for EVERY plan. execute-phase.md step 2.5 computes
|
||||
// USE_WORKTREES_FOR_PLAN per plan specifically to keep submodule-touching
|
||||
// plans OUT of worktree isolation (the executor commit protocol cannot
|
||||
// correctly handle submodule commits inside an isolated worktree). The
|
||||
// Workflow backend must honor the SAME per-plan decision via `use_worktree`.
|
||||
|
||||
function waveWithSubmodulePlan() {
|
||||
return {
|
||||
phaseDir: '.planning/phases/01-foo',
|
||||
runId: 'run-2285-submodule',
|
||||
waves: [
|
||||
{
|
||||
id: 'w1',
|
||||
plans: [
|
||||
{ id: 'p1', brief: 'normal plan', files_modified: ['src/a.ts'] },
|
||||
{ id: 'p2', brief: 'submodule plan', files_modified: ['vendor/lib.c'], use_worktree: false },
|
||||
],
|
||||
},
|
||||
],
|
||||
};
|
||||
}
|
||||
|
||||
describe('F. Workflow backend never forces worktree isolation on a submodule / use_worktree:false plan', () => {
|
||||
test('[happy] resolveWaveDispatch (pure seam): the submodule plan\'s agent() call carries NO isolation, the normal plan\'s does', () => {
|
||||
const result = resolveWaveDispatch(baseInput({ ...waveWithSubmodulePlan() }));
|
||||
assert.strictEqual(result.backend, 'workflow');
|
||||
assert.match(result.script, /agent\("normal plan", \{ agentType: "gsd-executor", isolation: "worktree" \}\)/);
|
||||
// #2686: the options object legitimately gained an optional `model` key, so assert
|
||||
// the invariant this test exists to protect — agentType present, isolation absent —
|
||||
// rather than a frozen literal that any future additive key would break.
|
||||
assert.match(result.script, /agent\("submodule plan", \{ agentType: "gsd-executor"[^}]*\}\)/);
|
||||
assert.ok(
|
||||
!/agent\("submodule plan"[^)]*isolation/.test(result.script),
|
||||
'the submodule-touching plan must NEVER be emitted with forced worktree isolation',
|
||||
);
|
||||
});
|
||||
|
||||
test('[happy] CLI `resolve-wave-dispatch`: same per-plan guarantee end-to-end through the subprocess', (t) => {
|
||||
const tmp = createTempDir('fix-2285-submodule-');
|
||||
t.after(() => cleanup(tmp));
|
||||
fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(
|
||||
path.join(tmp, '.planning', 'config.json'),
|
||||
JSON.stringify({ claude_orchestration: { enabled: true, execution_backend: 'auto' } }),
|
||||
);
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, JSON.stringify({ waves: waveWithSubmodulePlan().waves }));
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', wavesPath,
|
||||
'--run-id', 'run-2285-submodule',
|
||||
'--phase-dir', '.planning/phases/01-foo',
|
||||
'--runtime', 'claude',
|
||||
'--agent-sdk-version', ABOVE_FLOOR_SDK,
|
||||
'--raw',
|
||||
], tmp);
|
||||
assert.strictEqual(res.success, true, 'CLI command must succeed; stderr: ' + (res.error || ''));
|
||||
const parsed = JSON.parse(res.output);
|
||||
assert.strictEqual(parsed.backend, 'workflow');
|
||||
// #2686: additive `model` key — see the note on the pure-seam test above.
|
||||
assert.match(parsed.script, /agent\("submodule plan", \{ agentType: "gsd-executor"[^}]*\}\)/);
|
||||
assert.ok(!/agent\("submodule plan"[^)]*isolation/.test(parsed.script));
|
||||
});
|
||||
|
||||
test('[negative] use_worktree defaults to true when omitted — a manifest with NO submodule info stays backward-compatible', () => {
|
||||
const result = resolveWaveDispatch(baseInput());
|
||||
assert.strictEqual(result.backend, 'workflow');
|
||||
assert.match(result.script, /isolation: "worktree"/, 'default (no use_worktree field) must still isolate — backward compatible');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section G: orthogonal-review finding 2 — missing top-level `waves` key must never silently exit 0 ──
|
||||
//
|
||||
// readWavesManifest previously collapsed "read/parse threw" and "parsed OK but
|
||||
// no top-level `waves` key" into the same `undefined` sentinel. The call sites'
|
||||
// `if (waves === undefined) return;` made the missing-key case exit 0 with ZERO
|
||||
// output — fail-silent, breaking the "exit 0 => parseable JSON verdict" contract.
|
||||
// A missing key must now flow through to emitWorkflowScript's own validation,
|
||||
// exactly like an explicit `{"waves": null}` manifest already does.
|
||||
|
||||
describe('G. missing top-level `waves` key never silently exits 0 with no output', () => {
|
||||
function projectWithEnabledCapability(prefix) {
|
||||
const tmp = createTempDir(prefix);
|
||||
fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(
|
||||
path.join(tmp, '.planning', 'config.json'),
|
||||
JSON.stringify({ claude_orchestration: { enabled: true, execution_backend: 'auto' } }),
|
||||
);
|
||||
return tmp;
|
||||
}
|
||||
|
||||
test('[negative] resolve-wave-dispatch with a {"notwaves":[]} manifest → non-empty JSON verdict (NOT silent exit 0)', (t) => {
|
||||
const tmp = projectWithEnabledCapability('fix-2285-missingkey-resolve-');
|
||||
t.after(() => cleanup(tmp));
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, JSON.stringify({ notwaves: [] }));
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', wavesPath, '--run-id', 'run-x',
|
||||
'--runtime', 'claude', '--agent-sdk-version', ABOVE_FLOOR_SDK,
|
||||
'--raw',
|
||||
], tmp);
|
||||
|
||||
assert.strictEqual(res.success, true, 'command must exit 0 (fail-closed to inline, not error); stderr: ' + (res.error || ''));
|
||||
assert.ok(res.output.length > 0, 'FAIL-SILENT REGRESSION: missing waves key must NOT produce empty stdout on exit 0');
|
||||
const parsed = JSON.parse(res.output);
|
||||
assert.strictEqual(parsed.backend, 'inline');
|
||||
assert.match(parsed.reason, /waves must be a non-empty array/, 'reason must surface emitWorkflowScript\'s own validation message');
|
||||
});
|
||||
|
||||
test('[negative] resolve-wave-dispatch: {"notwaves":[]} and {"waves": null} produce the IDENTICAL verdict (parity)', (t) => {
|
||||
const tmp = projectWithEnabledCapability('fix-2285-missingkey-parity-');
|
||||
t.after(() => cleanup(tmp));
|
||||
const missingKeyPath = path.join(tmp, 'missing.json');
|
||||
fs.writeFileSync(missingKeyPath, JSON.stringify({ notwaves: [] }));
|
||||
const nullWavesPath = path.join(tmp, 'null.json');
|
||||
fs.writeFileSync(nullWavesPath, JSON.stringify({ waves: null }));
|
||||
|
||||
const argsFor = (p) => [
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', p, '--run-id', 'run-x',
|
||||
'--runtime', 'claude', '--agent-sdk-version', ABOVE_FLOOR_SDK,
|
||||
'--raw',
|
||||
];
|
||||
const missingRes = runGsdTools(argsFor(missingKeyPath), tmp);
|
||||
const nullRes = runGsdTools(argsFor(nullWavesPath), tmp);
|
||||
assert.strictEqual(missingRes.success, true);
|
||||
assert.strictEqual(nullRes.success, true);
|
||||
assert.deepStrictEqual(JSON.parse(missingRes.output), JSON.parse(nullRes.output), 'a missing `waves` key must behave identically to an explicit `waves: null`');
|
||||
});
|
||||
|
||||
test('[negative] emit-workflow with a {"notwaves":[]} manifest → loud non-zero exit (NOT silent exit 0)', (t) => {
|
||||
const tmp = createTempDir('fix-2285-missingkey-emit-');
|
||||
t.after(() => cleanup(tmp));
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, JSON.stringify({ notwaves: [] }));
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'emit-workflow',
|
||||
'--waves', wavesPath, '--run-id', 'run-x',
|
||||
], tmp);
|
||||
|
||||
assert.strictEqual(res.success, false, 'FAIL-SILENT REGRESSION: missing waves key must produce a loud, non-zero-exit error, not a silent success');
|
||||
assert.ok(res.exitCode !== 0, 'non-zero exit');
|
||||
assert.match(res.error || '', /waves must be a non-empty array/);
|
||||
});
|
||||
|
||||
test('[happy] a genuinely malformed (unparseable) --waves file still fails loudly, unaffected by the fix', (t) => {
|
||||
const tmp = createTempDir('fix-2285-badjson-');
|
||||
t.after(() => cleanup(tmp));
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, 'not json at all');
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', wavesPath, '--run-id', 'run-x', '--raw',
|
||||
], tmp);
|
||||
assert.strictEqual(res.success, false, 'a real parse failure must still error');
|
||||
assert.match(res.error || '', /could not read\/parse --waves file/);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section H: orthogonal-review finding 3 — manifest construction guidance is concrete ──
|
||||
|
||||
describe('H. the execute:wave:pre fragment documents concrete manifest construction (finding 3)', () => {
|
||||
test('[happy] the fragment explains how to build WAVE_MANIFEST_PATH, PHASE_RUN_ID, and per-plan use_worktree', () => {
|
||||
const fragPath = path.join(ROOT, 'capabilities', 'claude-orchestration', 'fragments', 'execute-wave-pre.md');
|
||||
const content = fs.readFileSync(fragPath, 'utf8');
|
||||
assert.match(content, /Manifest construction/, 'fragment must have concrete manifest-construction guidance, not just reference undefined vars');
|
||||
assert.match(content, /PHASE_RUN_ID/);
|
||||
assert.match(content, /WAVE_MANIFEST_PATH/);
|
||||
assert.match(content, /use_worktree/);
|
||||
assert.match(content, /USE_WORKTREES_FOR_PLAN/, 'must tie use_worktree back to step 2.5\'s per-plan decision');
|
||||
});
|
||||
|
||||
test('[happy] execute-phase.md step 2.75 stays minimal — manifest/use_worktree detail lives ONLY in the fragment (#1168 byte-budget conformance)', () => {
|
||||
// Per the ADR-857 Phase 6 conformance gate (tests/phase6-capstone-conformance.test.cjs),
|
||||
// the host loop must stay small — optional-feature detail (manifest construction,
|
||||
// per-plan use_worktree carry-through) belongs in the capability fragment, not the
|
||||
// host workflow. Step 2.75 is intentionally just a render-hooks call + a one-line
|
||||
// "follow the contribution or fall through to step 3" instruction.
|
||||
const doc = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
||||
const stepStart = doc.indexOf('2.75. **Execute:wave:pre capability dispatch:**');
|
||||
const stepEnd = doc.indexOf('\n3. **Spawn executor agents:**', stepStart);
|
||||
assert.ok(stepStart !== -1 && stepEnd !== -1, 'step 2.75 must exist and precede step 3');
|
||||
const stepBody = doc.slice(stepStart, stepEnd);
|
||||
assert.match(stepBody, /loop render-hooks execute:wave:pre/, 'step 2.75 must still render the hook point');
|
||||
assert.doesNotMatch(stepBody, /use_worktree/, 'manifest-construction detail (use_worktree) must live in the fragment, not the host step');
|
||||
assert.doesNotMatch(stepBody, /USE_WORKTREES_FOR_PLAN/, 'per-plan worktree gate detail must live in the fragment, not the host step');
|
||||
});
|
||||
|
||||
test('[happy] execute-phase.md is below the ADR-857 Phase 6 pre-phase-6 byte ceiling (#1168), with margin', () => {
|
||||
const { lfByteCount } = require('../scripts/workflow-size.cjs');
|
||||
const bytes = lfByteCount(WORKFLOW_PATH);
|
||||
assert.ok(bytes < 93600, `execute-phase.md must stay below the frozen pre-phase-6 ceiling (93600); got ${bytes}`);
|
||||
assert.ok(bytes <= 93400, `execute-phase.md should carry a comfortable margin (<=93400) so minor future edits don't re-trip the gate; got ${bytes}`);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section I: orthogonal-review finding 4 — stale doc fixed ───────────────
|
||||
|
||||
describe('I. docs/explanation/claude-orchestration-capability.md reflects the execute:wave:pre move (finding 4)', () => {
|
||||
test('[happy] the doc no longer claims the capability registers at execute:wave:post', () => {
|
||||
const docPath = path.join(ROOT, 'docs', 'explanation', 'claude-orchestration-capability.md');
|
||||
const content = fs.readFileSync(docPath, 'utf8');
|
||||
assert.match(content, /execute:wave:pre/, 'doc must mention execute:wave:pre as the wired point');
|
||||
assert.ok(
|
||||
!/execute:wave:post.*\(into the executor\)/.test(content),
|
||||
'doc must not still claim the wired point is execute:wave:post',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── #2590 — emitted Workflow scripts satisfy the Workflow tool contract
|
||||
// (folded from fix-2590-workflow-script-contract.test.cjs) ────────────────────
|
||||
//
|
||||
// `emitWorkflowScript` generated four constructs the tool does not accept. The
|
||||
// first was fatal on its own, so the backend could never dispatch a wave:
|
||||
//
|
||||
// 1. no `export const meta = {…}` first statement -> whole script rejected
|
||||
// 2. `resumeFromRunId("<id>")` -> "resumeFromRunId is not defined"
|
||||
// (it is a Workflow TOOL INPUT parameter, not a script function)
|
||||
// 3. `budget(<n>)` -> "budget is not a function"
|
||||
// (`budget` is a read-only object { total, spent(), remaining() })
|
||||
// 4. `parallel(agent(…), agent(…))` -> "parallel() expects an array of functions"
|
||||
//
|
||||
// Plus two secondary defects that kept the emitted script from ever being
|
||||
// REACHED, which is why this shipped undetected:
|
||||
//
|
||||
// 5. nothing resolved the Agent SDK version, so gate 5 returned
|
||||
// `agent_sdk_version_unknown` on every automated run
|
||||
// 6. the runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging
|
||||
// from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any
|
||||
// invocation without --runtime reported `runtime_not_claude`
|
||||
//
|
||||
// The script assertions parse the emitted text as a real ES module rather than
|
||||
// pattern-matching it, so a syntactically invalid script fails outright.
|
||||
|
||||
const TOOLS_2590 = path.join(ROOT, 'gsd-core', 'bin', 'gsd-tools.cjs');
|
||||
|
||||
function emit(overrides) {
|
||||
const input = Object.assign({
|
||||
phaseDir: '.planning/phases/01',
|
||||
runId: 'execute-1',
|
||||
waves: [{ id: 'wave-1', plans: [{ id: '01-01', brief: 'noop', files_modified: ['a.ts'] }] }],
|
||||
}, overrides || {});
|
||||
const r = emitWorkflowScript(input);
|
||||
assert.ok(r.ok, `emit failed: ${JSON.stringify(r)}`);
|
||||
return r;
|
||||
}
|
||||
|
||||
/** First non-comment, non-blank line — the script's first actual statement. */
|
||||
function firstStatement(script) {
|
||||
return script.split('\n').map((l) => l.trim())
|
||||
.find((l) => l.length > 0 && !l.startsWith('//')) || '';
|
||||
}
|
||||
|
||||
describe('#2590: emitted Workflow scripts satisfy the Workflow tool contract', () => {
|
||||
test('the emitted script is syntactically valid as an ES module', (t) => {
|
||||
// `export const meta` + top-level `await` only parse in module context —
|
||||
// which is exactly the context the Workflow tool runs the script in.
|
||||
const { script } = emit({
|
||||
waves: [
|
||||
{ id: 'w1', plans: [
|
||||
{ id: 'a', brief: 'one', files_modified: ['a.ts'] },
|
||||
{ id: 'b', brief: 'two', files_modified: ['b.ts'] },
|
||||
] },
|
||||
{ id: 'w2', plans: [{ id: 'c', brief: 'three', files_modified: ['c.ts'] }] },
|
||||
],
|
||||
});
|
||||
// .mjs so node parses it in module context, inside a helper temp dir so
|
||||
// cleanup() carries the Windows-EBUSY retry budget.
|
||||
const dir = createTempDir('gsd-2590-parse-');
|
||||
t.after(() => cleanup(dir));
|
||||
const f = path.join(dir, 'emitted.mjs');
|
||||
fs.writeFileSync(f, script);
|
||||
const result = runNode(['--check', f], { timeoutMs: PROBE_TIMEOUT_MS });
|
||||
throwIfFailed(result, `node --check ${f} (emitted script must parse)`);
|
||||
});
|
||||
|
||||
test('1. `export const meta` is the first statement', () => {
|
||||
const { script } = emit();
|
||||
assert.match(
|
||||
firstStatement(script),
|
||||
/^export const meta = \{/,
|
||||
'the Workflow tool rejects any script whose first statement is not the meta block',
|
||||
);
|
||||
});
|
||||
|
||||
test('meta.phases titles match the emitted phase() calls exactly', () => {
|
||||
// The tool matches phase titles by exact string; a mismatch silently splits
|
||||
// progress into an unnamed group.
|
||||
const { script } = emit({
|
||||
waves: [
|
||||
{ id: 'alpha', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] },
|
||||
{ id: 'beta', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] },
|
||||
],
|
||||
});
|
||||
const metaTitles = [...script.matchAll(/\{ title: "([^"]+)"/g)].map((m) => m[1]);
|
||||
const phaseTitles = [...script.matchAll(/^phase\("([^"]+)"\)/gm)].map((m) => m[1]);
|
||||
assert.deepEqual(metaTitles, ['Wave alpha', 'Wave beta']);
|
||||
assert.deepEqual(phaseTitles, metaTitles, 'phase() titles must match meta.phases exactly');
|
||||
});
|
||||
|
||||
test('duplicate wave ids are rejected (phase titles must map 1:1)', () => {
|
||||
// Two waves sharing an id emit two identical `phase("Wave x")` calls and two
|
||||
// identical meta.phases entries; the tool matches titles by exact string, so
|
||||
// the second wave's agents would be attributed to the first's progress group.
|
||||
const r = emitWorkflowScript({
|
||||
phaseDir: '.planning/phases/01',
|
||||
runId: 'execute-1',
|
||||
waves: [
|
||||
{ id: 'dup', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] },
|
||||
{ id: 'dup', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] },
|
||||
],
|
||||
});
|
||||
assert.equal(r.ok, false, 'duplicate wave ids must be rejected, not silently merged');
|
||||
assert.match(String(r.reason), /duplicate wave id/);
|
||||
});
|
||||
|
||||
test('distinct wave ids are still accepted (the boundary either side)', () => {
|
||||
const r = emitWorkflowScript({
|
||||
phaseDir: '.planning/phases/01',
|
||||
runId: 'execute-1',
|
||||
waves: [
|
||||
{ id: 'w1', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] },
|
||||
{ id: 'w2', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] },
|
||||
],
|
||||
});
|
||||
assert.equal(r.ok, true, `distinct wave ids must pass: ${JSON.stringify(r)}`);
|
||||
});
|
||||
|
||||
test('2. resumeFromRunId is never CALLED (it is a tool input, not a function)', () => {
|
||||
const { script, summary } = emit({ runId: 'execute-7' });
|
||||
assert.ok(
|
||||
!/^\s*resumeFromRunId\s*\(/m.test(script),
|
||||
'calling resumeFromRunId() throws "resumeFromRunId is not defined"',
|
||||
);
|
||||
// The run id must still reach the caller, which passes it as the tool input.
|
||||
assert.equal(summary.resumeRunId, 'execute-7');
|
||||
});
|
||||
|
||||
test('3. budget is never CALLED, at and around the boundary', () => {
|
||||
// budgetTokens is floored at > 0; check 0 (rejected), 1 (accepted), and a
|
||||
// large value — none may produce a budget(...) call.
|
||||
for (const tokens of [0, 1, 500000]) {
|
||||
const { script, summary } = emit({ budgetTokens: tokens });
|
||||
assert.ok(
|
||||
!/^\s*budget\s*\(/m.test(script),
|
||||
`budgetTokens=${tokens}: calling budget() throws "budget is not a function"`,
|
||||
);
|
||||
assert.equal(summary.budgetTokens, tokens > 0 ? tokens : null);
|
||||
}
|
||||
});
|
||||
|
||||
test('4. parallel() receives an array of thunks, not agent() results', () => {
|
||||
const { script } = emit({
|
||||
waves: [{ id: 'w', plans: [
|
||||
{ id: 'a', brief: 'one', files_modified: ['a.ts'] },
|
||||
{ id: 'b', brief: 'two', files_modified: ['b.ts'] },
|
||||
] }],
|
||||
});
|
||||
assert.ok(/parallel\(\[/.test(script), 'parallel() expects an array of functions');
|
||||
assert.ok(
|
||||
!/parallel\(\s*agent\(/.test(script),
|
||||
'passing agent() results directly both throws and starts every agent eagerly',
|
||||
);
|
||||
// Each agent must be wrapped in a thunk so parallel() can bound concurrency.
|
||||
const agents = [...script.matchAll(/agent\("/g)].length;
|
||||
const thunks = [...script.matchAll(/\(\) => agent\("/g)].length;
|
||||
assert.equal(thunks, agents, 'every agent() must be wrapped in a () => thunk');
|
||||
});
|
||||
|
||||
test('single-plan stages also emit an array (regression: the 1-plan branch)', () => {
|
||||
// The pre-fix code had a SEPARATE single-plan branch that emitted
|
||||
// `parallel(\n agent(...)\n)` — valid-looking but the same defect.
|
||||
const { script } = emit();
|
||||
assert.ok(/parallel\(\[/.test(script));
|
||||
assert.equal([...script.matchAll(/\(\) => agent\("/g)].length, 1);
|
||||
});
|
||||
|
||||
test('per-plan worktree isolation still mirrors use_worktree', () => {
|
||||
const { script } = emit({
|
||||
waves: [{ id: 'w', plans: [
|
||||
{ id: 'a', brief: 'iso', files_modified: ['a.ts'] },
|
||||
{ id: 'b', brief: 'noiso', files_modified: ['b.ts'], use_worktree: false },
|
||||
] }],
|
||||
});
|
||||
assert.match(script, /agent\("iso", \{ agentType: "gsd-executor", isolation: "worktree" \}\)/);
|
||||
assert.match(script, /agent\("noiso", \{ agentType: "gsd-executor" \}\)/);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2590: the backend is reachable without hand-passed flags', () => {
|
||||
function repro() {
|
||||
const dir = createTempDir('gsd-2590-repro-');
|
||||
fs.mkdirSync(path.join(dir, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(
|
||||
path.join(dir, '.planning', 'config.json'),
|
||||
JSON.stringify({ claude_orchestration: { enabled: true } }),
|
||||
);
|
||||
fs.writeFileSync(
|
||||
path.join(dir, 'waves.json'),
|
||||
JSON.stringify({ waves: [{ id: 'wave-1', plans: [{ id: '01-01', brief: 'noop', files_modified: ['a.ts'] }] }] }),
|
||||
);
|
||||
return dir;
|
||||
}
|
||||
|
||||
function resolve(dir, extraArgs) {
|
||||
const result = runNode([
|
||||
TOOLS_2590, 'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', 'waves.json', '--run-id', 'execute-1',
|
||||
'--phase-dir', '.planning/phases/01', '--raw',
|
||||
...(extraArgs || []),
|
||||
], { cwd: dir, timeoutMs: PROBE_TIMEOUT_MS });
|
||||
throwIfFailed(result, 'gsd-tools claude-orchestration resolve-wave-dispatch');
|
||||
return JSON.parse(result.stdout);
|
||||
}
|
||||
|
||||
test('5+6. no --runtime and no --agent-sdk-version still reaches the version gate', (t) => {
|
||||
const dir = repro();
|
||||
t.after(() => cleanup(dir));
|
||||
const r = resolve(dir);
|
||||
// Pre-fix this was `agent_sdk_version_unknown` (nothing resolved a
|
||||
// version) or `runtime_not_claude` (the divergent fallback). Either is a
|
||||
// regression; the version gate must now be reached and answer truthfully.
|
||||
assert.notEqual(r.reason, 'agent_sdk_version_unknown',
|
||||
'the router must resolve the installed SDK version itself');
|
||||
assert.notEqual(r.reason, 'runtime_not_claude',
|
||||
'runtime must fall back to the canonical config.runtime > claude chain');
|
||||
});
|
||||
|
||||
test('an SDK version above the floor activates the workflow backend end to end', (t) => {
|
||||
const dir = repro();
|
||||
t.after(() => cleanup(dir));
|
||||
const r = resolve(dir, ['--agent-sdk-version', '0.3.149']);
|
||||
assert.equal(r.backend, 'workflow', `expected workflow backend, got ${JSON.stringify(r)}`);
|
||||
assert.ok(typeof r.script === 'string' && r.script.length > 0);
|
||||
assert.match(firstStatement(r.script), /^export const meta = \{/);
|
||||
});
|
||||
|
||||
test('an explicit --agent-sdk-version still wins over the installed one', (t) => {
|
||||
const dir = repro();
|
||||
t.after(() => cleanup(dir));
|
||||
// A deliberately ancient pin must be honored (and decline), proving the
|
||||
// flag is not ignored now that a fallback exists.
|
||||
const r = resolve(dir, ['--agent-sdk-version', '0.0.1']);
|
||||
assert.equal(r.backend, 'inline');
|
||||
assert.equal(r.reason, 'agent_sdk_version_below_floor');
|
||||
});
|
||||
|
||||
test('GSD_AGENT_SDK_VERSION is honored between the flag and the installed version', (t) => {
|
||||
const dir = repro();
|
||||
t.after(() => cleanup(dir));
|
||||
const result = runNode([
|
||||
TOOLS_2590, 'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', 'waves.json', '--run-id', 'execute-1',
|
||||
'--phase-dir', '.planning/phases/01', '--raw',
|
||||
], { cwd: dir, env: { ...process.env, GSD_AGENT_SDK_VERSION: '0.3.149' }, timeoutMs: PROBE_TIMEOUT_MS });
|
||||
throwIfFailed(result, 'gsd-tools claude-orchestration resolve-wave-dispatch (GSD_AGENT_SDK_VERSION)');
|
||||
assert.equal(JSON.parse(result.stdout).backend, 'workflow');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -15,6 +15,8 @@ const { describe, test, beforeEach, afterEach } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
const os = require('os');
|
||||
const { spawnSync } = require('child_process');
|
||||
const { createTempGitProject, cleanup, runGsdTools } = require('./helpers.cjs');
|
||||
const { gitOrThrow } = require('./helpers/git-fixture.cjs');
|
||||
// #3145: class-norm timeout, not a per-suite value — see helpers/timeouts.cjs.
|
||||
@@ -251,3 +253,522 @@ describe('commit --files: pathspec honors declared scope (#2112)', () => {
|
||||
assert.strictEqual(status, '', `index must be clean (no pollution): ${status}`);
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* ── #2608: commit --files / commit-to-subrepo fail closed when `git add`
|
||||
* fails ──────────────────────────────────────────────────────────────────
|
||||
*
|
||||
* A `git add` that fails (unwritable index in a linked worktree whose git
|
||||
* dir is outside the managed writable root, permissions, timeout) was not
|
||||
* surfaced. #2523 above had already stopped a failed path entering the
|
||||
* commit pathspec, but skipping it silently left two bad outcomes, both
|
||||
* reproduced by this suite against the pre-fix build:
|
||||
*
|
||||
* - SOME paths fail -> `{"committed":true}`. `git commit` still ran and
|
||||
* PARTIALLY committed the subset that happened to
|
||||
* stage, under a message describing the full
|
||||
* requested scope.
|
||||
* - EVERY path fails -> `{"reason":"nothing_to_commit"}`, which is not
|
||||
* what happened and points the operator nowhere.
|
||||
*
|
||||
* In both cases the original `git add` stderr was discarded, so the user
|
||||
* saw a downstream `commit_failed` / pathspec error naming an innocent
|
||||
* file.
|
||||
*
|
||||
* The fix collects staging failures and fails closed BEFORE `git commit`
|
||||
* runs, returning `staging_failed` (or `staging_timeout`) with the
|
||||
* offending file and the original stderr preserved.
|
||||
*
|
||||
* ── INJECTION SEAM ──────────────────────────────────────────────────────
|
||||
* `execGit` is monkeypatched on the shell-command-projection module object.
|
||||
* The compiled call site is `(0, mod.execGit)(...)` — a property lookup at
|
||||
* call time — so the override takes effect. Per CLAUDE.md this is required
|
||||
* over `chmod 0o000` permission tricks, which do not fault under root (root
|
||||
* Docker/CI) and would make these tests silently vacuous.
|
||||
*
|
||||
* The patched call runs in a short-lived `node -e` CHILD rather than
|
||||
* in-process, for two reasons: `output()` writes with `fs.writeSync(1, …)`,
|
||||
* which neither `process.stdout.write` nor `console.log` interception can
|
||||
* capture; and a child keeps the patch from leaking into sibling suites. It
|
||||
* is a plain `process.execPath` spawn — no PATH stub and no exec bit, so it
|
||||
* is not subject to DEFECT.WINDOWS-TEST-PORTABILITY and runs on every
|
||||
* platform.
|
||||
*/
|
||||
|
||||
// Git plumbing (add/commit/status/rev-parse/rev-list/diff) on a small
|
||||
// mkdtemp fixture repo, for the #2608 suite below only. Kept as its own
|
||||
// local constant (distinct from the shared GIT_TIMEOUT_MS imported above)
|
||||
// per helpers/timeouts.cjs's own guidance: a call site whose class
|
||||
// genuinely differs keeps its own justified value rather than forcing a
|
||||
// shared norm that doesn't describe it — 5000ms here vs. 15000ms for the
|
||||
// shared DEFAULT_GIT_TIMEOUT_MS norm.
|
||||
const STAGING_GIT_TIMEOUT_MS = 5000;
|
||||
|
||||
const LIB = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib');
|
||||
|
||||
/**
|
||||
* Run cmdCommit with `git add <file>` forced to fail for the paths in `failFor`,
|
||||
* returning the parsed JSON result and the git argv list that was actually
|
||||
* executed (so "git commit never ran" is asserted directly, not inferred).
|
||||
*/
|
||||
function commitWithFailingAdd({ cwd, files, failFor = [], stderr = 'fatal: injected staging failure', timeout = false, amend = false, gitVerb = 'add' }) {
|
||||
const callsOut = path.join(fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-2608-')), 'calls.json');
|
||||
// `timeout` is `false` | `true` (alias for `'posix'`) | `'posix'` | `'windows'` —
|
||||
// #3050: the shared isSpawnTimeout predicate only requires `error.code ===
|
||||
// 'ETIMEDOUT'`, NOT `signal === 'SIGTERM'` (Windows does not reliably report
|
||||
// SIGTERM), so both shapes must be proven to still read as a timeout.
|
||||
const timeoutShape = timeout === true ? 'posix' : timeout;
|
||||
const script = `
|
||||
const path = require('path');
|
||||
const LIB = ${JSON.stringify(LIB)};
|
||||
const projection = require(path.join(LIB, 'shell-command-projection.cjs'));
|
||||
const { cmdCommit } = require(path.join(LIB, 'commands.cjs'));
|
||||
const failFor = ${JSON.stringify(failFor)};
|
||||
const stderrText = ${JSON.stringify(stderr)};
|
||||
const timeoutShape = ${JSON.stringify(timeoutShape)};
|
||||
const gitVerb = ${JSON.stringify(gitVerb)};
|
||||
const real = projection.execGit;
|
||||
const calls = [];
|
||||
projection.execGit = (args, opts) => {
|
||||
calls.push(args);
|
||||
if (args[0] === gitVerb && failFor.includes(args[args.length - 1])) {
|
||||
if (timeoutShape === 'posix') {
|
||||
// The exact shape spawnSync produces on a POSIX timeout, which
|
||||
// shell-command-projection surfaces as signal + error.code.
|
||||
const e = new Error('spawnSync git ETIMEDOUT');
|
||||
e.code = 'ETIMEDOUT';
|
||||
return { exitCode: 1, stdout: '', stderr: stderrText, signal: 'SIGTERM', error: e };
|
||||
}
|
||||
if (timeoutShape === 'windows') {
|
||||
// Windows shape: spawnSync's timeout kill does not reliably report
|
||||
// signal:'SIGTERM' — only error.code:'ETIMEDOUT' is guaranteed (#3050).
|
||||
const e = new Error('spawnSync git ETIMEDOUT');
|
||||
e.code = 'ETIMEDOUT';
|
||||
return { exitCode: 1, stdout: '', stderr: stderrText, signal: null, error: e };
|
||||
}
|
||||
return { exitCode: 128, stdout: '', stderr: stderrText, signal: null, error: null };
|
||||
}
|
||||
return real(args, opts);
|
||||
};
|
||||
process.on('exit', () => {
|
||||
require('fs').writeFileSync(${JSON.stringify(callsOut)}, JSON.stringify(calls));
|
||||
});
|
||||
cmdCommit(${JSON.stringify(cwd)}, 'docs: map existing codebase', ${JSON.stringify(files)}, false, ${JSON.stringify(amend)}, false);
|
||||
`;
|
||||
|
||||
const run = spawnSync(process.execPath, ['-e', script], {
|
||||
encoding: 'utf8',
|
||||
timeout: 30000,
|
||||
killSignal: 'SIGKILL',
|
||||
env: { ...process.env, GSD_TEST_MODE: '1' },
|
||||
});
|
||||
|
||||
assert.ok(
|
||||
run.stdout && run.stdout.trim(),
|
||||
`cmdCommit child produced no stdout (status=${run.status}): ${run.stderr}`,
|
||||
);
|
||||
return {
|
||||
result: JSON.parse(run.stdout),
|
||||
gitCalls: JSON.parse(fs.readFileSync(callsOut, 'utf8')),
|
||||
};
|
||||
}
|
||||
|
||||
function headCount(cwd) {
|
||||
return Number(gitOrThrow(['rev-list', '--count', 'HEAD'], { cwd, timeoutMs: STAGING_GIT_TIMEOUT_MS }).trim());
|
||||
}
|
||||
|
||||
function committedFiles(cwd) {
|
||||
return gitOrThrow(['diff', 'HEAD~1', 'HEAD', '--name-only'], { cwd, timeoutMs: STAGING_GIT_TIMEOUT_MS })
|
||||
.trim().split('\n').filter(Boolean).sort();
|
||||
}
|
||||
|
||||
/**
|
||||
* Same harness for `cmdCommitToSubrepo` — the sub-repo twin of the staging loop,
|
||||
* which carried the identical defect (failed `git add` dropped, commit proceeds
|
||||
* with the subset that staged).
|
||||
*/
|
||||
function subrepoCommitWithFailingAdd({ cwd, files, failFor = [], timeout = false }) {
|
||||
// See commitWithFailingAdd above for the timeoutShape rationale (#3050).
|
||||
const timeoutShape = timeout === true ? 'posix' : timeout;
|
||||
const script = `
|
||||
const path = require('path');
|
||||
const LIB = ${JSON.stringify(LIB)};
|
||||
const projection = require(path.join(LIB, 'shell-command-projection.cjs'));
|
||||
const { cmdCommitToSubrepo } = require(path.join(LIB, 'commands.cjs'));
|
||||
const failFor = ${JSON.stringify(failFor)};
|
||||
const timeoutShape = ${JSON.stringify(timeoutShape)};
|
||||
const real = projection.execGit;
|
||||
projection.execGit = (args, opts) => {
|
||||
if (args[0] === 'add' && failFor.includes(args[args.length - 1])) {
|
||||
if (timeoutShape === 'posix') {
|
||||
const e = new Error('spawnSync git ETIMEDOUT');
|
||||
e.code = 'ETIMEDOUT';
|
||||
return { exitCode: 1, stdout: '', stderr: 'fatal: injected subrepo staging failure', signal: 'SIGTERM', error: e };
|
||||
}
|
||||
if (timeoutShape === 'windows') {
|
||||
const e = new Error('spawnSync git ETIMEDOUT');
|
||||
e.code = 'ETIMEDOUT';
|
||||
return { exitCode: 1, stdout: '', stderr: 'fatal: injected subrepo staging failure', signal: null, error: e };
|
||||
}
|
||||
return { exitCode: 128, stdout: '', stderr: 'fatal: injected subrepo staging failure', signal: null, error: null };
|
||||
}
|
||||
return real(args, opts);
|
||||
};
|
||||
cmdCommitToSubrepo(${JSON.stringify(cwd)}, 'feat: subrepo change', ${JSON.stringify(files)}, false);
|
||||
`;
|
||||
const run = spawnSync(process.execPath, ['-e', script], {
|
||||
encoding: 'utf8',
|
||||
timeout: 30000,
|
||||
killSignal: 'SIGKILL',
|
||||
env: { ...process.env, GSD_TEST_MODE: '1' },
|
||||
});
|
||||
assert.ok(run.stdout && run.stdout.trim(),
|
||||
`cmdCommitToSubrepo child produced no stdout (status=${run.status}): ${run.stderr}`);
|
||||
return JSON.parse(run.stdout);
|
||||
}
|
||||
|
||||
describe('#2608: commit-to-subrepo fails closed when git add fails', () => {
|
||||
let rootDir;
|
||||
let subDir;
|
||||
|
||||
beforeEach(() => {
|
||||
rootDir = createTempGitProject();
|
||||
fs.writeFileSync(
|
||||
path.join(rootDir, '.planning', 'config.json'),
|
||||
JSON.stringify({ planning: { sub_repos: ['backend'] } }, null, 2),
|
||||
);
|
||||
subDir = path.join(rootDir, 'backend');
|
||||
fs.mkdirSync(subDir, { recursive: true });
|
||||
for (const [cmd, args] of [['init', []], ['config', ['user.email', 'test@example.com']], ['config', ['user.name', 'Test']]]) {
|
||||
gitOrThrow([cmd, ...args], { cwd: subDir, timeoutMs: STAGING_GIT_TIMEOUT_MS });
|
||||
}
|
||||
fs.writeFileSync(path.join(subDir, 'seed.js'), '// seed\n');
|
||||
gitOrThrow(['add', 'seed.js'], { cwd: subDir, timeoutMs: STAGING_GIT_TIMEOUT_MS });
|
||||
gitOrThrow(['commit', '-m', 'seed'], { cwd: subDir, timeoutMs: STAGING_GIT_TIMEOUT_MS });
|
||||
fs.writeFileSync(path.join(subDir, 'a.js'), '// a\n');
|
||||
fs.writeFileSync(path.join(subDir, 'b.js'), '// b\n');
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(rootDir);
|
||||
});
|
||||
|
||||
test('a failed sub-repo git add reports staging_failed and commits nothing', () => {
|
||||
const before = headCount(subDir);
|
||||
const result = subrepoCommitWithFailingAdd({
|
||||
cwd: rootDir,
|
||||
files: ['backend/a.js', 'backend/b.js'],
|
||||
failFor: ['b.js'],
|
||||
});
|
||||
|
||||
assert.equal(result.repos.backend.reason, 'staging_failed',
|
||||
`expected staging_failed for the sub-repo, got ${JSON.stringify(result)}`);
|
||||
assert.equal(result.repos.backend.committed, false);
|
||||
assert.match(result.repos.backend.error, /injected subrepo staging failure/,
|
||||
"git's original stderr must be preserved");
|
||||
assert.equal(headCount(subDir), before, 'no partial sub-repo commit may be created');
|
||||
|
||||
const status = gitOrThrow(['status', '--porcelain'], { cwd: subDir, timeoutMs: STAGING_GIT_TIMEOUT_MS });
|
||||
assert.deepEqual(status.split('\n').filter((l) => /^A[ \t]/.test(l)), [],
|
||||
`the sub-repo index must be rolled back, status:\n${status}`);
|
||||
});
|
||||
|
||||
test('successful sub-repo staging still commits', () => {
|
||||
const before = headCount(subDir);
|
||||
const result = subrepoCommitWithFailingAdd({
|
||||
cwd: rootDir,
|
||||
files: ['backend/a.js', 'backend/b.js'],
|
||||
failFor: [],
|
||||
});
|
||||
|
||||
assert.equal(result.repos.backend.committed, true, `expected a commit, got ${JSON.stringify(result)}`);
|
||||
assert.equal(headCount(subDir), before + 1);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2608: commit --files fails closed when git add fails', () => {
|
||||
let tmpDir;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempGitProject();
|
||||
for (const name of ['ARCHITECTURE', 'CONCERNS', 'CONVENTIONS']) {
|
||||
fs.writeFileSync(path.join(tmpDir, '.planning', `${name}.md`), `# ${name}\n`);
|
||||
}
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(tmpDir);
|
||||
});
|
||||
|
||||
// ── AC1 + AC3: the failure is reported, with its original stderr ──────────
|
||||
|
||||
test('a failed git add returns staging_failed with the file and original stderr', () => {
|
||||
const before = headCount(tmpDir);
|
||||
const { result, gitCalls } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md'],
|
||||
stderr: 'fatal: Unable to create index.lock: Permission denied',
|
||||
});
|
||||
|
||||
assert.equal(result.committed, false);
|
||||
assert.equal(result.hash, null);
|
||||
assert.equal(result.reason, 'staging_failed',
|
||||
'the staging cause must be reported, not a downstream commit_failed/pathspec error');
|
||||
assert.equal(result.file, '.planning/ARCHITECTURE.md', 'the offending file must be named');
|
||||
assert.match(result.error, /Unable to create index\.lock/,
|
||||
'the original git add stderr must be preserved');
|
||||
|
||||
// AC2: git commit must never have been invoked.
|
||||
assert.ok(
|
||||
!gitCalls.some((a) => a[0] === 'commit'),
|
||||
`git commit must not run after a staging failure, calls: ${JSON.stringify(gitCalls)}`,
|
||||
);
|
||||
assert.equal(headCount(tmpDir), before, 'no commit may be created');
|
||||
});
|
||||
|
||||
// ── AC4: no partial commit of a multi-file explicit scope ─────────────────
|
||||
|
||||
test('when the second of three paths fails to stage, nothing is committed', () => {
|
||||
// Pre-fix this returned {"committed":true} — the two paths that DID stage
|
||||
// were committed under a message describing all three.
|
||||
const before = headCount(tmpDir);
|
||||
const { result, gitCalls } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md', '.planning/CONVENTIONS.md'],
|
||||
failFor: ['.planning/CONCERNS.md'],
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_failed');
|
||||
assert.equal(result.file, '.planning/CONCERNS.md');
|
||||
assert.ok(
|
||||
!gitCalls.some((a) => a[0] === 'commit'),
|
||||
'a partial commit of the paths that DID stage must not happen',
|
||||
);
|
||||
assert.equal(headCount(tmpDir), before, 'no partial commit may be created');
|
||||
});
|
||||
|
||||
test('every failing path is reported, not just the first', () => {
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md', '.planning/CONVENTIONS.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md', '.planning/CONVENTIONS.md'],
|
||||
});
|
||||
|
||||
assert.equal(result.failures.length, 2);
|
||||
assert.deepEqual(
|
||||
result.failures.map((f) => f.file).sort(),
|
||||
['.planning/ARCHITECTURE.md', '.planning/CONVENTIONS.md'],
|
||||
);
|
||||
});
|
||||
|
||||
// ── An all-paths-fail run must not masquerade as nothing_to_commit ────────
|
||||
|
||||
test('when every path fails to stage, the reason is staging_failed not nothing_to_commit', () => {
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md'],
|
||||
});
|
||||
|
||||
assert.notEqual(result.reason, 'nothing_to_commit',
|
||||
'every path failing to stage is a staging failure, not an empty changeset');
|
||||
assert.equal(result.reason, 'staging_failed');
|
||||
});
|
||||
|
||||
// ── AC5: a staging timeout is distinguishable from an ordinary failure ────
|
||||
|
||||
test('a staging timeout is reported as staging_timeout, not staging_failed', () => {
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md'],
|
||||
stderr: '',
|
||||
timeout: true,
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_timeout',
|
||||
'the projection exposes SIGTERM+ETIMEDOUT; a timeout must not read as an ordinary failure');
|
||||
assert.equal(result.failures[0].timed_out, true);
|
||||
});
|
||||
|
||||
// #3050 item 4: this site (commands.cts's `git add` staging loop) now routes
|
||||
// through the shared isSpawnTimeout predicate, which drops the `signal ===
|
||||
// 'SIGTERM'` requirement — a Windows-shaped timeout (no signal, only
|
||||
// error.code === 'ETIMEDOUT') must still be detected.
|
||||
test('a staging timeout is reported as staging_timeout even without SIGTERM (Windows shape, #3050)', () => {
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md'],
|
||||
stderr: '',
|
||||
timeout: 'windows',
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_timeout');
|
||||
assert.equal(result.failures[0].timed_out, true);
|
||||
});
|
||||
|
||||
// #3050 item 4: the `git rm --cached` branch of the same staging loop (the
|
||||
// default-mode "stage the deletion" path, distinct from `git add` above)
|
||||
// carries its own inline copy of the timeout check pre-fix. Drive it
|
||||
// directly: default mode (no explicit --files) stages '.planning/', and
|
||||
// when that path is absent on disk the loop takes the `git rm --cached`
|
||||
// branch instead of `git add`.
|
||||
test('a `git rm --cached` timeout in default mode is reported as staging_timeout, POSIX and Windows shapes (#3050)', () => {
|
||||
// Mid-test fixture mutation (simulating an absent '.planning/' on disk),
|
||||
// not teardown; the outer afterEach still runs helpers.cleanup(tmpDir) on
|
||||
// the whole tmpDir.
|
||||
// eslint-disable-next-line local/no-raw-rmsync-in-tests -- see comment above
|
||||
fs.rmSync(path.join(tmpDir, '.planning'), { recursive: true, force: true });
|
||||
|
||||
for (const shape of ['posix', 'windows']) {
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: undefined,
|
||||
failFor: ['.planning/'],
|
||||
gitVerb: 'rm',
|
||||
stderr: '',
|
||||
timeout: shape,
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_timeout', `shape=${shape}`);
|
||||
assert.equal(result.failures[0].timed_out, true, `shape=${shape}`);
|
||||
}
|
||||
});
|
||||
|
||||
test('an ordinary non-zero git add is NOT reported as a timeout', () => {
|
||||
// Boundary: the timeout carve-out must not swallow the ordinary case.
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md'],
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_failed');
|
||||
assert.equal(result.failures[0].timed_out, false);
|
||||
});
|
||||
|
||||
// ── Successful staging preserves the current scoped-commit behaviour ──────
|
||||
|
||||
test('successful staging still commits exactly the declared scope', () => {
|
||||
const before = headCount(tmpDir);
|
||||
fs.writeFileSync(path.join(tmpDir, 'unrelated-wip.txt'), 'wip\n');
|
||||
gitOrThrow(['add', 'unrelated-wip.txt'], { cwd: tmpDir, timeoutMs: STAGING_GIT_TIMEOUT_MS });
|
||||
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md'],
|
||||
failFor: [],
|
||||
});
|
||||
|
||||
assert.equal(result.committed, true, `expected a commit, got ${JSON.stringify(result)}`);
|
||||
assert.equal(result.reason, 'committed');
|
||||
assert.equal(headCount(tmpDir), before + 1);
|
||||
assert.deepEqual(
|
||||
committedFiles(tmpDir),
|
||||
['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md'],
|
||||
'the declared scope must still be honoured, and the unrelated staged file left alone',
|
||||
);
|
||||
});
|
||||
|
||||
// ── Missing explicit files keep their existing documented handling ────────
|
||||
|
||||
test('a missing explicit file is still skipped, not reported as a staging failure', () => {
|
||||
// #2014/#2523 behaviour: an explicitly-named file that does not exist is
|
||||
// skipped rather than staged as a deletion. It never reaches `git add`, so
|
||||
// it is not a staging failure and must not become one.
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/DOES-NOT-EXIST.md'],
|
||||
failFor: [],
|
||||
});
|
||||
|
||||
assert.equal(result.committed, true, `expected a commit, got ${JSON.stringify(result)}`);
|
||||
assert.deepEqual(committedFiles(tmpDir), ['.planning/ARCHITECTURE.md']);
|
||||
});
|
||||
|
||||
// ── The index must be left clean, not partially staged ───────────────────
|
||||
|
||||
test('a staging failure rolls back the paths this call had already staged', () => {
|
||||
// Without the rollback the paths that DID stage stay in the index with no
|
||||
// commit made, so the next bare `git commit` sweeps them up — the same
|
||||
// silent partial commit this fix exists to prevent, just deferred a step.
|
||||
commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md', '.planning/CONVENTIONS.md'],
|
||||
failFor: ['.planning/CONCERNS.md'],
|
||||
});
|
||||
|
||||
const status = gitOrThrow(['status', '--porcelain'], { cwd: tmpDir, timeoutMs: STAGING_GIT_TIMEOUT_MS });
|
||||
const stagedAdds = status.split('\n').filter((l) => /^A[ \t]/.test(l));
|
||||
assert.deepEqual(stagedAdds, [],
|
||||
`no path may remain staged after a staging failure, status:\n${status}`);
|
||||
});
|
||||
|
||||
test('the rollback does not unstage work the caller had staged before the call', () => {
|
||||
// Boundary: the reset must touch only what THIS call staged. Unstaging a
|
||||
// path the caller staged themselves would destroy their work.
|
||||
fs.writeFileSync(path.join(tmpDir, 'caller-staged.txt'), 'mine\n');
|
||||
gitOrThrow(['add', 'caller-staged.txt'], { cwd: tmpDir, timeoutMs: STAGING_GIT_TIMEOUT_MS });
|
||||
|
||||
commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md'],
|
||||
failFor: ['.planning/CONCERNS.md'],
|
||||
});
|
||||
|
||||
const status = gitOrThrow(['status', '--porcelain'], { cwd: tmpDir, timeoutMs: STAGING_GIT_TIMEOUT_MS });
|
||||
assert.match(status, /^A[ \t]+caller-staged\.txt$/m,
|
||||
`the caller's own staged file must survive the rollback, status:\n${status}`);
|
||||
});
|
||||
|
||||
// ── The default (non---files) staging path is guarded too ─────────────────
|
||||
|
||||
test('a failed default-mode git add fails closed instead of committing the index', () => {
|
||||
// Default mode stages `.planning/`. Pre-fix a failure there also fell
|
||||
// through to an unguarded `git commit`.
|
||||
const before = headCount(tmpDir);
|
||||
const { result, gitCalls } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: undefined,
|
||||
failFor: ['.planning/'],
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_failed');
|
||||
assert.ok(!gitCalls.some((a) => a[0] === 'commit'), 'git commit must not run');
|
||||
assert.equal(headCount(tmpDir), before);
|
||||
});
|
||||
|
||||
test('a failed default-mode git add blocks --amend too', () => {
|
||||
// --amend has no carve-out: amending on top of a failed staging would
|
||||
// rewrite the tip without the changes the caller asked for.
|
||||
const before = gitOrThrow(['rev-parse', 'HEAD'], { cwd: tmpDir, timeoutMs: STAGING_GIT_TIMEOUT_MS }).trim();
|
||||
const { result, gitCalls } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: undefined,
|
||||
failFor: ['.planning/'],
|
||||
amend: true,
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_failed');
|
||||
assert.ok(!gitCalls.some((a) => a[0] === 'commit'), 'git commit --amend must not run');
|
||||
assert.equal(
|
||||
gitOrThrow(['rev-parse', 'HEAD'], { cwd: tmpDir, timeoutMs: STAGING_GIT_TIMEOUT_MS }).trim(),
|
||||
before,
|
||||
'HEAD must not be rewritten when staging failed',
|
||||
);
|
||||
});
|
||||
|
||||
test('when all explicit files are missing the reason is still nothing_to_commit', () => {
|
||||
// The nothing_to_commit path must survive: no `git add` ran, so there is no
|
||||
// staging failure to report.
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/GONE-A.md', '.planning/GONE-B.md'],
|
||||
failFor: [],
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'nothing_to_commit');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -244,3 +244,172 @@ describe('debug skill dispatch and sub-orchestrator (#2148, #2151)', () => {
|
||||
'debug.md session_params must not hardcode .planning/debug/... as debug_file_path (#2376)');
|
||||
});
|
||||
});
|
||||
|
||||
// Tests for #2196 (folded from tests/fix-2196-debug-agent-handoff.test.cjs):
|
||||
// the /gsd-debug orchestrator misused the foreground session-manager Agent()
|
||||
// spawn as a background task, then queried the returned agent ID via
|
||||
// TaskOutput (which expects a task ID) — yielding "No task found with ID"
|
||||
// and leaving the workflow waiting on a handoff that was never queryable,
|
||||
// with no recovery. The fix makes debug.md state explicitly that the spawn
|
||||
// is foreground/blocking, that an agent ID must never be passed to
|
||||
// TaskOutput, and that a lost handoff must be recovered (preserve
|
||||
// checkpoint + resume).
|
||||
describe('#2196 debug.md session-manager spawn contract', () => {
|
||||
// allow-test-rule: source-text-is-the-product (#2196)
|
||||
const DEBUG_MD_2196 = path.join(__dirname, '..', 'gsd-core', 'workflows', 'debug.md');
|
||||
const content = fs.readFileSync(DEBUG_MD_2196, 'utf-8');
|
||||
const sectionStart = content.indexOf('Session Management');
|
||||
const section = sectionStart !== -1 ? content.slice(sectionStart) : '';
|
||||
|
||||
test('debug.md has the Session Management section', () => {
|
||||
assert.notEqual(sectionStart, -1, 'debug.md must contain the Session Management section');
|
||||
});
|
||||
|
||||
test('the session-manager spawn is declared foreground/blocking (not backgrounded)', () => {
|
||||
assert.ok(/foreground/i.test(section) && /blocking/i.test(section),
|
||||
'the Agent() spawn must be declared foreground and blocking so it is not polled');
|
||||
});
|
||||
|
||||
test('an agent ID must not be passed to TaskOutput', () => {
|
||||
assert.ok(/TaskOutput/.test(section),
|
||||
'the contract must mention TaskOutput by name');
|
||||
assert.ok(/agent ID is NOT a task ID|agent ID is not a task ID/i.test(section),
|
||||
'the contract must state an agent ID is not a task ID');
|
||||
});
|
||||
|
||||
test('a lost handoff has a recovery path (preserve checkpoint + resume)', () => {
|
||||
// Pin the CANONICAL colon form — the retired /gsd-debug hyphen syntax is
|
||||
// rejected by the slash-command-namespace guard, so this must be /gsd:debug.
|
||||
assert.ok(/\/gsd:debug continue \{slug\}/.test(section),
|
||||
'the contract must point to /gsd:debug continue {slug} (canonical colon form) as the resume path');
|
||||
assert.ok(/do not claim|do NOT claim/i.test(section),
|
||||
'the contract must forbid claiming a lost-handoff session is still running');
|
||||
});
|
||||
});
|
||||
|
||||
// Tests for #2257 (folded from tests/fix-2257-debug-nonterminal-resume.test.cjs):
|
||||
// the /gsd-debug orchestrator had no contract for a foreground
|
||||
// gsd-debug-session-manager return that is usable but non-terminal. Section 4
|
||||
// "Session Management" (and the `continue` subcommand's return handling in
|
||||
// Section 1c) recognized only two literal-string returns — `DEBUG SESSION
|
||||
// COMPLETE` and `ABANDONED` — with no else branch. Any other return (e.g. a
|
||||
// mid-investigation progress summary emitted when the manager's own
|
||||
// turn/context budget runs out) matched neither and fell through to the user
|
||||
// as if the debug session were complete, silently abandoning the
|
||||
// investigation mid-flight.
|
||||
//
|
||||
// The fix defines an explicit non-terminal marker, `CONTINUE_REQUIRED`, that
|
||||
// the session manager emits when it must stop before reaching a terminal
|
||||
// state (distinct from the two terminal returns and from a genuine
|
||||
// user-input/approval `CHECKPOINT REACHED`, which already correctly pauses
|
||||
// via AskUserQuestion). The orchestrator treats anything that is not one of
|
||||
// the two terminal markers as non-terminal and auto-resumes by re-spawning
|
||||
// the session manager from the on-disk checkpoint, bounded by an anti-loop
|
||||
// guard.
|
||||
//
|
||||
// Correction (orthogonal review): the first cut of the anti-loop guard
|
||||
// required BOTH `next_action` AND `updated` to be unchanged across two
|
||||
// resumes to detect no-progress — but `agents/gsd-debugger.md` overwrites
|
||||
// `updated` on every checkpoint write ("Update the file BEFORE taking
|
||||
// action"), so `updated` changes every cycle and the AND-condition could
|
||||
// never be true, making the guard dead (unbounded auto-resume / DoS
|
||||
// regression). The corrected guard keys no-progress detection off
|
||||
// `next_action` ALONE and adds an absolute, content-independent hard cap of
|
||||
// 3 total auto-resumes per slug per `/gsd:debug` invocation as the real
|
||||
// termination bound.
|
||||
describe('#2257 debug non-terminal session-manager return contract', () => {
|
||||
const DEBUG_MD_2257 = path.join(__dirname, '..', 'gsd-core', 'workflows', 'debug.md');
|
||||
const SESSION_MANAGER_MD = path.join(__dirname, '..', 'agents', 'gsd-debug-session-manager.md');
|
||||
|
||||
// allow-test-rule: source-text-is-the-product (#2257)
|
||||
// workflow/agent prose IS the runtime contract under test
|
||||
const debugContent = fs.readFileSync(DEBUG_MD_2257, 'utf-8');
|
||||
// allow-test-rule: source-text-is-the-product (#2257)
|
||||
// workflow/agent prose IS the runtime contract under test
|
||||
const managerContent = fs.readFileSync(SESSION_MANAGER_MD, 'utf-8');
|
||||
|
||||
const section4Start = debugContent.indexOf('## 4. Session Management');
|
||||
const section4 = section4Start !== -1 ? debugContent.slice(section4Start) : '';
|
||||
|
||||
const section1cStart = debugContent.indexOf('## 1c. CONTINUE subcommand');
|
||||
const section1dStart = debugContent.indexOf('## 1d. Check Active Sessions');
|
||||
const section1c =
|
||||
section1cStart !== -1 && section1dStart !== -1
|
||||
? debugContent.slice(section1cStart, section1dStart)
|
||||
: '';
|
||||
|
||||
test('debug.md has Section 4 (Session Management) and Section 1c (CONTINUE subcommand)', () => {
|
||||
assert.notEqual(section4Start, -1, 'debug.md must contain Section 4 Session Management');
|
||||
assert.notEqual(section1cStart, -1, 'debug.md must contain Section 1c CONTINUE subcommand');
|
||||
});
|
||||
|
||||
test('Section 4 has an exhaustive non-terminal branch that auto-resumes from the checkpoint', () => {
|
||||
assert.ok(/CONTINUE_REQUIRED/.test(section4),
|
||||
'Section 4 must reference the CONTINUE_REQUIRED non-terminal marker');
|
||||
assert.ok(/ANYTHING ELSE/i.test(section4),
|
||||
'Section 4 must exhaustively catch any return that is not one of the two terminal markers');
|
||||
assert.ok(/AUTO-RESUME/i.test(section4) && /re-spawning/i.test(section4),
|
||||
'Section 4 must auto-resume by re-spawning the session manager, not return control to the user');
|
||||
assert.ok(/same.{0,20}slug/i.test(section4),
|
||||
'Section 4 auto-resume must use the SAME slug/checkpoint as the original spawn');
|
||||
});
|
||||
|
||||
test('Section 1c has the same exhaustive non-terminal auto-resume branch (not just the two literals)', () => {
|
||||
assert.ok(/CONTINUE_REQUIRED/.test(section1c),
|
||||
'Section 1c must reference the CONTINUE_REQUIRED non-terminal marker');
|
||||
assert.ok(/ANYTHING ELSE/i.test(section1c),
|
||||
'Section 1c must exhaustively catch any return that is not one of the two terminal markers');
|
||||
assert.ok(/AUTO-RESUME/i.test(section1c) && /re-spawning/i.test(section1c),
|
||||
'Section 1c must auto-resume by re-spawning the session manager, not return control to the user');
|
||||
assert.ok(/same.{0,20}slug/i.test(section1c),
|
||||
'Section 1c auto-resume must use the SAME slug/checkpoint as the original spawn (symmetric with Section 4)');
|
||||
});
|
||||
|
||||
test('gsd-debug-session-manager.md defines CONTINUE_REQUIRED distinct from the two terminal formats', () => {
|
||||
assert.ok(/## CONTINUE_REQUIRED/.test(managerContent),
|
||||
'the agent must define an explicit ## CONTINUE_REQUIRED return heading');
|
||||
assert.ok(/## DEBUG SESSION COMPLETE/.test(managerContent),
|
||||
'the terminal DEBUG SESSION COMPLETE format must still be present');
|
||||
assert.ok(/ABANDONED/.test(managerContent),
|
||||
'the terminal ABANDONED format must still be present');
|
||||
assert.ok(/non-terminal/i.test(managerContent),
|
||||
'the agent must characterize CONTINUE_REQUIRED as non-terminal');
|
||||
assert.ok(/CHECKPOINT REACHED/.test(managerContent) && /distinct from/i.test(managerContent),
|
||||
'CONTINUE_REQUIRED must be explicitly distinguished from the genuine user-input CHECKPOINT REACHED shape');
|
||||
assert.ok(/\.planning\/debug\/\{slug\}\.md/.test(managerContent),
|
||||
'CONTINUE_REQUIRED must reference the on-disk checkpoint path');
|
||||
assert.ok(/next_action/.test(managerContent) && /status/.test(managerContent),
|
||||
'CONTINUE_REQUIRED must reference the checkpoint status/next_action fields');
|
||||
});
|
||||
|
||||
test('an anti-loop bound exists so repeated no-progress auto-resumes do not loop indefinitely', () => {
|
||||
assert.ok(/anti-loop guard/i.test(section4),
|
||||
'Section 4 must name an anti-loop guard');
|
||||
assert.ok(/blocker report/i.test(section4),
|
||||
'Section 4 must emit a blocker report to the user once the bound is exceeded, instead of looping forever');
|
||||
|
||||
assert.ok(/anti-loop guard/i.test(section1c),
|
||||
'Section 1c must name an anti-loop guard');
|
||||
});
|
||||
|
||||
test('the anti-loop guard has an absolute hard cap independent of no-progress detection (#2257 correction)', () => {
|
||||
for (const [label, section] of [['Section 4', section4], ['Section 1c', section1c]]) {
|
||||
assert.ok(/hard cap/i.test(section), `${label} must name an absolute hard cap`);
|
||||
assert.ok(/\b3\b/.test(section) && /total auto-resumes/i.test(section),
|
||||
`${label} must encode a concrete numeric cap of 3 total auto-resumes`);
|
||||
assert.ok(/regardless/i.test(section),
|
||||
`${label} hard cap must trip regardless of whether next_action changed (content-independent)`);
|
||||
}
|
||||
});
|
||||
|
||||
test('no-progress detection keys off next_action alone, never the always-changing updated timestamp (#2257 correction)', () => {
|
||||
for (const [label, section] of [['Section 4', section4], ['Section 1c', section1c]]) {
|
||||
assert.ok(/next_action/.test(section),
|
||||
`${label} no-progress heuristic must reference next_action`);
|
||||
assert.ok(/(do not|never).{0,40}updated/i.test(section),
|
||||
`${label} must explicitly forbid keying no-progress detection off updated`);
|
||||
assert.ok(/changes every cycle/i.test(section),
|
||||
`${label} must state WHY updated cannot be used: it is overwritten/changes every checkpoint cycle`);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
@@ -412,3 +412,265 @@ describe('bug #2002: offer_next checks CONTEXT.md before suggesting next step',
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
|
||||
// ────────────────────────────────────────────────────────────────────────
|
||||
// Folded from tests/fix-3177-execute-phase-dispatch-claim.test.cjs — test hygiene #3334 (H3)
|
||||
// ────────────────────────────────────────────────────────────────────────
|
||||
{
|
||||
const { describe: __foldDescribe } = require('node:test');
|
||||
__foldDescribe("folded:fix-3177-execute-phase-dispatch-claim (test hygiene #3334 H3)", () => {
|
||||
// allow-test-rule: runtime-contract-is-the-product #3177 — the workflow markdown is loaded
|
||||
// verbatim into the agent's context and the matrix/descriptor ARE the negotiated host
|
||||
// contract; asserting agreement between those documents is behavioral, not source-grep.
|
||||
|
||||
'use strict';
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const fc = require('fast-check');
|
||||
|
||||
const ROOT = path.join(__dirname, '..');
|
||||
const WORKFLOW = path.join(ROOT, 'gsd-core', 'workflows', 'execute-phase.md');
|
||||
const MATRIX = path.join(ROOT, 'docs', 'reference', 'host-integration-capability-matrix.md');
|
||||
const DESCRIPTOR = path.join(ROOT, 'capabilities', 'claude', 'capability.json');
|
||||
|
||||
const workflowText = () => fs.readFileSync(WORKFLOW, 'utf8');
|
||||
|
||||
/**
|
||||
* Value of `| <field> | <value> | …` inside the `## <host>` section of the matrix.
|
||||
*
|
||||
* Anchored on a whole heading LINE, not a substring: `## claude` must never match
|
||||
* `## claude-local`, and the section must end at the next `## ` heading so a field
|
||||
* absent from this host can never be answered from the next host's table.
|
||||
*
|
||||
* @param {string} matrix - full matrix document text
|
||||
* @param {string} host - section name, e.g. `claude`
|
||||
* @param {string} field - row label, e.g. `dispatch.background`
|
||||
* @returns {string|null} trimmed cell value, or null when the section or row is absent
|
||||
*/
|
||||
function matrixField(matrix, host, field) {
|
||||
const lines = matrix.split('\n');
|
||||
const start = lines.findIndex((l) => l.trim() === `## ${host}`);
|
||||
if (start === -1) return null;
|
||||
let end = lines.length;
|
||||
for (let i = start + 1; i < lines.length; i += 1) {
|
||||
if (lines[i].startsWith('## ')) { end = i; break; }
|
||||
}
|
||||
const row = lines.slice(start + 1, end).find((l) => l.startsWith(`| ${field} |`));
|
||||
if (!row) return null;
|
||||
return row.split('|')[2].trim();
|
||||
}
|
||||
|
||||
describe('#3177: execute-phase.md states Claude Code dispatch truthfully', () => {
|
||||
test('execute-phase.md never claims Claude Code Agent() blocks or returns synchronously', () => {
|
||||
// Row 1 — the failing-first regression. Both stale sentences, by their own text.
|
||||
const text = workflowText();
|
||||
const stale = ['blocks until complete', 'returns synchronously'];
|
||||
const present = stale.filter((phrase) => text.includes(phrase));
|
||||
assert.deepEqual(
|
||||
present, [],
|
||||
'execute-phase.md still asserts synchronous Claude Code dispatch. Claude Code backgrounds '
|
||||
+ 'subagents by default (v2.1.198+); `run_in_background: false` is the opt-out.',
|
||||
);
|
||||
});
|
||||
|
||||
test('the Claude Code dispatch bullet states the background-by-default model', () => {
|
||||
// Row 2 — the corrected sentence must actually SAY the true thing, not merely
|
||||
// omit the false one. A deletion would pass row 1 while teaching nothing.
|
||||
const bullet = workflowText()
|
||||
.split('\n')
|
||||
.find((l) => l.startsWith('- **Claude Code:**'));
|
||||
assert.ok(bullet, 'the <runtime_compatibility> Claude Code bullet must exist');
|
||||
assert.match(
|
||||
bullet, /backgrounded by default/,
|
||||
'the bullet must state that dispatch is backgrounded by default',
|
||||
);
|
||||
assert.match(
|
||||
bullet, /verify completion/,
|
||||
'the bullet must point at completion verification. This workflow deliberately backgrounds '
|
||||
+ 'its executors (the multi-plan path prescribes run_in_background: true), so the blocking '
|
||||
+ 'opt-out is not the guidance here — confirming completion is.',
|
||||
);
|
||||
assert.ok(
|
||||
bullet.includes('Agent(subagent_type="gsd-executor"'),
|
||||
'the bullet must still carry the dispatch mechanism the rest of the file depends on',
|
||||
);
|
||||
});
|
||||
|
||||
test('the workflow prose and the capability matrix agree on claude dispatch.background', () => {
|
||||
// Row 3 — the parity guard. This is the assertion that outlives the wording:
|
||||
// flip the matrix to `false` without touching the prose (or vice versa) and this reds.
|
||||
const declared = matrixField(fs.readFileSync(MATRIX, 'utf8'), 'claude', 'dispatch.background');
|
||||
assert.equal(declared, 'true', 'matrix must document claude dispatch.background');
|
||||
|
||||
const bullet = workflowText()
|
||||
.split('\n')
|
||||
.find((l) => l.startsWith('- **Claude Code:**'));
|
||||
const proseSaysBackground = /backgrounded by default/.test(bullet ?? '');
|
||||
assert.equal(
|
||||
proseSaysBackground, declared === 'true',
|
||||
'execute-phase.md and the host-integration matrix disagree about whether claude backgrounds '
|
||||
+ 'its subagent dispatch. They describe the same runtime; exactly one of them is wrong (#3177).',
|
||||
);
|
||||
});
|
||||
|
||||
test('the claude descriptor and the matrix agree on dispatch.background', () => {
|
||||
// Row 4 — the other half of the divergence class. The matrix is generated from
|
||||
// the descriptor, so this pins the generator's output to its input.
|
||||
const descriptor = JSON.parse(fs.readFileSync(DESCRIPTOR, 'utf8'));
|
||||
const declared = matrixField(fs.readFileSync(MATRIX, 'utf8'), 'claude', 'dispatch.background');
|
||||
assert.equal(
|
||||
String(descriptor.runtime.hostIntegration.dispatch.background), declared,
|
||||
'capabilities/claude/capability.json and the rendered matrix disagree',
|
||||
);
|
||||
});
|
||||
|
||||
test('the Codex orchestrator rule is not swept by the Claude Code correction', () => {
|
||||
// Row 5 — negative space. Codex dispatch IS synchronous. A regex sweep for
|
||||
// "return its result" would introduce a NEW falsehood here; this catches that.
|
||||
const text = workflowText();
|
||||
const codexRules = text
|
||||
.split('\n')
|
||||
.filter((l) => l.includes('ORCHESTRATOR RULE — CODEX RUNTIME'));
|
||||
assert.ok(codexRules.length >= 2, 'both Codex orchestrator rules must survive');
|
||||
for (const rule of codexRules) {
|
||||
assert.ok(
|
||||
rule.includes('Wait for the subagent to return its result'),
|
||||
'Codex dispatch is genuinely synchronous — its wait rule must not be corrected away',
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test('the Copilot and multi-plan dispatch rules survive the correction', () => {
|
||||
// Row 6 — negative space. These were already true and are adjacent to the edit.
|
||||
const text = workflowText();
|
||||
assert.ok(
|
||||
text.includes('- **Copilot:** Subagent spawning does not reliably return completion signals.'),
|
||||
'the Copilot bullet must survive verbatim',
|
||||
);
|
||||
assert.ok(
|
||||
text.includes('one at a time with `run_in_background: true`'),
|
||||
'the multi-plan wave prescription must survive verbatim',
|
||||
);
|
||||
assert.ok(
|
||||
text.includes('If `Agent` IS available (top-level Claude'),
|
||||
'the spawn mandate is derived from TOOL AVAILABILITY, not from blocking — it must survive',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#3177: matrix section extraction is bounded by its heading', () => {
|
||||
const matrix = () => fs.readFileSync(MATRIX, 'utf8');
|
||||
|
||||
test('section extraction — first row of the section', () => {
|
||||
// limit-1: the row immediately after the `## claude` heading is INSIDE.
|
||||
assert.equal(matrixField(matrix(), 'claude', 'embeddingMode'), 'imperative');
|
||||
});
|
||||
|
||||
test('section extraction — last row before the next heading', () => {
|
||||
// limit: the final row of `## claude` is still INSIDE.
|
||||
assert.ok(matrixField(matrix(), 'claude', 'dispatch.isolation').startsWith('harness-worktree'));
|
||||
});
|
||||
|
||||
test('section extraction — a row in the next section never leaks in', () => {
|
||||
// limit+1: a field absent from claude must be null rather than silently
|
||||
// resolved from `## codex` below it, and two hosts with different values
|
||||
// must never resolve to the same cell.
|
||||
const doc = matrix();
|
||||
assert.equal(matrixField(doc, 'claude', 'dispatch.background'), 'true');
|
||||
assert.equal(matrixField(doc, 'claude', '__definitely_not_a_field__'), null);
|
||||
assert.notEqual(
|
||||
matrixField(doc, 'claude', 'effortSurface'),
|
||||
matrixField(doc, 'kilo', 'effortSurface'),
|
||||
'two hosts with different values must not resolve to the same cell',
|
||||
);
|
||||
});
|
||||
|
||||
test('fc property: field extraction never leaks across ## boundaries', () => {
|
||||
const hostArb = fc.stringMatching(/^[a-z][a-z0-9-]{0,12}$/);
|
||||
const valueArb = fc.stringMatching(/^[a-z0-9]{1,10}$/);
|
||||
fc.assert(
|
||||
fc.property(
|
||||
fc.uniqueArray(fc.tuple(hostArb, valueArb), {
|
||||
minLength: 2, maxLength: 6, selector: ([h]) => h,
|
||||
}),
|
||||
fc.nat(),
|
||||
(sections, pick) => {
|
||||
const doc = sections
|
||||
.map(([host, value]) => `## ${host}\n\n| axis | value |\n| f | ${value} |\n`)
|
||||
.join('\n');
|
||||
const [host, value] = sections[pick % sections.length];
|
||||
// Exactly the requested section's value, never a neighbor's.
|
||||
assert.equal(matrixField(doc, host, 'f'), value);
|
||||
// A longer name that merely EXTENDS a real heading resolves to nothing.
|
||||
// `_` is outside hostArb's alphabet, so this probe can NEVER collide with
|
||||
// another generated section — the `-local` form could, and did.
|
||||
assert.equal(matrixField(doc, `${host}_x`, 'f'), null);
|
||||
},
|
||||
),
|
||||
{ numRuns: 200 },
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#3177: debug.md dispatches its session manager in the foreground', () => {
|
||||
const DEBUG_WF = path.join(ROOT, 'gsd-core', 'workflows', 'debug.md');
|
||||
const debugText = () => fs.readFileSync(DEBUG_WF, 'utf8');
|
||||
|
||||
/**
|
||||
* Every fenced `Agent( … )` block in debug.md that dispatches the session manager.
|
||||
*
|
||||
* Anchored on a line that is exactly `Agent(` so the PROSE mention of
|
||||
* `Agent(subagent_type="gsd-debug-session-manager", …)` inside the blockquote at
|
||||
* :206 is not mistaken for a dispatch. Positional selection (first match wins) was
|
||||
* the original bug here: it silently checked the continue path while the
|
||||
* new-session path went unexamined.
|
||||
*/
|
||||
function sessionManagerDispatches(text) {
|
||||
const blocks = [];
|
||||
for (const m of text.matchAll(/^Agent\($/gm)) {
|
||||
const close = text.indexOf('\n)', m.index);
|
||||
if (close === -1) continue;
|
||||
const block = text.slice(m.index, close);
|
||||
if (block.includes('subagent_type="gsd-debug-session-manager"')) blocks.push(block);
|
||||
}
|
||||
return blocks;
|
||||
}
|
||||
|
||||
test('every session-manager spawn carries the run_in_background: false opt-out', () => {
|
||||
// #2196 required this dispatch be foreground and blocking so the orchestrator
|
||||
// receives the session summary inline; debug.md still says "Wait for it; do not
|
||||
// background it" and "Display the compact summary returned by the session
|
||||
// manager". Claude Code backgrounds subagents by DEFAULT, so that intent only
|
||||
// holds if each call states the opt-out explicitly — prose alone silently
|
||||
// reinstated the exact lost-handoff failure #2196 was filed to fix.
|
||||
const blocks = sessionManagerDispatches(debugText());
|
||||
assert.equal(
|
||||
blocks.length, 2,
|
||||
'debug.md dispatches the session manager on BOTH the new-session and continue paths; '
|
||||
+ 'a change to that count means a dispatch was added or removed and must be re-checked.',
|
||||
);
|
||||
for (const block of blocks) {
|
||||
assert.match(
|
||||
block, /run_in_background\s*=\s*false/,
|
||||
'every gsd-debug-session-manager dispatch must pass run_in_background=false — without '
|
||||
+ 'it Claude Code backgrounds the spawn and the compact summary never returns (#2196).',
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test('debug.md does not assert the spawn is inherently foreground', () => {
|
||||
// The old premise ("is FOREGROUND and BLOCKING") was a property claim about the
|
||||
// host, not an instruction — and it was false for the same reason as #3177.
|
||||
assert.ok(
|
||||
!debugText().includes('is FOREGROUND and BLOCKING'),
|
||||
'debug.md must not claim the Agent() call is inherently foreground; it must name the '
|
||||
+ 'run_in_background: false opt-out that actually makes it so.',
|
||||
);
|
||||
});
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
@@ -1,52 +0,0 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* #2194… no — #2196: the /gsd-debug orchestrator misused the foreground
|
||||
* session-manager Agent() spawn as a background task, then queried the returned
|
||||
* agent ID via TaskOutput (which expects a task ID) — yielding "No task found
|
||||
* with ID" and leaving the workflow waiting on a handoff that was never
|
||||
* queryable, with no recovery.
|
||||
*
|
||||
* The fix makes debug.md state explicitly that the spawn is foreground/blocking,
|
||||
* that an agent ID must never be passed to TaskOutput, and that a lost handoff
|
||||
* must be recovered (preserve checkpoint + resume). debug.md IS the product the
|
||||
* runtime loads, so this asserts the deployed text carries that contract.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const DEBUG_MD = path.join(__dirname, '..', 'gsd-core', 'workflows', 'debug.md');
|
||||
|
||||
describe('#2196 debug.md session-manager spawn contract', () => {
|
||||
const content = fs.readFileSync(DEBUG_MD, 'utf-8');
|
||||
const sectionStart = content.indexOf('Session Management');
|
||||
const section = sectionStart !== -1 ? content.slice(sectionStart) : '';
|
||||
|
||||
test('debug.md has the Session Management section', () => {
|
||||
assert.notEqual(sectionStart, -1, 'debug.md must contain the Session Management section');
|
||||
});
|
||||
|
||||
test('the session-manager spawn is declared foreground/blocking (not backgrounded)', () => {
|
||||
assert.ok(/foreground/i.test(section) && /blocking/i.test(section),
|
||||
'the Agent() spawn must be declared foreground and blocking so it is not polled');
|
||||
});
|
||||
|
||||
test('an agent ID must not be passed to TaskOutput', () => {
|
||||
assert.ok(/TaskOutput/.test(section),
|
||||
'the contract must mention TaskOutput by name');
|
||||
assert.ok(/agent ID is NOT a task ID|agent ID is not a task ID/i.test(section),
|
||||
'the contract must state an agent ID is not a task ID');
|
||||
});
|
||||
|
||||
test('a lost handoff has a recovery path (preserve checkpoint + resume)', () => {
|
||||
// Pin the CANONICAL colon form — the retired /gsd-debug hyphen syntax is
|
||||
// rejected by the slash-command-namespace guard, so this must be /gsd:debug.
|
||||
assert.ok(/\/gsd:debug continue \{slug\}/.test(section),
|
||||
'the contract must point to /gsd:debug continue {slug} (canonical colon form) as the resume path');
|
||||
assert.ok(/do not claim|do NOT claim/i.test(section),
|
||||
'the contract must forbid claiming a lost-handoff session is still running');
|
||||
});
|
||||
});
|
||||
@@ -1,140 +0,0 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* #2257: the /gsd-debug orchestrator had no contract for a foreground
|
||||
* gsd-debug-session-manager return that is usable but non-terminal. Section 4
|
||||
* "Session Management" (and the `continue` subcommand's return handling in
|
||||
* Section 1c) recognized only two literal-string returns — `DEBUG SESSION
|
||||
* COMPLETE` and `ABANDONED` — with no else branch. Any other return (e.g. a
|
||||
* mid-investigation progress summary emitted when the manager's own
|
||||
* turn/context budget runs out) matched neither and fell through to the user
|
||||
* as if the debug session were complete, silently abandoning the
|
||||
* investigation mid-flight.
|
||||
*
|
||||
* The fix defines an explicit non-terminal marker, `CONTINUE_REQUIRED`, that
|
||||
* the session manager emits when it must stop before reaching a terminal
|
||||
* state (distinct from the two terminal returns and from a genuine
|
||||
* user-input/approval `CHECKPOINT REACHED`, which already correctly pauses
|
||||
* via AskUserQuestion). The orchestrator treats anything that is not one of
|
||||
* the two terminal markers as non-terminal and auto-resumes by re-spawning
|
||||
* the session manager from the on-disk checkpoint, bounded by an anti-loop
|
||||
* guard.
|
||||
*
|
||||
* Correction (orthogonal review): the first cut of the anti-loop guard
|
||||
* required BOTH `next_action` AND `updated` to be unchanged across two
|
||||
* resumes to detect no-progress — but `agents/gsd-debugger.md` overwrites
|
||||
* `updated` on every checkpoint write ("Update the file BEFORE taking
|
||||
* action"), so `updated` changes every cycle and the AND-condition could
|
||||
* never be true, making the guard dead (unbounded auto-resume / DoS
|
||||
* regression). The corrected guard keys no-progress detection off
|
||||
* `next_action` ALONE and adds an absolute, content-independent hard cap of
|
||||
* 3 total auto-resumes per slug per `/gsd:debug` invocation as the real
|
||||
* termination bound.
|
||||
*
|
||||
* debug.md and gsd-debug-session-manager.md ARE the product the runtime
|
||||
* loads, so this asserts the deployed text carries the contract — the
|
||||
* sanctioned source-text/contract-guard idiom (see
|
||||
* tests/fix-2196-debug-agent-handoff.test.cjs).
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const DEBUG_MD = path.join(__dirname, '..', 'gsd-core', 'workflows', 'debug.md');
|
||||
const SESSION_MANAGER_MD = path.join(__dirname, '..', 'agents', 'gsd-debug-session-manager.md');
|
||||
|
||||
describe('#2257 debug non-terminal session-manager return contract', () => {
|
||||
// allow-test-rule: source-text-is-the-product (#2257)
|
||||
// workflow/agent prose IS the runtime contract under test
|
||||
const debugContent = fs.readFileSync(DEBUG_MD, 'utf-8');
|
||||
// allow-test-rule: source-text-is-the-product (#2257)
|
||||
// workflow/agent prose IS the runtime contract under test
|
||||
const managerContent = fs.readFileSync(SESSION_MANAGER_MD, 'utf-8');
|
||||
|
||||
const section4Start = debugContent.indexOf('## 4. Session Management');
|
||||
const section4 = section4Start !== -1 ? debugContent.slice(section4Start) : '';
|
||||
|
||||
const section1cStart = debugContent.indexOf('## 1c. CONTINUE subcommand');
|
||||
const section1dStart = debugContent.indexOf('## 1d. Check Active Sessions');
|
||||
const section1c =
|
||||
section1cStart !== -1 && section1dStart !== -1
|
||||
? debugContent.slice(section1cStart, section1dStart)
|
||||
: '';
|
||||
|
||||
test('debug.md has Section 4 (Session Management) and Section 1c (CONTINUE subcommand)', () => {
|
||||
assert.notEqual(section4Start, -1, 'debug.md must contain Section 4 Session Management');
|
||||
assert.notEqual(section1cStart, -1, 'debug.md must contain Section 1c CONTINUE subcommand');
|
||||
});
|
||||
|
||||
test('Section 4 has an exhaustive non-terminal branch that auto-resumes from the checkpoint', () => {
|
||||
assert.ok(/CONTINUE_REQUIRED/.test(section4),
|
||||
'Section 4 must reference the CONTINUE_REQUIRED non-terminal marker');
|
||||
assert.ok(/ANYTHING ELSE/i.test(section4),
|
||||
'Section 4 must exhaustively catch any return that is not one of the two terminal markers');
|
||||
assert.ok(/AUTO-RESUME/i.test(section4) && /re-spawning/i.test(section4),
|
||||
'Section 4 must auto-resume by re-spawning the session manager, not return control to the user');
|
||||
assert.ok(/same.{0,20}slug/i.test(section4),
|
||||
'Section 4 auto-resume must use the SAME slug/checkpoint as the original spawn');
|
||||
});
|
||||
|
||||
test('Section 1c has the same exhaustive non-terminal auto-resume branch (not just the two literals)', () => {
|
||||
assert.ok(/CONTINUE_REQUIRED/.test(section1c),
|
||||
'Section 1c must reference the CONTINUE_REQUIRED non-terminal marker');
|
||||
assert.ok(/ANYTHING ELSE/i.test(section1c),
|
||||
'Section 1c must exhaustively catch any return that is not one of the two terminal markers');
|
||||
assert.ok(/AUTO-RESUME/i.test(section1c) && /re-spawning/i.test(section1c),
|
||||
'Section 1c must auto-resume by re-spawning the session manager, not return control to the user');
|
||||
assert.ok(/same.{0,20}slug/i.test(section1c),
|
||||
'Section 1c auto-resume must use the SAME slug/checkpoint as the original spawn (symmetric with Section 4)');
|
||||
});
|
||||
|
||||
test('gsd-debug-session-manager.md defines CONTINUE_REQUIRED distinct from the two terminal formats', () => {
|
||||
assert.ok(/## CONTINUE_REQUIRED/.test(managerContent),
|
||||
'the agent must define an explicit ## CONTINUE_REQUIRED return heading');
|
||||
assert.ok(/## DEBUG SESSION COMPLETE/.test(managerContent),
|
||||
'the terminal DEBUG SESSION COMPLETE format must still be present');
|
||||
assert.ok(/ABANDONED/.test(managerContent),
|
||||
'the terminal ABANDONED format must still be present');
|
||||
assert.ok(/non-terminal/i.test(managerContent),
|
||||
'the agent must characterize CONTINUE_REQUIRED as non-terminal');
|
||||
assert.ok(/CHECKPOINT REACHED/.test(managerContent) && /distinct from/i.test(managerContent),
|
||||
'CONTINUE_REQUIRED must be explicitly distinguished from the genuine user-input CHECKPOINT REACHED shape');
|
||||
assert.ok(/\.planning\/debug\/\{slug\}\.md/.test(managerContent),
|
||||
'CONTINUE_REQUIRED must reference the on-disk checkpoint path');
|
||||
assert.ok(/next_action/.test(managerContent) && /status/.test(managerContent),
|
||||
'CONTINUE_REQUIRED must reference the checkpoint status/next_action fields');
|
||||
});
|
||||
|
||||
test('an anti-loop bound exists so repeated no-progress auto-resumes do not loop indefinitely', () => {
|
||||
assert.ok(/anti-loop guard/i.test(section4),
|
||||
'Section 4 must name an anti-loop guard');
|
||||
assert.ok(/blocker report/i.test(section4),
|
||||
'Section 4 must emit a blocker report to the user once the bound is exceeded, instead of looping forever');
|
||||
|
||||
assert.ok(/anti-loop guard/i.test(section1c),
|
||||
'Section 1c must name an anti-loop guard');
|
||||
});
|
||||
|
||||
test('the anti-loop guard has an absolute hard cap independent of no-progress detection (#2257 correction)', () => {
|
||||
for (const [label, section] of [['Section 4', section4], ['Section 1c', section1c]]) {
|
||||
assert.ok(/hard cap/i.test(section), `${label} must name an absolute hard cap`);
|
||||
assert.ok(/\b3\b/.test(section) && /total auto-resumes/i.test(section),
|
||||
`${label} must encode a concrete numeric cap of 3 total auto-resumes`);
|
||||
assert.ok(/regardless/i.test(section),
|
||||
`${label} hard cap must trip regardless of whether next_action changed (content-independent)`);
|
||||
}
|
||||
});
|
||||
|
||||
test('no-progress detection keys off next_action alone, never the always-changing updated timestamp (#2257 correction)', () => {
|
||||
for (const [label, section] of [['Section 4', section4], ['Section 1c', section1c]]) {
|
||||
assert.ok(/next_action/.test(section),
|
||||
`${label} no-progress heuristic must reference next_action`);
|
||||
assert.ok(/(do not|never).{0,40}updated/i.test(section),
|
||||
`${label} must explicitly forbid keying no-progress detection off updated`);
|
||||
assert.ok(/changes every cycle/i.test(section),
|
||||
`${label} must state WHY updated cannot be used: it is overwritten/changes every checkpoint cycle`);
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -1,662 +0,0 @@
|
||||
'use strict';
|
||||
|
||||
// allow-test-rule: source-text-is-the-product, see #2285 — reads gsd-core/workflows/execute-phase.md
|
||||
// prose to verify the render-hooks call site + ordering. The workflow markdown IS the runtime
|
||||
// contract executed by the orchestrator; there is no behavioral seam to drive this assertion
|
||||
// through other than the rendered prose itself.
|
||||
|
||||
/**
|
||||
* fix-2285-claude-orchestration-wiring.test.cjs
|
||||
*
|
||||
* #2285 — the `claude-orchestration` capability (Workflow backend, #1143) was
|
||||
* registered `active` but fully INERT: `detectWorkflowBackend`/`emitWorkflowScript`
|
||||
* had zero callers outside their own CLI router, and execute-phase.md declared
|
||||
* `execute:wave:pre` as a hook point in its frontmatter but never rendered it —
|
||||
* the wave loop only ever dispatched `execute:pre`, `execute:wave:post`, and
|
||||
* `execute:post`. `claude_orchestration.enabled:true` therefore had no effect on
|
||||
* a real execute-phase run.
|
||||
*
|
||||
* Fix (Approach B):
|
||||
* 1. execute-phase.md now renders `execute:wave:pre` immediately before each
|
||||
* wave's agents are dispatched (step 2.75, before step 3's Agent() loop).
|
||||
* 2. The claude-orchestration contribution moved from `execute:wave:post`
|
||||
* (fires too late — after the wave already dispatched inline) to
|
||||
* `execute:wave:pre` (fires before dispatch, where a backend selector
|
||||
* actually has to run to matter).
|
||||
* 3. `resolveWaveDispatch` in src/claude-orchestration.cts composes
|
||||
* `detectWorkflowBackend` + `emitWorkflowScript` into ONE decision seam,
|
||||
* giving both functions a real caller outside their CLI router and outside
|
||||
* tests. It is also exposed via `gsd-tools claude-orchestration
|
||||
* resolve-wave-dispatch`.
|
||||
*
|
||||
* This file drives the real seam (no source-grep on implementation files) and
|
||||
* asserts the fail-closed contract: disabled or any gate miss => inline,
|
||||
* byte-identical to today's dispatch shape.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const fc = require('fast-check');
|
||||
|
||||
const {
|
||||
detectWorkflowBackend,
|
||||
emitWorkflowScript,
|
||||
resolveWaveDispatch,
|
||||
WORKFLOW_TOOL_FLOOR_VERSION,
|
||||
} = require('../gsd-core/bin/lib/claude-orchestration.cjs');
|
||||
|
||||
const { runGsdTools, createTempDir, cleanup } = require('./helpers.cjs');
|
||||
|
||||
const ROOT = path.resolve(__dirname, '..');
|
||||
const WORKFLOW_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'execute-phase.md');
|
||||
const CAP_PATH = path.join(ROOT, 'capabilities', 'claude-orchestration', 'capability.json');
|
||||
|
||||
// ─── Fixtures ───────────────────────────────────────────────────────────────
|
||||
|
||||
/** A host-integration descriptor whose dispatch axis signals Workflow-tool capability. */
|
||||
const CAPABLE_HOST = { dispatch: { nested: true, background: true } };
|
||||
/** A descriptor that fails the nested/background dispatch gate. */
|
||||
const INCAPABLE_HOST = { dispatch: { nested: false, background: true } };
|
||||
|
||||
const ABOVE_FLOOR_SDK = '0.3.150';
|
||||
const AT_FLOOR_SDK = WORKFLOW_TOOL_FLOOR_VERSION; // '0.3.149'
|
||||
const BELOW_FLOOR_SDK = '0.3.148';
|
||||
|
||||
function enabledConfig(overrides = {}) {
|
||||
return {
|
||||
'claude_orchestration.enabled': true,
|
||||
'claude_orchestration.execution_backend': 'auto',
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
function singleWave() {
|
||||
return {
|
||||
phaseDir: '.planning/phases/01-foo',
|
||||
runId: 'run-2285-1',
|
||||
waves: [
|
||||
{
|
||||
id: 'w1',
|
||||
plans: [
|
||||
{ id: 'p1', brief: 'Implement the foo module', files_modified: ['src/foo.cts'] },
|
||||
],
|
||||
},
|
||||
],
|
||||
};
|
||||
}
|
||||
|
||||
function baseInput(overrides = {}) {
|
||||
return {
|
||||
runtimeId: 'claude',
|
||||
hostIntegration: CAPABLE_HOST,
|
||||
agentSdkVersion: ABOVE_FLOOR_SDK,
|
||||
config: enabledConfig(),
|
||||
...singleWave(),
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
// ─── Section A: happy path — every gate satisfied → workflow backend ────────
|
||||
|
||||
describe('A. resolveWaveDispatch — enabled + all gates satisfied → workflow backend with emitted script', () => {
|
||||
test('[happy] enabled, claude runtime, capable host, SDK above floor, auto backend → backend:"workflow"', () => {
|
||||
const result = resolveWaveDispatch(baseInput());
|
||||
assert.strictEqual(result.backend, 'workflow');
|
||||
assert.strictEqual(result.reason, 'workflow_backend_active');
|
||||
assert.ok(typeof result.script === 'string' && result.script.length > 0, 'script must be a non-empty string');
|
||||
// #2590: resumeFromRunId is a Workflow TOOL INPUT, not a script function —
|
||||
// calling it threw "resumeFromRunId is not defined". The run id must still
|
||||
// reach the caller (which passes it as that input), but never as a call.
|
||||
assert.ok(!/^\s*resumeFromRunId\s*\(/m.test(result.script), 'must not CALL resumeFromRunId');
|
||||
assert.strictEqual(result.summary.resumeRunId, 'run-2285-1');
|
||||
assert.match(result.script, /agentType: "gsd-executor", isolation: "worktree"/);
|
||||
assert.ok(result.summary && result.summary.plans === 1, 'summary.plans must reflect the manifest');
|
||||
});
|
||||
|
||||
test('[happy] execution_backend explicitly "workflow" (not just "auto") also activates', () => {
|
||||
const result = resolveWaveDispatch(baseInput({
|
||||
config: enabledConfig({ 'claude_orchestration.execution_backend': 'workflow' }),
|
||||
}));
|
||||
assert.strictEqual(result.backend, 'workflow');
|
||||
});
|
||||
|
||||
test('[bva] SDK version boundary: floor-1 → inline, floor exact → workflow, floor+1 → workflow', () => {
|
||||
const below = resolveWaveDispatch(baseInput({ agentSdkVersion: BELOW_FLOOR_SDK }));
|
||||
assert.strictEqual(below.backend, 'inline', 'below floor must be inline');
|
||||
assert.strictEqual(below.reason, 'agent_sdk_version_below_floor');
|
||||
|
||||
const at = resolveWaveDispatch(baseInput({ agentSdkVersion: AT_FLOOR_SDK }));
|
||||
assert.strictEqual(at.backend, 'workflow', 'exactly at floor must activate workflow');
|
||||
|
||||
const above = resolveWaveDispatch(baseInput({ agentSdkVersion: ABOVE_FLOOR_SDK }));
|
||||
assert.strictEqual(above.backend, 'workflow', 'above floor must activate workflow');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section B: fail-closed contract — disabled / each gate individually failing → inline ──
|
||||
|
||||
describe('B. resolveWaveDispatch — fail-closed contract: disabled or any gate miss → inline, matches detectWorkflowBackend 1:1', () => {
|
||||
const GATE_MISS_CASES = [
|
||||
{
|
||||
label: 'capability disabled',
|
||||
overrides: { config: {} },
|
||||
expectedReason: 'capability_disabled',
|
||||
},
|
||||
{
|
||||
label: 'capability explicitly disabled',
|
||||
overrides: { config: enabledConfig({ 'claude_orchestration.enabled': false }) },
|
||||
expectedReason: 'capability_disabled',
|
||||
},
|
||||
{
|
||||
label: 'runtime is not claude',
|
||||
overrides: { runtimeId: 'codex' },
|
||||
expectedReason: 'runtime_not_claude',
|
||||
},
|
||||
{
|
||||
label: 'execution_backend explicitly "inline"',
|
||||
overrides: { config: enabledConfig({ 'claude_orchestration.execution_backend': 'inline' }) },
|
||||
expectedReason: 'backend_inline',
|
||||
},
|
||||
{
|
||||
label: 'host descriptor incapable (nested:false)',
|
||||
overrides: { hostIntegration: INCAPABLE_HOST },
|
||||
expectedReason: 'workflow_tool_unavailable',
|
||||
},
|
||||
{
|
||||
label: 'host descriptor missing entirely',
|
||||
overrides: { hostIntegration: null },
|
||||
expectedReason: 'workflow_tool_unavailable',
|
||||
},
|
||||
{
|
||||
label: 'agent SDK version missing',
|
||||
overrides: { agentSdkVersion: undefined },
|
||||
expectedReason: 'agent_sdk_version_unknown',
|
||||
},
|
||||
{
|
||||
label: 'agent SDK version malformed (not semver)',
|
||||
overrides: { agentSdkVersion: 'not-a-version' },
|
||||
expectedReason: 'agent_sdk_version_unknown',
|
||||
},
|
||||
{
|
||||
label: 'agent SDK version below floor',
|
||||
overrides: { agentSdkVersion: BELOW_FLOOR_SDK },
|
||||
expectedReason: 'agent_sdk_version_below_floor',
|
||||
},
|
||||
];
|
||||
|
||||
for (const { label, overrides, expectedReason } of GATE_MISS_CASES) {
|
||||
test(`[negative] ${label} → backend:"inline", reason:"${expectedReason}"`, () => {
|
||||
const input = baseInput(overrides);
|
||||
const result = resolveWaveDispatch(input);
|
||||
|
||||
assert.strictEqual(result.backend, 'inline', `${label}: must resolve to inline`);
|
||||
assert.strictEqual(result.reason, expectedReason, `${label}: reason mismatch`);
|
||||
|
||||
// Fail-closed CONTRACT: today's (byte-identical) inline dispatch carries no
|
||||
// script/summary. Verify the shape never leaks emitter fields on a gate miss.
|
||||
assert.deepStrictEqual(
|
||||
Object.keys(result).sort(),
|
||||
['backend', 'reason'],
|
||||
`${label}: inline result must be exactly {backend, reason}, got keys: ${Object.keys(result).join(',')}`,
|
||||
);
|
||||
|
||||
// Parity: resolveWaveDispatch must not reimplement the gate ladder — its
|
||||
// reason for a detect-side miss must be IDENTICAL to calling
|
||||
// detectWorkflowBackend directly with the same gate-relevant fields.
|
||||
const direct = detectWorkflowBackend({
|
||||
runtimeId: input.runtimeId,
|
||||
hostIntegration: input.hostIntegration,
|
||||
config: input.config,
|
||||
agentSdkVersion: input.agentSdkVersion,
|
||||
});
|
||||
assert.strictEqual(direct.backend, 'inline', `${label}: detectWorkflowBackend parity check must also be inline`);
|
||||
assert.strictEqual(result.reason, direct.reason, `${label}: resolveWaveDispatch must surface detectWorkflowBackend's own reason verbatim`);
|
||||
});
|
||||
}
|
||||
|
||||
test('[negative] null/undefined/non-object input → inline, reason:"invalid_input" (never throws)', () => {
|
||||
assert.deepStrictEqual(resolveWaveDispatch(null), { backend: 'inline', reason: 'invalid_input' });
|
||||
assert.deepStrictEqual(resolveWaveDispatch(undefined), { backend: 'inline', reason: 'invalid_input' });
|
||||
assert.deepStrictEqual(resolveWaveDispatch('not-an-object'), { backend: 'inline', reason: 'invalid_input' });
|
||||
});
|
||||
|
||||
test('[happy] a valid, dispatch-ready waves manifest never flips a gate-missed decision to workflow', () => {
|
||||
// Prove the gate ladder short-circuits BEFORE emitWorkflowScript ever runs:
|
||||
// even with a perfectly valid wave manifest, a disabled capability stays inline.
|
||||
const result = resolveWaveDispatch(baseInput({ config: {}, ...singleWave() }));
|
||||
assert.strictEqual(result.backend, 'inline');
|
||||
assert.strictEqual(result.reason, 'capability_disabled');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section C: composition correctness — detect + emit have a real, non-CLI, non-test caller ──
|
||||
|
||||
describe('C. resolveWaveDispatch composes detectWorkflowBackend + emitWorkflowScript (the seam itself)', () => {
|
||||
test('[happy] resolveWaveDispatch is exported as a function from the core module', () => {
|
||||
assert.strictEqual(typeof resolveWaveDispatch, 'function');
|
||||
});
|
||||
|
||||
test('[happy] on a workflow-hit, the emitted script/summary are IDENTICAL to calling emitWorkflowScript directly with the same wave data', () => {
|
||||
const input = baseInput();
|
||||
const composed = resolveWaveDispatch(input);
|
||||
assert.strictEqual(composed.backend, 'workflow');
|
||||
|
||||
const directEmit = emitWorkflowScript({
|
||||
phaseDir: input.phaseDir,
|
||||
waves: input.waves,
|
||||
runId: input.runId,
|
||||
});
|
||||
assert.strictEqual(directEmit.ok, true);
|
||||
assert.strictEqual(composed.script, directEmit.script, 'resolveWaveDispatch must not re-implement emission — script must match emitWorkflowScript byte-for-byte');
|
||||
assert.deepStrictEqual(composed.summary, directEmit.summary);
|
||||
});
|
||||
|
||||
test('[negative] detect-hit but a malformed wave manifest (emit failure) → inline, carrying emitWorkflowScript\'s own failure reason', () => {
|
||||
const input = baseInput({ waves: [] }); // emitWorkflowScript rejects empty waves
|
||||
const result = resolveWaveDispatch(input);
|
||||
assert.strictEqual(result.backend, 'inline');
|
||||
|
||||
const directEmit = emitWorkflowScript({ phaseDir: input.phaseDir, waves: input.waves, runId: input.runId });
|
||||
assert.strictEqual(directEmit.ok, false);
|
||||
assert.strictEqual(result.reason, 'emit_failed: ' + directEmit.reason, 'the emit failure reason must be surfaced verbatim, prefixed');
|
||||
|
||||
// Still byte-identical inline shape — no partial/broken script ever leaks.
|
||||
assert.deepStrictEqual(Object.keys(result).sort(), ['backend', 'reason']);
|
||||
});
|
||||
|
||||
test('[happy] the CLI subcommand `claude-orchestration resolve-wave-dispatch` is ALSO a caller and matches the pure function output', () => {
|
||||
const tmp = createTempDir('fix-2285-');
|
||||
try {
|
||||
fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(
|
||||
path.join(tmp, '.planning', 'config.json'),
|
||||
JSON.stringify({ claude_orchestration: { enabled: true, execution_backend: 'auto' } }),
|
||||
);
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, JSON.stringify({ waves: singleWave().waves }));
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', wavesPath,
|
||||
'--run-id', 'run-2285-1',
|
||||
'--phase-dir', '.planning/phases/01-foo',
|
||||
'--runtime', 'claude',
|
||||
'--agent-sdk-version', ABOVE_FLOOR_SDK,
|
||||
// #2686: the router now DEFAULTS the executor model from project config,
|
||||
// so pin it on both sides — otherwise this compares a config-resolved CLI
|
||||
// run against a pure call that was given no model, and the equality this
|
||||
// test exists to prove would be testing the default instead of the seam.
|
||||
'--executor-model', 'sonnet',
|
||||
'--raw',
|
||||
], tmp);
|
||||
assert.strictEqual(res.success, true, 'CLI command must succeed; stderr: ' + (res.error || ''));
|
||||
const parsed = JSON.parse(res.output);
|
||||
|
||||
const direct = resolveWaveDispatch(baseInput({ executorModel: 'sonnet' }));
|
||||
assert.strictEqual(parsed.backend, direct.backend);
|
||||
assert.strictEqual(parsed.script, direct.script);
|
||||
assert.deepStrictEqual(parsed.summary, direct.summary);
|
||||
} finally {
|
||||
cleanup(tmp);
|
||||
}
|
||||
});
|
||||
|
||||
test('[negative] CLI subcommand fails closed to inline exactly like the pure function when disabled', () => {
|
||||
const tmp = createTempDir('fix-2285-off-');
|
||||
try {
|
||||
fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(path.join(tmp, '.planning', 'config.json'), '{}');
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, JSON.stringify({ waves: singleWave().waves }));
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', wavesPath, '--run-id', 'run-x', '--raw',
|
||||
], tmp);
|
||||
assert.strictEqual(res.success, true, 'CLI command must succeed (fail-closed, not error); stderr: ' + (res.error || ''));
|
||||
const parsed = JSON.parse(res.output);
|
||||
assert.strictEqual(parsed.backend, 'inline');
|
||||
assert.strictEqual(parsed.reason, 'capability_disabled');
|
||||
assert.deepStrictEqual(Object.keys(parsed).sort(), ['backend', 'reason']);
|
||||
} finally {
|
||||
cleanup(tmp);
|
||||
}
|
||||
});
|
||||
|
||||
test('property: for ANY input, resolveWaveDispatch never throws, backend is always "inline"|"workflow", and "inline" results carry exactly {backend, reason}', () => {
|
||||
fc.assert(fc.property(
|
||||
fc.record({
|
||||
runtimeId: fc.oneof(fc.constant('claude'), fc.constant('codex'), fc.constant(undefined), fc.string()),
|
||||
agentSdkVersion: fc.oneof(fc.constant(ABOVE_FLOOR_SDK), fc.constant(BELOW_FLOOR_SDK), fc.constant(undefined), fc.string()),
|
||||
enabled: fc.boolean(),
|
||||
capableHost: fc.boolean(),
|
||||
backendPref: fc.constantFrom('auto', 'workflow', 'inline'),
|
||||
}),
|
||||
({ runtimeId, agentSdkVersion, enabled, capableHost, backendPref }) => {
|
||||
const input = {
|
||||
runtimeId,
|
||||
hostIntegration: capableHost ? CAPABLE_HOST : INCAPABLE_HOST,
|
||||
agentSdkVersion,
|
||||
config: {
|
||||
'claude_orchestration.enabled': enabled,
|
||||
'claude_orchestration.execution_backend': backendPref,
|
||||
},
|
||||
...singleWave(),
|
||||
};
|
||||
const result = resolveWaveDispatch(input);
|
||||
assert.ok(result.backend === 'inline' || result.backend === 'workflow');
|
||||
if (result.backend === 'inline') {
|
||||
assert.deepStrictEqual(Object.keys(result).sort(), ['backend', 'reason']);
|
||||
} else {
|
||||
assert.ok(typeof result.script === 'string' && result.script.length > 0);
|
||||
}
|
||||
},
|
||||
));
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section D: capability declaration now targets execute:wave:pre ─────────
|
||||
|
||||
describe('D. capability.json declares the contribution at execute:wave:pre (#2285)', () => {
|
||||
test('[happy] contribution point is execute:wave:pre, not execute:wave:post', () => {
|
||||
const cap = JSON.parse(fs.readFileSync(CAP_PATH, 'utf8'));
|
||||
const wavePreContrib = cap.contributions.find((c) => c.point === 'execute:wave:pre');
|
||||
assert.ok(wavePreContrib, 'capability.json must declare a contribution at execute:wave:pre');
|
||||
assert.strictEqual(wavePreContrib.into, 'executor');
|
||||
assert.strictEqual(wavePreContrib.when, 'claude_orchestration.enabled');
|
||||
assert.strictEqual(wavePreContrib.onError, 'skip');
|
||||
assert.strictEqual(wavePreContrib.fragment.path, 'fragments/execute-wave-pre.md');
|
||||
|
||||
const wavePostContrib = cap.contributions.find((c) => c.point === 'execute:wave:post');
|
||||
assert.strictEqual(wavePostContrib, undefined, 'the capability must no longer contribute at execute:wave:post');
|
||||
});
|
||||
|
||||
test('[happy] the declared fragment file exists on disk', () => {
|
||||
const fragPath = path.join(ROOT, 'capabilities', 'claude-orchestration', 'fragments', 'execute-wave-pre.md');
|
||||
assert.ok(fs.existsSync(fragPath), 'fragments/execute-wave-pre.md must exist');
|
||||
const content = fs.readFileSync(fragPath, 'utf8');
|
||||
assert.match(content, /execute:wave:pre/);
|
||||
assert.match(content, /resolve-wave-dispatch/);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section E: source-contract guard — execute-phase.md renders execute:wave:pre BEFORE dispatch ──
|
||||
|
||||
describe('E. execute-phase.md actually renders execute:wave:pre (the dead hook is now live)', () => {
|
||||
test('[happy] execute-phase.md invokes `loop render-hooks execute:wave:pre`', () => {
|
||||
const doc = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
||||
assert.ok(
|
||||
/loop render-hooks execute:wave:pre/.test(doc),
|
||||
'execute-phase.md must dispatch execute:wave:pre hooks (was declared in frontmatter but never rendered — #2285)',
|
||||
);
|
||||
});
|
||||
|
||||
test('[happy] the execute:wave:pre render-hooks call site appears BEFORE the wave\'s Agent() dispatch (pre-wave, not post)', () => {
|
||||
const doc = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
||||
const preHooksIdx = doc.indexOf('loop render-hooks execute:wave:pre');
|
||||
// Anchor on the actual per-wave dispatch call (step 3), not the generic
|
||||
// `subagent_type="gsd-executor"` mention in <runtime_compatibility> near the
|
||||
// top of the file — that mention predates the wave loop entirely and would
|
||||
// give a false "before" reading.
|
||||
const agentDispatchIdx = doc.indexOf('description="Execute plan {plan_number}');
|
||||
assert.ok(preHooksIdx !== -1, 'execute:wave:pre render-hooks call site must exist');
|
||||
assert.ok(agentDispatchIdx !== -1, 'the gsd-executor Agent() dispatch call site (step 3) must exist');
|
||||
assert.ok(
|
||||
preHooksIdx < agentDispatchIdx,
|
||||
`execute:wave:pre render-hooks (idx ${preHooksIdx}) must appear BEFORE the wave's Agent() dispatch (idx ${agentDispatchIdx}) — it is a pre-wave hook`,
|
||||
);
|
||||
});
|
||||
|
||||
test('[happy] the frontmatter still declares all four execute:* points (regression guard)', () => {
|
||||
const doc = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
||||
const frontmatterMatch = doc.match(/points:\s*(.+)/);
|
||||
assert.ok(frontmatterMatch, 'frontmatter must declare a points: line');
|
||||
for (const point of ['execute:pre', 'execute:wave:pre', 'execute:wave:post', 'execute:post']) {
|
||||
assert.ok(frontmatterMatch[1].includes(point), `frontmatter points: line must include ${point}`);
|
||||
}
|
||||
});
|
||||
|
||||
test('[happy] execute:wave:post is still rendered too (regression guard — did not accidentally remove the post-wave gate dispatch)', () => {
|
||||
const doc = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
||||
assert.ok(
|
||||
/loop render-hooks execute:wave:post/.test(doc),
|
||||
'execute-phase.md must still dispatch execute:wave:post hooks (drift/ui gates unaffected by #2285)',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section F: orthogonal-review finding 1 — submodule plans never forced into worktree isolation ──
|
||||
//
|
||||
// #2772 / #2285 finding 1: emitWorkflowScript previously hardcoded
|
||||
// `isolation: "worktree"` for EVERY plan. execute-phase.md step 2.5 computes
|
||||
// USE_WORKTREES_FOR_PLAN per plan specifically to keep submodule-touching
|
||||
// plans OUT of worktree isolation (the executor commit protocol cannot
|
||||
// correctly handle submodule commits inside an isolated worktree). The
|
||||
// Workflow backend must honor the SAME per-plan decision via `use_worktree`.
|
||||
|
||||
function waveWithSubmodulePlan() {
|
||||
return {
|
||||
phaseDir: '.planning/phases/01-foo',
|
||||
runId: 'run-2285-submodule',
|
||||
waves: [
|
||||
{
|
||||
id: 'w1',
|
||||
plans: [
|
||||
{ id: 'p1', brief: 'normal plan', files_modified: ['src/a.ts'] },
|
||||
{ id: 'p2', brief: 'submodule plan', files_modified: ['vendor/lib.c'], use_worktree: false },
|
||||
],
|
||||
},
|
||||
],
|
||||
};
|
||||
}
|
||||
|
||||
describe('F. Workflow backend never forces worktree isolation on a submodule / use_worktree:false plan', () => {
|
||||
test('[happy] resolveWaveDispatch (pure seam): the submodule plan\'s agent() call carries NO isolation, the normal plan\'s does', () => {
|
||||
const result = resolveWaveDispatch(baseInput({ ...waveWithSubmodulePlan() }));
|
||||
assert.strictEqual(result.backend, 'workflow');
|
||||
assert.match(result.script, /agent\("normal plan", \{ agentType: "gsd-executor", isolation: "worktree" \}\)/);
|
||||
// #2686: the options object legitimately gained an optional `model` key, so assert
|
||||
// the invariant this test exists to protect — agentType present, isolation absent —
|
||||
// rather than a frozen literal that any future additive key would break.
|
||||
assert.match(result.script, /agent\("submodule plan", \{ agentType: "gsd-executor"[^}]*\}\)/);
|
||||
assert.ok(
|
||||
!/agent\("submodule plan"[^)]*isolation/.test(result.script),
|
||||
'the submodule-touching plan must NEVER be emitted with forced worktree isolation',
|
||||
);
|
||||
});
|
||||
|
||||
test('[happy] CLI `resolve-wave-dispatch`: same per-plan guarantee end-to-end through the subprocess', () => {
|
||||
const tmp = createTempDir('fix-2285-submodule-');
|
||||
try {
|
||||
fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(
|
||||
path.join(tmp, '.planning', 'config.json'),
|
||||
JSON.stringify({ claude_orchestration: { enabled: true, execution_backend: 'auto' } }),
|
||||
);
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, JSON.stringify({ waves: waveWithSubmodulePlan().waves }));
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', wavesPath,
|
||||
'--run-id', 'run-2285-submodule',
|
||||
'--phase-dir', '.planning/phases/01-foo',
|
||||
'--runtime', 'claude',
|
||||
'--agent-sdk-version', ABOVE_FLOOR_SDK,
|
||||
'--raw',
|
||||
], tmp);
|
||||
assert.strictEqual(res.success, true, 'CLI command must succeed; stderr: ' + (res.error || ''));
|
||||
const parsed = JSON.parse(res.output);
|
||||
assert.strictEqual(parsed.backend, 'workflow');
|
||||
// #2686: additive `model` key — see the note on the pure-seam test above.
|
||||
assert.match(parsed.script, /agent\("submodule plan", \{ agentType: "gsd-executor"[^}]*\}\)/);
|
||||
assert.ok(!/agent\("submodule plan"[^)]*isolation/.test(parsed.script));
|
||||
} finally {
|
||||
cleanup(tmp);
|
||||
}
|
||||
});
|
||||
|
||||
test('[negative] use_worktree defaults to true when omitted — a manifest with NO submodule info stays backward-compatible', () => {
|
||||
const result = resolveWaveDispatch(baseInput());
|
||||
assert.strictEqual(result.backend, 'workflow');
|
||||
assert.match(result.script, /isolation: "worktree"/, 'default (no use_worktree field) must still isolate — backward compatible');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section G: orthogonal-review finding 2 — missing top-level `waves` key must never silently exit 0 ──
|
||||
//
|
||||
// readWavesManifest previously collapsed "read/parse threw" and "parsed OK but
|
||||
// no top-level `waves` key" into the same `undefined` sentinel. The call sites'
|
||||
// `if (waves === undefined) return;` made the missing-key case exit 0 with ZERO
|
||||
// output — fail-silent, breaking the "exit 0 => parseable JSON verdict" contract.
|
||||
// A missing key must now flow through to emitWorkflowScript's own validation,
|
||||
// exactly like an explicit `{"waves": null}` manifest already does.
|
||||
|
||||
describe('G. missing top-level `waves` key never silently exits 0 with no output', () => {
|
||||
function projectWithEnabledCapability(prefix) {
|
||||
const tmp = createTempDir(prefix);
|
||||
fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(
|
||||
path.join(tmp, '.planning', 'config.json'),
|
||||
JSON.stringify({ claude_orchestration: { enabled: true, execution_backend: 'auto' } }),
|
||||
);
|
||||
return tmp;
|
||||
}
|
||||
|
||||
test('[negative] resolve-wave-dispatch with a {"notwaves":[]} manifest → non-empty JSON verdict (NOT silent exit 0)', () => {
|
||||
const tmp = projectWithEnabledCapability('fix-2285-missingkey-resolve-');
|
||||
try {
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, JSON.stringify({ notwaves: [] }));
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', wavesPath, '--run-id', 'run-x',
|
||||
'--runtime', 'claude', '--agent-sdk-version', ABOVE_FLOOR_SDK,
|
||||
'--raw',
|
||||
], tmp);
|
||||
|
||||
assert.strictEqual(res.success, true, 'command must exit 0 (fail-closed to inline, not error); stderr: ' + (res.error || ''));
|
||||
assert.ok(res.output.length > 0, 'FAIL-SILENT REGRESSION: missing waves key must NOT produce empty stdout on exit 0');
|
||||
const parsed = JSON.parse(res.output);
|
||||
assert.strictEqual(parsed.backend, 'inline');
|
||||
assert.match(parsed.reason, /waves must be a non-empty array/, 'reason must surface emitWorkflowScript\'s own validation message');
|
||||
} finally {
|
||||
cleanup(tmp);
|
||||
}
|
||||
});
|
||||
|
||||
test('[negative] resolve-wave-dispatch: {"notwaves":[]} and {"waves": null} produce the IDENTICAL verdict (parity)', () => {
|
||||
const tmp = projectWithEnabledCapability('fix-2285-missingkey-parity-');
|
||||
try {
|
||||
const missingKeyPath = path.join(tmp, 'missing.json');
|
||||
fs.writeFileSync(missingKeyPath, JSON.stringify({ notwaves: [] }));
|
||||
const nullWavesPath = path.join(tmp, 'null.json');
|
||||
fs.writeFileSync(nullWavesPath, JSON.stringify({ waves: null }));
|
||||
|
||||
const argsFor = (p) => [
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', p, '--run-id', 'run-x',
|
||||
'--runtime', 'claude', '--agent-sdk-version', ABOVE_FLOOR_SDK,
|
||||
'--raw',
|
||||
];
|
||||
const missingRes = runGsdTools(argsFor(missingKeyPath), tmp);
|
||||
const nullRes = runGsdTools(argsFor(nullWavesPath), tmp);
|
||||
assert.strictEqual(missingRes.success, true);
|
||||
assert.strictEqual(nullRes.success, true);
|
||||
assert.deepStrictEqual(JSON.parse(missingRes.output), JSON.parse(nullRes.output), 'a missing `waves` key must behave identically to an explicit `waves: null`');
|
||||
} finally {
|
||||
cleanup(tmp);
|
||||
}
|
||||
});
|
||||
|
||||
test('[negative] emit-workflow with a {"notwaves":[]} manifest → loud non-zero exit (NOT silent exit 0)', () => {
|
||||
const tmp = createTempDir('fix-2285-missingkey-emit-');
|
||||
try {
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, JSON.stringify({ notwaves: [] }));
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'emit-workflow',
|
||||
'--waves', wavesPath, '--run-id', 'run-x',
|
||||
], tmp);
|
||||
|
||||
assert.strictEqual(res.success, false, 'FAIL-SILENT REGRESSION: missing waves key must produce a loud, non-zero-exit error, not a silent success');
|
||||
assert.ok(res.exitCode !== 0, 'non-zero exit');
|
||||
assert.match(res.error || '', /waves must be a non-empty array/);
|
||||
} finally {
|
||||
cleanup(tmp);
|
||||
}
|
||||
});
|
||||
|
||||
test('[happy] a genuinely malformed (unparseable) --waves file still fails loudly, unaffected by the fix', () => {
|
||||
const tmp = createTempDir('fix-2285-badjson-');
|
||||
try {
|
||||
const wavesPath = path.join(tmp, 'waves.json');
|
||||
fs.writeFileSync(wavesPath, 'not json at all');
|
||||
|
||||
const res = runGsdTools([
|
||||
'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', wavesPath, '--run-id', 'run-x', '--raw',
|
||||
], tmp);
|
||||
assert.strictEqual(res.success, false, 'a real parse failure must still error');
|
||||
assert.match(res.error || '', /could not read\/parse --waves file/);
|
||||
} finally {
|
||||
cleanup(tmp);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section H: orthogonal-review finding 3 — manifest construction guidance is concrete ──
|
||||
|
||||
describe('H. the execute:wave:pre fragment documents concrete manifest construction (finding 3)', () => {
|
||||
test('[happy] the fragment explains how to build WAVE_MANIFEST_PATH, PHASE_RUN_ID, and per-plan use_worktree', () => {
|
||||
const fragPath = path.join(ROOT, 'capabilities', 'claude-orchestration', 'fragments', 'execute-wave-pre.md');
|
||||
const content = fs.readFileSync(fragPath, 'utf8');
|
||||
assert.match(content, /Manifest construction/, 'fragment must have concrete manifest-construction guidance, not just reference undefined vars');
|
||||
assert.match(content, /PHASE_RUN_ID/);
|
||||
assert.match(content, /WAVE_MANIFEST_PATH/);
|
||||
assert.match(content, /use_worktree/);
|
||||
assert.match(content, /USE_WORKTREES_FOR_PLAN/, 'must tie use_worktree back to step 2.5\'s per-plan decision');
|
||||
});
|
||||
|
||||
test('[happy] execute-phase.md step 2.75 stays minimal — manifest/use_worktree detail lives ONLY in the fragment (#1168 byte-budget conformance)', () => {
|
||||
// Per the ADR-857 Phase 6 conformance gate (tests/phase6-capstone-conformance.test.cjs),
|
||||
// the host loop must stay small — optional-feature detail (manifest construction,
|
||||
// per-plan use_worktree carry-through) belongs in the capability fragment, not the
|
||||
// host workflow. Step 2.75 is intentionally just a render-hooks call + a one-line
|
||||
// "follow the contribution or fall through to step 3" instruction.
|
||||
const doc = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
||||
const stepStart = doc.indexOf('2.75. **Execute:wave:pre capability dispatch:**');
|
||||
const stepEnd = doc.indexOf('\n3. **Spawn executor agents:**', stepStart);
|
||||
assert.ok(stepStart !== -1 && stepEnd !== -1, 'step 2.75 must exist and precede step 3');
|
||||
const stepBody = doc.slice(stepStart, stepEnd);
|
||||
assert.match(stepBody, /loop render-hooks execute:wave:pre/, 'step 2.75 must still render the hook point');
|
||||
assert.doesNotMatch(stepBody, /use_worktree/, 'manifest-construction detail (use_worktree) must live in the fragment, not the host step');
|
||||
assert.doesNotMatch(stepBody, /USE_WORKTREES_FOR_PLAN/, 'per-plan worktree gate detail must live in the fragment, not the host step');
|
||||
});
|
||||
|
||||
test('[happy] execute-phase.md is below the ADR-857 Phase 6 pre-phase-6 byte ceiling (#1168), with margin', () => {
|
||||
const { lfByteCount } = require('../scripts/workflow-size.cjs');
|
||||
const bytes = lfByteCount(WORKFLOW_PATH);
|
||||
assert.ok(bytes < 93600, `execute-phase.md must stay below the frozen pre-phase-6 ceiling (93600); got ${bytes}`);
|
||||
assert.ok(bytes <= 93400, `execute-phase.md should carry a comfortable margin (<=93400) so minor future edits don't re-trip the gate; got ${bytes}`);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Section I: orthogonal-review finding 4 — stale doc fixed ───────────────
|
||||
|
||||
describe('I. docs/explanation/claude-orchestration-capability.md reflects the execute:wave:pre move (finding 4)', () => {
|
||||
test('[happy] the doc no longer claims the capability registers at execute:wave:post', () => {
|
||||
const docPath = path.join(ROOT, 'docs', 'explanation', 'claude-orchestration-capability.md');
|
||||
const content = fs.readFileSync(docPath, 'utf8');
|
||||
assert.match(content, /execute:wave:pre/, 'doc must mention execute:wave:pre as the wired point');
|
||||
assert.ok(
|
||||
!/execute:wave:post.*\(into the executor\)/.test(content),
|
||||
'doc must not still claim the wired point is execute:wave:post',
|
||||
);
|
||||
});
|
||||
});
|
||||
@@ -1,345 +0,0 @@
|
||||
/**
|
||||
* #2287 — deferred-items.md has no reader anywhere in gsd-core.
|
||||
*
|
||||
* The SCOPE BOUNDARY convention (`agents/gsd-executor.md`) instructs the
|
||||
* executor to log out-of-scope discoveries to `deferred-items.md` inside the
|
||||
* phase directory. Nothing read that file back: `cmdAuditUat` (src/uat.cts)
|
||||
* filtered phase-directory files down to `*-UAT.md` / `*-VERIFICATION.md`
|
||||
* only, and the `forensic_audit` workflow step (gsd-core/workflows/
|
||||
* progress.md) ran 6 checks, none of which globbed the phase-directory
|
||||
* `deferred-items.md` path. An entry written there was permanently invisible.
|
||||
*
|
||||
* This fix:
|
||||
* - `cmdAuditUat` gains a `deferred-items.md` scan per phase directory,
|
||||
* surfacing every UNRESOLVED entry as a `type: 'deferred'` result. An
|
||||
* entry is resolved only when it carries an explicit `status: resolved`
|
||||
* field (mirroring the established `## Gaps` convention from #2286) — a
|
||||
* missing/garbled status fails safe and is surfaced.
|
||||
* - `forensic_audit` gains a 7th check that globs the same path and reports
|
||||
* unresolved entries with the same ✓/⚠ semantics as the other 6 checks.
|
||||
*
|
||||
* `deferred-items.md` remains the single source of truth — no duplicate
|
||||
* `.planning/todos/pending/*.md` entry is required.
|
||||
*/
|
||||
|
||||
'use strict';
|
||||
|
||||
const { test, describe, beforeEach, afterEach } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const fc = require('./helpers/fast-check-setup.cjs');
|
||||
const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs');
|
||||
const { parseDeferredItems } = require('../gsd-core/bin/lib/uat.cjs');
|
||||
|
||||
// ─── cmdAuditUat behavioral coverage ───────────────────────────────────────
|
||||
|
||||
describe('#2287 cmdAuditUat: deferred-items.md awareness', () => {
|
||||
let tmpDir;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempProject();
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(tmpDir);
|
||||
});
|
||||
|
||||
test('no deferred-items.md present (0 entries) → no results, no false positive', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
fs.writeFileSync(path.join(phaseDir, '.gitkeep'), '');
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.deepStrictEqual(output.results, []);
|
||||
assert.strictEqual(output.summary.total_items, 0);
|
||||
assert.strictEqual(output.summary.total_files, 0);
|
||||
});
|
||||
|
||||
test('deferred-items.md with only a resolved entry (0 unresolved) → no result surfaced', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- Already handled unrelated lint warning.',
|
||||
' status: resolved',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.deepStrictEqual(output.results, [],
|
||||
'a fully-resolved deferred-items.md must not surface any result');
|
||||
assert.strictEqual(output.summary.total_items, 0);
|
||||
});
|
||||
|
||||
test('deferred-items.md with 1 unresolved entry → surfaced in structured JSON output', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- Found an unrelated pre-existing test failure in `some-other-module` while working on',
|
||||
' this phase\'s task. Out of scope for this task — logged here per SCOPE BOUNDARY.',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.strictEqual(output.summary.total_items, 1);
|
||||
assert.strictEqual(output.summary.total_files, 1);
|
||||
assert.strictEqual(output.summary.by_category.deferred, 1);
|
||||
assert.strictEqual(output.summary.by_phase['01'], 1);
|
||||
|
||||
const deferredResult = output.results.find(r => r.type === 'deferred');
|
||||
assert.ok(deferredResult, 'a deferred-typed result must be present');
|
||||
assert.strictEqual(deferredResult.phase, '01');
|
||||
assert.strictEqual(deferredResult.file, 'deferred-items.md');
|
||||
assert.strictEqual(
|
||||
deferredResult.file_path,
|
||||
'.planning/phases/01-foundation/deferred-items.md',
|
||||
);
|
||||
assert.strictEqual(deferredResult.items.length, 1);
|
||||
assert.match(deferredResult.items[0].name, /unrelated pre-existing test failure/);
|
||||
assert.strictEqual(deferredResult.items[0].result, 'unresolved');
|
||||
assert.strictEqual(deferredResult.items[0].category, 'deferred');
|
||||
});
|
||||
|
||||
test('deferred-items.md with 2+ entries (mixed resolved/unresolved) → only unresolved surfaced', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- First unrelated finding, still open.',
|
||||
'- Second unrelated finding, also still open.',
|
||||
'- Third finding, already fixed separately.',
|
||||
' status: resolved',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
const deferredResult = output.results.find(r => r.type === 'deferred');
|
||||
assert.ok(deferredResult);
|
||||
assert.strictEqual(deferredResult.items.length, 2,
|
||||
'exactly the 2 unresolved entries must surface; the resolved 3rd must not');
|
||||
const names = deferredResult.items.map(i => i.name);
|
||||
assert.ok(names.some(n => n.includes('First unrelated finding')));
|
||||
assert.ok(names.some(n => n.includes('Second unrelated finding')));
|
||||
assert.ok(!names.some(n => n.includes('Third finding')));
|
||||
});
|
||||
|
||||
test('deferred entries surface across multiple phase directories', () => {
|
||||
const phase1 = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
const phase2 = path.join(tmpDir, '.planning', 'phases', '02-auth');
|
||||
fs.mkdirSync(phase1, { recursive: true });
|
||||
fs.mkdirSync(phase2, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phase1, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- Phase 1 unrelated finding.',
|
||||
].join('\n'));
|
||||
fs.writeFileSync(path.join(phase2, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- Phase 2 unrelated finding.',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
const deferredResults = output.results.filter(r => r.type === 'deferred');
|
||||
assert.strictEqual(deferredResults.length, 2);
|
||||
assert.strictEqual(output.summary.total_items, 2);
|
||||
assert.strictEqual(output.summary.by_phase['01'], 1);
|
||||
assert.strictEqual(output.summary.by_phase['02'], 1);
|
||||
});
|
||||
|
||||
test('an entry with a garbled/missing status fails safe and is surfaced (not silently dropped)', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- An entry with no status field at all.',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.strictEqual(output.summary.total_items, 1,
|
||||
'missing status must SURFACE the entry, not silently drop it');
|
||||
});
|
||||
|
||||
test('existing UAT/VERIFICATION scanning is unchanged when a deferred-items.md is also present', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, '01-UAT.md'), [
|
||||
'---',
|
||||
'status: testing',
|
||||
'phase: 01-foundation',
|
||||
'started: 2025-01-01T00:00:00Z',
|
||||
'updated: 2025-01-01T00:00:00Z',
|
||||
'---',
|
||||
'',
|
||||
'## Tests',
|
||||
'',
|
||||
'### 1. Login Form',
|
||||
'expected: Form displays with email and password fields',
|
||||
'result: pending',
|
||||
].join('\n'));
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- An unrelated out-of-scope finding.',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.strictEqual(output.results.length, 2, 'both the UAT file and deferred-items.md must surface as separate results');
|
||||
const uatResult = output.results.find(r => r.type === 'uat');
|
||||
const deferredResult = output.results.find(r => r.type === 'deferred');
|
||||
assert.ok(uatResult, 'existing uat-type result must still be present');
|
||||
assert.strictEqual(uatResult.items.length, 1);
|
||||
assert.strictEqual(uatResult.items[0].result, 'pending');
|
||||
assert.ok(deferredResult, 'new deferred-type result must be present');
|
||||
assert.strictEqual(deferredResult.items.length, 1);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── forensic_audit workflow-prose source-contract guard ──────────────────
|
||||
|
||||
// #2994 fragmentization moved the --forensic-gated forensic_audit step out of
|
||||
// progress.md into gsd-core/workflows/progress/steps/forensic-audit.md behind
|
||||
// a section marker. Read that step file directly — it is the sole remaining
|
||||
// source of the forensic_audit step body these guards assert on.
|
||||
const PROGRESS_MD = path.join(__dirname, '..', 'gsd-core', 'workflows', 'progress', 'steps', 'forensic-audit.md');
|
||||
|
||||
describe('#2287 progress.md forensic_audit: deferred-items.md contract', () => {
|
||||
const content = fs.readFileSync(PROGRESS_MD, 'utf-8');
|
||||
const stepStart = content.indexOf('<step name="forensic_audit">');
|
||||
const stepEnd = content.indexOf('</step>', stepStart);
|
||||
const section = stepStart !== -1 && stepEnd !== -1 ? content.slice(stepStart, stepEnd) : '';
|
||||
|
||||
test('forensic_audit step exists', () => {
|
||||
assert.notEqual(stepStart, -1, 'progress.md (or its extracted progress/steps/forensic-audit.md) must contain the forensic_audit step');
|
||||
});
|
||||
|
||||
test('forensic_audit now runs 7 checks (was 6) and globs deferred-items.md', () => {
|
||||
assert.ok(/running 7 deep checks/i.test(section),
|
||||
'forensic_audit must advertise 7 deep checks (was 6) now that deferred-items.md is read');
|
||||
assert.ok(/\.planning\/phases\/\*\/deferred-items\.md/.test(section),
|
||||
'forensic_audit must glob .planning/phases/*/deferred-items.md');
|
||||
});
|
||||
|
||||
test('the new check reports unresolved deferred items with the same ✓/⚠ semantics as the other checks', () => {
|
||||
assert.ok(/check\s*7/i.test(section),
|
||||
'a 7th check must be present');
|
||||
assert.ok(/unresolved deferred items/i.test(section),
|
||||
'the check must be framed around unresolved deferred items');
|
||||
assert.ok(/✓[^\n]*no unresolved deferred items/i.test(section),
|
||||
'the check must emit a ✓ pass line when no unresolved deferred items exist');
|
||||
assert.ok(/⚠[^\n]*unresolved deferred items found/i.test(section),
|
||||
'the check must emit a ⚠ warning line when unresolved deferred items exist');
|
||||
});
|
||||
|
||||
test('an entry is resolved only via an explicit status: resolved field (fail-safe otherwise)', () => {
|
||||
assert.ok(/status:\s*resolved/i.test(section),
|
||||
'the resolved/unresolved parsing rule must be documented in the step prose');
|
||||
});
|
||||
|
||||
test('the verdict summary now gates on 7 checks (was 6)', () => {
|
||||
assert.ok(/after all 7 checks/i.test(section),
|
||||
'the verdict section must say "after all 7 checks"');
|
||||
assert.ok(/if all 7 checks passed/i.test(section),
|
||||
'the verdict section must say "if all 7 checks passed"');
|
||||
assert.ok(!/after all 6 checks/i.test(section) && !/if all 6 checks passed/i.test(section),
|
||||
'stale "6 checks" phrasing must not remain in the step');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── parseDeferredItems property test ──────────────────────────────────────
|
||||
|
||||
describe('#2287 parseDeferredItems: property (status: resolved fail-safe)', () => {
|
||||
// Single-line entry text: no newlines (would break bullet-entry splitting),
|
||||
// non-empty after trim, and never itself SHAPED like a `status:` field line
|
||||
// (that would be indistinguishable from a real field regardless of intent).
|
||||
const plainText = fc.string({ minLength: 1, maxLength: 40 })
|
||||
.map((s) => s.replace(/[\r\n]/g, ' ').trim())
|
||||
.filter((s) => s.length > 0 && !/^status:/i.test(s));
|
||||
|
||||
// Decoy: entry text that CONTAINS a `status: resolved`-shaped substring
|
||||
// mid-line (not at line start) — must never be misread as a resolved
|
||||
// marker, since extractGapEntryFields only recognises a field anchored to
|
||||
// the START of its own trimmed line (see parseDeferredItems' doc comment).
|
||||
const decoyText = plainText.map((s) => `${s} status: resolved trailing note`);
|
||||
|
||||
const textArb = fc.oneof(plainText, decoyText);
|
||||
const entryArb = fc.record({ text: textArb, resolved: fc.boolean() });
|
||||
|
||||
test('property: an entry is surfaced iff it is NOT marked status: resolved; surfaced count == non-resolved count', () => {
|
||||
fc.assert(
|
||||
fc.property(
|
||||
fc.array(entryArb, { maxLength: 20 }),
|
||||
(rawEntries) => {
|
||||
// Index-prefix for uniqueness so surfaced items can be mapped back
|
||||
// to their source entry unambiguously even with colliding random text.
|
||||
const entries = rawEntries.map((e, i) => ({ text: `E${i}_${e.text}`, resolved: e.resolved }));
|
||||
|
||||
const lines = ['## Deferred Items', ''];
|
||||
for (const e of entries) {
|
||||
lines.push(`- ${e.text}`);
|
||||
if (e.resolved) lines.push(' status: resolved');
|
||||
}
|
||||
const content = lines.join('\n');
|
||||
|
||||
const items = parseDeferredItems(content);
|
||||
const surfacedNames = new Set(items.map((it) => it.name));
|
||||
|
||||
const expectedUnresolved = entries.filter((e) => !e.resolved);
|
||||
const expectedResolved = entries.filter((e) => e.resolved);
|
||||
|
||||
// Total surfaced count equals the count of non-resolved entries.
|
||||
assert.strictEqual(items.length, expectedUnresolved.length);
|
||||
|
||||
// Every non-resolved entry IS surfaced (including status:-shaped
|
||||
// decoy substrings embedded mid-line — those must not flip the
|
||||
// outcome).
|
||||
for (const e of expectedUnresolved) {
|
||||
assert.ok(surfacedNames.has(e.text), `expected unresolved entry to surface: ${e.text}`);
|
||||
}
|
||||
|
||||
// No status:-resolved entry is EVER surfaced.
|
||||
for (const e of expectedResolved) {
|
||||
assert.ok(!surfacedNames.has(e.text), `status: resolved entry must never surface: ${e.text}`);
|
||||
}
|
||||
|
||||
// Every returned item carries the fixed deferred category/result shape.
|
||||
for (const item of items) {
|
||||
assert.strictEqual(item.result, 'unresolved');
|
||||
assert.strictEqual(item.category, 'deferred');
|
||||
}
|
||||
}
|
||||
)
|
||||
);
|
||||
});
|
||||
});
|
||||
@@ -1,204 +0,0 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* #2289 — gsd-context-monitor lifecycle-event output allowlist.
|
||||
*
|
||||
* The context monitor emits a `hookSpecificOutput.additionalContext` envelope
|
||||
* to inject context warnings. That shape is only valid for the context-injection
|
||||
* events (PostToolUse, and AfterTool for the Gemini dialect). Codex also wires
|
||||
* this hook to Stop / SubagentStart / SubagentStop / PreCompact (#772), and
|
||||
* Codex's Stop schema REJECTS the envelope ("hook returned invalid stop hook
|
||||
* JSON output"). The fix uses a positive allowlist: emit only for
|
||||
* injection-capable events; every other event — and a missing/unknown name —
|
||||
* exits 0 with NO stdout, while side effects (debounce, critical-session
|
||||
* recording) still run.
|
||||
*
|
||||
* These tests drive the real hook script end-to-end (spawn + stdin + a fresh
|
||||
* metrics bridge file), asserting behavior, not source text.
|
||||
*/
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const os = require('node:os');
|
||||
const path = require('node:path');
|
||||
const { execFileSync } = require('node:child_process');
|
||||
|
||||
const HOOK_PATH = path.join(__dirname, '..', 'hooks', 'gsd-context-monitor.js');
|
||||
|
||||
// Run the monitor with a synthetic, fresh metrics bridge file.
|
||||
// Returns { stdout, warnData } and cleans up the bridge + sentinel files.
|
||||
// opts: { event, remaining, used = 80, gemini = false, gsdActive = false }
|
||||
function runMonitor(opts) {
|
||||
const {
|
||||
event,
|
||||
remaining,
|
||||
used = 80,
|
||||
gemini = false,
|
||||
gsdActive = false,
|
||||
} = opts;
|
||||
|
||||
const sessionId = `fix-2289-${Date.now()}-${Math.random().toString(36).slice(2)}`;
|
||||
const tmpDir = os.tmpdir();
|
||||
const metricsPath = path.join(tmpDir, `claude-ctx-${sessionId}.json`);
|
||||
const warnPath = path.join(tmpDir, `claude-ctx-${sessionId}-warned.json`);
|
||||
|
||||
// Fresh (non-stale) metrics: timestamp is "now" in seconds.
|
||||
fs.writeFileSync(metricsPath, JSON.stringify({
|
||||
timestamp: Math.floor(Date.now() / 1000),
|
||||
remaining_percentage: remaining,
|
||||
used_pct: used,
|
||||
}));
|
||||
|
||||
// Optional GSD-active project dir (STATE.md present) so the critical-session
|
||||
// recording side effect is reachable.
|
||||
let cwd = tmpDir;
|
||||
let projDir = null;
|
||||
if (gsdActive) {
|
||||
projDir = fs.mkdtempSync(path.join(tmpDir, 'fix-2289-proj-'));
|
||||
fs.mkdirSync(path.join(projDir, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(path.join(projDir, '.planning', 'STATE.md'), '# State\n');
|
||||
cwd = projDir;
|
||||
}
|
||||
|
||||
const payload = { session_id: sessionId, cwd };
|
||||
if (event !== undefined) payload.hook_event_name = event;
|
||||
|
||||
const env = { ...process.env };
|
||||
if (gemini) env.GEMINI_API_KEY = 'test-key';
|
||||
else delete env.GEMINI_API_KEY;
|
||||
|
||||
let stdout = '';
|
||||
try {
|
||||
stdout = execFileSync(process.execPath, [HOOK_PATH], {
|
||||
input: JSON.stringify(payload),
|
||||
env,
|
||||
encoding: 'utf8',
|
||||
timeout: 8000,
|
||||
});
|
||||
} catch (e) {
|
||||
stdout = e.stdout || '';
|
||||
}
|
||||
|
||||
let warnData = null;
|
||||
try {
|
||||
warnData = JSON.parse(fs.readFileSync(warnPath, 'utf8'));
|
||||
} catch { /* sentinel may not exist */ }
|
||||
|
||||
// Cleanup
|
||||
for (const p of [metricsPath, warnPath]) {
|
||||
try { fs.unlinkSync(p); } catch { /* ignore */ }
|
||||
}
|
||||
if (projDir) {
|
||||
// Retry-tolerant teardown: the critical path fires a detached, unref()'d
|
||||
// `state record-session` grandchild against projDir, and execFileSync does
|
||||
// not wait for it. maxRetries/retryDelay absorbs the transient
|
||||
// EBUSY/ENOTEMPTY window while that process exits, so cleanup can neither
|
||||
// flake nor leak the temp dir (mirrors tests/helpers.cjs cleanup(); see the
|
||||
// #2289 review and the prior fix in perf-317-context-monitor-fs.test.cjs).
|
||||
// eslint-disable-next-line local/no-raw-rmsync-in-tests -- test fixture teardown of a unique mkdtemp dir
|
||||
try { fs.rmSync(projDir, { recursive: true, force: true, maxRetries: 20, retryDelay: 100 }); } catch { /* ignore */ }
|
||||
}
|
||||
|
||||
return { stdout, warnData };
|
||||
}
|
||||
|
||||
describe('#2289 context-monitor: non-injection events exit silently', () => {
|
||||
// Boundary coverage around WARNING (35) and CRITICAL (25) — Stop must stay
|
||||
// silent at limit-1 / limit / limit+1 for BOTH thresholds.
|
||||
for (const remaining of [40, 36, 35, 34, 26, 25, 24, 20]) {
|
||||
test(`Stop event at remaining=${remaining}% → exit 0, empty stdout`, () => {
|
||||
const { stdout } = runMonitor({ event: 'Stop', remaining });
|
||||
assert.strictEqual(stdout, '', `Stop must emit nothing at remaining=${remaining}% (Codex rejects the envelope)`);
|
||||
});
|
||||
}
|
||||
|
||||
test('missing hook_event_name (no Gemini) at 30% → empty stdout', () => {
|
||||
const { stdout } = runMonitor({ event: undefined, remaining: 30 });
|
||||
assert.strictEqual(stdout, '', 'a missing event name must not fall through to the injection envelope');
|
||||
});
|
||||
|
||||
test('empty-string hook_event_name (no Gemini) at 30% → empty stdout', () => {
|
||||
const { stdout } = runMonitor({ event: ' ', remaining: 30 });
|
||||
assert.strictEqual(stdout, '', 'a blank event name must be treated as missing → silent');
|
||||
});
|
||||
|
||||
for (const event of ['SubagentStart', 'SubagentStop', 'PreCompact', 'SessionStart', 'BeforeTool']) {
|
||||
test(`unknown/non-injection event "${event}" at 30% → empty stdout`, () => {
|
||||
const { stdout } = runMonitor({ event, remaining: 30 });
|
||||
assert.strictEqual(stdout, '', `${event} is not injection-capable and must emit nothing`);
|
||||
});
|
||||
}
|
||||
});
|
||||
|
||||
describe('#2289 context-monitor: injection events still warn (unchanged)', () => {
|
||||
test('PostToolUse at 30% → WARNING envelope with hookEventName PostToolUse', () => {
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 30, used: 70 });
|
||||
assert.notStrictEqual(stdout, '', 'PostToolUse must still emit a warning envelope');
|
||||
const parsed = JSON.parse(stdout);
|
||||
assert.strictEqual(parsed.hookSpecificOutput.hookEventName, 'PostToolUse');
|
||||
assert.match(parsed.hookSpecificOutput.additionalContext, /CONTEXT WARNING/);
|
||||
});
|
||||
|
||||
test('PostToolUse at 20% → CRITICAL envelope', () => {
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 20, used: 80 });
|
||||
const parsed = JSON.parse(stdout);
|
||||
assert.strictEqual(parsed.hookSpecificOutput.hookEventName, 'PostToolUse');
|
||||
assert.match(parsed.hookSpecificOutput.additionalContext, /CONTEXT CRITICAL/);
|
||||
});
|
||||
|
||||
test('AfterTool at 30% → WARNING envelope with hookEventName AfterTool', () => {
|
||||
const { stdout } = runMonitor({ event: 'AfterTool', remaining: 30 });
|
||||
const parsed = JSON.parse(stdout);
|
||||
assert.strictEqual(parsed.hookSpecificOutput.hookEventName, 'AfterTool');
|
||||
assert.match(parsed.hookSpecificOutput.additionalContext, /CONTEXT WARNING/);
|
||||
});
|
||||
|
||||
test('missing event name WITH Gemini env at 30% → AfterTool envelope (fallback preserved)', () => {
|
||||
const { stdout } = runMonitor({ event: undefined, remaining: 30, gemini: true });
|
||||
assert.notStrictEqual(stdout, '', 'Gemini AfterTool fallback must still emit when the event name is absent');
|
||||
const parsed = JSON.parse(stdout);
|
||||
assert.strictEqual(parsed.hookSpecificOutput.hookEventName, 'AfterTool');
|
||||
});
|
||||
|
||||
test('explicit PostToolUse WITH Gemini env → explicit name wins over the AfterTool fallback', () => {
|
||||
// Precedence guard: the Gemini fallback only applies to a MISSING name; an
|
||||
// explicit PostToolUse must still report as PostToolUse even under GEMINI_API_KEY.
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 30, gemini: true });
|
||||
const parsed = JSON.parse(stdout);
|
||||
assert.strictEqual(parsed.hookSpecificOutput.hookEventName, 'PostToolUse');
|
||||
assert.match(parsed.hookSpecificOutput.additionalContext, /CONTEXT WARNING/);
|
||||
});
|
||||
|
||||
// Threshold boundaries on the emit path: 36 = no warn, 35 = warn, 25 = critical, 26 = warn.
|
||||
test('PostToolUse at 36% (above WARNING) → empty stdout', () => {
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 36 });
|
||||
assert.strictEqual(stdout, '', 'no warning above the 35% threshold');
|
||||
});
|
||||
|
||||
test('PostToolUse at 35% (WARNING boundary) → WARNING envelope', () => {
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 35 });
|
||||
assert.match(JSON.parse(stdout).hookSpecificOutput.additionalContext, /CONTEXT WARNING/);
|
||||
});
|
||||
|
||||
test('PostToolUse at 25% (CRITICAL boundary) → CRITICAL envelope', () => {
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 25 });
|
||||
assert.match(JSON.parse(stdout).hookSpecificOutput.additionalContext, /CONTEXT CRITICAL/);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2289 context-monitor: side effects still fire on silent events (no output ≠ no side effect)', () => {
|
||||
test('Stop at 30% still writes the debounce sentinel (bookkeeping runs)', () => {
|
||||
const { stdout, warnData } = runMonitor({ event: 'Stop', remaining: 30 });
|
||||
assert.strictEqual(stdout, '', 'Stop emits nothing');
|
||||
assert.ok(warnData, 'the debounce sentinel must still be written on a silenced Stop event');
|
||||
assert.strictEqual(warnData.lastLevel, 'warning', 'debounce level bookkeeping runs regardless of output');
|
||||
});
|
||||
|
||||
test('Stop at 20% in a GSD project still records the critical-session sentinel', () => {
|
||||
const { stdout, warnData } = runMonitor({ event: 'Stop', remaining: 20, used: 80, gsdActive: true });
|
||||
assert.strictEqual(stdout, '', 'Stop emits nothing even at critical context');
|
||||
assert.ok(warnData, 'sentinel must be written');
|
||||
assert.strictEqual(warnData.criticalRecorded, true, 'critical-session recording side effect fires on the silent Stop event');
|
||||
});
|
||||
});
|
||||
@@ -1,150 +0,0 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* #2358 — review.md (and ship.md's external peer-review step) wrote every
|
||||
* temp file to a hardcoded, phase-number-only path under /tmp
|
||||
* (`/tmp/gsd-review-prompt-{phase}.md`, `/tmp/gsd-review-<reviewer>-{phase}.*`,
|
||||
* `/tmp/gsd-review-stderr.log`). Two GSD projects with a phase sharing the
|
||||
* same small integer number collide on the exact same path; a crashed prior
|
||||
* run's leftover file is bait a later, unrelated run can silently read (the
|
||||
* reporter forensically confirmed agy read a 3-week-old stale prompt from a
|
||||
* DIFFERENT project). Neither file used the portable ${TMPDIR:-/tmp} seam,
|
||||
* and review.md had no cleanup.
|
||||
*
|
||||
* The fix threads a single `mktemp -d "${TMPDIR:-/tmp}/gsd-review-XXXXXX"`
|
||||
* run directory (RUN_DIR / {run_dir}) through every review.md temp path, and
|
||||
* ship.md's stderr capture through a per-run `mktemp` file — eliminating the
|
||||
* shared-path collision by construction rather than by convention.
|
||||
*
|
||||
* review.md and ship.md ARE the product the runtime loads (an AI agent reads
|
||||
* and executes these workflow instructions verbatim), so this is a
|
||||
* static-content regression against the deployed text, mirroring
|
||||
* fix-2194-review-timeout-guidance.test.cjs.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const os = require('node:os');
|
||||
const path = require('node:path');
|
||||
|
||||
const { cleanup } = require('./helpers.cjs');
|
||||
|
||||
const REVIEW_MD = path.join(__dirname, '..', 'gsd-core', 'workflows', 'review.md');
|
||||
|
||||
describe('#2358 review.md temp paths are run-scoped, not phase-only', () => {
|
||||
const content = fs.readFileSync(REVIEW_MD, 'utf-8');
|
||||
|
||||
test('no bare, unscoped /tmp/gsd-review-* path remains', () => {
|
||||
assert.ok(
|
||||
!content.includes('/tmp/gsd-review'),
|
||||
'review.md must not contain any hardcoded /tmp/gsd-review* literal — ' +
|
||||
'every review temp path must be rooted under the run-scoped mktemp directory'
|
||||
);
|
||||
});
|
||||
|
||||
test('creates exactly one run-scoped directory via the portable ${TMPDIR:-/tmp} seam', () => {
|
||||
const mktempAssignments = content.match(/RUN_DIR=\$\(mktemp -d "\$\{TMPDIR:-\/tmp\}\/gsd-review-XXXXXX"\)/g) || [];
|
||||
assert.equal(
|
||||
mktempAssignments.length, 1,
|
||||
'review.md must create the run directory with exactly one `mktemp -d "${TMPDIR:-/tmp}/gsd-review-XXXXXX"` — ' +
|
||||
'a hardcoded /tmp (no ${TMPDIR:-/tmp} seam) breaks on Windows, and re-mktemp-ing per block would break the ' +
|
||||
'write/read pairing between build_prompt and the local-reviewer budget-trimming reads'
|
||||
);
|
||||
});
|
||||
|
||||
test('every downstream temp path is threaded through {run_dir} / $RUN_DIR, not re-derived from {phase}', () => {
|
||||
assert.ok(
|
||||
/\{run_dir\}\/gsd-review-/.test(content),
|
||||
'reviewer blocks must reference {run_dir}/gsd-review-... (the run-scoped placeholder)'
|
||||
);
|
||||
assert.ok(
|
||||
/\$\{RUN_DIR\}\/gsd-review-/.test(content),
|
||||
'the build_prompt section-file writes must reference ${RUN_DIR}/gsd-review-... (the run-scoped shell var)'
|
||||
);
|
||||
// The old isolation key must be gone entirely from path construction.
|
||||
assert.ok(
|
||||
!/\/tmp\/gsd-review[^\r\n]*\{phase\}/.test(content),
|
||||
'no temp path may still be keyed on a bare {phase} placeholder'
|
||||
);
|
||||
assert.ok(
|
||||
!/\$\{PHASE\}-(?:instructions|roadmap|plan|project|context|research|requirements)\.md/.test(content),
|
||||
'no temp path may still be keyed on the ${PHASE} shell var'
|
||||
);
|
||||
});
|
||||
|
||||
// Phase 5b (#2799) moved these strings out of review.md's bash and into the resolver and the
|
||||
// antigravity handler, so the assertions follow them. The invariant is unchanged and is what
|
||||
// #2358 was about: every reviewer artifact must live under the run-scoped mktemp directory, never
|
||||
// a bare `{phase}`-keyed /tmp path that a concurrent review could collide with.
|
||||
test('every lane anchors its prompt and artifacts under the run dir', () => {
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const RUN = '/run-scoped';
|
||||
for (const lane of REVIEWER_LANES) {
|
||||
const r = resolveLanePlan({
|
||||
lane, configGet: () => undefined, runDir: RUN, repoRoot: '/repo',
|
||||
});
|
||||
assert.equal(r.ok, true, `${lane.slug} failed to resolve`);
|
||||
const p = r.plan;
|
||||
assert.ok(p.reviewPath.startsWith(`${RUN}/`), `${lane.slug} review path escapes the run dir`);
|
||||
assert.ok(p.errPath.startsWith(`${RUN}/`), `${lane.slug} err path escapes the run dir`);
|
||||
assert.ok(p.promptPath.startsWith(`${RUN}/`), `${lane.slug} prompt path escapes the run dir`);
|
||||
}
|
||||
});
|
||||
|
||||
test('the argv-borne prompt instruction references the run-scoped path', () => {
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const RUN = '/run-scoped';
|
||||
const fileRefLanes = REVIEWER_LANES.filter(
|
||||
(l) => l.transport === 'spawn' && l.invoke.promptChannel === 'argv-file-ref',
|
||||
);
|
||||
assert.ok(fileRefLanes.length > 0, 'expected at least one argv-file-ref lane');
|
||||
for (const lane of fileRefLanes) {
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: RUN, repoRoot: '/repo' });
|
||||
const arg = r.plan.argv[r.plan.argv.length - 1];
|
||||
assert.ok(arg.includes(`${RUN}/gsd-review-prompt.md`), `${lane.slug} prompt not run-scoped`);
|
||||
}
|
||||
});
|
||||
|
||||
test('an instance writes under the run dir, keyed by its own identity', () => {
|
||||
// Two instances of one adapter must not overwrite each other, and neither may escape the run
|
||||
// dir — the identity is sanitized to a flat filename.
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === 'opencode');
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: '/run-scoped', repoRoot: '/repo' });
|
||||
assert.ok(r.plan.reviewPath.startsWith('/run-scoped/'));
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2358 design principle: run-scoped temp dirs never collide across projects/phases', () => {
|
||||
// review.md and ship.md are markdown instructions an AI agent executes, not
|
||||
// node-executable code, so this does not shell out to the literal snippet —
|
||||
// it validates the underlying guarantee the fix relies on (mktemp-style
|
||||
// randomized-suffix isolation) using Node's built-in equivalent, which is
|
||||
// cross-platform (Windows included) unlike shelling out to `mktemp`/bash.
|
||||
test('two runs — even for the same phase number, same or different project — get distinct run dirs', () => {
|
||||
const prefix = path.join(os.tmpdir(), 'gsd-review-');
|
||||
const runDirA = fs.mkdtempSync(prefix);
|
||||
const runDirB = fs.mkdtempSync(prefix);
|
||||
try {
|
||||
assert.notEqual(
|
||||
runDirA, runDirB,
|
||||
'two review runs sharing the same phase number must never resolve to the same run-scoped directory'
|
||||
);
|
||||
const phase = '10'; // same phase number in both "projects" — the historical collision case
|
||||
const staleProjectAPath = path.join(runDirA, `gsd-review-prompt.md`);
|
||||
const laterProjectBPath = path.join(runDirB, `gsd-review-prompt.md`);
|
||||
assert.notEqual(
|
||||
staleProjectAPath, laterProjectBPath,
|
||||
`phase ${phase} in two different runs must not resolve to the same prompt path`
|
||||
);
|
||||
} finally {
|
||||
// helpers.cleanup (not raw fs.rmSync) carries the Windows-EBUSY retry budget.
|
||||
cleanup(runDirA);
|
||||
cleanup(runDirB);
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -1,121 +0,0 @@
|
||||
/**
|
||||
* #2494 — a failed reviewer lane must be diagnosable, never a silent drop.
|
||||
*
|
||||
* Before the fix, gemini and claude sent stderr to `/dev/null` and wrote nothing on failure. A
|
||||
* failed lane — CLI missing, unauthenticated, rate-limited, crashed, any exit that writes no
|
||||
* stdout — left a zero-byte file that `write_reviews` rendered as "a reviewer that ran cleanly with
|
||||
* nothing to report", silently dropping a lane from the cross-AI consensus while `present_results`
|
||||
* reported success.
|
||||
*
|
||||
* The invariant is unchanged; what moved is where it lives. This suite used to extract the two
|
||||
* shell blocks verbatim from `review.md` and run them under a real bash against a failing stub.
|
||||
* Phase 5b (#2799) deleted those blocks, so the tests now drive the real runner with stubbed
|
||||
* dependencies — the same behavioural altitude (a lane is actually run and its artifacts
|
||||
* inspected), against the surface that ships today. Nothing here reads source text any more, so
|
||||
* the source-text exemption this file used to carry is gone.
|
||||
*
|
||||
* The guarantee also got STRONGER in one way worth locking: the policy is uniform across every
|
||||
* lane now rather than fixed per-leg, so these assertions run over the whole spawn roster instead
|
||||
* of the two legs the issue named.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const { runLane } = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
|
||||
const RUN = '/run';
|
||||
const ROOT = '/repo';
|
||||
|
||||
/** Lanes whose empty-output policy is the shared stub (antigravity owns its own diagnostics). */
|
||||
const STUB_LANES = REVIEWER_LANES.filter(
|
||||
(l) => l.transport === 'spawn' && l.emptyOutput === 'stub-with-stderr',
|
||||
);
|
||||
|
||||
function planFor(slug) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true, `${slug} failed to resolve`);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
function deps(spawnResult, files = {}) {
|
||||
return {
|
||||
files,
|
||||
// `kimi-code` declares a `command-capability` probe, so the runner spawns `--help` BEFORE the
|
||||
// review. Answer that separately or the probe fails and the lane never reaches the invocation
|
||||
// this test is about.
|
||||
spawn: (binary, argv) =>
|
||||
argv && argv.length === 1 && argv[0] === '--help'
|
||||
? { status: 0, stdout: '--output-format', stderr: '' }
|
||||
: spawnResult,
|
||||
httpJson: async () => ({ ok: true, status: 200, body: '{}' }),
|
||||
readFile: (p) => { if (!(p in files)) throw new Error(`ENOENT ${p}`); return files[p]; },
|
||||
writeFile: (p, c) => { files[p] = c; },
|
||||
exists: (p) => p in files,
|
||||
hasBinary: () => true,
|
||||
configGet: () => undefined,
|
||||
homeDir: '/home/u',
|
||||
warn: () => {},
|
||||
};
|
||||
}
|
||||
|
||||
describe('#2494 — a failed lane writes a diagnosable stub, not a zero-byte file', () => {
|
||||
for (const lane of STUB_LANES) {
|
||||
test(`${lane.slug}: a lane that exits non-zero with no stdout is stubbed`, async () => {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps({ status: 127, stdout: '', stderr: 'command not found' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
|
||||
assert.equal(r.stubbed, true, 'a failed lane must be reported as stubbed');
|
||||
const review = d.files[p.reviewPath];
|
||||
assert.ok(review !== undefined, 'a review file must exist after a failed lane');
|
||||
assert.notStrictEqual(review.trim(), '', 'the review file must not be empty');
|
||||
assert.ok(
|
||||
review.includes('failed or returned empty output'),
|
||||
'the stub must be distinguishable from a real review',
|
||||
);
|
||||
});
|
||||
|
||||
test(`${lane.slug}: stderr is captured to a .err sidecar, never discarded`, async () => {
|
||||
// The sidecar is the difference between "this lane failed" and "this lane failed BECAUSE…".
|
||||
// Without it every failure mode looks identical to every other.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps({ status: 1, stdout: '', stderr: 'HTTP 429 rate limited' });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
|
||||
assert.equal(d.files[p.errPath], 'HTTP 429 rate limited', 'stderr must reach the sidecar');
|
||||
assert.ok(
|
||||
d.files[p.reviewPath].includes('HTTP 429 rate limited'),
|
||||
'and must be surfaced in the stub, where a reader will actually see it',
|
||||
);
|
||||
assert.ok(p.reviewPath.endsWith('.md'), 'review output path unchanged');
|
||||
});
|
||||
}
|
||||
|
||||
test('a successful review passes through untouched', async () => {
|
||||
const p = planFor('gemini');
|
||||
const d = deps({ status: 0, stdout: 'Looks good.\n', stderr: '' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('Looks good.'));
|
||||
assert.ok(
|
||||
!d.files[p.reviewPath].includes('failed or returned empty output'),
|
||||
'a real review must never carry the failure header',
|
||||
);
|
||||
});
|
||||
|
||||
test('no lane sends stderr to /dev/null — the sidecar is unconditional', async () => {
|
||||
// The original defect in one line, asserted over the whole roster rather than the two legs the
|
||||
// issue named: the policy is uniform now, and a future lane must not be able to opt out.
|
||||
for (const lane of STUB_LANES) {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps({ status: 0, stdout: 'ok', stderr: 'a warning' });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(d.files[p.errPath], 'a warning', `${lane.slug} discarded stderr`);
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -1,280 +0,0 @@
|
||||
/**
|
||||
* #2590 — every emitted Workflow script was rejected by the Workflow tool.
|
||||
*
|
||||
* `emitWorkflowScript` generated four constructs the tool does not accept. The
|
||||
* first was fatal on its own, so the backend could never dispatch a wave:
|
||||
*
|
||||
* 1. no `export const meta = {…}` first statement -> whole script rejected
|
||||
* 2. `resumeFromRunId("<id>")` -> "resumeFromRunId is not defined"
|
||||
* (it is a Workflow TOOL INPUT parameter, not a script function)
|
||||
* 3. `budget(<n>)` -> "budget is not a function"
|
||||
* (`budget` is a read-only object { total, spent(), remaining() })
|
||||
* 4. `parallel(agent(…), agent(…))` -> "parallel() expects an array of functions"
|
||||
*
|
||||
* Plus two secondary defects that kept the emitted script from ever being
|
||||
* REACHED, which is why this shipped undetected:
|
||||
*
|
||||
* 5. nothing resolved the Agent SDK version, so gate 5 returned
|
||||
* `agent_sdk_version_unknown` on every automated run
|
||||
* 6. the runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging
|
||||
* from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any
|
||||
* invocation without --runtime reported `runtime_not_claude`
|
||||
*
|
||||
* The script assertions parse the emitted text as a real ES module rather than
|
||||
* pattern-matching it, so a syntactically invalid script fails outright.
|
||||
*/
|
||||
|
||||
'use strict';
|
||||
|
||||
process.env.GSD_TEST_MODE = '1';
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const { createTempDir, cleanup } = require('./helpers.cjs');
|
||||
const { runNode } = require('./helpers/process-seam.cjs');
|
||||
const { throwIfFailed } = require('./helpers/git-fixture.cjs');
|
||||
const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs');
|
||||
|
||||
const core = require('../gsd-core/bin/lib/claude-orchestration.cjs');
|
||||
const TOOLS = path.join(__dirname, '..', 'gsd-core', 'bin', 'gsd-tools.cjs');
|
||||
|
||||
function emit(overrides) {
|
||||
const input = Object.assign({
|
||||
phaseDir: '.planning/phases/01',
|
||||
runId: 'execute-1',
|
||||
waves: [{ id: 'wave-1', plans: [{ id: '01-01', brief: 'noop', files_modified: ['a.ts'] }] }],
|
||||
}, overrides || {});
|
||||
const r = core.emitWorkflowScript(input);
|
||||
assert.ok(r.ok, `emit failed: ${JSON.stringify(r)}`);
|
||||
return r;
|
||||
}
|
||||
|
||||
/** First non-comment, non-blank line — the script's first actual statement. */
|
||||
function firstStatement(script) {
|
||||
return script.split('\n').map((l) => l.trim())
|
||||
.find((l) => l.length > 0 && !l.startsWith('//')) || '';
|
||||
}
|
||||
|
||||
describe('#2590: emitted Workflow scripts satisfy the Workflow tool contract', () => {
|
||||
test('the emitted script is syntactically valid as an ES module', () => {
|
||||
// `export const meta` + top-level `await` only parse in module context —
|
||||
// which is exactly the context the Workflow tool runs the script in.
|
||||
const { script } = emit({
|
||||
waves: [
|
||||
{ id: 'w1', plans: [
|
||||
{ id: 'a', brief: 'one', files_modified: ['a.ts'] },
|
||||
{ id: 'b', brief: 'two', files_modified: ['b.ts'] },
|
||||
] },
|
||||
{ id: 'w2', plans: [{ id: 'c', brief: 'three', files_modified: ['c.ts'] }] },
|
||||
],
|
||||
});
|
||||
// .mjs so node parses it in module context, inside a helper temp dir so
|
||||
// cleanup() carries the Windows-EBUSY retry budget.
|
||||
const dir = createTempDir('gsd-2590-parse-');
|
||||
const f = path.join(dir, 'emitted.mjs');
|
||||
fs.writeFileSync(f, script);
|
||||
try {
|
||||
const result = runNode(['--check', f], { timeoutMs: PROBE_TIMEOUT_MS });
|
||||
throwIfFailed(result, `node --check ${f} (emitted script must parse)`);
|
||||
} finally {
|
||||
cleanup(dir);
|
||||
}
|
||||
});
|
||||
|
||||
test('1. `export const meta` is the first statement', () => {
|
||||
const { script } = emit();
|
||||
assert.match(
|
||||
firstStatement(script),
|
||||
/^export const meta = \{/,
|
||||
'the Workflow tool rejects any script whose first statement is not the meta block',
|
||||
);
|
||||
});
|
||||
|
||||
test('meta.phases titles match the emitted phase() calls exactly', () => {
|
||||
// The tool matches phase titles by exact string; a mismatch silently splits
|
||||
// progress into an unnamed group.
|
||||
const { script } = emit({
|
||||
waves: [
|
||||
{ id: 'alpha', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] },
|
||||
{ id: 'beta', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] },
|
||||
],
|
||||
});
|
||||
const metaTitles = [...script.matchAll(/\{ title: "([^"]+)"/g)].map((m) => m[1]);
|
||||
const phaseTitles = [...script.matchAll(/^phase\("([^"]+)"\)/gm)].map((m) => m[1]);
|
||||
assert.deepEqual(metaTitles, ['Wave alpha', 'Wave beta']);
|
||||
assert.deepEqual(phaseTitles, metaTitles, 'phase() titles must match meta.phases exactly');
|
||||
});
|
||||
|
||||
test('duplicate wave ids are rejected (phase titles must map 1:1)', () => {
|
||||
// Two waves sharing an id emit two identical `phase("Wave x")` calls and two
|
||||
// identical meta.phases entries; the tool matches titles by exact string, so
|
||||
// the second wave's agents would be attributed to the first's progress group.
|
||||
const r = core.emitWorkflowScript({
|
||||
phaseDir: '.planning/phases/01',
|
||||
runId: 'execute-1',
|
||||
waves: [
|
||||
{ id: 'dup', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] },
|
||||
{ id: 'dup', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] },
|
||||
],
|
||||
});
|
||||
assert.equal(r.ok, false, 'duplicate wave ids must be rejected, not silently merged');
|
||||
assert.match(String(r.reason), /duplicate wave id/);
|
||||
});
|
||||
|
||||
test('distinct wave ids are still accepted (the boundary either side)', () => {
|
||||
const r = core.emitWorkflowScript({
|
||||
phaseDir: '.planning/phases/01',
|
||||
runId: 'execute-1',
|
||||
waves: [
|
||||
{ id: 'w1', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] },
|
||||
{ id: 'w2', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] },
|
||||
],
|
||||
});
|
||||
assert.equal(r.ok, true, `distinct wave ids must pass: ${JSON.stringify(r)}`);
|
||||
});
|
||||
|
||||
test('2. resumeFromRunId is never CALLED (it is a tool input, not a function)', () => {
|
||||
const { script, summary } = emit({ runId: 'execute-7' });
|
||||
assert.ok(
|
||||
!/^\s*resumeFromRunId\s*\(/m.test(script),
|
||||
'calling resumeFromRunId() throws "resumeFromRunId is not defined"',
|
||||
);
|
||||
// The run id must still reach the caller, which passes it as the tool input.
|
||||
assert.equal(summary.resumeRunId, 'execute-7');
|
||||
});
|
||||
|
||||
test('3. budget is never CALLED, at and around the boundary', () => {
|
||||
// budgetTokens is floored at > 0; check 0 (rejected), 1 (accepted), and a
|
||||
// large value — none may produce a budget(...) call.
|
||||
for (const tokens of [0, 1, 500000]) {
|
||||
const { script, summary } = emit({ budgetTokens: tokens });
|
||||
assert.ok(
|
||||
!/^\s*budget\s*\(/m.test(script),
|
||||
`budgetTokens=${tokens}: calling budget() throws "budget is not a function"`,
|
||||
);
|
||||
assert.equal(summary.budgetTokens, tokens > 0 ? tokens : null);
|
||||
}
|
||||
});
|
||||
|
||||
test('4. parallel() receives an array of thunks, not agent() results', () => {
|
||||
const { script } = emit({
|
||||
waves: [{ id: 'w', plans: [
|
||||
{ id: 'a', brief: 'one', files_modified: ['a.ts'] },
|
||||
{ id: 'b', brief: 'two', files_modified: ['b.ts'] },
|
||||
] }],
|
||||
});
|
||||
assert.ok(/parallel\(\[/.test(script), 'parallel() expects an array of functions');
|
||||
assert.ok(
|
||||
!/parallel\(\s*agent\(/.test(script),
|
||||
'passing agent() results directly both throws and starts every agent eagerly',
|
||||
);
|
||||
// Each agent must be wrapped in a thunk so parallel() can bound concurrency.
|
||||
const agents = [...script.matchAll(/agent\("/g)].length;
|
||||
const thunks = [...script.matchAll(/\(\) => agent\("/g)].length;
|
||||
assert.equal(thunks, agents, 'every agent() must be wrapped in a () => thunk');
|
||||
});
|
||||
|
||||
test('single-plan stages also emit an array (regression: the 1-plan branch)', () => {
|
||||
// The pre-fix code had a SEPARATE single-plan branch that emitted
|
||||
// `parallel(\n agent(...)\n)` — valid-looking but the same defect.
|
||||
const { script } = emit();
|
||||
assert.ok(/parallel\(\[/.test(script));
|
||||
assert.equal([...script.matchAll(/\(\) => agent\("/g)].length, 1);
|
||||
});
|
||||
|
||||
test('per-plan worktree isolation still mirrors use_worktree', () => {
|
||||
const { script } = emit({
|
||||
waves: [{ id: 'w', plans: [
|
||||
{ id: 'a', brief: 'iso', files_modified: ['a.ts'] },
|
||||
{ id: 'b', brief: 'noiso', files_modified: ['b.ts'], use_worktree: false },
|
||||
] }],
|
||||
});
|
||||
assert.match(script, /agent\("iso", \{ agentType: "gsd-executor", isolation: "worktree" \}\)/);
|
||||
assert.match(script, /agent\("noiso", \{ agentType: "gsd-executor" \}\)/);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2590: the backend is reachable without hand-passed flags', () => {
|
||||
function repro() {
|
||||
const dir = createTempDir('gsd-2590-repro-');
|
||||
fs.mkdirSync(path.join(dir, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(
|
||||
path.join(dir, '.planning', 'config.json'),
|
||||
JSON.stringify({ claude_orchestration: { enabled: true } }),
|
||||
);
|
||||
fs.writeFileSync(
|
||||
path.join(dir, 'waves.json'),
|
||||
JSON.stringify({ waves: [{ id: 'wave-1', plans: [{ id: '01-01', brief: 'noop', files_modified: ['a.ts'] }] }] }),
|
||||
);
|
||||
return dir;
|
||||
}
|
||||
|
||||
function resolve(dir, extraArgs) {
|
||||
const result = runNode([
|
||||
TOOLS, 'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', 'waves.json', '--run-id', 'execute-1',
|
||||
'--phase-dir', '.planning/phases/01', '--raw',
|
||||
...(extraArgs || []),
|
||||
], { cwd: dir, timeoutMs: PROBE_TIMEOUT_MS });
|
||||
throwIfFailed(result, 'gsd-tools claude-orchestration resolve-wave-dispatch');
|
||||
return JSON.parse(result.stdout);
|
||||
}
|
||||
|
||||
test('5+6. no --runtime and no --agent-sdk-version still reaches the version gate', () => {
|
||||
const dir = repro();
|
||||
try {
|
||||
const r = resolve(dir);
|
||||
// Pre-fix this was `agent_sdk_version_unknown` (nothing resolved a
|
||||
// version) or `runtime_not_claude` (the divergent fallback). Either is a
|
||||
// regression; the version gate must now be reached and answer truthfully.
|
||||
assert.notEqual(r.reason, 'agent_sdk_version_unknown',
|
||||
'the router must resolve the installed SDK version itself');
|
||||
assert.notEqual(r.reason, 'runtime_not_claude',
|
||||
'runtime must fall back to the canonical config.runtime > claude chain');
|
||||
} finally {
|
||||
cleanup(dir);
|
||||
}
|
||||
});
|
||||
|
||||
test('an SDK version above the floor activates the workflow backend end to end', () => {
|
||||
const dir = repro();
|
||||
try {
|
||||
const r = resolve(dir, ['--agent-sdk-version', '0.3.149']);
|
||||
assert.equal(r.backend, 'workflow', `expected workflow backend, got ${JSON.stringify(r)}`);
|
||||
assert.ok(typeof r.script === 'string' && r.script.length > 0);
|
||||
assert.match(firstStatement(r.script), /^export const meta = \{/);
|
||||
} finally {
|
||||
cleanup(dir);
|
||||
}
|
||||
});
|
||||
|
||||
test('an explicit --agent-sdk-version still wins over the installed one', () => {
|
||||
const dir = repro();
|
||||
try {
|
||||
// A deliberately ancient pin must be honored (and decline), proving the
|
||||
// flag is not ignored now that a fallback exists.
|
||||
const r = resolve(dir, ['--agent-sdk-version', '0.0.1']);
|
||||
assert.equal(r.backend, 'inline');
|
||||
assert.equal(r.reason, 'agent_sdk_version_below_floor');
|
||||
} finally {
|
||||
cleanup(dir);
|
||||
}
|
||||
});
|
||||
|
||||
test('GSD_AGENT_SDK_VERSION is honored between the flag and the installed version', () => {
|
||||
const dir = repro();
|
||||
try {
|
||||
const result = runNode([
|
||||
TOOLS, 'claude-orchestration', 'resolve-wave-dispatch',
|
||||
'--waves', 'waves.json', '--run-id', 'execute-1',
|
||||
'--phase-dir', '.planning/phases/01', '--raw',
|
||||
], { cwd: dir, env: { ...process.env, GSD_AGENT_SDK_VERSION: '0.3.149' }, timeoutMs: PROBE_TIMEOUT_MS });
|
||||
throwIfFailed(result, 'gsd-tools claude-orchestration resolve-wave-dispatch (GSD_AGENT_SDK_VERSION)');
|
||||
assert.equal(JSON.parse(result.stdout).backend, 'workflow');
|
||||
} finally {
|
||||
cleanup(dir);
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -1,147 +0,0 @@
|
||||
/**
|
||||
* #2605 — the local OpenAI-compatible lanes (ollama / lm_studio / llama.cpp) dropped silently.
|
||||
*
|
||||
* The original defects, all of which made a failed lane indistinguishable from a clean empty
|
||||
* review: bare `curl -s` suppressed curl's own error text; the response was piped straight into
|
||||
* `jq` so the BODY — where an OpenAI-compatible server puts its error JSON on an HTTP 4xx/5xx while
|
||||
* curl still exits 0 — was discarded unread; nothing was written when content was empty, so the
|
||||
* file never existed and `write_reviews` omitted the section entirely; and a whitespace-only reply
|
||||
* passed the byte-counting `[ ! -s … ]` guard as a successful review.
|
||||
*
|
||||
* Phase 5b (#2799) replaced the curl/jq pipeline with an in-process HTTP call and `JSON.parse`, so
|
||||
* this suite drives the runner instead of extracting and executing shell. Two of the original
|
||||
* defects are now structurally impossible rather than merely guarded: there is no pipe to discard
|
||||
* the body, and no `echo` to swallow a `-n`-shaped reply.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const { runLane, runOpenAiCompatible } = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
|
||||
const RUN = '/run';
|
||||
const ROOT = '/repo';
|
||||
const HTTP_LANES = REVIEWER_LANES.filter((l) => l.transport === 'openai-http');
|
||||
|
||||
function planFor(slug, config = {}) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({ lane, configGet: (k) => config[k], runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
/**
|
||||
* These lanes declare an `http-reachable` probe, so the runner performs a GET on /v1/models BEFORE
|
||||
* the chat call. The stub must answer that separately — otherwise the lane is reported unreachable
|
||||
* and never reaches the invocation these tests are actually about.
|
||||
*/
|
||||
function reachableThen(chatResponse) {
|
||||
return async (url, opts) =>
|
||||
opts.method === 'GET'
|
||||
? { ok: true, status: 200, body: JSON.stringify({ data: [{ id: 'stub-model' }] }) }
|
||||
: (typeof chatResponse === 'function' ? chatResponse(url, opts) : chatResponse);
|
||||
}
|
||||
|
||||
function deps(httpJson, files = { [`${RUN}/gsd-review-prompt.md`]: 'PLAN' }) {
|
||||
const warnings = [];
|
||||
return {
|
||||
files,
|
||||
warnings,
|
||||
spawn: () => ({ status: 0, stdout: '', stderr: '' }),
|
||||
httpJson,
|
||||
readFile: (p) => { if (!(p in files)) throw new Error(`ENOENT ${p}`); return files[p]; },
|
||||
writeFile: (p, c) => { files[p] = c; },
|
||||
exists: (p) => p in files,
|
||||
hasBinary: () => true,
|
||||
configGet: () => undefined,
|
||||
homeDir: '/home/u',
|
||||
warn: (m) => warnings.push(m),
|
||||
};
|
||||
}
|
||||
|
||||
const okBody = (content) => ({
|
||||
ok: true, status: 200, body: JSON.stringify({ choices: [{ message: { content } }] }),
|
||||
});
|
||||
|
||||
describe('#2605 local OpenAI-compatible lanes produce diagnosable output', () => {
|
||||
for (const lane of HTTP_LANES) {
|
||||
test(`${lane.slug}: an unreachable endpoint produces a stub carrying the transport error`, async () => {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen({ ok: false, status: 0, body: '', error: 'ECONNREFUSED' }));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
assert.ok(d.files[p.reviewPath].includes('ECONNREFUSED'),
|
||||
'the transport error must be visible — bare `curl -s` used to swallow it');
|
||||
});
|
||||
|
||||
test(`${lane.slug}: an HTTP error body is preserved in the stub`, async () => {
|
||||
// The body is the ONLY evidence on a 4xx/5xx: such a server returns its error JSON there and
|
||||
// curl still exits 0, so stderr is empty. The old pipe into jq discarded it.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen({ ok: false, status: 404, body: '{"error":"model not found"}' }));
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.ok(d.files[p.reviewPath].includes('Raw response body:'));
|
||||
assert.ok(d.files[p.reviewPath].includes('model not found'));
|
||||
});
|
||||
|
||||
test(`${lane.slug}: an empty 200 response still produces a file`, async () => {
|
||||
// Previously nothing was written, so the file never existed, write_reviews omitted the
|
||||
// section, and the result was indistinguishable from the reviewer never being selected.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody('')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
assert.ok(d.files[p.reviewPath] !== undefined, 'a file must exist even on an empty reply');
|
||||
assert.ok(d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
|
||||
test(`${lane.slug}: a whitespace-only response is empty, not a successful review`, async () => {
|
||||
// `[ ! -s … ]` counted BYTES, so " " passed as a real review.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody(' \n\t ')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
});
|
||||
|
||||
test(`${lane.slug}: a reply that is exactly an echo option is NOT misclassified`, async () => {
|
||||
// `echo "$VAR"` would write 0 bytes for `-n`/`-e`/`-E`. Nothing here goes through echo, so
|
||||
// this is structurally impossible now — locked anyway.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody('-n')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('-n'));
|
||||
assert.ok(!d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
|
||||
test(`${lane.slug}: a successful review passes through untouched`, async () => {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody('## Findings\nreal review')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('## Findings'));
|
||||
assert.ok(!d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
}
|
||||
|
||||
test('a served-model mismatch is warned about, not silently accepted', async () => {
|
||||
const p = planFor('lm_studio', { 'review.models.lm_studio': 'asked-for' });
|
||||
const d = deps(async () => ({
|
||||
ok: true, status: 200,
|
||||
body: JSON.stringify({ model: 'actually-served', choices: [{ message: { content: 'R' } }] }),
|
||||
}));
|
||||
const out = await runOpenAiCompatible(p, 'PLAN', d);
|
||||
assert.equal(out.review, 'R');
|
||||
assert.ok(d.warnings.some((w) => w.includes('actually-served') && w.includes('asked-for')));
|
||||
});
|
||||
|
||||
test('neither jq nor curl is required by any of these lanes', () => {
|
||||
// The dependency is gone, not merely satisfied: parsing is JSON.parse and the request is
|
||||
// in-process. `jq` is absent on stock Windows/Git-Bash (#2589), which gated these lanes.
|
||||
for (const lane of HTTP_LANES) {
|
||||
assert.deepStrictEqual([...lane.requiresBinaries], [], `${lane.slug} still declares a binary`);
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -1,522 +0,0 @@
|
||||
/**
|
||||
* Regression tests for #2608 — `query commit --files` ignored `git add` failures
|
||||
* and misreported them.
|
||||
*
|
||||
* A `git add` that fails (unwritable index in a linked worktree whose git dir is
|
||||
* outside the managed writable root, permissions, timeout) was not surfaced.
|
||||
* #2523 had already stopped a failed path entering the commit pathspec, but
|
||||
* skipping it silently left two bad outcomes, both reproduced by this suite
|
||||
* against the pre-fix build:
|
||||
*
|
||||
* - SOME paths fail -> `{"committed":true}`. `git commit` still ran and
|
||||
* PARTIALLY committed the subset that happened to stage,
|
||||
* under a message describing the full requested scope.
|
||||
* - EVERY path fails -> `{"reason":"nothing_to_commit"}`, which is not what
|
||||
* happened and points the operator nowhere.
|
||||
*
|
||||
* In both cases the original `git add` stderr was discarded, so the user saw a
|
||||
* downstream `commit_failed` / pathspec error naming an innocent file.
|
||||
*
|
||||
* The fix collects staging failures and fails closed BEFORE `git commit` runs,
|
||||
* returning `staging_failed` (or `staging_timeout`) with the offending file and
|
||||
* the original stderr preserved.
|
||||
*
|
||||
* ── INJECTION SEAM ────────────────────────────────────────────────────────────
|
||||
* `execGit` is monkeypatched on the shell-command-projection module object. The
|
||||
* compiled call site is `(0, mod.execGit)(...)` — a property lookup at call time
|
||||
* — so the override takes effect. Per CLAUDE.md this is required over
|
||||
* `chmod 0o000` permission tricks, which do not fault under root (root
|
||||
* Docker/CI) and would make these tests silently vacuous.
|
||||
*
|
||||
* The patched call runs in a short-lived `node -e` CHILD rather than in-process,
|
||||
* for two reasons: `output()` writes with `fs.writeSync(1, …)`, which neither
|
||||
* `process.stdout.write` nor `console.log` interception can capture; and a child
|
||||
* keeps the patch from leaking into sibling suites. It is a plain
|
||||
* `process.execPath` spawn — no PATH stub and no exec bit, so it is not subject
|
||||
* to DEFECT.WINDOWS-TEST-PORTABILITY and runs on every platform.
|
||||
*/
|
||||
|
||||
'use strict';
|
||||
|
||||
const { describe, test, beforeEach, afterEach } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const os = require('node:os');
|
||||
const path = require('node:path');
|
||||
const { spawnSync } = require('node:child_process');
|
||||
|
||||
const { createTempGitProject, cleanup } = require('./helpers.cjs');
|
||||
const { gitOrThrow } = require('./helpers/git-fixture.cjs');
|
||||
|
||||
// 5000ms: git plumbing (add/commit/status/rev-parse/rev-list/diff) on a small
|
||||
// mkdtemp fixture repo — well over any observed duration for that class of call.
|
||||
const GIT_TIMEOUT_MS = 5000;
|
||||
|
||||
const LIB = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib');
|
||||
|
||||
/**
|
||||
* Run cmdCommit with `git add <file>` forced to fail for the paths in `failFor`,
|
||||
* returning the parsed JSON result and the git argv list that was actually
|
||||
* executed (so "git commit never ran" is asserted directly, not inferred).
|
||||
*/
|
||||
function commitWithFailingAdd({ cwd, files, failFor = [], stderr = 'fatal: injected staging failure', timeout = false, amend = false, gitVerb = 'add' }) {
|
||||
const callsOut = path.join(fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-2608-')), 'calls.json');
|
||||
// `timeout` is `false` | `true` (alias for `'posix'`) | `'posix'` | `'windows'` —
|
||||
// #3050: the shared isSpawnTimeout predicate only requires `error.code ===
|
||||
// 'ETIMEDOUT'`, NOT `signal === 'SIGTERM'` (Windows does not reliably report
|
||||
// SIGTERM), so both shapes must be proven to still read as a timeout.
|
||||
const timeoutShape = timeout === true ? 'posix' : timeout;
|
||||
const script = `
|
||||
const path = require('path');
|
||||
const LIB = ${JSON.stringify(LIB)};
|
||||
const projection = require(path.join(LIB, 'shell-command-projection.cjs'));
|
||||
const { cmdCommit } = require(path.join(LIB, 'commands.cjs'));
|
||||
const failFor = ${JSON.stringify(failFor)};
|
||||
const stderrText = ${JSON.stringify(stderr)};
|
||||
const timeoutShape = ${JSON.stringify(timeoutShape)};
|
||||
const gitVerb = ${JSON.stringify(gitVerb)};
|
||||
const real = projection.execGit;
|
||||
const calls = [];
|
||||
projection.execGit = (args, opts) => {
|
||||
calls.push(args);
|
||||
if (args[0] === gitVerb && failFor.includes(args[args.length - 1])) {
|
||||
if (timeoutShape === 'posix') {
|
||||
// The exact shape spawnSync produces on a POSIX timeout, which
|
||||
// shell-command-projection surfaces as signal + error.code.
|
||||
const e = new Error('spawnSync git ETIMEDOUT');
|
||||
e.code = 'ETIMEDOUT';
|
||||
return { exitCode: 1, stdout: '', stderr: stderrText, signal: 'SIGTERM', error: e };
|
||||
}
|
||||
if (timeoutShape === 'windows') {
|
||||
// Windows shape: spawnSync's timeout kill does not reliably report
|
||||
// signal:'SIGTERM' — only error.code:'ETIMEDOUT' is guaranteed (#3050).
|
||||
const e = new Error('spawnSync git ETIMEDOUT');
|
||||
e.code = 'ETIMEDOUT';
|
||||
return { exitCode: 1, stdout: '', stderr: stderrText, signal: null, error: e };
|
||||
}
|
||||
return { exitCode: 128, stdout: '', stderr: stderrText, signal: null, error: null };
|
||||
}
|
||||
return real(args, opts);
|
||||
};
|
||||
process.on('exit', () => {
|
||||
require('fs').writeFileSync(${JSON.stringify(callsOut)}, JSON.stringify(calls));
|
||||
});
|
||||
cmdCommit(${JSON.stringify(cwd)}, 'docs: map existing codebase', ${JSON.stringify(files)}, false, ${JSON.stringify(amend)}, false);
|
||||
`;
|
||||
|
||||
const run = spawnSync(process.execPath, ['-e', script], {
|
||||
encoding: 'utf8',
|
||||
timeout: 30000,
|
||||
killSignal: 'SIGKILL',
|
||||
env: { ...process.env, GSD_TEST_MODE: '1' },
|
||||
});
|
||||
|
||||
assert.ok(
|
||||
run.stdout && run.stdout.trim(),
|
||||
`cmdCommit child produced no stdout (status=${run.status}): ${run.stderr}`,
|
||||
);
|
||||
return {
|
||||
result: JSON.parse(run.stdout),
|
||||
gitCalls: JSON.parse(fs.readFileSync(callsOut, 'utf8')),
|
||||
};
|
||||
}
|
||||
|
||||
function headCount(cwd) {
|
||||
return Number(gitOrThrow(['rev-list', '--count', 'HEAD'], { cwd, timeoutMs: GIT_TIMEOUT_MS }).trim());
|
||||
}
|
||||
|
||||
function committedFiles(cwd) {
|
||||
return gitOrThrow(['diff', 'HEAD~1', 'HEAD', '--name-only'], { cwd, timeoutMs: GIT_TIMEOUT_MS })
|
||||
.trim().split('\n').filter(Boolean).sort();
|
||||
}
|
||||
|
||||
/**
|
||||
* Same harness for `cmdCommitToSubrepo` — the sub-repo twin of the staging loop,
|
||||
* which carried the identical defect (failed `git add` dropped, commit proceeds
|
||||
* with the subset that staged).
|
||||
*/
|
||||
function subrepoCommitWithFailingAdd({ cwd, files, failFor = [], timeout = false }) {
|
||||
// See commitWithFailingAdd above for the timeoutShape rationale (#3050).
|
||||
const timeoutShape = timeout === true ? 'posix' : timeout;
|
||||
const script = `
|
||||
const path = require('path');
|
||||
const LIB = ${JSON.stringify(LIB)};
|
||||
const projection = require(path.join(LIB, 'shell-command-projection.cjs'));
|
||||
const { cmdCommitToSubrepo } = require(path.join(LIB, 'commands.cjs'));
|
||||
const failFor = ${JSON.stringify(failFor)};
|
||||
const timeoutShape = ${JSON.stringify(timeoutShape)};
|
||||
const real = projection.execGit;
|
||||
projection.execGit = (args, opts) => {
|
||||
if (args[0] === 'add' && failFor.includes(args[args.length - 1])) {
|
||||
if (timeoutShape === 'posix') {
|
||||
const e = new Error('spawnSync git ETIMEDOUT');
|
||||
e.code = 'ETIMEDOUT';
|
||||
return { exitCode: 1, stdout: '', stderr: 'fatal: injected subrepo staging failure', signal: 'SIGTERM', error: e };
|
||||
}
|
||||
if (timeoutShape === 'windows') {
|
||||
const e = new Error('spawnSync git ETIMEDOUT');
|
||||
e.code = 'ETIMEDOUT';
|
||||
return { exitCode: 1, stdout: '', stderr: 'fatal: injected subrepo staging failure', signal: null, error: e };
|
||||
}
|
||||
return { exitCode: 128, stdout: '', stderr: 'fatal: injected subrepo staging failure', signal: null, error: null };
|
||||
}
|
||||
return real(args, opts);
|
||||
};
|
||||
cmdCommitToSubrepo(${JSON.stringify(cwd)}, 'feat: subrepo change', ${JSON.stringify(files)}, false);
|
||||
`;
|
||||
const run = spawnSync(process.execPath, ['-e', script], {
|
||||
encoding: 'utf8',
|
||||
timeout: 30000,
|
||||
killSignal: 'SIGKILL',
|
||||
env: { ...process.env, GSD_TEST_MODE: '1' },
|
||||
});
|
||||
assert.ok(run.stdout && run.stdout.trim(),
|
||||
`cmdCommitToSubrepo child produced no stdout (status=${run.status}): ${run.stderr}`);
|
||||
return JSON.parse(run.stdout);
|
||||
}
|
||||
|
||||
describe('#2608: commit-to-subrepo fails closed when git add fails', () => {
|
||||
let rootDir;
|
||||
let subDir;
|
||||
|
||||
beforeEach(() => {
|
||||
rootDir = createTempGitProject();
|
||||
fs.writeFileSync(
|
||||
path.join(rootDir, '.planning', 'config.json'),
|
||||
JSON.stringify({ planning: { sub_repos: ['backend'] } }, null, 2),
|
||||
);
|
||||
subDir = path.join(rootDir, 'backend');
|
||||
fs.mkdirSync(subDir, { recursive: true });
|
||||
for (const [cmd, args] of [['init', []], ['config', ['user.email', 'test@example.com']], ['config', ['user.name', 'Test']]]) {
|
||||
gitOrThrow([cmd, ...args], { cwd: subDir, timeoutMs: GIT_TIMEOUT_MS });
|
||||
}
|
||||
fs.writeFileSync(path.join(subDir, 'seed.js'), '// seed\n');
|
||||
gitOrThrow(['add', 'seed.js'], { cwd: subDir, timeoutMs: GIT_TIMEOUT_MS });
|
||||
gitOrThrow(['commit', '-m', 'seed'], { cwd: subDir, timeoutMs: GIT_TIMEOUT_MS });
|
||||
fs.writeFileSync(path.join(subDir, 'a.js'), '// a\n');
|
||||
fs.writeFileSync(path.join(subDir, 'b.js'), '// b\n');
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(rootDir);
|
||||
});
|
||||
|
||||
test('a failed sub-repo git add reports staging_failed and commits nothing', () => {
|
||||
const before = headCount(subDir);
|
||||
const result = subrepoCommitWithFailingAdd({
|
||||
cwd: rootDir,
|
||||
files: ['backend/a.js', 'backend/b.js'],
|
||||
failFor: ['b.js'],
|
||||
});
|
||||
|
||||
assert.equal(result.repos.backend.reason, 'staging_failed',
|
||||
`expected staging_failed for the sub-repo, got ${JSON.stringify(result)}`);
|
||||
assert.equal(result.repos.backend.committed, false);
|
||||
assert.match(result.repos.backend.error, /injected subrepo staging failure/,
|
||||
"git's original stderr must be preserved");
|
||||
assert.equal(headCount(subDir), before, 'no partial sub-repo commit may be created');
|
||||
|
||||
const status = gitOrThrow(['status', '--porcelain'], { cwd: subDir, timeoutMs: GIT_TIMEOUT_MS });
|
||||
assert.deepEqual(status.split('\n').filter((l) => /^A[ \t]/.test(l)), [],
|
||||
`the sub-repo index must be rolled back, status:\n${status}`);
|
||||
});
|
||||
|
||||
test('successful sub-repo staging still commits', () => {
|
||||
const before = headCount(subDir);
|
||||
const result = subrepoCommitWithFailingAdd({
|
||||
cwd: rootDir,
|
||||
files: ['backend/a.js', 'backend/b.js'],
|
||||
failFor: [],
|
||||
});
|
||||
|
||||
assert.equal(result.repos.backend.committed, true, `expected a commit, got ${JSON.stringify(result)}`);
|
||||
assert.equal(headCount(subDir), before + 1);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2608: commit --files fails closed when git add fails', () => {
|
||||
let tmpDir;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempGitProject();
|
||||
for (const name of ['ARCHITECTURE', 'CONCERNS', 'CONVENTIONS']) {
|
||||
fs.writeFileSync(path.join(tmpDir, '.planning', `${name}.md`), `# ${name}\n`);
|
||||
}
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(tmpDir);
|
||||
});
|
||||
|
||||
// ── AC1 + AC3: the failure is reported, with its original stderr ──────────
|
||||
|
||||
test('a failed git add returns staging_failed with the file and original stderr', () => {
|
||||
const before = headCount(tmpDir);
|
||||
const { result, gitCalls } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md'],
|
||||
stderr: 'fatal: Unable to create index.lock: Permission denied',
|
||||
});
|
||||
|
||||
assert.equal(result.committed, false);
|
||||
assert.equal(result.hash, null);
|
||||
assert.equal(result.reason, 'staging_failed',
|
||||
'the staging cause must be reported, not a downstream commit_failed/pathspec error');
|
||||
assert.equal(result.file, '.planning/ARCHITECTURE.md', 'the offending file must be named');
|
||||
assert.match(result.error, /Unable to create index\.lock/,
|
||||
'the original git add stderr must be preserved');
|
||||
|
||||
// AC2: git commit must never have been invoked.
|
||||
assert.ok(
|
||||
!gitCalls.some((a) => a[0] === 'commit'),
|
||||
`git commit must not run after a staging failure, calls: ${JSON.stringify(gitCalls)}`,
|
||||
);
|
||||
assert.equal(headCount(tmpDir), before, 'no commit may be created');
|
||||
});
|
||||
|
||||
// ── AC4: no partial commit of a multi-file explicit scope ─────────────────
|
||||
|
||||
test('when the second of three paths fails to stage, nothing is committed', () => {
|
||||
// Pre-fix this returned {"committed":true} — the two paths that DID stage
|
||||
// were committed under a message describing all three.
|
||||
const before = headCount(tmpDir);
|
||||
const { result, gitCalls } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md', '.planning/CONVENTIONS.md'],
|
||||
failFor: ['.planning/CONCERNS.md'],
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_failed');
|
||||
assert.equal(result.file, '.planning/CONCERNS.md');
|
||||
assert.ok(
|
||||
!gitCalls.some((a) => a[0] === 'commit'),
|
||||
'a partial commit of the paths that DID stage must not happen',
|
||||
);
|
||||
assert.equal(headCount(tmpDir), before, 'no partial commit may be created');
|
||||
});
|
||||
|
||||
test('every failing path is reported, not just the first', () => {
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md', '.planning/CONVENTIONS.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md', '.planning/CONVENTIONS.md'],
|
||||
});
|
||||
|
||||
assert.equal(result.failures.length, 2);
|
||||
assert.deepEqual(
|
||||
result.failures.map((f) => f.file).sort(),
|
||||
['.planning/ARCHITECTURE.md', '.planning/CONVENTIONS.md'],
|
||||
);
|
||||
});
|
||||
|
||||
// ── An all-paths-fail run must not masquerade as nothing_to_commit ────────
|
||||
|
||||
test('when every path fails to stage, the reason is staging_failed not nothing_to_commit', () => {
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md'],
|
||||
});
|
||||
|
||||
assert.notEqual(result.reason, 'nothing_to_commit',
|
||||
'every path failing to stage is a staging failure, not an empty changeset');
|
||||
assert.equal(result.reason, 'staging_failed');
|
||||
});
|
||||
|
||||
// ── AC5: a staging timeout is distinguishable from an ordinary failure ────
|
||||
|
||||
test('a staging timeout is reported as staging_timeout, not staging_failed', () => {
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md'],
|
||||
stderr: '',
|
||||
timeout: true,
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_timeout',
|
||||
'the projection exposes SIGTERM+ETIMEDOUT; a timeout must not read as an ordinary failure');
|
||||
assert.equal(result.failures[0].timed_out, true);
|
||||
});
|
||||
|
||||
// #3050 item 4: this site (commands.cts's `git add` staging loop) now routes
|
||||
// through the shared isSpawnTimeout predicate, which drops the `signal ===
|
||||
// 'SIGTERM'` requirement — a Windows-shaped timeout (no signal, only
|
||||
// error.code === 'ETIMEDOUT') must still be detected.
|
||||
test('a staging timeout is reported as staging_timeout even without SIGTERM (Windows shape, #3050)', () => {
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md'],
|
||||
stderr: '',
|
||||
timeout: 'windows',
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_timeout');
|
||||
assert.equal(result.failures[0].timed_out, true);
|
||||
});
|
||||
|
||||
// #3050 item 4: the `git rm --cached` branch of the same staging loop (the
|
||||
// default-mode "stage the deletion" path, distinct from `git add` above)
|
||||
// carries its own inline copy of the timeout check pre-fix. Drive it
|
||||
// directly: default mode (no explicit --files) stages '.planning/', and
|
||||
// when that path is absent on disk the loop takes the `git rm --cached`
|
||||
// branch instead of `git add`.
|
||||
test('a `git rm --cached` timeout in default mode is reported as staging_timeout, POSIX and Windows shapes (#3050)', () => {
|
||||
// Mid-test fixture mutation (simulating an absent '.planning/' on disk),
|
||||
// not teardown; the outer afterEach still runs helpers.cleanup(tmpDir) on
|
||||
// the whole tmpDir.
|
||||
// eslint-disable-next-line local/no-raw-rmsync-in-tests -- see comment above
|
||||
fs.rmSync(path.join(tmpDir, '.planning'), { recursive: true, force: true });
|
||||
|
||||
for (const shape of ['posix', 'windows']) {
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: undefined,
|
||||
failFor: ['.planning/'],
|
||||
gitVerb: 'rm',
|
||||
stderr: '',
|
||||
timeout: shape,
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_timeout', `shape=${shape}`);
|
||||
assert.equal(result.failures[0].timed_out, true, `shape=${shape}`);
|
||||
}
|
||||
});
|
||||
|
||||
test('an ordinary non-zero git add is NOT reported as a timeout', () => {
|
||||
// Boundary: the timeout carve-out must not swallow the ordinary case.
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md'],
|
||||
failFor: ['.planning/ARCHITECTURE.md'],
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_failed');
|
||||
assert.equal(result.failures[0].timed_out, false);
|
||||
});
|
||||
|
||||
// ── Successful staging preserves the current scoped-commit behaviour ──────
|
||||
|
||||
test('successful staging still commits exactly the declared scope', () => {
|
||||
const before = headCount(tmpDir);
|
||||
fs.writeFileSync(path.join(tmpDir, 'unrelated-wip.txt'), 'wip\n');
|
||||
gitOrThrow(['add', 'unrelated-wip.txt'], { cwd: tmpDir, timeoutMs: GIT_TIMEOUT_MS });
|
||||
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md'],
|
||||
failFor: [],
|
||||
});
|
||||
|
||||
assert.equal(result.committed, true, `expected a commit, got ${JSON.stringify(result)}`);
|
||||
assert.equal(result.reason, 'committed');
|
||||
assert.equal(headCount(tmpDir), before + 1);
|
||||
assert.deepEqual(
|
||||
committedFiles(tmpDir),
|
||||
['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md'],
|
||||
'the declared scope must still be honoured, and the unrelated staged file left alone',
|
||||
);
|
||||
});
|
||||
|
||||
// ── Missing explicit files keep their existing documented handling ────────
|
||||
|
||||
test('a missing explicit file is still skipped, not reported as a staging failure', () => {
|
||||
// #2014/#2523 behaviour: an explicitly-named file that does not exist is
|
||||
// skipped rather than staged as a deletion. It never reaches `git add`, so
|
||||
// it is not a staging failure and must not become one.
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/DOES-NOT-EXIST.md'],
|
||||
failFor: [],
|
||||
});
|
||||
|
||||
assert.equal(result.committed, true, `expected a commit, got ${JSON.stringify(result)}`);
|
||||
assert.deepEqual(committedFiles(tmpDir), ['.planning/ARCHITECTURE.md']);
|
||||
});
|
||||
|
||||
// ── The index must be left clean, not partially staged ───────────────────
|
||||
|
||||
test('a staging failure rolls back the paths this call had already staged', () => {
|
||||
// Without the rollback the paths that DID stage stay in the index with no
|
||||
// commit made, so the next bare `git commit` sweeps them up — the same
|
||||
// silent partial commit this fix exists to prevent, just deferred a step.
|
||||
commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md', '.planning/CONVENTIONS.md'],
|
||||
failFor: ['.planning/CONCERNS.md'],
|
||||
});
|
||||
|
||||
const status = gitOrThrow(['status', '--porcelain'], { cwd: tmpDir, timeoutMs: GIT_TIMEOUT_MS });
|
||||
const stagedAdds = status.split('\n').filter((l) => /^A[ \t]/.test(l));
|
||||
assert.deepEqual(stagedAdds, [],
|
||||
`no path may remain staged after a staging failure, status:\n${status}`);
|
||||
});
|
||||
|
||||
test('the rollback does not unstage work the caller had staged before the call', () => {
|
||||
// Boundary: the reset must touch only what THIS call staged. Unstaging a
|
||||
// path the caller staged themselves would destroy their work.
|
||||
fs.writeFileSync(path.join(tmpDir, 'caller-staged.txt'), 'mine\n');
|
||||
gitOrThrow(['add', 'caller-staged.txt'], { cwd: tmpDir, timeoutMs: GIT_TIMEOUT_MS });
|
||||
|
||||
commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/ARCHITECTURE.md', '.planning/CONCERNS.md'],
|
||||
failFor: ['.planning/CONCERNS.md'],
|
||||
});
|
||||
|
||||
const status = gitOrThrow(['status', '--porcelain'], { cwd: tmpDir, timeoutMs: GIT_TIMEOUT_MS });
|
||||
assert.match(status, /^A[ \t]+caller-staged\.txt$/m,
|
||||
`the caller's own staged file must survive the rollback, status:\n${status}`);
|
||||
});
|
||||
|
||||
// ── The default (non---files) staging path is guarded too ─────────────────
|
||||
|
||||
test('a failed default-mode git add fails closed instead of committing the index', () => {
|
||||
// Default mode stages `.planning/`. Pre-fix a failure there also fell
|
||||
// through to an unguarded `git commit`.
|
||||
const before = headCount(tmpDir);
|
||||
const { result, gitCalls } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: undefined,
|
||||
failFor: ['.planning/'],
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_failed');
|
||||
assert.ok(!gitCalls.some((a) => a[0] === 'commit'), 'git commit must not run');
|
||||
assert.equal(headCount(tmpDir), before);
|
||||
});
|
||||
|
||||
test('a failed default-mode git add blocks --amend too', () => {
|
||||
// --amend has no carve-out: amending on top of a failed staging would
|
||||
// rewrite the tip without the changes the caller asked for.
|
||||
const before = gitOrThrow(['rev-parse', 'HEAD'], { cwd: tmpDir, timeoutMs: GIT_TIMEOUT_MS }).trim();
|
||||
const { result, gitCalls } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: undefined,
|
||||
failFor: ['.planning/'],
|
||||
amend: true,
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'staging_failed');
|
||||
assert.ok(!gitCalls.some((a) => a[0] === 'commit'), 'git commit --amend must not run');
|
||||
assert.equal(
|
||||
gitOrThrow(['rev-parse', 'HEAD'], { cwd: tmpDir, timeoutMs: GIT_TIMEOUT_MS }).trim(),
|
||||
before,
|
||||
'HEAD must not be rewritten when staging failed',
|
||||
);
|
||||
});
|
||||
|
||||
test('when all explicit files are missing the reason is still nothing_to_commit', () => {
|
||||
// The nothing_to_commit path must survive: no `git add` ran, so there is no
|
||||
// staging failure to report.
|
||||
const { result } = commitWithFailingAdd({
|
||||
cwd: tmpDir,
|
||||
files: ['.planning/GONE-A.md', '.planning/GONE-B.md'],
|
||||
failFor: [],
|
||||
});
|
||||
|
||||
assert.equal(result.reason, 'nothing_to_commit');
|
||||
});
|
||||
});
|
||||
@@ -1,391 +0,0 @@
|
||||
/**
|
||||
* #2766 — three audit-uat false negatives.
|
||||
*
|
||||
* All three fail SILENTLY and in the reassuring direction: outstanding UAT work
|
||||
* is under-reported, so the command whose job is to be the backstop reports a
|
||||
* clean bill of health over real unresolved items.
|
||||
*
|
||||
* 1. `cmdAuditUat` scanned only `.planning/phases/`. On milestone completion
|
||||
* `milestone.cts` MOVES each phase dir into
|
||||
* `.planning/milestones/<version>-phases/` (archive-by-default since #1871),
|
||||
* so a partly-archived project silently omitted the archived phases and a
|
||||
* fully-archived one hard-errored with "No phases directory found",
|
||||
* indistinguishable from a broken install.
|
||||
* 2. `parseDeferredItems` delegates entry splitting to `splitGapsEntries`, which
|
||||
* keys entirely on `- ` bullet openers — so a `deferred-items.md` recording
|
||||
* entries as a GFM table produced ZERO items. The SCOPE BOUNDARY convention
|
||||
* mandates no shape, and a table is natural for "test → failing seeds".
|
||||
* 3. `parseGapsItems` has the identical blindness for the same shared reason, so
|
||||
* a table-shaped `## Gaps` section surfaced nothing.
|
||||
*
|
||||
* Same false-negative family as #2286 and #2287, which fixed the first two
|
||||
* shapes; these are the next ones out.
|
||||
*
|
||||
* This fix:
|
||||
* - `cmdAuditUat` collects scan targets from the active tree AND
|
||||
* `getArchivedPhaseDirs` (the canonical seam `findPhaseInternal` already uses),
|
||||
* erroring only when both are empty. Archived dirs deliberately bypass
|
||||
* `getMilestonePhaseFilter` — it derives the CURRENT milestone's phase numbers
|
||||
* from ROADMAP.md, so applying it to archived dirs discards every one and
|
||||
* silently reinstates the bug. Results gain `archived_milestone` for provenance.
|
||||
* - One shared `collectTableRows` walker (header/delimiter/boundary handling in
|
||||
* one place) feeds table scans in BOTH parsers as a UNION with the existing
|
||||
* bullet scan. A `|`-leading line is never a `- ` bullet opener, so mixed files
|
||||
* surface both with no double-counting. Resolution semantics are unchanged
|
||||
* from the bullet path: suppress only on an explicit marker.
|
||||
*/
|
||||
|
||||
'use strict';
|
||||
|
||||
const { test, describe, beforeEach, afterEach } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const { runGsdTools, createTempProject, createTempDir, cleanup } = require('./helpers.cjs');
|
||||
const { parseDeferredItems } = require('../gsd-core/bin/lib/uat.cjs');
|
||||
|
||||
const UAT_ONE_PENDING = [
|
||||
'---',
|
||||
'status: partial',
|
||||
'phase: 01-foundation',
|
||||
'---',
|
||||
'',
|
||||
'## Current Test',
|
||||
'',
|
||||
'[awaiting human testing]',
|
||||
'',
|
||||
'## Tests',
|
||||
'',
|
||||
'### 1. A scenario nobody ever ran',
|
||||
'expected: something observable happens',
|
||||
'result: [pending]',
|
||||
'',
|
||||
'## Summary',
|
||||
'',
|
||||
'total: 1',
|
||||
'pending: 1',
|
||||
'',
|
||||
'## Gaps',
|
||||
'',
|
||||
].join('\n');
|
||||
|
||||
/** Write a UAT file whose `## Gaps` section holds `gapsBody`. */
|
||||
function uatWithGaps(gapsBody) {
|
||||
return [
|
||||
'---',
|
||||
'status: complete',
|
||||
'phase: 50-gaps',
|
||||
'---',
|
||||
'',
|
||||
'## Current Test',
|
||||
'',
|
||||
'[testing complete]',
|
||||
'',
|
||||
'## Tests',
|
||||
'',
|
||||
'### 1. A passing scenario',
|
||||
'expected: this one is fine',
|
||||
'result: pass',
|
||||
'',
|
||||
'## Summary',
|
||||
'',
|
||||
'total: 1',
|
||||
'passed: 1',
|
||||
'',
|
||||
'## Gaps',
|
||||
'',
|
||||
gapsBody,
|
||||
'',
|
||||
].join('\n');
|
||||
}
|
||||
|
||||
// ─── Bug 1: archived phase dirs ───────────────────────────────────────────────
|
||||
|
||||
describe('#2766 cmdAuditUat: archived phase directories', () => {
|
||||
let tmpDir;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempProject();
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(tmpDir);
|
||||
});
|
||||
|
||||
test('phases ONLY in the archive → items surfaced, not a hard error', () => {
|
||||
const archiveDir = path.join(
|
||||
tmpDir, '.planning', 'milestones', 'v1.0-phases', '01-foundation',
|
||||
);
|
||||
fs.mkdirSync(archiveDir, { recursive: true });
|
||||
fs.writeFileSync(path.join(archiveDir, '01-UAT.md'), UAT_ONE_PENDING);
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.strictEqual(output.summary.total_items, 1);
|
||||
assert.strictEqual(output.results.length, 1);
|
||||
assert.strictEqual(output.results[0].phase, '01');
|
||||
assert.strictEqual(output.results[0].archived_milestone, 'v1.0');
|
||||
assert.match(output.results[0].file_path, /milestones\/v1\.0-phases\//);
|
||||
});
|
||||
|
||||
test('active and archived trees are both scanned', () => {
|
||||
const activeDir = path.join(tmpDir, '.planning', 'phases', '40-current');
|
||||
fs.mkdirSync(activeDir, { recursive: true });
|
||||
fs.writeFileSync(path.join(activeDir, '40-UAT.md'), UAT_ONE_PENDING);
|
||||
|
||||
const archiveDir = path.join(
|
||||
tmpDir, '.planning', 'milestones', 'v1.0-phases', '01-foundation',
|
||||
);
|
||||
fs.mkdirSync(archiveDir, { recursive: true });
|
||||
fs.writeFileSync(path.join(archiveDir, '01-UAT.md'), UAT_ONE_PENDING);
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
const byPhase = new Map(output.results.map(r => [r.phase, r]));
|
||||
assert.ok(byPhase.has('01'), `archived phase missing: ${JSON.stringify([...byPhase.keys()])}`);
|
||||
assert.ok(byPhase.has('40'), `active phase missing: ${JSON.stringify([...byPhase.keys()])}`);
|
||||
assert.strictEqual(byPhase.get('01').archived_milestone, 'v1.0');
|
||||
assert.strictEqual(byPhase.get('40').archived_milestone, undefined);
|
||||
});
|
||||
|
||||
test('multiple archived milestones are all scanned', () => {
|
||||
for (const [version, phase] of [['v1.0', '01-foundation'], ['v2.0', '07-later']]) {
|
||||
const dir = path.join(tmpDir, '.planning', 'milestones', `${version}-phases`, phase);
|
||||
fs.mkdirSync(dir, { recursive: true });
|
||||
fs.writeFileSync(path.join(dir, `${phase.slice(0, 2)}-UAT.md`), UAT_ONE_PENDING);
|
||||
}
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.strictEqual(output.summary.total_items, 2);
|
||||
assert.deepStrictEqual(
|
||||
output.results.map(r => r.archived_milestone).sort(),
|
||||
['v1.0', 'v2.0'],
|
||||
);
|
||||
});
|
||||
|
||||
test('an empty active phases dir still succeeds with no items (pre-existing behavior)', () => {
|
||||
// createTempProject() ships an empty `.planning/phases/`, so this is the
|
||||
// shape the existing uat.test.cjs "no UAT files" case covers — the archive
|
||||
// change must not turn it into an error.
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.deepStrictEqual(output.results, []);
|
||||
assert.strictEqual(output.summary.total_items, 0);
|
||||
});
|
||||
|
||||
test('no phases dir AND no archive still errors — no false all-clear', () => {
|
||||
// A bare temp dir with a .planning/ that has NO phases subdir and no
|
||||
// milestones archive — built from createTempDir rather than by deleting
|
||||
// createTempProject's phases dir, so nothing is torn down mid-test.
|
||||
const bare = createTempDir();
|
||||
try {
|
||||
fs.mkdirSync(path.join(bare, '.planning'), { recursive: true });
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', bare);
|
||||
assert.strictEqual(result.success, false, 'expected a failure when no phases exist at all');
|
||||
} finally {
|
||||
cleanup(bare);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Bug 2: table-shaped deferred-items.md ────────────────────────────────────
|
||||
|
||||
describe('#2766 parseDeferredItems: GFM table shape', () => {
|
||||
const names = (md) => parseDeferredItems(md).map(i => i.name);
|
||||
|
||||
test('header + delimiter → header dropped, data rows surfaced', () => {
|
||||
assert.deepStrictEqual(
|
||||
names([
|
||||
'## Discovered during 01-03',
|
||||
'',
|
||||
'| Test | Failing seeds |',
|
||||
'|------|---------------|',
|
||||
'| test_a | 0, 1 |',
|
||||
'| test_b | 424242 |',
|
||||
].join('\n')),
|
||||
['test_a — 0, 1', 'test_b — 424242'],
|
||||
);
|
||||
});
|
||||
|
||||
test('later columns are preserved, not truncated to the first cell', () => {
|
||||
const [name] = names('| T | seeds |\n|---|---|\n| test_a | 0, 1, 424242 |');
|
||||
assert.match(name, /0, 1, 424242/);
|
||||
});
|
||||
|
||||
test('headerless table → every row surfaced', () => {
|
||||
assert.deepStrictEqual(
|
||||
names('| test_a | 0 |\n| test_b | 1 |'),
|
||||
['test_a — 0', 'test_b — 1'],
|
||||
);
|
||||
});
|
||||
|
||||
test('row marked resolved/done/pass is suppressed', () => {
|
||||
assert.deepStrictEqual(
|
||||
names([
|
||||
'| Test | Seeds | Status |',
|
||||
'|---|---|---|',
|
||||
'| test_open | 0 | open |',
|
||||
'| test_fixed | 1 | resolved |',
|
||||
'| test_done | 2 | DONE |',
|
||||
].join('\n')),
|
||||
['test_open — 0 — open'],
|
||||
);
|
||||
});
|
||||
|
||||
test('two prose-separated tables → each drops its own header', () => {
|
||||
assert.deepStrictEqual(
|
||||
names([
|
||||
'| T1 | x |', '|---|---|', '| one | 1 |',
|
||||
'',
|
||||
'some prose in between',
|
||||
'',
|
||||
'| T2 | y |', '|---|---|', '| two | 2 |',
|
||||
].join('\n')),
|
||||
['one — 1', 'two — 2'],
|
||||
);
|
||||
});
|
||||
|
||||
test('bullets and a table in one file → union, no double-counting', () => {
|
||||
const got = names([
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- a bullet-shaped deferred entry',
|
||||
'',
|
||||
'| Test | Seeds |',
|
||||
'|---|---|',
|
||||
'| test_a | 0 |',
|
||||
].join('\n'));
|
||||
assert.strictEqual(got.length, 2, JSON.stringify(got));
|
||||
assert.ok(got.some(n => n.includes('bullet-shaped')));
|
||||
assert.ok(got.some(n => n.startsWith('test_a')));
|
||||
});
|
||||
|
||||
test('bullet-only file unchanged (no regression on #2287)', () => {
|
||||
assert.deepStrictEqual(
|
||||
names('## Deferred Items\n\n- entry one\n- entry two\n'),
|
||||
['entry one', 'entry two'],
|
||||
);
|
||||
});
|
||||
|
||||
test('explicit status: resolved bullet still suppressed (no regression on #2287)', () => {
|
||||
const got = names(
|
||||
'## Deferred Items\n\n- truth: "closed thing"\n status: resolved\n- truth: "open thing"\n',
|
||||
);
|
||||
assert.strictEqual(got.length, 1, JSON.stringify(got));
|
||||
assert.match(got[0], /open thing/);
|
||||
});
|
||||
|
||||
test('no table and no bullets → zero items, no throw', () => {
|
||||
assert.deepStrictEqual(names('# Notes\n\njust prose, nothing actionable.\n'), []);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Bug 3: table-shaped ## Gaps section ──────────────────────────────────────
|
||||
|
||||
describe('#2766 parseGapsItems: GFM table shape', () => {
|
||||
let tmpDir;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempProject();
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(tmpDir);
|
||||
});
|
||||
|
||||
/** Run audit-uat over a phase whose UAT file has `gapsBody` as its Gaps section. */
|
||||
function gapsItems(gapsBody) {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '50-gaps');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
fs.writeFileSync(path.join(phaseDir, '50-UAT.md'), uatWithGaps(gapsBody));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
const output = JSON.parse(result.output);
|
||||
const uat = output.results.find(r => r.type === 'uat');
|
||||
return uat ? uat.items : [];
|
||||
}
|
||||
|
||||
test('header-mapped table → truth/status/reason/test extracted', () => {
|
||||
const items = gapsItems([
|
||||
'| Truth | Status | Reason | Test |',
|
||||
'|-------|--------|--------|------|',
|
||||
'| Login should redirect | failed | User reported a 500 | 1 |',
|
||||
].join('\n'));
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items));
|
||||
assert.strictEqual(items[0].name, 'Login should redirect');
|
||||
assert.strictEqual(items[0].result, 'failed');
|
||||
assert.strictEqual(items[0].reason, 'User reported a 500');
|
||||
assert.strictEqual(items[0].test, 1);
|
||||
});
|
||||
|
||||
test('status: resolved row suppressed, open row kept', () => {
|
||||
const items = gapsItems([
|
||||
'| Truth | Status |',
|
||||
'|-------|--------|',
|
||||
'| closed thing | resolved |',
|
||||
'| open thing | failed |',
|
||||
].join('\n'));
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items.map(i => i.name)));
|
||||
assert.strictEqual(items[0].name, 'open thing');
|
||||
});
|
||||
|
||||
test('no status column → surfaced as unknown, not dropped', () => {
|
||||
const items = gapsItems('| Truth | Note |\n|---|---|\n| something is off | see logs |');
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items));
|
||||
assert.strictEqual(items[0].result, 'unknown');
|
||||
assert.strictEqual(items[0].name, 'something is off');
|
||||
});
|
||||
|
||||
test('unrecognizable header → joined cells + unknown status', () => {
|
||||
const items = gapsItems('| Alpha | Beta |\n|---|---|\n| xxx | yyy |');
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items));
|
||||
assert.strictEqual(items[0].result, 'unknown');
|
||||
assert.match(items[0].name, /xxx/);
|
||||
assert.match(items[0].name, /yyy/);
|
||||
});
|
||||
|
||||
test('headerless table → explicit resolved cell still suppressed', () => {
|
||||
const items = gapsItems('| open thing | failed |\n| closed thing | resolved |');
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items.map(i => i.name)));
|
||||
assert.match(items[0].name, /open thing/);
|
||||
});
|
||||
|
||||
test('bullets and a table in one Gaps section → union, no double-counting', () => {
|
||||
const items = gapsItems([
|
||||
'- truth: "a bullet gap"',
|
||||
' status: failed',
|
||||
'',
|
||||
'| Truth | Status |',
|
||||
'|---|---|',
|
||||
'| a table gap | failed |',
|
||||
].join('\n'));
|
||||
|
||||
assert.strictEqual(items.length, 2, JSON.stringify(items.map(i => i.name)));
|
||||
assert.ok(items.some(i => i.name === 'a bullet gap'));
|
||||
assert.ok(items.some(i => i.name === 'a table gap'));
|
||||
});
|
||||
|
||||
test('bullet-only Gaps unchanged (no regression on #2286)', () => {
|
||||
const items = gapsItems('- truth: "only a bullet"\n status: failed\n reason: "because"\n');
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items));
|
||||
assert.strictEqual(items[0].name, 'only a bullet');
|
||||
assert.strictEqual(items[0].reason, 'because');
|
||||
});
|
||||
});
|
||||
@@ -1,93 +0,0 @@
|
||||
/**
|
||||
* #2794 — the qwen reviewer leg was the last one still sending stderr to /dev/null.
|
||||
*
|
||||
* Every other lane captured stderr to a `.err` sidecar and appended it to the stub (#2494/#2605);
|
||||
* qwen wrote a bare "failed or returned empty output." with no diagnostic at all, so a missing
|
||||
* binary, an auth prompt and a rate-limit were indistinguishable from each other AND from a clean
|
||||
* empty review.
|
||||
*
|
||||
* Phase 5b (#2799) deleted the per-CLI bash this suite used to extract and run under a real bash.
|
||||
* The invariant is now structural rather than per-leg — the sidecar is written by the shared runner
|
||||
* for every lane — so these tests drive the runner directly. That also means the defect this issue
|
||||
* describes can no longer recur for ONE lane: there is no longer a per-lane place to get it wrong.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const { runLane } = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
|
||||
const RUN = '/run';
|
||||
const ROOT = '/repo';
|
||||
|
||||
function planFor(slug) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
function deps(spawnResult, files = {}, hasBinary = () => true) {
|
||||
return {
|
||||
files,
|
||||
spawn: () => spawnResult,
|
||||
httpJson: async () => ({ ok: true, status: 200, body: '{}' }),
|
||||
readFile: (p) => { if (!(p in files)) throw new Error(`ENOENT ${p}`); return files[p]; },
|
||||
writeFile: (p, c) => { files[p] = c; },
|
||||
exists: (p) => p in files,
|
||||
hasBinary,
|
||||
configGet: () => undefined,
|
||||
homeDir: '/home/u',
|
||||
warn: () => {},
|
||||
};
|
||||
}
|
||||
|
||||
describe('#2794 qwen reviewer stderr capture', () => {
|
||||
test('writes the review on success', async () => {
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 0, stdout: '## Qwen findings\nall good\n', stderr: '' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('## Qwen findings'));
|
||||
});
|
||||
|
||||
test('a failed lane surfaces its stderr in the review stub', async () => {
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 1, stdout: '', stderr: 'auth required: run `qwen login`' });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.ok(d.files[p.reviewPath].includes('auth required'),
|
||||
'the diagnostic must reach the review, not just the sidecar');
|
||||
assert.equal(d.files[p.errPath], 'auth required: run `qwen login`');
|
||||
});
|
||||
|
||||
test('a silently empty lane still produces a diagnosable stub', async () => {
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 0, stdout: '', stderr: '' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
assert.ok(d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
|
||||
test('a missing qwen binary reports unavailable rather than an empty review', async () => {
|
||||
// Stronger than the original: the lane is now reported with a TYPED reason before it is ever
|
||||
// spawned, instead of producing a stub that looked the same as every other failure.
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 0, stdout: '', stderr: '' }, {}, () => false);
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.ok, false);
|
||||
assert.equal(r.reason, 'missing_binary');
|
||||
assert.equal(d.files[p.reviewPath], undefined, 'no review file for a lane that never ran');
|
||||
});
|
||||
|
||||
test('the sidecar is structural — no lane can opt out of it', async () => {
|
||||
// The #2794 defect was one lane diverging from a convention every other lane followed. There is
|
||||
// no per-lane place to diverge any more; this asserts that directly.
|
||||
for (const lane of REVIEWER_LANES.filter((l) => l.transport === 'spawn')) {
|
||||
const p = planFor(lane.slug);
|
||||
assert.ok(p.errPath.endsWith('.err'), `${lane.slug} must declare a stderr sidecar`);
|
||||
assert.notEqual(p.errPath, '/dev/null');
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -1,278 +0,0 @@
|
||||
/**
|
||||
* #3177 — `gsd-core/workflows/execute-phase.md` asserted that Claude Code's
|
||||
* `Agent(...)` "blocks until complete, returns result", and restated the same
|
||||
* claim further down as "Claude Code's Agent() normally returns synchronously".
|
||||
*
|
||||
* Both were stale. Claude Code runs subagents in the BACKGROUND by default as of
|
||||
* v2.1.198; `run_in_background: false` is the explicit opt-out when an immediate
|
||||
* result is needed (code.claude.com/docs/en/sub-agents,
|
||||
* /agent-sdk/subagents, /tools-reference).
|
||||
*
|
||||
* This package already recorded the true value: `capabilities/claude/capability.json`
|
||||
* declares `dispatch.background: true` and `docs/reference/host-integration-capability-matrix.md`
|
||||
* renders it. So the SHIPPED PROSE and the SHIPPED MATRIX disagreed about the same
|
||||
* runtime — and the prose is the copy loaded verbatim into the orchestrator's context
|
||||
* on every `/gsd-execute-phase` run, where it is the premise a spawn-safety conclusion
|
||||
* leans on (#3159 documents a host where that difference cost a wave of executor work).
|
||||
*
|
||||
* The durable guard is therefore NOT a snapshot of today's wording: it is a PARITY
|
||||
* assertion (tests 3 and 4) that reds whenever the prose, the matrix, and the descriptor
|
||||
* stop agreeing about `dispatch.background` — in either direction. Tests 5 and 6 pin the
|
||||
* negative space so a future correction cannot over-sweep: Codex dispatch genuinely IS
|
||||
* synchronous, and its orchestrator rules must survive a fix aimed at Claude Code.
|
||||
*/
|
||||
|
||||
// allow-test-rule: runtime-contract-is-the-product #3177 — the workflow markdown is loaded
|
||||
// verbatim into the agent's context and the matrix/descriptor ARE the negotiated host
|
||||
// contract; asserting agreement between those documents is behavioral, not source-grep.
|
||||
|
||||
'use strict';
|
||||
|
||||
process.env.GSD_TEST_MODE = '1';
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const fc = require('fast-check');
|
||||
|
||||
const ROOT = path.join(__dirname, '..');
|
||||
const WORKFLOW = path.join(ROOT, 'gsd-core', 'workflows', 'execute-phase.md');
|
||||
const MATRIX = path.join(ROOT, 'docs', 'reference', 'host-integration-capability-matrix.md');
|
||||
const DESCRIPTOR = path.join(ROOT, 'capabilities', 'claude', 'capability.json');
|
||||
|
||||
const workflowText = () => fs.readFileSync(WORKFLOW, 'utf8');
|
||||
|
||||
/**
|
||||
* Value of `| <field> | <value> | …` inside the `## <host>` section of the matrix.
|
||||
*
|
||||
* Anchored on a whole heading LINE, not a substring: `## claude` must never match
|
||||
* `## claude-local`, and the section must end at the next `## ` heading so a field
|
||||
* absent from this host can never be answered from the next host's table.
|
||||
*
|
||||
* @param {string} matrix - full matrix document text
|
||||
* @param {string} host - section name, e.g. `claude`
|
||||
* @param {string} field - row label, e.g. `dispatch.background`
|
||||
* @returns {string|null} trimmed cell value, or null when the section or row is absent
|
||||
*/
|
||||
function matrixField(matrix, host, field) {
|
||||
const lines = matrix.split('\n');
|
||||
const start = lines.findIndex((l) => l.trim() === `## ${host}`);
|
||||
if (start === -1) return null;
|
||||
let end = lines.length;
|
||||
for (let i = start + 1; i < lines.length; i += 1) {
|
||||
if (lines[i].startsWith('## ')) { end = i; break; }
|
||||
}
|
||||
const row = lines.slice(start + 1, end).find((l) => l.startsWith(`| ${field} |`));
|
||||
if (!row) return null;
|
||||
return row.split('|')[2].trim();
|
||||
}
|
||||
|
||||
describe('#3177: execute-phase.md states Claude Code dispatch truthfully', () => {
|
||||
test('execute-phase.md never claims Claude Code Agent() blocks or returns synchronously', () => {
|
||||
// Row 1 — the failing-first regression. Both stale sentences, by their own text.
|
||||
const text = workflowText();
|
||||
const stale = ['blocks until complete', 'returns synchronously'];
|
||||
const present = stale.filter((phrase) => text.includes(phrase));
|
||||
assert.deepEqual(
|
||||
present, [],
|
||||
'execute-phase.md still asserts synchronous Claude Code dispatch. Claude Code backgrounds '
|
||||
+ 'subagents by default (v2.1.198+); `run_in_background: false` is the opt-out.',
|
||||
);
|
||||
});
|
||||
|
||||
test('the Claude Code dispatch bullet states the background-by-default model', () => {
|
||||
// Row 2 — the corrected sentence must actually SAY the true thing, not merely
|
||||
// omit the false one. A deletion would pass row 1 while teaching nothing.
|
||||
const bullet = workflowText()
|
||||
.split('\n')
|
||||
.find((l) => l.startsWith('- **Claude Code:**'));
|
||||
assert.ok(bullet, 'the <runtime_compatibility> Claude Code bullet must exist');
|
||||
assert.match(
|
||||
bullet, /backgrounded by default/,
|
||||
'the bullet must state that dispatch is backgrounded by default',
|
||||
);
|
||||
assert.match(
|
||||
bullet, /verify completion/,
|
||||
'the bullet must point at completion verification. This workflow deliberately backgrounds '
|
||||
+ 'its executors (the multi-plan path prescribes run_in_background: true), so the blocking '
|
||||
+ 'opt-out is not the guidance here — confirming completion is.',
|
||||
);
|
||||
assert.ok(
|
||||
bullet.includes('Agent(subagent_type="gsd-executor"'),
|
||||
'the bullet must still carry the dispatch mechanism the rest of the file depends on',
|
||||
);
|
||||
});
|
||||
|
||||
test('the workflow prose and the capability matrix agree on claude dispatch.background', () => {
|
||||
// Row 3 — the parity guard. This is the assertion that outlives the wording:
|
||||
// flip the matrix to `false` without touching the prose (or vice versa) and this reds.
|
||||
const declared = matrixField(fs.readFileSync(MATRIX, 'utf8'), 'claude', 'dispatch.background');
|
||||
assert.equal(declared, 'true', 'matrix must document claude dispatch.background');
|
||||
|
||||
const bullet = workflowText()
|
||||
.split('\n')
|
||||
.find((l) => l.startsWith('- **Claude Code:**'));
|
||||
const proseSaysBackground = /backgrounded by default/.test(bullet ?? '');
|
||||
assert.equal(
|
||||
proseSaysBackground, declared === 'true',
|
||||
'execute-phase.md and the host-integration matrix disagree about whether claude backgrounds '
|
||||
+ 'its subagent dispatch. They describe the same runtime; exactly one of them is wrong (#3177).',
|
||||
);
|
||||
});
|
||||
|
||||
test('the claude descriptor and the matrix agree on dispatch.background', () => {
|
||||
// Row 4 — the other half of the divergence class. The matrix is generated from
|
||||
// the descriptor, so this pins the generator's output to its input.
|
||||
const descriptor = JSON.parse(fs.readFileSync(DESCRIPTOR, 'utf8'));
|
||||
const declared = matrixField(fs.readFileSync(MATRIX, 'utf8'), 'claude', 'dispatch.background');
|
||||
assert.equal(
|
||||
String(descriptor.runtime.hostIntegration.dispatch.background), declared,
|
||||
'capabilities/claude/capability.json and the rendered matrix disagree',
|
||||
);
|
||||
});
|
||||
|
||||
test('the Codex orchestrator rule is not swept by the Claude Code correction', () => {
|
||||
// Row 5 — negative space. Codex dispatch IS synchronous. A regex sweep for
|
||||
// "return its result" would introduce a NEW falsehood here; this catches that.
|
||||
const text = workflowText();
|
||||
const codexRules = text
|
||||
.split('\n')
|
||||
.filter((l) => l.includes('ORCHESTRATOR RULE — CODEX RUNTIME'));
|
||||
assert.ok(codexRules.length >= 2, 'both Codex orchestrator rules must survive');
|
||||
for (const rule of codexRules) {
|
||||
assert.ok(
|
||||
rule.includes('Wait for the subagent to return its result'),
|
||||
'Codex dispatch is genuinely synchronous — its wait rule must not be corrected away',
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test('the Copilot and multi-plan dispatch rules survive the correction', () => {
|
||||
// Row 6 — negative space. These were already true and are adjacent to the edit.
|
||||
const text = workflowText();
|
||||
assert.ok(
|
||||
text.includes('- **Copilot:** Subagent spawning does not reliably return completion signals.'),
|
||||
'the Copilot bullet must survive verbatim',
|
||||
);
|
||||
assert.ok(
|
||||
text.includes('one at a time with `run_in_background: true`'),
|
||||
'the multi-plan wave prescription must survive verbatim',
|
||||
);
|
||||
assert.ok(
|
||||
text.includes('If `Agent` IS available (top-level Claude'),
|
||||
'the spawn mandate is derived from TOOL AVAILABILITY, not from blocking — it must survive',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#3177: matrix section extraction is bounded by its heading', () => {
|
||||
const matrix = () => fs.readFileSync(MATRIX, 'utf8');
|
||||
|
||||
test('section extraction — first row of the section', () => {
|
||||
// limit-1: the row immediately after the `## claude` heading is INSIDE.
|
||||
assert.equal(matrixField(matrix(), 'claude', 'embeddingMode'), 'imperative');
|
||||
});
|
||||
|
||||
test('section extraction — last row before the next heading', () => {
|
||||
// limit: the final row of `## claude` is still INSIDE.
|
||||
assert.ok(matrixField(matrix(), 'claude', 'dispatch.isolation').startsWith('harness-worktree'));
|
||||
});
|
||||
|
||||
test('section extraction — a row in the next section never leaks in', () => {
|
||||
// limit+1: a field absent from claude must be null rather than silently
|
||||
// resolved from `## codex` below it, and two hosts with different values
|
||||
// must never resolve to the same cell.
|
||||
const doc = matrix();
|
||||
assert.equal(matrixField(doc, 'claude', 'dispatch.background'), 'true');
|
||||
assert.equal(matrixField(doc, 'claude', '__definitely_not_a_field__'), null);
|
||||
assert.notEqual(
|
||||
matrixField(doc, 'claude', 'effortSurface'),
|
||||
matrixField(doc, 'kilo', 'effortSurface'),
|
||||
'two hosts with different values must not resolve to the same cell',
|
||||
);
|
||||
});
|
||||
|
||||
test('fc property: field extraction never leaks across ## boundaries', () => {
|
||||
const hostArb = fc.stringMatching(/^[a-z][a-z0-9-]{0,12}$/);
|
||||
const valueArb = fc.stringMatching(/^[a-z0-9]{1,10}$/);
|
||||
fc.assert(
|
||||
fc.property(
|
||||
fc.uniqueArray(fc.tuple(hostArb, valueArb), {
|
||||
minLength: 2, maxLength: 6, selector: ([h]) => h,
|
||||
}),
|
||||
fc.nat(),
|
||||
(sections, pick) => {
|
||||
const doc = sections
|
||||
.map(([host, value]) => `## ${host}\n\n| axis | value |\n| f | ${value} |\n`)
|
||||
.join('\n');
|
||||
const [host, value] = sections[pick % sections.length];
|
||||
// Exactly the requested section's value, never a neighbor's.
|
||||
assert.equal(matrixField(doc, host, 'f'), value);
|
||||
// A longer name that merely EXTENDS a real heading resolves to nothing.
|
||||
// `_` is outside hostArb's alphabet, so this probe can NEVER collide with
|
||||
// another generated section — the `-local` form could, and did.
|
||||
assert.equal(matrixField(doc, `${host}_x`, 'f'), null);
|
||||
},
|
||||
),
|
||||
{ numRuns: 200 },
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#3177: debug.md dispatches its session manager in the foreground', () => {
|
||||
const DEBUG_WF = path.join(ROOT, 'gsd-core', 'workflows', 'debug.md');
|
||||
const debugText = () => fs.readFileSync(DEBUG_WF, 'utf8');
|
||||
|
||||
/**
|
||||
* Every fenced `Agent( … )` block in debug.md that dispatches the session manager.
|
||||
*
|
||||
* Anchored on a line that is exactly `Agent(` so the PROSE mention of
|
||||
* `Agent(subagent_type="gsd-debug-session-manager", …)` inside the blockquote at
|
||||
* :206 is not mistaken for a dispatch. Positional selection (first match wins) was
|
||||
* the original bug here: it silently checked the continue path while the
|
||||
* new-session path went unexamined.
|
||||
*/
|
||||
function sessionManagerDispatches(text) {
|
||||
const blocks = [];
|
||||
for (const m of text.matchAll(/^Agent\($/gm)) {
|
||||
const close = text.indexOf('\n)', m.index);
|
||||
if (close === -1) continue;
|
||||
const block = text.slice(m.index, close);
|
||||
if (block.includes('subagent_type="gsd-debug-session-manager"')) blocks.push(block);
|
||||
}
|
||||
return blocks;
|
||||
}
|
||||
|
||||
test('every session-manager spawn carries the run_in_background: false opt-out', () => {
|
||||
// #2196 required this dispatch be foreground and blocking so the orchestrator
|
||||
// receives the session summary inline; debug.md still says "Wait for it; do not
|
||||
// background it" and "Display the compact summary returned by the session
|
||||
// manager". Claude Code backgrounds subagents by DEFAULT, so that intent only
|
||||
// holds if each call states the opt-out explicitly — prose alone silently
|
||||
// reinstated the exact lost-handoff failure #2196 was filed to fix.
|
||||
const blocks = sessionManagerDispatches(debugText());
|
||||
assert.equal(
|
||||
blocks.length, 2,
|
||||
'debug.md dispatches the session manager on BOTH the new-session and continue paths; '
|
||||
+ 'a change to that count means a dispatch was added or removed and must be re-checked.',
|
||||
);
|
||||
for (const block of blocks) {
|
||||
assert.match(
|
||||
block, /run_in_background\s*=\s*false/,
|
||||
'every gsd-debug-session-manager dispatch must pass run_in_background=false — without '
|
||||
+ 'it Claude Code backgrounds the spawn and the compact summary never returns (#2196).',
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test('debug.md does not assert the spawn is inherently foreground', () => {
|
||||
// The old premise ("is FOREGROUND and BLOCKING") was a property claim about the
|
||||
// host, not an instruction — and it was false for the same reason as #3177.
|
||||
assert.ok(
|
||||
!debugText().includes('is FOREGROUND and BLOCKING'),
|
||||
'debug.md must not claim the Agent() call is inherently foreground; it must name the '
|
||||
+ 'run_in_background: false opt-out that actually makes it so.',
|
||||
);
|
||||
});
|
||||
});
|
||||
@@ -948,3 +948,211 @@ describe('bug #925: critical threshold warning also uses correct hookEventName',
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
|
||||
// ────────────────────────────────────────────────────────────────────────
|
||||
// Folded from tests/fix-2289-context-monitor-event-allowlist.test.cjs — H3 test-hygiene (#3315/#3334)
|
||||
//
|
||||
// Dropped as exact duplicates already covered by the "folded:bug-925-context-
|
||||
// monitor-hook-event-name" section above:
|
||||
// - "missing hook_event_name (no Gemini) at 30% → empty stdout" (dupe of
|
||||
// "absent hook_event_name (non-Gemini) is now silent (#2289)")
|
||||
// - "empty-string hook_event_name (no Gemini) at 30% → empty stdout" (this
|
||||
// test actually used a whitespace-only event name ' '; dupe of
|
||||
// "whitespace-only hook_event_name (non-Gemini) is now silent (#2289)")
|
||||
// - "missing event name WITH Gemini env at 30% → AfterTool envelope
|
||||
// (fallback preserved)" (dupe of "falls back to \"AfterTool\" when
|
||||
// hook_event_name is absent and GEMINI_API_KEY is set")
|
||||
// ────────────────────────────────────────────────────────────────────────
|
||||
{
|
||||
const { describe: __foldDescribe } = require('node:test');
|
||||
__foldDescribe("folded:fix-2289-context-monitor-event-allowlist (#3315/#3334)", () => {
|
||||
/**
|
||||
* #2289 — gsd-context-monitor lifecycle-event output allowlist.
|
||||
*
|
||||
* The context monitor emits a `hookSpecificOutput.additionalContext` envelope
|
||||
* to inject context warnings. That shape is only valid for the context-injection
|
||||
* events (PostToolUse, and AfterTool for the Gemini dialect). Codex also wires
|
||||
* this hook to Stop / SubagentStart / SubagentStop / PreCompact (#772), and
|
||||
* Codex's Stop schema REJECTS the envelope ("hook returned invalid stop hook
|
||||
* JSON output"). The fix uses a positive allowlist: emit only for
|
||||
* injection-capable events; every other event — and a missing/unknown name —
|
||||
* exits 0 with NO stdout, while side effects (debounce, critical-session
|
||||
* recording) still run.
|
||||
*
|
||||
* These tests drive the real hook script end-to-end (spawn + stdin + a fresh
|
||||
* metrics bridge file), asserting behavior, not source text.
|
||||
*/
|
||||
|
||||
'use strict';
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const os = require('node:os');
|
||||
const path = require('node:path');
|
||||
const { execFileSync } = require('node:child_process');
|
||||
|
||||
const HOOK_PATH = path.join(__dirname, '..', 'hooks', 'gsd-context-monitor.js');
|
||||
|
||||
// Run the monitor with a synthetic, fresh metrics bridge file.
|
||||
// Returns { stdout, warnData } and cleans up the bridge + sentinel files.
|
||||
// opts: { event, remaining, used = 80, gemini = false, gsdActive = false }
|
||||
function runMonitor(opts) {
|
||||
const {
|
||||
event,
|
||||
remaining,
|
||||
used = 80,
|
||||
gemini = false,
|
||||
gsdActive = false,
|
||||
} = opts;
|
||||
|
||||
const sessionId = `fix-2289-${Date.now()}-${Math.random().toString(36).slice(2)}`;
|
||||
const tmpDir = os.tmpdir();
|
||||
const metricsPath = path.join(tmpDir, `claude-ctx-${sessionId}.json`);
|
||||
const warnPath = path.join(tmpDir, `claude-ctx-${sessionId}-warned.json`);
|
||||
|
||||
// Fresh (non-stale) metrics: timestamp is "now" in seconds.
|
||||
fs.writeFileSync(metricsPath, JSON.stringify({
|
||||
timestamp: Math.floor(Date.now() / 1000),
|
||||
remaining_percentage: remaining,
|
||||
used_pct: used,
|
||||
}));
|
||||
|
||||
// Optional GSD-active project dir (STATE.md present) so the critical-session
|
||||
// recording side effect is reachable.
|
||||
let cwd = tmpDir;
|
||||
let projDir = null;
|
||||
if (gsdActive) {
|
||||
projDir = fs.mkdtempSync(path.join(tmpDir, 'fix-2289-proj-'));
|
||||
fs.mkdirSync(path.join(projDir, '.planning'), { recursive: true });
|
||||
fs.writeFileSync(path.join(projDir, '.planning', 'STATE.md'), '# State\n');
|
||||
cwd = projDir;
|
||||
}
|
||||
|
||||
const payload = { session_id: sessionId, cwd };
|
||||
if (event !== undefined) payload.hook_event_name = event;
|
||||
|
||||
const env = { ...process.env };
|
||||
if (gemini) env.GEMINI_API_KEY = 'test-key';
|
||||
else delete env.GEMINI_API_KEY;
|
||||
|
||||
let stdout = '';
|
||||
try {
|
||||
stdout = execFileSync(process.execPath, [HOOK_PATH], {
|
||||
input: JSON.stringify(payload),
|
||||
env,
|
||||
encoding: 'utf8',
|
||||
timeout: 8000,
|
||||
});
|
||||
} catch (e) {
|
||||
stdout = e.stdout || '';
|
||||
}
|
||||
|
||||
let warnData = null;
|
||||
try {
|
||||
warnData = JSON.parse(fs.readFileSync(warnPath, 'utf8'));
|
||||
} catch { /* sentinel may not exist */ }
|
||||
|
||||
// Cleanup
|
||||
for (const p of [metricsPath, warnPath]) {
|
||||
try { fs.unlinkSync(p); } catch { /* ignore */ }
|
||||
}
|
||||
if (projDir) {
|
||||
// Retry-tolerant teardown: the critical path fires a detached, unref()'d
|
||||
// `state record-session` grandchild against projDir, and execFileSync does
|
||||
// not wait for it. maxRetries/retryDelay absorbs the transient
|
||||
// EBUSY/ENOTEMPTY window while that process exits, so cleanup can neither
|
||||
// flake nor leak the temp dir (mirrors tests/helpers.cjs cleanup(); see the
|
||||
// #2289 review and the prior fix in perf-317-context-monitor-fs.test.cjs).
|
||||
// eslint-disable-next-line local/no-raw-rmsync-in-tests -- test fixture teardown of a unique mkdtemp dir
|
||||
try { fs.rmSync(projDir, { recursive: true, force: true, maxRetries: 20, retryDelay: 100 }); } catch { /* ignore */ }
|
||||
}
|
||||
|
||||
return { stdout, warnData };
|
||||
}
|
||||
|
||||
describe('#2289 context-monitor: non-injection events exit silently', () => {
|
||||
// Boundary coverage around WARNING (35) and CRITICAL (25) — Stop must stay
|
||||
// silent at limit-1 / limit / limit+1 for BOTH thresholds.
|
||||
for (const remaining of [40, 36, 35, 34, 26, 25, 24, 20]) {
|
||||
test(`Stop event at remaining=${remaining}% → exit 0, empty stdout`, () => {
|
||||
const { stdout } = runMonitor({ event: 'Stop', remaining });
|
||||
assert.strictEqual(stdout, '', `Stop must emit nothing at remaining=${remaining}% (Codex rejects the envelope)`);
|
||||
});
|
||||
}
|
||||
|
||||
for (const event of ['SubagentStart', 'SubagentStop', 'PreCompact', 'SessionStart', 'BeforeTool']) {
|
||||
test(`unknown/non-injection event "${event}" at 30% → empty stdout`, () => {
|
||||
const { stdout } = runMonitor({ event, remaining: 30 });
|
||||
assert.strictEqual(stdout, '', `${event} is not injection-capable and must emit nothing`);
|
||||
});
|
||||
}
|
||||
});
|
||||
|
||||
describe('#2289 context-monitor: injection events still warn (unchanged)', () => {
|
||||
test('PostToolUse at 30% → WARNING envelope with hookEventName PostToolUse', () => {
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 30, used: 70 });
|
||||
assert.notStrictEqual(stdout, '', 'PostToolUse must still emit a warning envelope');
|
||||
const parsed = JSON.parse(stdout);
|
||||
assert.strictEqual(parsed.hookSpecificOutput.hookEventName, 'PostToolUse');
|
||||
assert.match(parsed.hookSpecificOutput.additionalContext, /CONTEXT WARNING/);
|
||||
});
|
||||
|
||||
test('PostToolUse at 20% → CRITICAL envelope', () => {
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 20, used: 80 });
|
||||
const parsed = JSON.parse(stdout);
|
||||
assert.strictEqual(parsed.hookSpecificOutput.hookEventName, 'PostToolUse');
|
||||
assert.match(parsed.hookSpecificOutput.additionalContext, /CONTEXT CRITICAL/);
|
||||
});
|
||||
|
||||
test('AfterTool at 30% → WARNING envelope with hookEventName AfterTool', () => {
|
||||
const { stdout } = runMonitor({ event: 'AfterTool', remaining: 30 });
|
||||
const parsed = JSON.parse(stdout);
|
||||
assert.strictEqual(parsed.hookSpecificOutput.hookEventName, 'AfterTool');
|
||||
assert.match(parsed.hookSpecificOutput.additionalContext, /CONTEXT WARNING/);
|
||||
});
|
||||
|
||||
test('explicit PostToolUse WITH Gemini env → explicit name wins over the AfterTool fallback', () => {
|
||||
// Precedence guard: the Gemini fallback only applies to a MISSING name; an
|
||||
// explicit PostToolUse must still report as PostToolUse even under GEMINI_API_KEY.
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 30, gemini: true });
|
||||
const parsed = JSON.parse(stdout);
|
||||
assert.strictEqual(parsed.hookSpecificOutput.hookEventName, 'PostToolUse');
|
||||
assert.match(parsed.hookSpecificOutput.additionalContext, /CONTEXT WARNING/);
|
||||
});
|
||||
|
||||
// Threshold boundaries on the emit path: 36 = no warn, 35 = warn, 25 = critical, 26 = warn.
|
||||
test('PostToolUse at 36% (above WARNING) → empty stdout', () => {
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 36 });
|
||||
assert.strictEqual(stdout, '', 'no warning above the 35% threshold');
|
||||
});
|
||||
|
||||
test('PostToolUse at 35% (WARNING boundary) → WARNING envelope', () => {
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 35 });
|
||||
assert.match(JSON.parse(stdout).hookSpecificOutput.additionalContext, /CONTEXT WARNING/);
|
||||
});
|
||||
|
||||
test('PostToolUse at 25% (CRITICAL boundary) → CRITICAL envelope', () => {
|
||||
const { stdout } = runMonitor({ event: 'PostToolUse', remaining: 25 });
|
||||
assert.match(JSON.parse(stdout).hookSpecificOutput.additionalContext, /CONTEXT CRITICAL/);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2289 context-monitor: side effects still fire on silent events (no output ≠ no side effect)', () => {
|
||||
test('Stop at 30% still writes the debounce sentinel (bookkeeping runs)', () => {
|
||||
const { stdout, warnData } = runMonitor({ event: 'Stop', remaining: 30 });
|
||||
assert.strictEqual(stdout, '', 'Stop emits nothing');
|
||||
assert.ok(warnData, 'the debounce sentinel must still be written on a silenced Stop event');
|
||||
assert.strictEqual(warnData.lastLevel, 'warning', 'debounce level bookkeeping runs regardless of output');
|
||||
});
|
||||
|
||||
test('Stop at 20% in a GSD project still records the critical-session sentinel', () => {
|
||||
const { stdout, warnData } = runMonitor({ event: 'Stop', remaining: 20, used: 80, gsdActive: true });
|
||||
assert.strictEqual(stdout, '', 'Stop emits nothing even at critical context');
|
||||
assert.ok(warnData, 'sentinel must be written');
|
||||
assert.strictEqual(warnData.criticalRecorded, true, 'critical-session recording side effect fires on the silent Stop event');
|
||||
});
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
@@ -15,6 +15,11 @@
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fc = require('fast-check');
|
||||
const fs = require('node:fs');
|
||||
const os = require('node:os');
|
||||
const path = require('node:path');
|
||||
|
||||
const { cleanup } = require('./helpers.cjs');
|
||||
|
||||
const {
|
||||
REVIEWER_LANES,
|
||||
@@ -29,6 +34,7 @@ const {
|
||||
|
||||
const RUN = '/run';
|
||||
const ROOT = '/repo';
|
||||
const REVIEW_MD = path.join(__dirname, '..', 'gsd-core', 'workflows', 'review.md');
|
||||
|
||||
/** Deterministic property runs — pinned seed, bounded, replay printed on failure. */
|
||||
const FC = { seed: 42, numRuns: 200 };
|
||||
@@ -450,3 +456,119 @@ describe('reviewer lane invocation — properties', () => {
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// Folded in from tests/fix-2358-review-temp-path-scoping.test.cjs (#3334 H3 test-hygiene fold-in).
|
||||
//
|
||||
// #2358 — review.md (and ship.md's external peer-review step) wrote every temp file to a
|
||||
// hardcoded, phase-number-only path under /tmp, which two GSD projects sharing a phase number
|
||||
// could collide on. The fix threads a single run-scoped `mktemp -d` directory through every
|
||||
// review.md temp path.
|
||||
describe('#2358 review.md temp paths are run-scoped, not phase-only', () => {
|
||||
const reviewMdContent = fs.readFileSync(REVIEW_MD, 'utf-8');
|
||||
|
||||
test('no bare, unscoped /tmp/gsd-review-* path remains', () => {
|
||||
assert.ok(
|
||||
!reviewMdContent.includes('/tmp/gsd-review'),
|
||||
'review.md must not contain any hardcoded /tmp/gsd-review* literal — ' +
|
||||
'every review temp path must be rooted under the run-scoped mktemp directory'
|
||||
);
|
||||
});
|
||||
|
||||
test('creates exactly one run-scoped directory via the portable ${TMPDIR:-/tmp} seam', () => {
|
||||
const mktempAssignments = reviewMdContent.match(/RUN_DIR=\$\(mktemp -d "\$\{TMPDIR:-\/tmp\}\/gsd-review-XXXXXX"\)/g) || [];
|
||||
assert.equal(
|
||||
mktempAssignments.length, 1,
|
||||
'review.md must create the run directory with exactly one `mktemp -d "${TMPDIR:-/tmp}/gsd-review-XXXXXX"` — ' +
|
||||
'a hardcoded /tmp (no ${TMPDIR:-/tmp} seam) breaks on Windows, and re-mktemp-ing per block would break the ' +
|
||||
'write/read pairing between build_prompt and the local-reviewer budget-trimming reads'
|
||||
);
|
||||
});
|
||||
|
||||
test('every downstream temp path is threaded through {run_dir} / $RUN_DIR, not re-derived from {phase}', () => {
|
||||
assert.ok(
|
||||
/\{run_dir\}\/gsd-review-/.test(reviewMdContent),
|
||||
'reviewer blocks must reference {run_dir}/gsd-review-... (the run-scoped placeholder)'
|
||||
);
|
||||
assert.ok(
|
||||
/\$\{RUN_DIR\}\/gsd-review-/.test(reviewMdContent),
|
||||
'the build_prompt section-file writes must reference ${RUN_DIR}/gsd-review-... (the run-scoped shell var)'
|
||||
);
|
||||
// The old isolation key must be gone entirely from path construction.
|
||||
assert.ok(
|
||||
!/\/tmp\/gsd-review[^\r\n]*\{phase\}/.test(reviewMdContent),
|
||||
'no temp path may still be keyed on a bare {phase} placeholder'
|
||||
);
|
||||
assert.ok(
|
||||
!/\$\{PHASE\}-(?:instructions|roadmap|plan|project|context|research|requirements)\.md/.test(reviewMdContent),
|
||||
'no temp path may still be keyed on the ${PHASE} shell var'
|
||||
);
|
||||
});
|
||||
|
||||
// Phase 5b (#2799) moved these strings out of review.md's bash and into the resolver and the
|
||||
// antigravity handler, so the assertions follow them. The invariant is unchanged and is what
|
||||
// #2358 was about: every reviewer artifact must live under the run-scoped mktemp directory, never
|
||||
// a bare `{phase}`-keyed /tmp path that a concurrent review could collide with.
|
||||
test('every lane anchors its prompt and artifacts under the run dir', () => {
|
||||
const RUN_2358 = '/run-scoped';
|
||||
for (const lane of REVIEWER_LANES) {
|
||||
const r = resolveLanePlan({
|
||||
lane, configGet: () => undefined, runDir: RUN_2358, repoRoot: '/repo',
|
||||
});
|
||||
assert.equal(r.ok, true, `${lane.slug} failed to resolve`);
|
||||
const p = r.plan;
|
||||
assert.ok(p.reviewPath.startsWith(`${RUN_2358}/`), `${lane.slug} review path escapes the run dir`);
|
||||
assert.ok(p.errPath.startsWith(`${RUN_2358}/`), `${lane.slug} err path escapes the run dir`);
|
||||
assert.ok(p.promptPath.startsWith(`${RUN_2358}/`), `${lane.slug} prompt path escapes the run dir`);
|
||||
}
|
||||
});
|
||||
|
||||
test('the argv-borne prompt instruction references the run-scoped path', () => {
|
||||
const RUN_2358 = '/run-scoped';
|
||||
const fileRefLanes = REVIEWER_LANES.filter(
|
||||
(l) => l.transport === 'spawn' && l.invoke.promptChannel === 'argv-file-ref',
|
||||
);
|
||||
assert.ok(fileRefLanes.length > 0, 'expected at least one argv-file-ref lane');
|
||||
for (const lane of fileRefLanes) {
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: RUN_2358, repoRoot: '/repo' });
|
||||
const arg = r.plan.argv[r.plan.argv.length - 1];
|
||||
assert.ok(arg.includes(`${RUN_2358}/gsd-review-prompt.md`), `${lane.slug} prompt not run-scoped`);
|
||||
}
|
||||
});
|
||||
|
||||
test('an instance writes under the run dir, keyed by its own identity', () => {
|
||||
// Two instances of one adapter must not overwrite each other, and neither may escape the run
|
||||
// dir — the identity is sanitized to a flat filename.
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === 'opencode');
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: '/run-scoped', repoRoot: '/repo' });
|
||||
assert.ok(r.plan.reviewPath.startsWith('/run-scoped/'));
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2358 design principle: run-scoped temp dirs never collide across projects/phases', () => {
|
||||
// review.md and ship.md are markdown instructions an AI agent executes, not
|
||||
// node-executable code, so this does not shell out to the literal snippet —
|
||||
// it validates the underlying guarantee the fix relies on (mktemp-style
|
||||
// randomized-suffix isolation) using Node's built-in equivalent, which is
|
||||
// cross-platform (Windows included) unlike shelling out to `mktemp`/bash.
|
||||
test('two runs — even for the same phase number, same or different project — get distinct run dirs', (t) => {
|
||||
const prefix = path.join(os.tmpdir(), 'gsd-review-');
|
||||
const runDirA = fs.mkdtempSync(prefix);
|
||||
const runDirB = fs.mkdtempSync(prefix);
|
||||
// helpers.cleanup (not raw fs.rmSync) carries the Windows-EBUSY retry budget.
|
||||
t.after(() => {
|
||||
cleanup(runDirA);
|
||||
cleanup(runDirB);
|
||||
});
|
||||
assert.notEqual(
|
||||
runDirA, runDirB,
|
||||
'two review runs sharing the same phase number must never resolve to the same run-scoped directory'
|
||||
);
|
||||
const phase = '10'; // same phase number in both "projects" — the historical collision case
|
||||
const staleProjectAPath = path.join(runDirA, `gsd-review-prompt.md`);
|
||||
const laterProjectBPath = path.join(runDirB, `gsd-review-prompt.md`);
|
||||
assert.notEqual(
|
||||
staleProjectAPath, laterProjectBPath,
|
||||
`phase ${phase} in two different runs must not resolve to the same prompt path`
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -704,3 +704,307 @@ describe('runner — #3086: spawn errorCode surfaced in err file', () => {
|
||||
'legitimate stderr should still be written');
|
||||
});
|
||||
});
|
||||
|
||||
// Folded from tests/fix-2494-review-claude-gemini-empty-guard.test.cjs (#3334/H3).
|
||||
//
|
||||
// #2494 — a failed reviewer lane must be diagnosable, never a silent drop. Before the fix, gemini
|
||||
// and claude sent stderr to `/dev/null` and wrote nothing on failure. A failed lane — CLI missing,
|
||||
// unauthenticated, rate-limited, crashed, any exit that writes no stdout — left a zero-byte file
|
||||
// that `write_reviews` rendered as "a reviewer that ran cleanly with nothing to report", silently
|
||||
// dropping a lane from the cross-AI consensus while `present_results` reported success. The policy
|
||||
// is uniform across every lane now rather than fixed per-leg, so these assertions run over the
|
||||
// whole spawn roster instead of the two legs the issue named.
|
||||
describe('#2494 — a failed lane writes a diagnosable stub, not a zero-byte file', () => {
|
||||
/** Lanes whose empty-output policy is the shared stub (antigravity owns its own diagnostics). */
|
||||
const STUB_LANES = REVIEWER_LANES.filter(
|
||||
(l) => l.transport === 'spawn' && l.emptyOutput === 'stub-with-stderr',
|
||||
);
|
||||
|
||||
function planFor(slug) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true, `${slug} failed to resolve`);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
function deps(spawnResult, files = {}) {
|
||||
return {
|
||||
files,
|
||||
// `kimi-code` declares a `command-capability` probe, so the runner spawns `--help` BEFORE the
|
||||
// review. Answer that separately or the probe fails and the lane never reaches the invocation
|
||||
// this test is about.
|
||||
spawn: (binary, argv) =>
|
||||
argv && argv.length === 1 && argv[0] === '--help'
|
||||
? { status: 0, stdout: '--output-format', stderr: '' }
|
||||
: spawnResult,
|
||||
httpJson: async () => ({ ok: true, status: 200, body: '{}' }),
|
||||
readFile: (p) => { if (!(p in files)) throw new Error(`ENOENT ${p}`); return files[p]; },
|
||||
writeFile: (p, c) => { files[p] = c; },
|
||||
exists: (p) => p in files,
|
||||
hasBinary: () => true,
|
||||
configGet: () => undefined,
|
||||
homeDir: '/home/u',
|
||||
warn: () => {},
|
||||
};
|
||||
}
|
||||
|
||||
for (const lane of STUB_LANES) {
|
||||
test(`${lane.slug}: a lane that exits non-zero with no stdout is stubbed`, async () => {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps({ status: 127, stdout: '', stderr: 'command not found' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
|
||||
assert.equal(r.stubbed, true, 'a failed lane must be reported as stubbed');
|
||||
const review = d.files[p.reviewPath];
|
||||
assert.ok(review !== undefined, 'a review file must exist after a failed lane');
|
||||
assert.notStrictEqual(review.trim(), '', 'the review file must not be empty');
|
||||
assert.ok(
|
||||
review.includes('failed or returned empty output'),
|
||||
'the stub must be distinguishable from a real review',
|
||||
);
|
||||
});
|
||||
|
||||
test(`${lane.slug}: stderr is captured to a .err sidecar, never discarded`, async () => {
|
||||
// The sidecar is the difference between "this lane failed" and "this lane failed BECAUSE…".
|
||||
// Without it every failure mode looks identical to every other.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps({ status: 1, stdout: '', stderr: 'HTTP 429 rate limited' });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
|
||||
assert.equal(d.files[p.errPath], 'HTTP 429 rate limited', 'stderr must reach the sidecar');
|
||||
assert.ok(
|
||||
d.files[p.reviewPath].includes('HTTP 429 rate limited'),
|
||||
'and must be surfaced in the stub, where a reader will actually see it',
|
||||
);
|
||||
assert.ok(p.reviewPath.endsWith('.md'), 'review output path unchanged');
|
||||
});
|
||||
}
|
||||
|
||||
test('a successful review passes through untouched', async () => {
|
||||
const p = planFor('gemini');
|
||||
const d = deps({ status: 0, stdout: 'Looks good.\n', stderr: '' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('Looks good.'));
|
||||
assert.ok(
|
||||
!d.files[p.reviewPath].includes('failed or returned empty output'),
|
||||
'a real review must never carry the failure header',
|
||||
);
|
||||
});
|
||||
|
||||
test('no lane sends stderr to /dev/null — the sidecar is unconditional', async () => {
|
||||
// The original defect in one line, asserted over the whole roster rather than the two legs the
|
||||
// issue named: the policy is uniform now, and a future lane must not be able to opt out.
|
||||
for (const lane of STUB_LANES) {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps({ status: 0, stdout: 'ok', stderr: 'a warning' });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(d.files[p.errPath], 'a warning', `${lane.slug} discarded stderr`);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
// Folded from tests/fix-2605-review-local-server-empty-guard.test.cjs (#3334/H3).
|
||||
//
|
||||
// #2605 — the local OpenAI-compatible lanes (ollama / lm_studio / llama.cpp) dropped silently. The
|
||||
// original defects, all of which made a failed lane indistinguishable from a clean empty review:
|
||||
// bare `curl -s` suppressed curl's own error text; the response was piped straight into `jq` so the
|
||||
// BODY — where an OpenAI-compatible server puts its error JSON on an HTTP 4xx/5xx while curl still
|
||||
// exits 0 — was discarded unread; nothing was written when content was empty, so the file never
|
||||
// existed and `write_reviews` omitted the section entirely; and a whitespace-only reply passed the
|
||||
// byte-counting `[ ! -s … ]` guard as a successful review.
|
||||
describe('#2605 local OpenAI-compatible lanes produce diagnosable output', () => {
|
||||
const HTTP_LANES = REVIEWER_LANES.filter((l) => l.transport === 'openai-http');
|
||||
|
||||
function planFor(slug, config = {}) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({ lane, configGet: (k) => config[k], runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
/**
|
||||
* These lanes declare an `http-reachable` probe, so the runner performs a GET on /v1/models BEFORE
|
||||
* the chat call. The stub must answer that separately — otherwise the lane is reported unreachable
|
||||
* and never reaches the invocation these tests are actually about.
|
||||
*/
|
||||
function reachableThen(chatResponse) {
|
||||
return async (url, opts) =>
|
||||
opts.method === 'GET'
|
||||
? { ok: true, status: 200, body: JSON.stringify({ data: [{ id: 'stub-model' }] }) }
|
||||
: (typeof chatResponse === 'function' ? chatResponse(url, opts) : chatResponse);
|
||||
}
|
||||
|
||||
function deps(httpJson, files = { [`${RUN}/gsd-review-prompt.md`]: 'PLAN' }) {
|
||||
const warnings = [];
|
||||
return {
|
||||
files,
|
||||
warnings,
|
||||
spawn: () => ({ status: 0, stdout: '', stderr: '' }),
|
||||
httpJson,
|
||||
readFile: (p) => { if (!(p in files)) throw new Error(`ENOENT ${p}`); return files[p]; },
|
||||
writeFile: (p, c) => { files[p] = c; },
|
||||
exists: (p) => p in files,
|
||||
hasBinary: () => true,
|
||||
configGet: () => undefined,
|
||||
homeDir: '/home/u',
|
||||
warn: (m) => warnings.push(m),
|
||||
};
|
||||
}
|
||||
|
||||
const okBody = (content) => ({
|
||||
ok: true, status: 200, body: JSON.stringify({ choices: [{ message: { content } }] }),
|
||||
});
|
||||
|
||||
for (const lane of HTTP_LANES) {
|
||||
test(`${lane.slug}: an unreachable endpoint produces a stub carrying the transport error`, async () => {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen({ ok: false, status: 0, body: '', error: 'ECONNREFUSED' }));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
assert.ok(d.files[p.reviewPath].includes('ECONNREFUSED'),
|
||||
'the transport error must be visible — bare `curl -s` used to swallow it');
|
||||
});
|
||||
|
||||
test(`${lane.slug}: an HTTP error body is preserved in the stub`, async () => {
|
||||
// The body is the ONLY evidence on a 4xx/5xx: such a server returns its error JSON there and
|
||||
// curl still exits 0, so stderr is empty. The old pipe into jq discarded it.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen({ ok: false, status: 404, body: '{"error":"model not found"}' }));
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.ok(d.files[p.reviewPath].includes('Raw response body:'));
|
||||
assert.ok(d.files[p.reviewPath].includes('model not found'));
|
||||
});
|
||||
|
||||
test(`${lane.slug}: an empty 200 response still produces a file`, async () => {
|
||||
// Previously nothing was written, so the file never existed, write_reviews omitted the
|
||||
// section, and the result was indistinguishable from the reviewer never being selected.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody('')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
assert.ok(d.files[p.reviewPath] !== undefined, 'a file must exist even on an empty reply');
|
||||
assert.ok(d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
|
||||
test(`${lane.slug}: a whitespace-only response is empty, not a successful review`, async () => {
|
||||
// `[ ! -s … ]` counted BYTES, so " " passed as a real review.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody(' \n\t ')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
});
|
||||
|
||||
test(`${lane.slug}: a reply that is exactly an echo option is NOT misclassified`, async () => {
|
||||
// `echo "$VAR"` would write 0 bytes for `-n`/`-e`/`-E`. Nothing here goes through echo, so
|
||||
// this is structurally impossible now — locked anyway.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody('-n')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('-n'));
|
||||
assert.ok(!d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
|
||||
test(`${lane.slug}: a successful review passes through untouched`, async () => {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody('## Findings\nreal review')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('## Findings'));
|
||||
assert.ok(!d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
}
|
||||
|
||||
// NOTE: 'a served-model mismatch is warned about, not silently accepted' (originally in
|
||||
// fix-2605-review-local-server-empty-guard.test.cjs) was DROPPED here as a genuine duplicate of
|
||||
// "runner — openai-compatible handler > a served-model mismatch warns without failing the review"
|
||||
// above: same lane (lm_studio), same call path (runOpenAiCompatible directly), same assertion
|
||||
// shape (review content + a warning naming both the served and asked-for model), differing only
|
||||
// in the literal string labels used for the mismatched model names.
|
||||
|
||||
test('neither jq nor curl is required by any of these lanes', () => {
|
||||
// The dependency is gone, not merely satisfied: parsing is JSON.parse and the request is
|
||||
// in-process. `jq` is absent on stock Windows/Git-Bash (#2589), which gated these lanes.
|
||||
for (const lane of HTTP_LANES) {
|
||||
assert.deepStrictEqual([...lane.requiresBinaries], [], `${lane.slug} still declares a binary`);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
// Folded from tests/fix-2794-review-qwen-empty-guard.test.cjs (#3334/H3).
|
||||
//
|
||||
// #2794 — the qwen reviewer leg was the last one still sending stderr to /dev/null. Every other lane
|
||||
// captured stderr to a `.err` sidecar and appended it to the stub (#2494/#2605); qwen wrote a bare
|
||||
// "failed or returned empty output." with no diagnostic at all, so a missing binary, an auth prompt
|
||||
// and a rate-limit were indistinguishable from each other AND from a clean empty review.
|
||||
describe('#2794 qwen reviewer stderr capture', () => {
|
||||
function planFor(slug) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
function deps(spawnResult, files = {}, hasBinary = () => true) {
|
||||
return {
|
||||
files,
|
||||
spawn: () => spawnResult,
|
||||
httpJson: async () => ({ ok: true, status: 200, body: '{}' }),
|
||||
readFile: (p) => { if (!(p in files)) throw new Error(`ENOENT ${p}`); return files[p]; },
|
||||
writeFile: (p, c) => { files[p] = c; },
|
||||
exists: (p) => p in files,
|
||||
hasBinary,
|
||||
configGet: () => undefined,
|
||||
homeDir: '/home/u',
|
||||
warn: () => {},
|
||||
};
|
||||
}
|
||||
|
||||
test('writes the review on success', async () => {
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 0, stdout: '## Qwen findings\nall good\n', stderr: '' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('## Qwen findings'));
|
||||
});
|
||||
|
||||
test('a failed lane surfaces its stderr in the review stub', async () => {
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 1, stdout: '', stderr: 'auth required: run `qwen login`' });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.ok(d.files[p.reviewPath].includes('auth required'),
|
||||
'the diagnostic must reach the review, not just the sidecar');
|
||||
assert.equal(d.files[p.errPath], 'auth required: run `qwen login`');
|
||||
});
|
||||
|
||||
test('a silently empty lane still produces a diagnosable stub', async () => {
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 0, stdout: '', stderr: '' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
assert.ok(d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
|
||||
test('a missing qwen binary reports unavailable rather than an empty review', async () => {
|
||||
// Stronger than the original: the lane is now reported with a TYPED reason before it is ever
|
||||
// spawned, instead of producing a stub that looked the same as every other failure.
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 0, stdout: '', stderr: '' }, {}, () => false);
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.ok, false);
|
||||
assert.equal(r.reason, 'missing_binary');
|
||||
assert.equal(d.files[p.reviewPath], undefined, 'no review file for a lane that never ran');
|
||||
});
|
||||
|
||||
test('the sidecar is structural — no lane can opt out of it', async () => {
|
||||
// The #2794 defect was one lane diverging from a convention every other lane followed. There is
|
||||
// no per-lane place to diverge any more; this asserts that directly.
|
||||
for (const lane of REVIEWER_LANES.filter((l) => l.transport === 'spawn')) {
|
||||
const p = planFor(lane.slug);
|
||||
assert.ok(p.errPath.endsWith('.err'), `${lane.slug} must declare a stderr sidecar`);
|
||||
assert.notEqual(p.errPath, '/dev/null');
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
@@ -9,13 +9,14 @@ const assert = require('node:assert/strict');
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
const fc = require('./helpers/fast-check-setup.cjs');
|
||||
const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs');
|
||||
const { runGsdTools, createTempProject, createTempDir, cleanup } = require('./helpers.cjs');
|
||||
const {
|
||||
buildCheckpoint,
|
||||
CHECKPOINT_FRAMES,
|
||||
CHECKPOINT_LANGUAGE_ALIASES,
|
||||
resolveCheckpointFrame,
|
||||
checkpointBoxLine,
|
||||
parseDeferredItems,
|
||||
} = require('../gsd-core/bin/lib/uat.cjs');
|
||||
|
||||
describe('audit-uat command', () => {
|
||||
@@ -1414,3 +1415,661 @@ awaiting: user response
|
||||
assert.strictEqual(result.output, expected);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── cmdAuditUat behavioral coverage (#2287 deferred-items.md) ─────────────
|
||||
|
||||
describe('#2287 cmdAuditUat: deferred-items.md awareness', () => {
|
||||
let tmpDir;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempProject();
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(tmpDir);
|
||||
});
|
||||
|
||||
test('no deferred-items.md present (0 entries) → no results, no false positive', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
fs.writeFileSync(path.join(phaseDir, '.gitkeep'), '');
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.deepStrictEqual(output.results, []);
|
||||
assert.strictEqual(output.summary.total_items, 0);
|
||||
assert.strictEqual(output.summary.total_files, 0);
|
||||
});
|
||||
|
||||
test('deferred-items.md with only a resolved entry (0 unresolved) → no result surfaced', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- Already handled unrelated lint warning.',
|
||||
' status: resolved',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.deepStrictEqual(output.results, [],
|
||||
'a fully-resolved deferred-items.md must not surface any result');
|
||||
assert.strictEqual(output.summary.total_items, 0);
|
||||
});
|
||||
|
||||
test('deferred-items.md with 1 unresolved entry → surfaced in structured JSON output', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- Found an unrelated pre-existing test failure in `some-other-module` while working on',
|
||||
' this phase\'s task. Out of scope for this task — logged here per SCOPE BOUNDARY.',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.strictEqual(output.summary.total_items, 1);
|
||||
assert.strictEqual(output.summary.total_files, 1);
|
||||
assert.strictEqual(output.summary.by_category.deferred, 1);
|
||||
assert.strictEqual(output.summary.by_phase['01'], 1);
|
||||
|
||||
const deferredResult = output.results.find(r => r.type === 'deferred');
|
||||
assert.ok(deferredResult, 'a deferred-typed result must be present');
|
||||
assert.strictEqual(deferredResult.phase, '01');
|
||||
assert.strictEqual(deferredResult.file, 'deferred-items.md');
|
||||
assert.strictEqual(
|
||||
deferredResult.file_path,
|
||||
'.planning/phases/01-foundation/deferred-items.md',
|
||||
);
|
||||
assert.strictEqual(deferredResult.items.length, 1);
|
||||
assert.match(deferredResult.items[0].name, /unrelated pre-existing test failure/);
|
||||
assert.strictEqual(deferredResult.items[0].result, 'unresolved');
|
||||
assert.strictEqual(deferredResult.items[0].category, 'deferred');
|
||||
});
|
||||
|
||||
test('deferred-items.md with 2+ entries (mixed resolved/unresolved) → only unresolved surfaced', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- First unrelated finding, still open.',
|
||||
'- Second unrelated finding, also still open.',
|
||||
'- Third finding, already fixed separately.',
|
||||
' status: resolved',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
const deferredResult = output.results.find(r => r.type === 'deferred');
|
||||
assert.ok(deferredResult);
|
||||
assert.strictEqual(deferredResult.items.length, 2,
|
||||
'exactly the 2 unresolved entries must surface; the resolved 3rd must not');
|
||||
const names = deferredResult.items.map(i => i.name);
|
||||
assert.ok(names.some(n => n.includes('First unrelated finding')));
|
||||
assert.ok(names.some(n => n.includes('Second unrelated finding')));
|
||||
assert.ok(!names.some(n => n.includes('Third finding')));
|
||||
});
|
||||
|
||||
test('deferred entries surface across multiple phase directories', () => {
|
||||
const phase1 = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
const phase2 = path.join(tmpDir, '.planning', 'phases', '02-auth');
|
||||
fs.mkdirSync(phase1, { recursive: true });
|
||||
fs.mkdirSync(phase2, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phase1, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- Phase 1 unrelated finding.',
|
||||
].join('\n'));
|
||||
fs.writeFileSync(path.join(phase2, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- Phase 2 unrelated finding.',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
const deferredResults = output.results.filter(r => r.type === 'deferred');
|
||||
assert.strictEqual(deferredResults.length, 2);
|
||||
assert.strictEqual(output.summary.total_items, 2);
|
||||
assert.strictEqual(output.summary.by_phase['01'], 1);
|
||||
assert.strictEqual(output.summary.by_phase['02'], 1);
|
||||
});
|
||||
|
||||
test('an entry with a garbled/missing status fails safe and is surfaced (not silently dropped)', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- An entry with no status field at all.',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.strictEqual(output.summary.total_items, 1,
|
||||
'missing status must SURFACE the entry, not silently drop it');
|
||||
});
|
||||
|
||||
test('existing UAT/VERIFICATION scanning is unchanged when a deferred-items.md is also present', () => {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, '01-UAT.md'), [
|
||||
'---',
|
||||
'status: testing',
|
||||
'phase: 01-foundation',
|
||||
'started: 2025-01-01T00:00:00Z',
|
||||
'updated: 2025-01-01T00:00:00Z',
|
||||
'---',
|
||||
'',
|
||||
'## Tests',
|
||||
'',
|
||||
'### 1. Login Form',
|
||||
'expected: Form displays with email and password fields',
|
||||
'result: pending',
|
||||
].join('\n'));
|
||||
|
||||
fs.writeFileSync(path.join(phaseDir, 'deferred-items.md'), [
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- An unrelated out-of-scope finding.',
|
||||
].join('\n'));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.strictEqual(output.results.length, 2, 'both the UAT file and deferred-items.md must surface as separate results');
|
||||
const uatResult = output.results.find(r => r.type === 'uat');
|
||||
const deferredResult = output.results.find(r => r.type === 'deferred');
|
||||
assert.ok(uatResult, 'existing uat-type result must still be present');
|
||||
assert.strictEqual(uatResult.items.length, 1);
|
||||
assert.strictEqual(uatResult.items[0].result, 'pending');
|
||||
assert.ok(deferredResult, 'new deferred-type result must be present');
|
||||
assert.strictEqual(deferredResult.items.length, 1);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── forensic_audit workflow-prose source-contract guard (#2287) ──────────
|
||||
|
||||
// #2994 fragmentization moved the --forensic-gated forensic_audit step out of
|
||||
// progress.md into gsd-core/workflows/progress/steps/forensic-audit.md behind
|
||||
// a section marker. Read that step file directly — it is the sole remaining
|
||||
// source of the forensic_audit step body these guards assert on.
|
||||
const PROGRESS_MD = path.join(__dirname, '..', 'gsd-core', 'workflows', 'progress', 'steps', 'forensic-audit.md');
|
||||
|
||||
describe('#2287 progress.md forensic_audit: deferred-items.md contract', () => {
|
||||
const content = fs.readFileSync(PROGRESS_MD, 'utf-8');
|
||||
const stepStart = content.indexOf('<step name="forensic_audit">');
|
||||
const stepEnd = content.indexOf('</step>', stepStart);
|
||||
const section = stepStart !== -1 && stepEnd !== -1 ? content.slice(stepStart, stepEnd) : '';
|
||||
|
||||
test('forensic_audit step exists', () => {
|
||||
assert.notEqual(stepStart, -1, 'progress.md (or its extracted progress/steps/forensic-audit.md) must contain the forensic_audit step');
|
||||
});
|
||||
|
||||
test('forensic_audit now runs 7 checks (was 6) and globs deferred-items.md', () => {
|
||||
assert.ok(/running 7 deep checks/i.test(section),
|
||||
'forensic_audit must advertise 7 deep checks (was 6) now that deferred-items.md is read');
|
||||
assert.ok(/\.planning\/phases\/\*\/deferred-items\.md/.test(section),
|
||||
'forensic_audit must glob .planning/phases/*/deferred-items.md');
|
||||
});
|
||||
|
||||
test('the new check reports unresolved deferred items with the same ✓/⚠ semantics as the other checks', () => {
|
||||
assert.ok(/check\s*7/i.test(section),
|
||||
'a 7th check must be present');
|
||||
assert.ok(/unresolved deferred items/i.test(section),
|
||||
'the check must be framed around unresolved deferred items');
|
||||
assert.ok(/✓[^\n]*no unresolved deferred items/i.test(section),
|
||||
'the check must emit a ✓ pass line when no unresolved deferred items exist');
|
||||
assert.ok(/⚠[^\n]*unresolved deferred items found/i.test(section),
|
||||
'the check must emit a ⚠ warning line when unresolved deferred items exist');
|
||||
});
|
||||
|
||||
test('an entry is resolved only via an explicit status: resolved field (fail-safe otherwise)', () => {
|
||||
assert.ok(/status:\s*resolved/i.test(section),
|
||||
'the resolved/unresolved parsing rule must be documented in the step prose');
|
||||
});
|
||||
|
||||
test('the verdict summary now gates on 7 checks (was 6)', () => {
|
||||
assert.ok(/after all 7 checks/i.test(section),
|
||||
'the verdict section must say "after all 7 checks"');
|
||||
assert.ok(/if all 7 checks passed/i.test(section),
|
||||
'the verdict section must say "if all 7 checks passed"');
|
||||
assert.ok(!/after all 6 checks/i.test(section) && !/if all 6 checks passed/i.test(section),
|
||||
'stale "6 checks" phrasing must not remain in the step');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── parseDeferredItems property test (#2287) ──────────────────────────────
|
||||
|
||||
describe('#2287 parseDeferredItems: property (status: resolved fail-safe)', () => {
|
||||
// Single-line entry text: no newlines (would break bullet-entry splitting),
|
||||
// non-empty after trim, and never itself SHAPED like a `status:` field line
|
||||
// (that would be indistinguishable from a real field regardless of intent).
|
||||
const plainText = fc.string({ minLength: 1, maxLength: 40 })
|
||||
.map((s) => s.replace(/[\r\n]/g, ' ').trim())
|
||||
.filter((s) => s.length > 0 && !/^status:/i.test(s));
|
||||
|
||||
// Decoy: entry text that CONTAINS a `status: resolved`-shaped substring
|
||||
// mid-line (not at line start) — must never be misread as a resolved
|
||||
// marker, since extractGapEntryFields only recognises a field anchored to
|
||||
// the START of its own trimmed line (see parseDeferredItems' doc comment).
|
||||
const decoyText = plainText.map((s) => `${s} status: resolved trailing note`);
|
||||
|
||||
const textArb = fc.oneof(plainText, decoyText);
|
||||
const entryArb = fc.record({ text: textArb, resolved: fc.boolean() });
|
||||
|
||||
test('property: an entry is surfaced iff it is NOT marked status: resolved; surfaced count == non-resolved count', () => {
|
||||
fc.assert(
|
||||
fc.property(
|
||||
fc.array(entryArb, { maxLength: 20 }),
|
||||
(rawEntries) => {
|
||||
// Index-prefix for uniqueness so surfaced items can be mapped back
|
||||
// to their source entry unambiguously even with colliding random text.
|
||||
const entries = rawEntries.map((e, i) => ({ text: `E${i}_${e.text}`, resolved: e.resolved }));
|
||||
|
||||
const lines = ['## Deferred Items', ''];
|
||||
for (const e of entries) {
|
||||
lines.push(`- ${e.text}`);
|
||||
if (e.resolved) lines.push(' status: resolved');
|
||||
}
|
||||
const content = lines.join('\n');
|
||||
|
||||
const items = parseDeferredItems(content);
|
||||
const surfacedNames = new Set(items.map((it) => it.name));
|
||||
|
||||
const expectedUnresolved = entries.filter((e) => !e.resolved);
|
||||
const expectedResolved = entries.filter((e) => e.resolved);
|
||||
|
||||
// Total surfaced count equals the count of non-resolved entries.
|
||||
assert.strictEqual(items.length, expectedUnresolved.length);
|
||||
|
||||
// Every non-resolved entry IS surfaced (including status:-shaped
|
||||
// decoy substrings embedded mid-line — those must not flip the
|
||||
// outcome).
|
||||
for (const e of expectedUnresolved) {
|
||||
assert.ok(surfacedNames.has(e.text), `expected unresolved entry to surface: ${e.text}`);
|
||||
}
|
||||
|
||||
// No status:-resolved entry is EVER surfaced.
|
||||
for (const e of expectedResolved) {
|
||||
assert.ok(!surfacedNames.has(e.text), `status: resolved entry must never surface: ${e.text}`);
|
||||
}
|
||||
|
||||
// Every returned item carries the fixed deferred category/result shape.
|
||||
for (const item of items) {
|
||||
assert.strictEqual(item.result, 'unresolved');
|
||||
assert.strictEqual(item.category, 'deferred');
|
||||
}
|
||||
}
|
||||
)
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── #2766: archived phase dirs, and GFM-table-shaped deferred/gaps ────────
|
||||
|
||||
const UAT_ONE_PENDING = [
|
||||
'---',
|
||||
'status: partial',
|
||||
'phase: 01-foundation',
|
||||
'---',
|
||||
'',
|
||||
'## Current Test',
|
||||
'',
|
||||
'[awaiting human testing]',
|
||||
'',
|
||||
'## Tests',
|
||||
'',
|
||||
'### 1. A scenario nobody ever ran',
|
||||
'expected: something observable happens',
|
||||
'result: [pending]',
|
||||
'',
|
||||
'## Summary',
|
||||
'',
|
||||
'total: 1',
|
||||
'pending: 1',
|
||||
'',
|
||||
'## Gaps',
|
||||
'',
|
||||
].join('\n');
|
||||
|
||||
/** Write a UAT file whose `## Gaps` section holds `gapsBody`. */
|
||||
function uatWithGaps(gapsBody) {
|
||||
return [
|
||||
'---',
|
||||
'status: complete',
|
||||
'phase: 50-gaps',
|
||||
'---',
|
||||
'',
|
||||
'## Current Test',
|
||||
'',
|
||||
'[testing complete]',
|
||||
'',
|
||||
'## Tests',
|
||||
'',
|
||||
'### 1. A passing scenario',
|
||||
'expected: this one is fine',
|
||||
'result: pass',
|
||||
'',
|
||||
'## Summary',
|
||||
'',
|
||||
'total: 1',
|
||||
'passed: 1',
|
||||
'',
|
||||
'## Gaps',
|
||||
'',
|
||||
gapsBody,
|
||||
'',
|
||||
].join('\n');
|
||||
}
|
||||
|
||||
// ─── Bug 1: archived phase dirs ───────────────────────────────────────────────
|
||||
|
||||
describe('#2766 cmdAuditUat: archived phase directories', () => {
|
||||
let tmpDir;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempProject();
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(tmpDir);
|
||||
});
|
||||
|
||||
test('phases ONLY in the archive → items surfaced, not a hard error', () => {
|
||||
const archiveDir = path.join(
|
||||
tmpDir, '.planning', 'milestones', 'v1.0-phases', '01-foundation',
|
||||
);
|
||||
fs.mkdirSync(archiveDir, { recursive: true });
|
||||
fs.writeFileSync(path.join(archiveDir, '01-UAT.md'), UAT_ONE_PENDING);
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.strictEqual(output.summary.total_items, 1);
|
||||
assert.strictEqual(output.results.length, 1);
|
||||
assert.strictEqual(output.results[0].phase, '01');
|
||||
assert.strictEqual(output.results[0].archived_milestone, 'v1.0');
|
||||
assert.match(output.results[0].file_path, /milestones\/v1\.0-phases\//);
|
||||
});
|
||||
|
||||
test('active and archived trees are both scanned', () => {
|
||||
const activeDir = path.join(tmpDir, '.planning', 'phases', '40-current');
|
||||
fs.mkdirSync(activeDir, { recursive: true });
|
||||
fs.writeFileSync(path.join(activeDir, '40-UAT.md'), UAT_ONE_PENDING);
|
||||
|
||||
const archiveDir = path.join(
|
||||
tmpDir, '.planning', 'milestones', 'v1.0-phases', '01-foundation',
|
||||
);
|
||||
fs.mkdirSync(archiveDir, { recursive: true });
|
||||
fs.writeFileSync(path.join(archiveDir, '01-UAT.md'), UAT_ONE_PENDING);
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
const byPhase = new Map(output.results.map(r => [r.phase, r]));
|
||||
assert.ok(byPhase.has('01'), `archived phase missing: ${JSON.stringify([...byPhase.keys()])}`);
|
||||
assert.ok(byPhase.has('40'), `active phase missing: ${JSON.stringify([...byPhase.keys()])}`);
|
||||
assert.strictEqual(byPhase.get('01').archived_milestone, 'v1.0');
|
||||
assert.strictEqual(byPhase.get('40').archived_milestone, undefined);
|
||||
});
|
||||
|
||||
test('multiple archived milestones are all scanned', () => {
|
||||
for (const [version, phase] of [['v1.0', '01-foundation'], ['v2.0', '07-later']]) {
|
||||
const dir = path.join(tmpDir, '.planning', 'milestones', `${version}-phases`, phase);
|
||||
fs.mkdirSync(dir, { recursive: true });
|
||||
fs.writeFileSync(path.join(dir, `${phase.slice(0, 2)}-UAT.md`), UAT_ONE_PENDING);
|
||||
}
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.strictEqual(output.summary.total_items, 2);
|
||||
assert.deepStrictEqual(
|
||||
output.results.map(r => r.archived_milestone).sort(),
|
||||
['v1.0', 'v2.0'],
|
||||
);
|
||||
});
|
||||
|
||||
test('an empty active phases dir still succeeds with no items (pre-existing behavior)', () => {
|
||||
// createTempProject() ships an empty `.planning/phases/`, so this is the
|
||||
// shape the existing uat.test.cjs "no UAT files" case covers — the archive
|
||||
// change must not turn it into an error.
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
|
||||
const output = JSON.parse(result.output);
|
||||
assert.deepStrictEqual(output.results, []);
|
||||
assert.strictEqual(output.summary.total_items, 0);
|
||||
});
|
||||
|
||||
test('no phases dir AND no archive still errors — no false all-clear', (t) => {
|
||||
// A bare temp dir with a .planning/ that has NO phases subdir and no
|
||||
// milestones archive — built from createTempDir rather than by deleting
|
||||
// createTempProject's phases dir, so nothing is torn down mid-test.
|
||||
const bare = createTempDir();
|
||||
t.after(() => cleanup(bare));
|
||||
|
||||
fs.mkdirSync(path.join(bare, '.planning'), { recursive: true });
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', bare);
|
||||
assert.strictEqual(result.success, false, 'expected a failure when no phases exist at all');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Bug 2: table-shaped deferred-items.md ────────────────────────────────────
|
||||
|
||||
describe('#2766 parseDeferredItems: GFM table shape', () => {
|
||||
const names = (md) => parseDeferredItems(md).map(i => i.name);
|
||||
|
||||
test('header + delimiter → header dropped, data rows surfaced', () => {
|
||||
assert.deepStrictEqual(
|
||||
names([
|
||||
'## Discovered during 01-03',
|
||||
'',
|
||||
'| Test | Failing seeds |',
|
||||
'|------|---------------|',
|
||||
'| test_a | 0, 1 |',
|
||||
'| test_b | 424242 |',
|
||||
].join('\n')),
|
||||
['test_a — 0, 1', 'test_b — 424242'],
|
||||
);
|
||||
});
|
||||
|
||||
test('later columns are preserved, not truncated to the first cell', () => {
|
||||
const [name] = names('| T | seeds |\n|---|---|\n| test_a | 0, 1, 424242 |');
|
||||
assert.match(name, /0, 1, 424242/);
|
||||
});
|
||||
|
||||
test('headerless table → every row surfaced', () => {
|
||||
assert.deepStrictEqual(
|
||||
names('| test_a | 0 |\n| test_b | 1 |'),
|
||||
['test_a — 0', 'test_b — 1'],
|
||||
);
|
||||
});
|
||||
|
||||
test('row marked resolved/done/pass is suppressed', () => {
|
||||
assert.deepStrictEqual(
|
||||
names([
|
||||
'| Test | Seeds | Status |',
|
||||
'|---|---|---|',
|
||||
'| test_open | 0 | open |',
|
||||
'| test_fixed | 1 | resolved |',
|
||||
'| test_done | 2 | DONE |',
|
||||
].join('\n')),
|
||||
['test_open — 0 — open'],
|
||||
);
|
||||
});
|
||||
|
||||
test('two prose-separated tables → each drops its own header', () => {
|
||||
assert.deepStrictEqual(
|
||||
names([
|
||||
'| T1 | x |', '|---|---|', '| one | 1 |',
|
||||
'',
|
||||
'some prose in between',
|
||||
'',
|
||||
'| T2 | y |', '|---|---|', '| two | 2 |',
|
||||
].join('\n')),
|
||||
['one — 1', 'two — 2'],
|
||||
);
|
||||
});
|
||||
|
||||
test('bullets and a table in one file → union, no double-counting', () => {
|
||||
const got = names([
|
||||
'## Deferred Items',
|
||||
'',
|
||||
'- a bullet-shaped deferred entry',
|
||||
'',
|
||||
'| Test | Seeds |',
|
||||
'|---|---|',
|
||||
'| test_a | 0 |',
|
||||
].join('\n'));
|
||||
assert.strictEqual(got.length, 2, JSON.stringify(got));
|
||||
assert.ok(got.some(n => n.includes('bullet-shaped')));
|
||||
assert.ok(got.some(n => n.startsWith('test_a')));
|
||||
});
|
||||
|
||||
test('bullet-only file unchanged (no regression on #2287)', () => {
|
||||
assert.deepStrictEqual(
|
||||
names('## Deferred Items\n\n- entry one\n- entry two\n'),
|
||||
['entry one', 'entry two'],
|
||||
);
|
||||
});
|
||||
|
||||
test('explicit status: resolved bullet still suppressed (no regression on #2287)', () => {
|
||||
const got = names(
|
||||
'## Deferred Items\n\n- truth: "closed thing"\n status: resolved\n- truth: "open thing"\n',
|
||||
);
|
||||
assert.strictEqual(got.length, 1, JSON.stringify(got));
|
||||
assert.match(got[0], /open thing/);
|
||||
});
|
||||
|
||||
test('no table and no bullets → zero items, no throw', () => {
|
||||
assert.deepStrictEqual(names('# Notes\n\njust prose, nothing actionable.\n'), []);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Bug 3: table-shaped ## Gaps section ──────────────────────────────────────
|
||||
|
||||
describe('#2766 parseGapsItems: GFM table shape', () => {
|
||||
let tmpDir;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempProject();
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(tmpDir);
|
||||
});
|
||||
|
||||
/** Run audit-uat over a phase whose UAT file has `gapsBody` as its Gaps section. */
|
||||
function gapsItems(gapsBody) {
|
||||
const phaseDir = path.join(tmpDir, '.planning', 'phases', '50-gaps');
|
||||
fs.mkdirSync(phaseDir, { recursive: true });
|
||||
fs.writeFileSync(path.join(phaseDir, '50-UAT.md'), uatWithGaps(gapsBody));
|
||||
|
||||
const result = runGsdTools('audit-uat --raw', tmpDir);
|
||||
assert.ok(result.success, `Command failed: ${result.error}`);
|
||||
const output = JSON.parse(result.output);
|
||||
const uat = output.results.find(r => r.type === 'uat');
|
||||
return uat ? uat.items : [];
|
||||
}
|
||||
|
||||
test('header-mapped table → truth/status/reason/test extracted', () => {
|
||||
const items = gapsItems([
|
||||
'| Truth | Status | Reason | Test |',
|
||||
'|-------|--------|--------|------|',
|
||||
'| Login should redirect | failed | User reported a 500 | 1 |',
|
||||
].join('\n'));
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items));
|
||||
assert.strictEqual(items[0].name, 'Login should redirect');
|
||||
assert.strictEqual(items[0].result, 'failed');
|
||||
assert.strictEqual(items[0].reason, 'User reported a 500');
|
||||
assert.strictEqual(items[0].test, 1);
|
||||
});
|
||||
|
||||
test('status: resolved row suppressed, open row kept', () => {
|
||||
const items = gapsItems([
|
||||
'| Truth | Status |',
|
||||
'|-------|--------|',
|
||||
'| closed thing | resolved |',
|
||||
'| open thing | failed |',
|
||||
].join('\n'));
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items.map(i => i.name)));
|
||||
assert.strictEqual(items[0].name, 'open thing');
|
||||
});
|
||||
|
||||
test('no status column → surfaced as unknown, not dropped', () => {
|
||||
const items = gapsItems('| Truth | Note |\n|---|---|\n| something is off | see logs |');
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items));
|
||||
assert.strictEqual(items[0].result, 'unknown');
|
||||
assert.strictEqual(items[0].name, 'something is off');
|
||||
});
|
||||
|
||||
test('unrecognizable header → joined cells + unknown status', () => {
|
||||
const items = gapsItems('| Alpha | Beta |\n|---|---|\n| xxx | yyy |');
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items));
|
||||
assert.strictEqual(items[0].result, 'unknown');
|
||||
assert.match(items[0].name, /xxx/);
|
||||
assert.match(items[0].name, /yyy/);
|
||||
});
|
||||
|
||||
test('headerless table → explicit resolved cell still suppressed', () => {
|
||||
const items = gapsItems('| open thing | failed |\n| closed thing | resolved |');
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items.map(i => i.name)));
|
||||
assert.match(items[0].name, /open thing/);
|
||||
});
|
||||
|
||||
test('bullets and a table in one Gaps section → union, no double-counting', () => {
|
||||
const items = gapsItems([
|
||||
'- truth: "a bullet gap"',
|
||||
' status: failed',
|
||||
'',
|
||||
'| Truth | Status |',
|
||||
'|---|---|',
|
||||
'| a table gap | failed |',
|
||||
].join('\n'));
|
||||
|
||||
assert.strictEqual(items.length, 2, JSON.stringify(items.map(i => i.name)));
|
||||
assert.ok(items.some(i => i.name === 'a bullet gap'));
|
||||
assert.ok(items.some(i => i.name === 'a table gap'));
|
||||
});
|
||||
|
||||
test('bullet-only Gaps unchanged (no regression on #2286)', () => {
|
||||
const items = gapsItems('- truth: "only a bullet"\n status: failed\n reason: "because"\n');
|
||||
|
||||
assert.strictEqual(items.length, 1, JSON.stringify(items));
|
||||
assert.strictEqual(items[0].name, 'only a bullet');
|
||||
assert.strictEqual(items[0].reason, 'because');
|
||||
});
|
||||
});
|
||||
|
||||
Reference in New Issue
Block a user