* test(#3409): failing-first regression tests for unreachable shell guard arms Drives the three live defects fail-first, executing the shipped workflow snippets rather than a re-typed copy: - G1/G2 plan-phase.md Walking Skeleton gate reads `--pick summaries_total`, a field that does not exist, so PRIOR_SUMMARIES is always "" and the gate has never fired (#3365). G2 is the load-bearing negative-space case: it rejects a fix that treats "no answer" as "zero" and fires unconditionally. - G3 plan-phase.md PHASE_REQ_IDS resolves "" instead of the TBD sentinel on a phase with zero requirements. - G4 complete-milestone.md's bare `cat <glob>` blocks on stdin under a nullglob left set by an earlier block (measured hang). Skipped on Windows for G4 only: the FIFO-blocked-stdin mechanism is POSIX only, and a weakened assertion there would pass vacuously. Refs #3409 * fix(#3409): make nine shell guards observe their own failure arm `--pick` coerces a missing field to empty string and exits 0, so the `|| echo <default>` fallback after it fires only on a verb typo, never on the field absence it was written for. Nine sites relied on that arm. - plan-phase.md walking-skeleton gate: `--pick summaries_total` names a field that does not exist under any flag combination, so the gate has never fired on any project (#3365). Repointed at the existing single owner, `phases.list --type summaries --pick count`, which returns a real integer in every case including a project with no `.planning` directory. No new counter is added: a second one would duplicate the ownership ADR-3180 Decision 1 forbids. The gate now fires only on a literal "0", so an unanswerable query fails safe instead of entering skeleton mode. - plan-phase.md phase_req_ids: now falls back to the documented TBD. - The remaining seven convert to an explicit empty test. - complete-milestone.md read all phase summaries through a bare `cat <glob>`; under a nullglob left set by an earlier block that is zero operands, so cat blocks on stdin. Guarded with the array shape the #3300 fix already established in review.md. Refs #3409 * fix(#3409): guard eleven more globs that defeat their own fallback arm The nullglob audit this issue asks for turned up the same class in files #3300 never touched. - Eight bare `cat <glob>` reads (transition, complete-milestone, planner x4, verifier, phase-researcher). With nullglob set that is zero operands, so cat reads stdin and blocks; measured rc=137 at 3s. - Three `ls <glob> || echo "<message>"` sites (session-report, review-backlog and its generated skill). nullglob makes ls succeed listing the cwd, so the message never prints and the user gets a directory listing instead. Guarded with `[ -e "${_ARR[0]}" ]` rather than `[ ${#_ARR[@]} -gt 0 ]`. The count form is correct only when nullglob is set, and six of these seven files never set it: without it the array holds the unmatched literal pattern, so the count is 1 and the guard passes wrongly. `-e` is correct in both worlds. review.md keeps its count guards — that block sets nullglob two lines above them. skills/gsd-review-backlog regenerated from commands/, never hand-edited. Refs #3409 * feat(#3409): add the unreachable-shell-guard drift lint A sibling of lint-planning-prompt-drift.cjs, consuming the shared scripts/lib/drift-scan.cjs rather than copying it, wired into lint:ci. Both detectors are one shape — a fallback arm defeated by a legitimate success-on-empty: - Detector A: `--pick` and `|| echo` on one line. `--pick` is the discriminator because "missing field renders empty at exit 0" is a documented CLI contract, not a heuristic. A rule keyed on gsd_run matched 111 lines, ~132 of them legitimate, and was rejected. - Detector B: `cat <glob>` in command position, and `ls <glob>` whose exit code feeds a real fallback or an if/while head. Informational `ls <glob>` whose stdout is consumed (97 sites) and `|| true` failure suppression (~15) are not guards and never fire. Shrink-only ratchet keyed on (file, trimmed text) with a per-pair count, POSIX-normalized unconditionally so Windows CI cannot report everything fresh and stale at once. Ships with a ZERO-entry baseline: every site it can find is fixed. Exemption is the per-line `# gsd-scan-ignore: #NNN` marker whose reason must name an issue or URL; a malformed reason reports a distinct error rather than silently exempting. No file allowlists. ADR-3409 records the invariant, the measurements behind both detectors, and why the upstream `--pick` contract fix belongs to #3473. Refs #3409 * fix(#3409): resolve review findings — typed surface, sanitized reports, tighter marker Standards axis (blocker): the guard's tests asserted on human-readable stdout/stderr and on free-form baseline-load prose, which CONTRIBUTING prohibits by name. Added the typed surface it prescribes instead of weakening the tests: a frozen REASON enum, a --json report mode, structured loadBaseline errors, and a test locking Object.keys(REASON) so a new reason stays three coordinated changes. Security axis: sanitizeForReport covered every violation field but not the baseline-load error path, which embeds raw JSON.stringify output -- that escapes nothing above 0x1f, so bidi and C1 controls reached CI logs unfiltered. Routed through the sanitizer at the output seam. Security axis: the scan-ignore marker accepted `#0` and a bare `http://`. Tightened to a positive issue number and a URL with a host. This diverges deliberately from the sibling in tests/commit-files-pathspec.test.cjs, whose looser form was copied verbatim; the header now records the divergence. Security axis: G4 built its FIFO with `mktemp -u`, reserving a name without creating it. Now created inside a `mktemp -d` directory. Spec axis: ADR-3409 claimed a ninth site landed after the issue was filed. git blame disproves it -- all nine predate it; the issue's hand count missed one. Corrected. The design and test matrix still specified B9 as a FLAG after implementation reversed it to PASS; both now record the reversal and why. Refs #3409 * docs(#3409): add the how-to for resolving unreachable-guard findings Reference and Explanation are carried by ADR-3409; this is the task-oriented quadrant CI cannot check for. The page exists mainly for one thing the lint structurally cannot catch: both `[ -e "${_ARR[0]}" ]` and `[ ${#_ARR[@]} -gt 0 ]` remove the glob from the command and therefore both pass, but the count form is correct only when nullglob is set — and nullglob is usually set in a different block of the same file. A reference table cannot carry that; a how-to can. Also documents the reason codes, so a reader can tell "nothing to report" from "could not look". No tutorial: this is a gate inside an existing CI loop, not a new entry point a newcomer starts from. Refs #3409 * fix(#3409): bring the touched prompt files back under their size gates The remote run was red on 14 tests, all size/attribution, none of them the regression suite. - agents/gsd-planner.md was 194 chars over a 49152 cap enforced by four separate tests, each of which says the remedy is extraction, not a bump. It had 41 chars of headroom before this branch. Its `## Checkpoint Types` section was an unlinked, condensed duplicate of references/checkpoints.md, which already carries all three types and their XML shapes; the section now points there and keeps the three names and percentages inline. Net -969, margin 1010. - gsd-core/workflows/execute-phase.md sat 2 chars under a comfortable margin assertion. Dropped the AUTO_MODE default: the `|| echo "false"` it replaced was unreachable, so the value was already sometimes empty on next, and its only consumer compares against `true`. Net -16. Left plan-phase.md's AUTO_CHAIN default alone -- that file names an explicit `false` branch, so empty would match neither branch. - Acknowledged the seven prompt files that genuinely grew, one specific reason each. Five of those paths were already claimed by spent fragments identical to next, which blocks a second source naming the same path; removed just the colliding key from each, deleting the two that this emptied. Refs #3409 * test(#3409): extract the whole PHASE_REQ_IDS block, not just its first line G3 failed on the remote runner with '' !== 'TBD'. The test was wrong, not the workflow. The shipped contract is now two consecutive lines -- the capture and the `${PHASE_REQ_IDS:-TBD}` default -- but the helper's `^PREFIX=.*$` regex returns only the first match, so the test executed half the contract and correctly observed the empty string. Renamed to extractAssignmentBlockFor and taught it to consume the contiguous run of lines sharing the prefix. The assertion is untouched: TBD is the right expectation, and weakening it to accept the empty string would have reinstated exactly the class this suite exists to catch -- a check that cannot observe the thing it is checking. extractFencedBashAfterAnchor is unaffected: it is fence-delimited rather than line-anchored, so G1/G2/G4 still capture their full blocks. Refs #3409 * chore(#3409): drop a spent ack fragment that collided on complete-milestone.md #3458 landed on next while this branch was in flight and its fragment claims complete-milestone.md, which this branch also grows. Two ack sources may never name the same path. Its entry is spent: the +9163 it explains is already absorbed at base, so it can no longer clear anything, and the checker's own guidance for spent entries is to delete them. Removing the key emptied the fragment, so the file goes too -- an empty one signals nothing. Refs #3409 * chore(#3409): backfill changeset pr number 3558 * test(#3409): hoist a regex subject out of exec() to clear the injection scan CI's prompt-injection scan flagged `MARKER_RE.exec('# gsd-scan-ignore: ...')`. The pattern `exec[[:space:]]*\(["']` is receiver-blind on purpose, so it catches `require('child_process').exec('...')` -- and the scanner's own header records that RegExp.prototype.exec is collateral, to be handled by its allowlist. Allowlisting the file would blind it to the real exec vector permanently, so the subject is hoisted into a const instead: same assertion, scanner left at full strength, no security surface widened. Refs #3409 --------- Co-authored-by: sim <sim@local>
380 lines
18 KiB
JavaScript
380 lines
18 KiB
JavaScript
// allow-test-rule: source-text-is-the-product — see #3409
|
|
// Workflow markdown is the installed orchestration contract; the snippets
|
|
// below are extracted from the shipped .md files and EXECUTED (not
|
|
// re-typed), so the test binds to the deployed contract rather than a copy
|
|
// that could silently drift from it.
|
|
|
|
'use strict';
|
|
|
|
/**
|
|
* Failing-first regression tests for #3409 (design:
|
|
* .gsd/phase/feat-3409-unreachable-shell-guard-lint/40-design.md; matrix:
|
|
* .gsd/phase/feat-3409-unreachable-shell-guard-lint/50-test-matrix.md,
|
|
* section "Regression — the three defects this PR fixes", rows G1-G4).
|
|
*
|
|
* Root cause (40-design.md): `gsd-tools.cjs`'s `--pick <field>` extractor
|
|
* coerces a missing/absent field to the empty string and exits 0. So
|
|
* `X=$(gsd_run query V --pick F 2>/dev/null || echo D)` can NEVER reach its
|
|
* `|| echo D` arm on field absence — only on a typo in the verb name. Three
|
|
* shipped shell guards silently rely on that unreachable arm:
|
|
*
|
|
* G1/G2 — plan-phase.md's Walking Skeleton gate reads a
|
|
* `phases.list --pick summaries_total` field that does not exist
|
|
* (#3365), so `PRIOR_SUMMARIES` is always `""`, never `"0"`, and
|
|
* the gate can never fire — not even for a genuinely fresh
|
|
* project (G1). G2 is the load-bearing negative-space case: it
|
|
* proves a bad fix that merely treats "no answer" as "zero"
|
|
* (making the gate fire unconditionally) is rejected, by pinning
|
|
* BOTH that the resolved count is a real nonzero integer AND that
|
|
* the gate stays off.
|
|
* G3 — plan-phase.md's `PHASE_REQ_IDS` site: on a phase with zero
|
|
* requirements, `query init.plan-phase <N> --pick phase_req_ids`
|
|
* exits 0 with empty stdout, so `|| echo TBD` never fires and
|
|
* `PHASE_REQ_IDS` resolves to `""` instead of the documented
|
|
* `TBD` sentinel (gate step reads "Skip if phase_req_ids is null
|
|
* or TBD").
|
|
* G4 — complete-milestone.md's bare `cat` over an unmatched-capable
|
|
* SUMMARY.md glob: under a `nullglob`
|
|
* left set by an earlier block in the SAME shell session (the
|
|
* `extract_accomplishments` step, a few hundred lines earlier in
|
|
* this same file), an unmatched glob expands to zero operands,
|
|
* so `cat` reads from stdin instead of erroring — and blocks
|
|
* forever if that stdin is not already at EOF.
|
|
*
|
|
* Each test below extracts the LIVE fenced-bash / single-line snippet out of
|
|
* the shipped workflow markdown (never a hand-typed copy — see
|
|
* `extractFencedBashAfterAnchor` / `extractAssignmentBlockFor`) and executes it
|
|
* with `runHook(..., { interpreter: 'bash' })`
|
|
* (`tests/helpers/process-seam.cjs`), against a temp project fixture, driving
|
|
* the real CLI at `gsd-core/bin/gsd-tools.cjs` through the real `gsd_run`
|
|
* shell function sourced from the shipped
|
|
* `gsd-core/workflows/_runtime-launcher.snippet.sh` preamble.
|
|
*/
|
|
|
|
const { test, describe } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
|
|
const { createTempDir, cleanup, readWorkflowCombined } = require('./helpers.cjs');
|
|
const { runHook, OUTCOME } = require('./helpers/process-seam.cjs');
|
|
const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs');
|
|
|
|
const REPO_ROOT = path.join(__dirname, '..');
|
|
const PLAN_PHASE_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'plan-phase.md');
|
|
const COMPLETE_MILESTONE_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'complete-milestone.md');
|
|
const LAUNCHER_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', '_runtime-launcher.snippet.sh');
|
|
const GSD_TOOLS_PATH = path.join(REPO_ROOT, 'gsd-core', 'bin', 'gsd-tools.cjs');
|
|
|
|
// A `cat`-under-blocked-stdin hang (G4) must be bounded well under this, but
|
|
// give the CI-shape headroom PROBE_TIMEOUT_MS documents for a short CLI call.
|
|
const G4_TIMEOUT_MS = 5000; // short and explicit per the test-matrix note (G4 must assert on
|
|
// `outcome`, never `signal` — a real timeout and a maxBuffer overflow both
|
|
// report SIGTERM; PROBE_TIMEOUT_MS (15000ms) would work too but a tight,
|
|
// named bound makes a genuine hang fail fast instead of eating the suite's
|
|
// time budget on every RED run.
|
|
|
|
// ─── extraction (source-text-is-the-product) ─────────────────────────────
|
|
|
|
/**
|
|
* Extract the first ```bash fence appearing AFTER `anchor` in `content`.
|
|
* Mirrors the extraction convention already established by
|
|
* tests/plan-phase-stall-detection.test.cjs's extractStallHelpersBash(): walk
|
|
* forward from the anchor to the next fence open, then to its close. Throws
|
|
* with a message naming the anchor and file so a relocated/renamed anchor
|
|
* fails loudly instead of silently extracting the wrong block.
|
|
*/
|
|
function extractFencedBashAfterAnchor(content, anchor, sourcePath) {
|
|
const anchorIdx = content.indexOf(anchor);
|
|
if (anchorIdx === -1) {
|
|
throw new Error(`extractFencedBashAfterAnchor: could not find anchor "${anchor}" in ${sourcePath}`);
|
|
}
|
|
const after = content.slice(anchorIdx);
|
|
const fenceOpen = after.match(/```bash\r?\n/);
|
|
if (!fenceOpen) {
|
|
throw new Error(`extractFencedBashAfterAnchor: no \`\`\`bash fence found after anchor "${anchor}" in ${sourcePath}`);
|
|
}
|
|
const bodyStart = anchorIdx + fenceOpen.index + fenceOpen[0].length;
|
|
const closeIdx = content.indexOf('```', bodyStart);
|
|
if (closeIdx === -1) {
|
|
throw new Error(`extractFencedBashAfterAnchor: unterminated \`\`\`bash fence after anchor "${anchor}" in ${sourcePath}`);
|
|
}
|
|
return content.slice(bodyStart, closeIdx);
|
|
}
|
|
|
|
/**
|
|
* Extract the CONTIGUOUS RUN of source lines beginning with `prefix` (e.g.
|
|
* `PHASE_REQ_IDS=`) — from the first matching line, keep consuming
|
|
* subsequent lines while they ALSO start with `prefix`, and join them with
|
|
* `\n`. A single-line extraction would silently test only half a
|
|
* multi-line contract (e.g. the capture line of `X=$(...)` / `X="${X:-D}"`
|
|
* without its fallback-default line), which is exactly the "guard that
|
|
* cannot observe its own failure" class this suite exists to catch. Throws
|
|
* with a message naming the prefix and file if no matching line is found,
|
|
* so a rename/relocation fails loudly rather than silently testing nothing.
|
|
*/
|
|
function extractAssignmentBlockFor(content, prefix, sourcePath) {
|
|
const lines = content.split('\n');
|
|
const startIdx = lines.findIndex((l) => l.startsWith(prefix));
|
|
if (startIdx === -1) {
|
|
throw new Error(`extractAssignmentBlockFor: no line starting with "${prefix}" found in ${sourcePath}`);
|
|
}
|
|
const block = [];
|
|
for (let i = startIdx; i < lines.length; i += 1) {
|
|
if (!lines[i].startsWith(prefix)) break;
|
|
block.push(lines[i]);
|
|
}
|
|
return block.join('\n');
|
|
}
|
|
|
|
// ─── shared bash-script runner ────────────────────────────────────────────
|
|
|
|
/**
|
|
* Write `script` to a fresh temp file and run it via the process seam's
|
|
* `runHook(..., { interpreter: 'bash' })` — a script PATH, not a `bash -c`
|
|
* argv string, matching tests/plan-phase-stall-detection.test.cjs's
|
|
* runBashScript() (#2650: a quote-dense multi-line script passed as a single
|
|
* `-c` argv element does not survive Windows argv serialization).
|
|
*
|
|
* @param {import('node:test').TestContext} t
|
|
* @param {string} script - full script body (a shebang + `set -e` are
|
|
* prepended).
|
|
* @param {object} [options] - forwarded to runHook (cwd, env, timeoutMs).
|
|
*/
|
|
function runBashScript(t, script, options = {}) {
|
|
const scriptDir = createTempDir('gsd-3409-sh-');
|
|
t.after(() => cleanup(scriptDir));
|
|
const scriptPath = path.join(scriptDir, 'script.sh');
|
|
fs.writeFileSync(scriptPath, `#!/usr/bin/env bash\nset -e\n${script}`, { mode: 0o755 });
|
|
return runHook(scriptPath, [], { interpreter: 'bash', ...options });
|
|
}
|
|
|
|
/**
|
|
* Parse `KEY=value` lines (one per line, as emitted by this file's own
|
|
* `echo "KEY=$VAR"` trailers) out of a script's stdout. Values may
|
|
* legitimately be the empty string (that IS the RED condition G1/G2/G3
|
|
* assert against), so this returns `''` rather than `undefined` when the key
|
|
* is present with nothing after `=`.
|
|
*/
|
|
function parseKeyValueStdout(stdout) {
|
|
const result = {};
|
|
for (const line of stdout.split('\n')) {
|
|
const eq = line.indexOf('=');
|
|
if (eq === -1) continue;
|
|
result[line.slice(0, eq)] = line.slice(eq + 1).replace(/\r$/, '');
|
|
}
|
|
return result;
|
|
}
|
|
|
|
// ─── fixtures ──────────────────────────────────────────────────────────────
|
|
|
|
/**
|
|
* A minimal `.planning/phases/01-foundation/` project fixture — enough for
|
|
* `phases.list` and `init.plan-phase` to resolve phase 01 without error.
|
|
* `withSummary` seeds one real `*-SUMMARY.md` file when the negative-space
|
|
* (G2) case needs a nonzero prior-summary count.
|
|
*/
|
|
function buildPhase01Fixture({ withSummary }) {
|
|
const root = createTempDir('gsd-3409-fixture-');
|
|
const phaseDir = path.join(root, '.planning', 'phases', '01-foundation');
|
|
fs.mkdirSync(phaseDir, { recursive: true });
|
|
if (withSummary) {
|
|
fs.writeFileSync(path.join(phaseDir, '01-01-SUMMARY.md'), '# Summary\n\nDone.\n');
|
|
}
|
|
return root;
|
|
}
|
|
|
|
// ─── G1 / G2 — Walking Skeleton gate (plan-phase.md, #3365) ───────────────
|
|
|
|
describe('#3409 G1/G2 — plan-phase.md Walking Skeleton gate observes a real summary count', () => {
|
|
const anchor = 'Walking Skeleton gate.';
|
|
|
|
function runWalkingSkeletonGate(t, projectRoot) {
|
|
const snippet = extractFencedBashAfterAnchor(
|
|
readWorkflowCombined(PLAN_PHASE_PATH),
|
|
anchor,
|
|
PLAN_PHASE_PATH,
|
|
);
|
|
const script = [
|
|
`. "${LAUNCHER_PATH}"`,
|
|
snippet,
|
|
'echo "GSD_TEST_WALKING_SKELETON=$WALKING_SKELETON"',
|
|
'echo "GSD_TEST_PRIOR_SUMMARIES=$PRIOR_SUMMARIES"',
|
|
].join('\n');
|
|
const result = runBashScript(t, script, {
|
|
cwd: projectRoot,
|
|
env: {
|
|
...process.env,
|
|
RUNTIME_DIR: REPO_ROOT,
|
|
MVP_MODE: 'true',
|
|
padded_phase: '01',
|
|
},
|
|
timeoutMs: PROBE_TIMEOUT_MS,
|
|
});
|
|
assert.equal(result.outcome, OUTCOME.EXITED, `gate script did not exit cleanly: ${result.stderr}`);
|
|
assert.equal(result.exitCode, 0, `gate script exited non-zero: ${result.stderr}`);
|
|
return parseKeyValueStdout(result.stdout);
|
|
}
|
|
|
|
test('G1: zero prior summaries — the gate observes a real integer 0, and fires', (t) => {
|
|
const root = buildPhase01Fixture({ withSummary: false });
|
|
t.after(() => cleanup(root));
|
|
const { GSD_TEST_WALKING_SKELETON, GSD_TEST_PRIOR_SUMMARIES } = runWalkingSkeletonGate(t, root);
|
|
|
|
// RED on the current tree: `--pick summaries_total` names a field that
|
|
// does not exist, so gsd-tools exits 0 with EMPTY stdout and the
|
|
// unreachable `|| echo "0"` arm never fires — PRIOR_SUMMARIES is `""`,
|
|
// not the integer `"0"` this asserts.
|
|
assert.equal(GSD_TEST_PRIOR_SUMMARIES, '0', 'prior-summary count must resolve to the integer 0, not empty string');
|
|
assert.equal(GSD_TEST_WALKING_SKELETON, 'true', 'a fresh phase-01 project must enter Walking Skeleton mode');
|
|
});
|
|
|
|
test('G2 (load-bearing negative space): a project WITH prior summaries does not enter skeleton mode', (t) => {
|
|
const root = buildPhase01Fixture({ withSummary: true });
|
|
t.after(() => cleanup(root));
|
|
const { GSD_TEST_WALKING_SKELETON, GSD_TEST_PRIOR_SUMMARIES } = runWalkingSkeletonGate(t, root);
|
|
|
|
// RED on the current tree: PRIOR_SUMMARIES is `""` here too (same
|
|
// unreachable-arm defect), which fails this integer check even though
|
|
// WALKING_SKELETON happens to read 'false' on the current, doubly-broken
|
|
// gate (it never fires for ANY input). This is what rejects a bad fix
|
|
// that treats "no answer" as "zero": such a fix would make
|
|
// WALKING_SKELETON fire unconditionally, which the second assertion
|
|
// below also catches.
|
|
assert.match(
|
|
GSD_TEST_PRIOR_SUMMARIES,
|
|
/^[1-9][0-9]*$/,
|
|
`prior-summary count must resolve to a nonzero integer, got ${JSON.stringify(GSD_TEST_PRIOR_SUMMARIES)}`,
|
|
);
|
|
assert.equal(GSD_TEST_WALKING_SKELETON, 'false', 'a project with prior summaries must NOT enter Walking Skeleton mode');
|
|
});
|
|
});
|
|
|
|
// ─── G3 — PHASE_REQ_IDS falls back to TBD (plan-phase.md) ─────────────────
|
|
|
|
test('#3409 G3: an empty phase_req_ids falls back to TBD, not the empty string', (t) => {
|
|
const root = createTempDir('gsd-3409-g3-');
|
|
t.after(() => cleanup(root));
|
|
fs.mkdirSync(path.join(root, '.planning', 'phases', '01-foundation'), { recursive: true });
|
|
// Deliberately no REQUIREMENTS.md / ROADMAP.md — phase 01 with zero
|
|
// requirements mapped to it, so `init.plan-phase --pick phase_req_ids`
|
|
// resolves `phase_req_ids: null` and `--pick` renders that as empty stdout
|
|
// (probe-confirmed: exit 0, empty stdout).
|
|
|
|
const block = extractAssignmentBlockFor(
|
|
readWorkflowCombined(PLAN_PHASE_PATH),
|
|
'PHASE_REQ_IDS=',
|
|
PLAN_PHASE_PATH,
|
|
);
|
|
const script = [
|
|
`. "${LAUNCHER_PATH}"`,
|
|
block,
|
|
'echo "GSD_TEST_PHASE_REQ_IDS=$PHASE_REQ_IDS"',
|
|
].join('\n');
|
|
const result = runBashScript(t, script, {
|
|
cwd: root,
|
|
env: { ...process.env, RUNTIME_DIR: REPO_ROOT, PHASE: '01' },
|
|
timeoutMs: PROBE_TIMEOUT_MS,
|
|
});
|
|
assert.equal(result.outcome, OUTCOME.EXITED, `PHASE_REQ_IDS script did not exit cleanly: ${result.stderr}`);
|
|
assert.equal(result.exitCode, 0, `PHASE_REQ_IDS script exited non-zero: ${result.stderr}`);
|
|
|
|
const { GSD_TEST_PHASE_REQ_IDS } = parseKeyValueStdout(result.stdout);
|
|
// RED on the current tree: `--pick phase_req_ids` exits 0 with empty
|
|
// stdout on a `null` field, so the unreachable `|| echo TBD` arm never
|
|
// fires and PHASE_REQ_IDS resolves to `""` instead of the documented
|
|
// `TBD` sentinel (plan-phase.md: "Skip if phase_req_ids is null or TBD").
|
|
assert.equal(GSD_TEST_PHASE_REQ_IDS, 'TBD');
|
|
});
|
|
|
|
// ─── G4 — complete-milestone.md bare `cat <glob>` does not block on stdin ──
|
|
|
|
test('#3409 G4: the milestone summary read does not hang with no summaries', (t) => {
|
|
// The blocked-stdin mechanism below is a read-write FIFO opened via
|
|
// `mkfifo` — POSIX-only, and unavailable/non-functional on the
|
|
// `windows-latest` CI lane. Under `set -e` an unsupported `mkfifo` fails
|
|
// the script during setup, before the `cat` under test ever runs, so a
|
|
// Windows run would exercise nothing and must be skipped, not weakened.
|
|
if (process.platform === 'win32') {
|
|
t.skip('mkfifo-blocked-stdin reproduction is POSIX-only; unreachable on Windows');
|
|
return;
|
|
}
|
|
|
|
const root = createTempDir('gsd-3409-g4-');
|
|
t.after(() => cleanup(root));
|
|
// Zero-summary milestone: a phase dir exists, but no *-SUMMARY.md file
|
|
// anywhere under it — the exact condition that makes the glob unmatched.
|
|
fs.mkdirSync(path.join(root, '.planning', 'phases', '01-foundation'), { recursive: true });
|
|
|
|
const snippet = extractFencedBashAfterAnchor(
|
|
readWorkflowCombined(COMPLETE_MILESTONE_PATH),
|
|
'Read all phase summaries:',
|
|
COMPLETE_MILESTONE_PATH,
|
|
);
|
|
|
|
// Two things this script must reproduce, both faithfully, neither
|
|
// confounded with the other:
|
|
//
|
|
// 1. `nullglob` set — not by this fenced block itself (it sets nothing),
|
|
// but by an EARLIER block in the SAME workflow file/shell session
|
|
// (`extract_accomplishments`'s `shopt -s nullglob`, a few hundred
|
|
// lines above this one). 40-design.md's B11 names this exact
|
|
// "latent option from a different block" hazard as why Detector B
|
|
// flags this site even though it never sets the option locally.
|
|
// 2. stdin genuinely blocked, not just closed. Node's spawnSync closes
|
|
// an unwritten stdin immediately (EOF) when no `input` option is
|
|
// given, which would make a zero-operand `cat` return instantly
|
|
// instead of reproducing the real hang — so this opens a FIFO
|
|
// read-write on fd 3 (a read-write open never sees EOF, because the
|
|
// process holds its own write end) and redirects fd 0 there. This
|
|
// avoids `<(process substitution)`, which would leave a background
|
|
// job holding the CAPTURED STDOUT pipe open instead — a different,
|
|
// confounding hang unrelated to the stdin defect under test. The FIFO
|
|
// lives inside a private `mktemp -d` directory (created atomically
|
|
// with mode 0700) rather than at a bare `mktemp -u` path: `-u` only
|
|
// RESERVES a name without creating it, leaving a window between the
|
|
// reservation and `mkfifo` in which another process on a shared /tmp
|
|
// could create that same path first (a symlink-race primitive) — the
|
|
// directory removes the race entirely.
|
|
const script = [
|
|
'shopt -s nullglob',
|
|
'FIFO_DIR=$(mktemp -d)',
|
|
'mkfifo "$FIFO_DIR/f"',
|
|
'exec 3<> "$FIFO_DIR/f"',
|
|
'rm -rf "$FIFO_DIR"',
|
|
'exec 0<&3',
|
|
snippet,
|
|
].join('\n');
|
|
|
|
const result = runBashScript(t, script, { cwd: root, timeoutMs: G4_TIMEOUT_MS });
|
|
|
|
// RED on the current tree: the bare `cat <glob>` reads from the blocked
|
|
// stdin and never returns within G4_TIMEOUT_MS, so `outcome` is
|
|
// TIMED_OUT. Asserting on `outcome` (never `signal`) per CONTRIBUTING.md's
|
|
// process-seam guidance — a timeout and a maxBuffer overflow both report
|
|
// SIGTERM, and only `outcome` discriminates them.
|
|
assert.equal(
|
|
result.outcome,
|
|
OUTCOME.EXITED,
|
|
`expected the summary read to complete, got outcome=${result.outcome} stderr=${result.stderr}`,
|
|
);
|
|
// A nonzero exit here means the script's own setup (mkfifo/exec/mktemp)
|
|
// failed under `set -e` and the process exited immediately — which also
|
|
// reports outcome=EXITED, so it would silently pass the assertion above
|
|
// without ever reaching the `cat` under test. Pinning exitCode===0
|
|
// distinguishes "setup failed" from "the blocked read actually completed".
|
|
assert.equal(
|
|
result.exitCode,
|
|
0,
|
|
`expected setup (mkfifo/exec/mktemp) to succeed and the read to complete cleanly, got exitCode=${result.exitCode} stderr=${result.stderr}`,
|
|
);
|
|
});
|
|
|
|
// Sanity: the module under test actually exists at the path every fixture
|
|
// above points `RUNTIME_DIR`/`gsd_run` at — a moved/renamed CLI would
|
|
// otherwise make every test above fail with a confusing "gsd-tools.cjs not
|
|
// found" error deep inside a bash script instead of a clear assertion here.
|
|
test('#3409: gsd-tools.cjs exists at the path this suite drives gsd_run through', () => {
|
|
assert.equal(fs.existsSync(GSD_TOOLS_PATH), true, `expected ${GSD_TOOLS_PATH} to exist`);
|
|
});
|