A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on dacae9273 while the PR that caused it (#3746) was green on every
check.
The docs-lint job in .github/workflows/docs-required.yml -- an ALREADY-REQUIRED
context -- now selects and runs the docs guards that read the specific docs files
the PR changed.
scripts/docs-guard-registry.cjs test file -> the docs paths it reads (63)
scripts/select-docs-guards.cjs pure (changedPaths, registry) -> test files
scripts/lint-docs-guard-registration.cjs drift guard, wired into lint:ci
scripts/ci-test-scope.cjs is NOT touched -- `git diff origin/next --` on it is
empty -- so #764's saving stands and its 21 pinning tests are untouched.
Selection: exact path; trailing-slash directory prefix (boundary-checked --
docs/adrenaline.md does NOT match docs/adr/, which a naive startsWith gets
wrong); and '*' for the 6 entries that walk docs/ generally or read a computed
path. Unknown maps to '*' -- guessing narrow is how a guard silently stops
running. Measured: a typo fix selects 6 of 63; docs/AGENTS.md selects 12;
docs/COMMANDS.md selects 18.
Four things this got wrong first, each found by an independent reviewer or by
probe, and each having been asserted safe in a comment:
1. The registry started as a RULE in ci-test-scope.cjs's RULES, on the theory
that classify()'s !codeChanged normalization made it inert. True for
docs-ONLY diffs; false for MIXED docs+code diffs, where codeChanged is true
and the normalization never runs:
node scripts/ci-test-scope.cjs --files "docs/a.md src/semver.cts"
with the RULE: 25 targeted_tests
origin/next: 3 targeted_tests
Category error: RULES is the scoped lane's input; a docs-guard registry is a
lane manifest for a consumer that never calls classify(). Extracted; pinned
by value.
2. The second attempt was a dedicated workflow with paths: [docs/**]. Such a
workflow never reports on a non-docs PR, so it can never be a required
context without hanging every non-docs PR -- and a non-required check does not
block a merge, so the guard would have been advisory and #3753 unfixed.
docs-required.yml already has no paths: filter, already supplies the required
docs-lint context, already computes docs_changed, and already ran one docs
guard gated on it. Generalizing that step needs no ruleset edit at all.
3. The registry and the drift lint were built from ONE path-segment heuristic, so
both were blind identically -- and blind at the guard that motivated the issue.
The reader-call regex required a character BEFORE its keyword, so a callee
named exactly read( / load( / parse( / doc( / file( / content( could never
match; and only an INLINE path.join(ROOT,'docs','X.md') argument was caught,
missing the two-step-via-variable form -- the MAJORITY spelling -- plus
template literals and concatenation. Detector 1 fired on 14 of ~450 files, so
35 genuine guards sat unregistered while the lint reported 0 violations,
including cursor-reviewer (reads docs/COMMANDS.md, asserts
.includes('--cursor')) and inventory-headings-countfree. The "accepted blind
spot" this shipped with was the common case, not a fringe.
4. With detection fixed the true population is 115 files: 63 genuine guards, 52
incidental. Running all 63 in a REQUIRED check on a one-line typo fix is the
cost #764 exists to avoid -- install.test.cjs is 7840 lines and reads exactly
one docs file, docs/AGENTS.md, for its frontmatter. Dropping it reproduces the
bug; running it for a typo elsewhere is waste. Hence the map.
Then a second review round found six more, all fixed here:
- fragment-single-edit-propagation.install.test.cjs was EXEMPTED as
"overlay fixture only". False: it reads the real docs/registries/eos.json and
asserts on a registry entry name, and reads the real ADR-0001 and asserts its
H1. A docs-only PR touching either would have gone green and red next -- #3753
shipping again, from inside the fix for it. Now registered against both paths,
and all 52 remaining exemptions were re-audited one by one.
- The SUITES-collision guard compared RAW registry keys, but run-tests.cjs strips
a leading `tests/` BEFORE its suite check. So it caught 'all' and missed
'tests/all' -- the only spelling that can actually occur, since every key
carries the prefix. One typo would have run all 824 test files inside the
required job. Now normalized the same way run-tests.cjs normalizes.
- The lint failed OPEN on an unreadable tests dir or candidate file: 0 violations,
ok:true. A guard that cannot read its input must never report success.
- The exemption ratchet gated identity only, so a baselined file that later
STARTED asserting on shipped docs stayed exempt silently -- 52 permanently blind
files. The baseline now fingerprints the docs paths each exempted file
references and fails when that set changes, naming what changed.
- The exemption marker was still honored inside a multi-line template literal in
the header window. The scanner now tracks template-literal and block-comment
state.
- `git diff --name-only | grep '^docs/'` silently dropped C-quoted non-ASCII docs
paths, making docs_changed=false a green zero-guard check. Both call sites now
pass -c core.quotepath=false.
- The run step was gated on hashFiles(), which a force-committed
.docs-guard-tests.txt would satisfy. The step now rm -f's both scratch files
first and gates on an output it sets itself.
Three empty states, deliberately distinct, because conflating them rebuilds
#3753: an empty or malformed registry HARD-FAILS; docs changed with no guard
covering them logs and skips; no docs change is already gated. The middle state
must never be expressed as an empty --files-from, which prints `no tests in suite
"all"` and exits 0 -- a green check that guarded nothing. With the current
registry that state is unreachable, because the six '*' entries always match;
the branch is kept as defensive handling for a future registry and says so.
timeout-minutes: 15 bounds the required job against a hanging fork-supplied test;
it had none. npm ci was added because the job never installed dependencies -- the
previous single-file step got away without it, the registry does not.
docs/contributing/docs-guard-registration.md documents the rule, following its
sibling cross-platform-portability-rules.md, and CONTRIBUTING.md's CI Test
Quality Checks table links to it. It is also load-bearing: without a docs/ file
in the diff this PR would not have triggered its own lane, shipping an
unexercised change to a required check.
One unrelated fix, included because this PR surfaced it and CLAUDE.md forbids
deferring a defect found while working. On this branch's first CI run,
`full test (windows-latest, 24, shard 3/3)` was CANCELLED at exactly 30 minutes;
tests were still passing 0.8s before the cancel, so it is a wall-clock timeout,
not a hang, and a cancelled job reddens `Required tests`.
The cause is not this PR's test file, which costs ~60ms. Shard composition is
unstable: adding ONE file to the unit suite reshuffled 115 of 268 files between
shards, and shard 3 drew a heavier mix. Underneath that is a real pre-existing
defect. tests/ci-test-job-timeout-budget.test.cjs requires every lane's budget to
be >= 1.5x its MEASURED cost -- "a lane that got slower must be re-budgeted, not
excused" -- and its test-full entry recorded 19m from a windows-22 shard. That is
stale. Measured on `next` with none of this PR's changes present: 26m18s (run
32614439702, windows-latest/24 shard 3/3), 23m36s and 23m17s on shard 2/3. So the
lane costs ~26m and the 30-minute cap carried 1.14x headroom, not 1.5x. The gate
had been out of compliance with its own rule; this PR was merely the file
addition that reshuffled shard 3 past the cliff.
Fixed as that file prescribes: measuredMinutes 19 -> 27 with fresh evidence, and
test-full timeout-minutes 30 -> 45. The rule's minimum for 27m is 41; 45 is
deliberately above it because the reshuffle means per-shard worst case moves run
to run, and a budget pinned to the exact minimum would be re-breached by the next
test file anyone adds. Only that one job's timeout changed; test.yml's scope,
matrix and steps are untouched, so #764's saving is unaffected.
Raising that cap let the Windows shard finish (28m45s, inside 45) and uncovered
a real failure the 30-minute cancel had been masking:
`new quick-task branch branches off origin/main (#2916)` died with
`outcome=timed_out exitCode=null`, SIGTERM, at the 15000ms bound.
tests/quick-branching.test.cjs:149 `runStep` runs a `#!/usr/bin/env bash` script
executing MULTIPLE git commands, but was bound to GIT_TIMEOUT_MS (15000) -- the
norm for a SINGLE git plumbing call. tests/helpers/timeouts.cjs already documents
this exact failure and exists to fix it: HOOK_FANOUT_TIMEOUT_MS was created after
PR #3285 recorded "outcome=timed_out exitCode=null at exactly the 15000ms probe
bound while every other lane passed the same commit", and calls that "a bound
sized for the wrong class, not a slow machine". Our failure is that case
verbatim, so both sites move to the class norm rather than to a bigger number.
The same class also failed on `next` itself 21 hours earlier -- run 32608945654,
windows-latest/24 shard 1/3, `plan touching only src/ in a submodule project
keeps worktree isolation ENABLED` -- where tests/worktree-safety.test.cjs:5845
`runGate` fans out to `git config --file .gitmodules` under a hardcoded 30000.
Fixed too, since it is a defect in the tree regardless of which branch surfaced
it.
A survey of the whole tests/ tree found the same class-mismatch at further
bash fan-out sites bound under 60000ms, and the maintainer approved sweeping
them rather than leaving them latent to surface the same way one at a time. 16
fan-out sites across 16 files now use the class norm.
The sweep is class-correctness, not raising numbers until things pass. Sites
were moved ONLY where the bash body demonstrably spawns something (git, node,
npm, a CLI); self-contained shell snippets were left where they are, and are
listed as deliberately unchanged: pure if/printf bodies (copilot-install), pure
array/case builtins (code-review-pipeline-regression:638), a documented
pure-shell gsd_run stub (host-integration), single-process hook calls
(workflow-guard:222/271/302), and a deliberately tight 5000ms fast-check hook
(gsd-write-guard.property). Nothing was lowered. process-seam.test.cjs:513
(literal 300) is untouched on purpose -- it tests timeout BEHAVIOR, so raising
it would destroy what it asserts.
Shared file-level constants were the trap here, and were handled per file rather
than by redefinition: GIT_TIMEOUT_MS has ~15 users in git-base-branch and only 1
is a fan-out; WORKTREE_TIMEOUT_MS has 16 users in worktree.test.cjs and 3 are;
PROBE_TIMEOUT_MS has several in three more files. In each the CALL SITE was
changed and the constant left alone, so no single-plumbing-call site silently
inherited a 60s bound. The one exception is hooks-opt-in.test.cjs, where
HOOK_TIMEOUT_MS has exactly one consumer -- spawnHook, the fan-out itself -- so
redefining it is identical in effect and reads better.
Only two of these sites have actually been observed failing. The rest cite that
shared class and those two run ids rather than inventing evidence of their own.
Co-authored-by: sim <sim@local>
391 lines
19 KiB
JavaScript
391 lines
19 KiB
JavaScript
// allow-test-rule: source-text-is-the-product — see #3409
|
|
// Workflow markdown is the installed orchestration contract; the snippets
|
|
// below are extracted from the shipped .md files and EXECUTED (not
|
|
// re-typed), so the test binds to the deployed contract rather than a copy
|
|
// that could silently drift from it.
|
|
|
|
'use strict';
|
|
|
|
/**
|
|
* Failing-first regression tests for #3409 (design:
|
|
* .gsd/phase/feat-3409-unreachable-shell-guard-lint/40-design.md; matrix:
|
|
* .gsd/phase/feat-3409-unreachable-shell-guard-lint/50-test-matrix.md,
|
|
* section "Regression — the three defects this PR fixes", rows G1-G4).
|
|
*
|
|
* Root cause (40-design.md): `gsd-tools.cjs`'s `--pick <field>` extractor
|
|
* coerces a missing/absent field to the empty string and exits 0. So
|
|
* `X=$(gsd_run query V --pick F 2>/dev/null || echo D)` can NEVER reach its
|
|
* `|| echo D` arm on field absence — only on a typo in the verb name. Three
|
|
* shipped shell guards silently rely on that unreachable arm:
|
|
*
|
|
* G1/G2 — plan-phase.md's Walking Skeleton gate reads a
|
|
* `phases.list --pick summaries_total` field that does not exist
|
|
* (#3365), so `PRIOR_SUMMARIES` is always `""`, never `"0"`, and
|
|
* the gate can never fire — not even for a genuinely fresh
|
|
* project (G1). G2 is the load-bearing negative-space case: it
|
|
* proves a bad fix that merely treats "no answer" as "zero"
|
|
* (making the gate fire unconditionally) is rejected, by pinning
|
|
* BOTH that the resolved count is a real nonzero integer AND that
|
|
* the gate stays off.
|
|
* G3 — plan-phase.md's `PHASE_REQ_IDS` site: on a phase with zero
|
|
* requirements, `query init.plan-phase <N> --pick phase_req_ids`
|
|
* exits 0 with empty stdout, so `|| echo TBD` never fires and
|
|
* `PHASE_REQ_IDS` resolves to `""` instead of the documented
|
|
* `TBD` sentinel (gate step reads "Skip if phase_req_ids is null
|
|
* or TBD").
|
|
* G4 — complete-milestone.md's bare `cat` over an unmatched-capable
|
|
* SUMMARY.md glob: under a `nullglob`
|
|
* left set by an earlier block in the SAME shell session (the
|
|
* `extract_accomplishments` step, a few hundred lines earlier in
|
|
* this same file), an unmatched glob expands to zero operands,
|
|
* so `cat` reads from stdin instead of erroring — and blocks
|
|
* forever if that stdin is not already at EOF.
|
|
*
|
|
* Each test below extracts the LIVE fenced-bash / single-line snippet out of
|
|
* the shipped workflow markdown (never a hand-typed copy — see
|
|
* `extractFencedBashAfterAnchor` / `extractAssignmentBlockFor`) and executes it
|
|
* with `runHook(..., { interpreter: 'bash' })`
|
|
* (`tests/helpers/process-seam.cjs`), against a temp project fixture, driving
|
|
* the real CLI at `gsd-core/bin/gsd-tools.cjs` through the real `gsd_run`
|
|
* shell function sourced from the shipped
|
|
* `gsd-core/workflows/_runtime-launcher.snippet.sh` preamble.
|
|
*/
|
|
|
|
const { test, describe } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
|
|
const { createTempDir, cleanup, readWorkflowCombined } = require('./helpers.cjs');
|
|
const { runHook, OUTCOME } = require('./helpers/process-seam.cjs');
|
|
const { HOOK_FANOUT_TIMEOUT_MS } = require('./helpers/timeouts.cjs');
|
|
|
|
const REPO_ROOT = path.join(__dirname, '..');
|
|
const PLAN_PHASE_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'plan-phase.md');
|
|
const COMPLETE_MILESTONE_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'complete-milestone.md');
|
|
const LAUNCHER_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', '_runtime-launcher.snippet.sh');
|
|
const GSD_TOOLS_PATH = path.join(REPO_ROOT, 'gsd-core', 'bin', 'gsd-tools.cjs');
|
|
|
|
// A `cat`-under-blocked-stdin hang (G4) must be bounded well under this, but
|
|
// give the CI-shape headroom PROBE_TIMEOUT_MS documents for a short CLI call.
|
|
const G4_TIMEOUT_MS = 5000; // short and explicit per the test-matrix note (G4 must assert on
|
|
// `outcome`, never `signal` — a real timeout and a maxBuffer overflow both
|
|
// report SIGTERM; PROBE_TIMEOUT_MS (15000ms) would work too but a tight,
|
|
// named bound makes a genuine hang fail fast instead of eating the suite's
|
|
// time budget on every RED run.
|
|
|
|
// ─── extraction (source-text-is-the-product) ─────────────────────────────
|
|
|
|
/**
|
|
* Extract the first ```bash fence appearing AFTER `anchor` in `content`.
|
|
* Mirrors the extraction convention already established by
|
|
* tests/plan-phase-stall-detection.test.cjs's extractStallHelpersBash(): walk
|
|
* forward from the anchor to the next fence open, then to its close. Throws
|
|
* with a message naming the anchor and file so a relocated/renamed anchor
|
|
* fails loudly instead of silently extracting the wrong block.
|
|
*/
|
|
function extractFencedBashAfterAnchor(content, anchor, sourcePath) {
|
|
const anchorIdx = content.indexOf(anchor);
|
|
if (anchorIdx === -1) {
|
|
throw new Error(`extractFencedBashAfterAnchor: could not find anchor "${anchor}" in ${sourcePath}`);
|
|
}
|
|
const after = content.slice(anchorIdx);
|
|
const fenceOpen = after.match(/```bash\r?\n/);
|
|
if (!fenceOpen) {
|
|
throw new Error(`extractFencedBashAfterAnchor: no \`\`\`bash fence found after anchor "${anchor}" in ${sourcePath}`);
|
|
}
|
|
const bodyStart = anchorIdx + fenceOpen.index + fenceOpen[0].length;
|
|
const closeIdx = content.indexOf('```', bodyStart);
|
|
if (closeIdx === -1) {
|
|
throw new Error(`extractFencedBashAfterAnchor: unterminated \`\`\`bash fence after anchor "${anchor}" in ${sourcePath}`);
|
|
}
|
|
return content.slice(bodyStart, closeIdx);
|
|
}
|
|
|
|
/**
|
|
* Extract the CONTIGUOUS RUN of source lines beginning with `prefix` (e.g.
|
|
* `PHASE_REQ_IDS=`) — from the first matching line, keep consuming
|
|
* subsequent lines while they ALSO start with `prefix`, and join them with
|
|
* `\n`. A single-line extraction would silently test only half a
|
|
* multi-line contract (e.g. the capture line of `X=$(...)` / `X="${X:-D}"`
|
|
* without its fallback-default line), which is exactly the "guard that
|
|
* cannot observe its own failure" class this suite exists to catch. Throws
|
|
* with a message naming the prefix and file if no matching line is found,
|
|
* so a rename/relocation fails loudly rather than silently testing nothing.
|
|
*/
|
|
function extractAssignmentBlockFor(content, prefix, sourcePath) {
|
|
const lines = content.split('\n');
|
|
const startIdx = lines.findIndex((l) => l.startsWith(prefix));
|
|
if (startIdx === -1) {
|
|
throw new Error(`extractAssignmentBlockFor: no line starting with "${prefix}" found in ${sourcePath}`);
|
|
}
|
|
const block = [];
|
|
for (let i = startIdx; i < lines.length; i += 1) {
|
|
if (!lines[i].startsWith(prefix)) break;
|
|
block.push(lines[i]);
|
|
}
|
|
return block.join('\n');
|
|
}
|
|
|
|
// ─── shared bash-script runner ────────────────────────────────────────────
|
|
|
|
/**
|
|
* Write `script` to a fresh temp file and run it via the process seam's
|
|
* `runHook(..., { interpreter: 'bash' })` — a script PATH, not a `bash -c`
|
|
* argv string, matching tests/plan-phase-stall-detection.test.cjs's
|
|
* runBashScript() (#2650: a quote-dense multi-line script passed as a single
|
|
* `-c` argv element does not survive Windows argv serialization).
|
|
*
|
|
* @param {import('node:test').TestContext} t
|
|
* @param {string} script - full script body (a shebang + `set -e` are
|
|
* prepended).
|
|
* @param {object} [options] - forwarded to runHook (cwd, env, timeoutMs).
|
|
*/
|
|
function runBashScript(t, script, options = {}) {
|
|
const scriptDir = createTempDir('gsd-3409-sh-');
|
|
t.after(() => cleanup(scriptDir));
|
|
const scriptPath = path.join(scriptDir, 'script.sh');
|
|
fs.writeFileSync(scriptPath, `#!/usr/bin/env bash\nset -e\n${script}`, { mode: 0o755 });
|
|
return runHook(scriptPath, [], { interpreter: 'bash', ...options });
|
|
}
|
|
|
|
/**
|
|
* Parse `KEY=value` lines (one per line, as emitted by this file's own
|
|
* `echo "KEY=$VAR"` trailers) out of a script's stdout. Values may
|
|
* legitimately be the empty string (that IS the RED condition G1/G2/G3
|
|
* assert against), so this returns `''` rather than `undefined` when the key
|
|
* is present with nothing after `=`.
|
|
*/
|
|
function parseKeyValueStdout(stdout) {
|
|
const result = {};
|
|
for (const line of stdout.split('\n')) {
|
|
const eq = line.indexOf('=');
|
|
if (eq === -1) continue;
|
|
result[line.slice(0, eq)] = line.slice(eq + 1).replace(/\r$/, '');
|
|
}
|
|
return result;
|
|
}
|
|
|
|
// ─── fixtures ──────────────────────────────────────────────────────────────
|
|
|
|
/**
|
|
* A minimal `.planning/phases/01-foundation/` project fixture — enough for
|
|
* `phases.list` and `init.plan-phase` to resolve phase 01 without error.
|
|
* `withSummary` seeds one real `*-SUMMARY.md` file when the negative-space
|
|
* (G2) case needs a nonzero prior-summary count.
|
|
*/
|
|
function buildPhase01Fixture({ withSummary }) {
|
|
const root = createTempDir('gsd-3409-fixture-');
|
|
const phaseDir = path.join(root, '.planning', 'phases', '01-foundation');
|
|
fs.mkdirSync(phaseDir, { recursive: true });
|
|
if (withSummary) {
|
|
fs.writeFileSync(path.join(phaseDir, '01-01-SUMMARY.md'), '# Summary\n\nDone.\n');
|
|
}
|
|
return root;
|
|
}
|
|
|
|
// ─── G1 / G2 — Walking Skeleton gate (plan-phase.md, #3365) ───────────────
|
|
|
|
describe('#3409 G1/G2 — plan-phase.md Walking Skeleton gate observes a real summary count', () => {
|
|
const anchor = 'Walking Skeleton gate.';
|
|
|
|
function runWalkingSkeletonGate(t, projectRoot) {
|
|
const snippet = extractFencedBashAfterAnchor(
|
|
readWorkflowCombined(PLAN_PHASE_PATH),
|
|
anchor,
|
|
PLAN_PHASE_PATH,
|
|
);
|
|
const script = [
|
|
`. "${LAUNCHER_PATH}"`,
|
|
snippet,
|
|
'echo "GSD_TEST_WALKING_SKELETON=$WALKING_SKELETON"',
|
|
'echo "GSD_TEST_PRIOR_SUMMARIES=$PRIOR_SUMMARIES"',
|
|
].join('\n');
|
|
const result = runBashScript(t, script, {
|
|
cwd: projectRoot,
|
|
env: {
|
|
...process.env,
|
|
RUNTIME_DIR: REPO_ROOT,
|
|
MVP_MODE: 'true',
|
|
padded_phase: '01',
|
|
},
|
|
// Bash FAN-OUT: the sourced `_runtime-launcher.snippet.sh` preamble
|
|
// defines the real `gsd_run` function, which the extracted snippet
|
|
// then calls — a bash + node invocation, not a single CLI probe. Same
|
|
// class as the observed CI failures in
|
|
// tests/quick-branching.test.cjs (PR #3787 run 32668773524) and
|
|
// tests/worktree-safety.test.cjs (`next` run 32608945654). See
|
|
// HOOK_FANOUT_TIMEOUT_MS in ./helpers/timeouts.cjs for the class
|
|
// rationale.
|
|
timeoutMs: HOOK_FANOUT_TIMEOUT_MS,
|
|
});
|
|
assert.equal(result.outcome, OUTCOME.EXITED, `gate script did not exit cleanly: ${result.stderr}`);
|
|
assert.equal(result.exitCode, 0, `gate script exited non-zero: ${result.stderr}`);
|
|
return parseKeyValueStdout(result.stdout);
|
|
}
|
|
|
|
test('G1: zero prior summaries — the gate observes a real integer 0, and fires', (t) => {
|
|
const root = buildPhase01Fixture({ withSummary: false });
|
|
t.after(() => cleanup(root));
|
|
const { GSD_TEST_WALKING_SKELETON, GSD_TEST_PRIOR_SUMMARIES } = runWalkingSkeletonGate(t, root);
|
|
|
|
// RED on the current tree: `--pick summaries_total` names a field that
|
|
// does not exist, so gsd-tools exits 0 with EMPTY stdout and the
|
|
// unreachable `|| echo "0"` arm never fires — PRIOR_SUMMARIES is `""`,
|
|
// not the integer `"0"` this asserts.
|
|
assert.equal(GSD_TEST_PRIOR_SUMMARIES, '0', 'prior-summary count must resolve to the integer 0, not empty string');
|
|
assert.equal(GSD_TEST_WALKING_SKELETON, 'true', 'a fresh phase-01 project must enter Walking Skeleton mode');
|
|
});
|
|
|
|
test('G2 (load-bearing negative space): a project WITH prior summaries does not enter skeleton mode', (t) => {
|
|
const root = buildPhase01Fixture({ withSummary: true });
|
|
t.after(() => cleanup(root));
|
|
const { GSD_TEST_WALKING_SKELETON, GSD_TEST_PRIOR_SUMMARIES } = runWalkingSkeletonGate(t, root);
|
|
|
|
// RED on the current tree: PRIOR_SUMMARIES is `""` here too (same
|
|
// unreachable-arm defect), which fails this integer check even though
|
|
// WALKING_SKELETON happens to read 'false' on the current, doubly-broken
|
|
// gate (it never fires for ANY input). This is what rejects a bad fix
|
|
// that treats "no answer" as "zero": such a fix would make
|
|
// WALKING_SKELETON fire unconditionally, which the second assertion
|
|
// below also catches.
|
|
assert.match(
|
|
GSD_TEST_PRIOR_SUMMARIES,
|
|
/^[1-9][0-9]*$/,
|
|
`prior-summary count must resolve to a nonzero integer, got ${JSON.stringify(GSD_TEST_PRIOR_SUMMARIES)}`,
|
|
);
|
|
assert.equal(GSD_TEST_WALKING_SKELETON, 'false', 'a project with prior summaries must NOT enter Walking Skeleton mode');
|
|
});
|
|
});
|
|
|
|
// ─── G3 — PHASE_REQ_IDS falls back to TBD (plan-phase.md) ─────────────────
|
|
|
|
test('#3409 G3: an empty phase_req_ids falls back to TBD, not the empty string', (t) => {
|
|
const root = createTempDir('gsd-3409-g3-');
|
|
t.after(() => cleanup(root));
|
|
fs.mkdirSync(path.join(root, '.planning', 'phases', '01-foundation'), { recursive: true });
|
|
// Deliberately no REQUIREMENTS.md / ROADMAP.md — phase 01 with zero
|
|
// requirements mapped to it, so `init.plan-phase --pick phase_req_ids`
|
|
// resolves `phase_req_ids: null` and `--pick` renders that as empty stdout
|
|
// (probe-confirmed: exit 0, empty stdout).
|
|
|
|
const block = extractAssignmentBlockFor(
|
|
readWorkflowCombined(PLAN_PHASE_PATH),
|
|
'PHASE_REQ_IDS=',
|
|
PLAN_PHASE_PATH,
|
|
);
|
|
const script = [
|
|
`. "${LAUNCHER_PATH}"`,
|
|
block,
|
|
'echo "GSD_TEST_PHASE_REQ_IDS=$PHASE_REQ_IDS"',
|
|
].join('\n');
|
|
// Bash FAN-OUT: same class as runWalkingSkeletonGate above — the sourced
|
|
// launcher's `gsd_run` shells out to node. See HOOK_FANOUT_TIMEOUT_MS in
|
|
// ./helpers/timeouts.cjs for the class rationale.
|
|
const result = runBashScript(t, script, {
|
|
cwd: root,
|
|
env: { ...process.env, RUNTIME_DIR: REPO_ROOT, PHASE: '01' },
|
|
timeoutMs: HOOK_FANOUT_TIMEOUT_MS,
|
|
});
|
|
assert.equal(result.outcome, OUTCOME.EXITED, `PHASE_REQ_IDS script did not exit cleanly: ${result.stderr}`);
|
|
assert.equal(result.exitCode, 0, `PHASE_REQ_IDS script exited non-zero: ${result.stderr}`);
|
|
|
|
const { GSD_TEST_PHASE_REQ_IDS } = parseKeyValueStdout(result.stdout);
|
|
// RED on the current tree: `--pick phase_req_ids` exits 0 with empty
|
|
// stdout on a `null` field, so the unreachable `|| echo TBD` arm never
|
|
// fires and PHASE_REQ_IDS resolves to `""` instead of the documented
|
|
// `TBD` sentinel (plan-phase.md: "Skip if phase_req_ids is null or TBD").
|
|
assert.equal(GSD_TEST_PHASE_REQ_IDS, 'TBD');
|
|
});
|
|
|
|
// ─── G4 — complete-milestone.md bare `cat <glob>` does not block on stdin ──
|
|
|
|
test('#3409 G4: the milestone summary read does not hang with no summaries', (t) => {
|
|
// The blocked-stdin mechanism below is a read-write FIFO opened via
|
|
// `mkfifo` — POSIX-only, and unavailable/non-functional on the
|
|
// `windows-latest` CI lane. Under `set -e` an unsupported `mkfifo` fails
|
|
// the script during setup, before the `cat` under test ever runs, so a
|
|
// Windows run would exercise nothing and must be skipped, not weakened.
|
|
if (process.platform === 'win32') {
|
|
t.skip('mkfifo-blocked-stdin reproduction is POSIX-only; unreachable on Windows');
|
|
return;
|
|
}
|
|
|
|
const root = createTempDir('gsd-3409-g4-');
|
|
t.after(() => cleanup(root));
|
|
// Zero-summary milestone: a phase dir exists, but no *-SUMMARY.md file
|
|
// anywhere under it — the exact condition that makes the glob unmatched.
|
|
fs.mkdirSync(path.join(root, '.planning', 'phases', '01-foundation'), { recursive: true });
|
|
|
|
const snippet = extractFencedBashAfterAnchor(
|
|
readWorkflowCombined(COMPLETE_MILESTONE_PATH),
|
|
'Read all phase summaries:',
|
|
COMPLETE_MILESTONE_PATH,
|
|
);
|
|
|
|
// Two things this script must reproduce, both faithfully, neither
|
|
// confounded with the other:
|
|
//
|
|
// 1. `nullglob` set — not by this fenced block itself (it sets nothing),
|
|
// but by an EARLIER block in the SAME workflow file/shell session
|
|
// (`extract_accomplishments`'s `shopt -s nullglob`, a few hundred
|
|
// lines above this one). 40-design.md's B11 names this exact
|
|
// "latent option from a different block" hazard as why Detector B
|
|
// flags this site even though it never sets the option locally.
|
|
// 2. stdin genuinely blocked, not just closed. Node's spawnSync closes
|
|
// an unwritten stdin immediately (EOF) when no `input` option is
|
|
// given, which would make a zero-operand `cat` return instantly
|
|
// instead of reproducing the real hang — so this opens a FIFO
|
|
// read-write on fd 3 (a read-write open never sees EOF, because the
|
|
// process holds its own write end) and redirects fd 0 there. This
|
|
// avoids `<(process substitution)`, which would leave a background
|
|
// job holding the CAPTURED STDOUT pipe open instead — a different,
|
|
// confounding hang unrelated to the stdin defect under test. The FIFO
|
|
// lives inside a private `mktemp -d` directory (created atomically
|
|
// with mode 0700) rather than at a bare `mktemp -u` path: `-u` only
|
|
// RESERVES a name without creating it, leaving a window between the
|
|
// reservation and `mkfifo` in which another process on a shared /tmp
|
|
// could create that same path first (a symlink-race primitive) — the
|
|
// directory removes the race entirely.
|
|
const script = [
|
|
'shopt -s nullglob',
|
|
'FIFO_DIR=$(mktemp -d)',
|
|
'mkfifo "$FIFO_DIR/f"',
|
|
'exec 3<> "$FIFO_DIR/f"',
|
|
'rm -rf "$FIFO_DIR"',
|
|
'exec 0<&3',
|
|
snippet,
|
|
].join('\n');
|
|
|
|
const result = runBashScript(t, script, { cwd: root, timeoutMs: G4_TIMEOUT_MS });
|
|
|
|
// RED on the current tree: the bare `cat <glob>` reads from the blocked
|
|
// stdin and never returns within G4_TIMEOUT_MS, so `outcome` is
|
|
// TIMED_OUT. Asserting on `outcome` (never `signal`) per CONTRIBUTING.md's
|
|
// process-seam guidance — a timeout and a maxBuffer overflow both report
|
|
// SIGTERM, and only `outcome` discriminates them.
|
|
assert.equal(
|
|
result.outcome,
|
|
OUTCOME.EXITED,
|
|
`expected the summary read to complete, got outcome=${result.outcome} stderr=${result.stderr}`,
|
|
);
|
|
// A nonzero exit here means the script's own setup (mkfifo/exec/mktemp)
|
|
// failed under `set -e` and the process exited immediately — which also
|
|
// reports outcome=EXITED, so it would silently pass the assertion above
|
|
// without ever reaching the `cat` under test. Pinning exitCode===0
|
|
// distinguishes "setup failed" from "the blocked read actually completed".
|
|
assert.equal(
|
|
result.exitCode,
|
|
0,
|
|
`expected setup (mkfifo/exec/mktemp) to succeed and the read to complete cleanly, got exitCode=${result.exitCode} stderr=${result.stderr}`,
|
|
);
|
|
});
|
|
|
|
// Sanity: the module under test actually exists at the path every fixture
|
|
// above points `RUNTIME_DIR`/`gsd_run` at — a moved/renamed CLI would
|
|
// otherwise make every test above fail with a confusing "gsd-tools.cjs not
|
|
// found" error deep inside a bash script instead of a clear assertion here.
|
|
test('#3409: gsd-tools.cjs exists at the path this suite drives gsd_run through', () => {
|
|
assert.equal(fs.existsSync(GSD_TOOLS_PATH), true, `expected ${GSD_TOOLS_PATH} to exist`);
|
|
});
|