Files
msd-core/tests/graphify-visualization.test.cjs
Tom Boucher 067a4d1c6c fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls (#3015)
* test(#2650): add failing-first regression for plan-phase stall detection

Regression test for gsd_stall_should_recover / gsd_stall_watch and the
planner.stall_* config keys, none of which exist yet — proves RED before
the fix lands in the next commit.

* fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls

Mirrors the already-shipped executor.stall_* pattern (execute-phase.md, bug
#3212) but with a dispatch change the executor's prose-only surveillance
lacks: the standard planner spawn, chunked-outline planner spawn,
chunked-per-plan planner spawn, plan-checker spawn, and revision-loop
planner respawn now dispatch with run_in_background=true and are followed
by a real, bounded bash poll (gsd_stall_watch) that returns control to the
orchestrator on its own schedule instead of waiting indefinitely on a
subagent that may never return. On stall, the existing accept-plans/retry/
stop recovery menu (9a/11a) is auto-surfaced instead of requiring a manual
interrupt.

New config keys planner.stall_detect_interval_minutes (default 5) /
planner.stall_threshold_minutes (default 10) mirror executor.stall_*.

The helper functions (gsd_stall_should_recover, gsd_stall_watch) live in a
new lazily-loaded gsd-core/workflows/plan-phase/steps/stall-detection-
helpers.md rather than inline, and per-site prose is kept minimal, because
plan-phase.md is frozen under the ADR-857 Phase 6 PRE_PHASE6 gate
(tests/phase6-capstone-conformance.test.cjs) with ~36 bytes of headroom at
baseline; the net effect is plan-phase.md.md ships slightly SMALLER than
before (the old unconditional-wait ORCHESTRATOR RULE sentences are gone at
the five touched sites, superseded by the bounded watcher).

Also fixes a stale doc comment in tests/workflow-size-budget.test.cjs that
still described the per-file workflow-size-baseline.json guard removed by
#2724 (ADR-2719 Phase 4) as if it were still the enforcement mechanism —
discovered while verifying this fix's own byte budget.

Researcher and pattern-mapper spawns are untouched (out of scope per the
issue's Agent Brief).

* fix(#2650): make gsd_stall_watch single-cycle; harden numeric config inputs

Two review findings addressed on top of the prior commit:

1. gsd_stall_watch previously looped internally for the full
   threshold+interval duration inside ONE Bash tool call (up to 15 min at
   defaults) — a single call blocking that long risks the host tool's own
   timeout killing it before it ever prints a result, silently defeating the
   fix. Redesigned to a single sleep-and-check cycle per call, taking an
   explicit dispatch_ts so the orchestrator prose can repeat the (short,
   default 5 min) call until it resolves; the outer threshold is now
   enforced by dispatch_ts accumulating across calls, not by one call's
   duration. Documented the resulting trade-off (up to one interval of
   added latency on the success path) in the changeset and reference doc.

2. PLANNER_STALL_INTERVAL_MINUTES/THRESHOLD_MINUTES are config-controlled
   values that flow into bash arithmetic ($(( ))). A review flagged this as
   command injection; empirically verified against both macOS bash 3.2.57
   and Docker bash:5 that this is NOT actually exploitable (bash hard-errors
   on a `$(cmd)`-shaped arithmetic operand rather than invoking it) — but an
   unvalidated malformed value WOULD abort the stall-watcher itself with
   that bash error, silently defeating the exact hang-recovery this issue
   ships. Added integer validation with safe-default fallback, both at the
   config-resolution point and defensively inside gsd_stall_should_recover.

Also adds the previously-missing integration coverage for gsd_stall_watch's
real execution (grep/find/date plumbing), not just the pure classifier.

* fix(#2650): correct AC2 self-test — helpers doc may name teams-status in prose

The AC2 regression test asserted the stall-detection-helpers.md step file
never contains the substring "teams-status" at all, but the file's own
prose explicitly documents its independence from that guard (containing
the word by design). Narrowed the assertion to what actually matters: no
second `query teams-status` call site and no gating on it, not a blanket
absence of the word.

* test(#2650): regenerate golden install-tree fixtures for the new step file

gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md is an
emitted file (installed for every runtime), so adding it changes the
install tree even though it is invisible to docs/INVENTORY.md and
docs/INVENTORY-MANIFEST.json (both explicitly scope to non-recursive
gsd-core/workflows/*.md — verified against the execute-phase #2930 and
pre-existing plan-phase step-file precedent, which are equally absent from
both inventory artifacts). The golden install tree snapshots the sorted
list of emitted relative paths per runtime, so a file invisible to the
inventory is still visible here. Regenerated via `npm run gen:install-tree`
— one line added per runtime fixture (19 files), no other drift.

* fix(#2650): restore 7 ORCHESTRATOR RULE labels; sync runtime-launcher preamble

Two more consequences of extracting helper bodies out of plan-phase.md,
both caught by verification (0017e1a78, 9 unique failures):

1. tests/plan-phase-drift-guard.test.cjs (#913) requires at least 7
   "ORCHESTRATOR RULE — ALL RUNTIMES" labels in plan-phase.md itself, one
   per agent spawn site. Moving the full explanatory blocks to
   plan-phase/steps/stall-detection-helpers.md carried 5 of the 7 labels
   out with them (only the untouched researcher/pattern-mapper sites kept
   theirs). Restored a short label at each of the 5 stall-watch sites,
   trimmed a few more redundant words ("Per 7.99, " — already established
   by the adjacent step-7.99 pointer) to stay under the frozen
   PRE_PHASE6 cap (94497 bytes, 21 bytes headroom).

2. tests/runtime-launcher-parity.test.cjs (#373) requires exactly one
   canonical gsd_run preamble, byte-equal to
   gsd-core/workflows/_runtime-launcher.snippet.sh, before the first
   gsd_run call in any workflow .md that calls it (recursive scan under
   gsd-core/workflows/, unlike the non-recursive inventory/step-tag-balance
   checks). The new step file's config-get calls use gsd_run without one.
   Fixed via `node scripts/sync-runtime-launcher.cjs`, verified: exactly 1
   preamble occurrence, before the first call, including the .claude/ and
   .codex/ home fallback arms.

Also verified (no fix needed, evidence recorded): the generic
`gsd-core-verbatim` identity rule in tests/helpers/emitted-provenance.cjs
(roots: ['gsd-core'], pattern matching workflows/.+) self-attributes any
new gsd-core/workflows/** path to itself, so the new step file needs no
drift-ack entry — consistent with plan-phase.md's own net shrinkage
requiring none either.

* test(#2650): acknowledge plan-phase.md's +14 byte drift

Restoring the 5 ORCHESTRATOR RULE — ALL RUNTIMES labels (#913) flipped
plan-phase.md from -142 bytes (post-extraction) to +14 bytes net growth
against baseline (94483 -> 94497), which the differential attribution
size ratchet (tests/emitted-attribution.test.cjs) correctly flags as
unacknowledged growth. Added tests/emitted-drift-acks/2650-plan-phase-
stall-detection.json, keyed on the bare filename plan-phase.md per the
existing fragment schema (see tests/emitted-drift-acks/2649-diagnose-
execute-plan-base-check.json), explaining the growth as exactly the 5
restored labels — still verified under the PRE_PHASE6 cap (94497 < 94519)
and satisfying #913's 7-label requirement.

* fix(#2650): bind {outputFile} from the real Agent() return — was dead code

Independent review blocker: PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE were
read by every gsd_stall_watch call but never assigned anywhere in the
diff. With the variable permanently empty, `[ -f "$output_file" ]` was
always false, marker_found could never become true, and marker_received
was unreachable — the marker-based detection path was permanently dead.

Worse for the plan-checker spawn specifically: a checker that PASSES
touches no *-PLAN.md files, so it had no working completion signal at
all without the marker path. A healthy plan-checker finishing cleanly in
two minutes would be declared stalled once planner.stall_threshold_minutes
elapsed and the recovery menu would fire on an already-succeeded agent —
worse than the original unbounded hang.

Fixed by replacing the dead bash variable with the `{outputFile}`
orchestrator-substitution token, the same convention docs-update.md:471
already uses for a real run_in_background=true Agent() return ("Read
tool: file_path: `{outputFile from README agent result}`"). This is a
net BYTE SAVING at each site (`"{outputFile}"` is shorter than
`"$PLANNER_OUTPUT_FILE"`), which funded moving the full binding
explanation — including why plan-checker's *-PLAN.md glob alone is not
a working completion signal — into the lazily-loaded reference file to
stay under the frozen PRE_PHASE6 cap (94496 bytes, 22 headroom; net +13
over baseline, acknowledged in tests/emitted-drift-acks/2650-plan-phase-
stall-detection.json).

Added a regression test asserting plan-phase.md itself binds {outputFile}
at all 5 spawn sites and contains no dangling $PLANNER_OUTPUT_FILE /
$CHECKER_OUTPUT_FILE reference — the previous test suite only exercised
gsd_stall_watch's behavior when handed a valid argument, which is why
the dead production wiring survived two rounds of review. Also fixed
tests/fix-2650-plan-phase-stall-detection.test.cjs:170-195's raw
try/finally to use t.after(), per CONTRIBUTING's test-cleanup convention.

* chore(#2650): backfill changeset PR number to 3015

* fix: normalize CRLF at the read boundary in all .md-bash-extraction tests

Maintainer-authorized scope expansion, folded into this PR rather than
deferred: the Windows CI lane on this PR's own tests/fix-2650-plan-phase-
stall-detection.test.cjs exposed DEFECT.TEST-SHELL-PIPELINE-NONPORTABLE
(CONTEXT.md; recurring since #1700) as a repo-wide latent class, not a
one-off. Ten test files parse a fenced ```bash block out of a workflow
.md file and execute it via spawnSync/execFileSync; a Windows checkout
can yield CRLF line endings despite .gitattributes eol=lf, and bash then
treats the trailing \r on every extracted line as part of the token —
"unexpected EOF while looking for matching `"'" or a bare syntax error,
partway through the script.

Added tests/helpers.cjs:readFileNormalized() — strips \r\n -> \n at the
read boundary, before any fence-slicing or regex runs, so every
downstream operation is correct by construction. Migrated all ten call
sites to it:

Previously broken (fs.readFileSync with no normalization anywhere
between read and spawn):
- tests/worktree-cleanup.test.cjs (extractCwdGuardBash) — also fixes a
  misleading comment claiming the fence regex alone was "CRLF-safe"; it
  protected only the fence delimiters, never the captured body.
- tests/new-milestone-clear-phases.test.cjs (extractFenceBetween,
  extractFenceContaining)
- tests/code-review-pipeline-regression.test.cjs (extractPostProcessingScript)
- tests/drift-detection.test.cjs (readGate/bashBlock, plus the snippet-file
  comparison read in the same test)
- tests/graphify-visualization.test.cjs (extractStep3Block)
- tests/pause-work-improvements.test.cjs (extractCheckBlock)
- tests/plan-review-convergence.test.cjs (extractReviewerFlagsParseBlock
  and the inline post-config-gate resolution-block slices)

Already correct (split(/\r?\n/) then join('\n')), migrated to the shared
helper for consistency rather than a fourth/fifth/sixth copy of the same
fix:
- tests/git-base-branch.test.cjs (extractHandleBranchingBash)
- tests/quick-branching.test.cjs (extractStep25Bash)
- tests/runtime-launcher-parity.test.cjs (extractResolverSnippet)

Verified against a simulated Windows CRLF checkout (not assumed): for
both the worktree-cleanup.test.cjs and new-milestone-clear-phases.test.cjs
extraction shapes, confirmed the pre-fix code produces a real bash syntax
error on CRLF input and the post-fix code does not.

One eslint follow-up: local/no-crlf-fragile-split statically flags any
bare `\n` inside a markdown-fence-shaped regex, regardless of whether the
receiver was already normalized — it cannot see the readFileNormalized()
data-flow. Kept `\r?\n` in extractCwdGuardBash's fence regex (redundant
but harmless on pre-normalized input) rather than fight the rule.

Scope note: this diff is broader than issue #2650's own change (plan-
phase.md stall detection) because the Windows lane surfaced a genuine
repo-wide defect class while verifying that fix, and the maintainer
authorized fixing it here rather than filing it separately and shipping
a known-broken pattern.

Runtime impact: none — this is a test-harness-only defect. The live
orchestrator (Claude Code or another runtime) does not do a byte-exact
extract-and-pipe of .md content into a shell the way these tests do; it
reads the instructions and generates its own bash invocation text, which
does not reproduce a raw CRLF pass-through the same way.

Not touched: tests/plan-review-convergence.test.cjs's separate, tracked
spawnSync ETIMEDOUT flake under bench load (#3005, reproduced on
unmodified next) — unrelated load-sensitivity, not a CRLF symptom.

* fix(#2650): remove stale drift-ack fragment — plan-phase.md is self-explaining

tests/emitted-drift-acks/2650-plan-phase-stall-detection.json acknowledged
plan-phase.md's own emitted-path hash move, but plan-phase.md is directly
edited in this diff. Per the emitted-attribution law (ADR-2719,
tests/emitted-attribution.test.cjs), a workflow's emitted key equals its
own source path (gsd-core-verbatim identity rule), so a direct edit to the
source is self-explaining and auto-attributed — no ack was ever needed.

Verified via the pre-merge lint (scripts/lint-emitted-drift-ack.cjs, run
through npm run lint:ci with a fully cleared eslint cache): it passes clean
with the fragment removed, confirming no contradiction between the lint and
the runtime attribution gate — this was simply an unnecessary fragment.

* fix(#2650): restore plan-phase.md drift-ack — size ratchet demands it against next

tests/emitted-drift-acks/2650-plan-phase-stall-detection.json was deleted in
the previous commit because, against an earlier verification base, it was
inert: it explained a moved emitted hash that a direct edit to plan-phase.md
already self-attributes. Against origin/next@f1af47766a the demand is
different: plan-phase.md is 13 bytes larger than the base copy, which trips
the emitted-attribution size ratchet — a job this same ack also performs.

Recreated in the documented shape, keyed on the bare filename plan-phase.md
(not the full path, and not restating the byte delta per review guidance),
describing the actual change: the {outputFile} binding fix for the dead
PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE variables and the 5 restored
ORCHESTRATOR RULE labels required by #913, both at the stall-watch spawn
sites, with explanatory bodies living in the lazily-loaded
gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md reference.

Confirmed no other fragment (on this branch or on next) claims the bare key
"plan-phase.md" before recreating — scripts/lint-emitted-drift-ack.cjs's
duplicate check is an exact string match, and the only other mention of
plan-phase.md in tests/emitted-drift-acks/ (2658-trae-instruction-file-path.json)
uses the full path as its key, so there is no collision.

* fix(#2650): real cause of Windows CI failure — bash -c argv-transport, not CRLF

The CRLF diagnosis for PR #3015's Windows failure was wrong. Proven wrong,
not assumed: .gitattributes' blanket `* text=auto eol=lf` means a Windows
checkout never receives CRLF for stall-detection-helpers.md, and the
extracted fence's line 64 is byte-identical and correctly balanced on every
platform. The real cause: runShouldRecover() passed a 70+ line, quote-dense
script as ONE argv element to `spawnSync('bash', ['-c', script, arg0, ...])`
PLUS four more positional args. Windows has no execve — Node serializes
that whole argv into a single CreateProcess command-line string, and Git
Bash's MSYS layer re-splits and unescapes it with its own rules. The
boundary between the script and the trailing args was not stable across
that round trip (live evidence: one failure's stderr was prefixed
`gsd_stall_should_recover_test:` — arg0 arrived — another `/usr/bin/bash:`
— arg0 did not).

Fixed by writing the script to a temp file and running `bash <file> <args>`
instead — the four values are now normal, quote-free positional args, and
the script itself never enters argv transport at all. Mirrors
tests/quick-branching.test.cjs's extractStep25Bash/runStep, which already
uses this exact shape and is green on Windows on `next`.
tests/worktree-cleanup.test.cjs's extractCwdGuardBash/runGuard stays on
`bash -c` but never appends extra positional args beyond the script itself,
so it never hits the same boundary — checked both siblings per review, not
assumed.

Corrected the now-actively-misleading CRLF comment in
extractStallHelpersBash(), and corrected the changeset's claim that the
repo-wide CRLF-normalization fix (folded into this branch, maintainer-
authorized) explains this PR's own Windows failure — it doesn't, though it
remains defensible on its own merits as general test-portability hardening.

Separately, while auditing the shipped (non-test) gsd_stall_watch for
Windows portability per review request, found and fixed a second, real
user-facing defect: the artifact-freshness check used GNU find's
`-newermt "@<epoch>"` shorthand, which the BSD find(1) actually shipped on
macOS does NOT understand ("Can't parse date/time: @<epoch>", verified live
against /usr/bin/find on both a stale and a genuinely fresh file). With the
adjacent `2>/dev/null`, that failed silently and permanently degraded
artifact_fresh to false on every macOS run — a plan-checker or planner
actively writing plan files could still be reported "stalled." Replaced
with `find $glob -mmin -N` ("modified less than N minutes ago"), which
needs no date-string parsing and is supported identically by GNU find and
BSD find; verified live that the old shape fails and the new shape passes
against the same real fresh file. Added a real-execution regression test
(gsd_stall_watch with `sleep` stubbed to a no-op so the test doesn't
actually wait, but the real `find ... -mmin` line still runs) proving the
fix, replacing the prior "not integration-tested" note for that path.

Note: the remote gsd-test runner is Linux-only, so it cannot itself confirm
the Windows fix — only the actual windows-latest CI lane can.

* fix(#2650): route the third bash -c call site through the same temp-file seam

runWatch() and a `-mmin` regression test still passed their script via
`bash -c <script>` after the previous commit only converted
runShouldRecover() — live Windows CI on 4b86cc57f confirmed the mechanism:
failures went 11 -> 4, and `full test (windows-latest, 22, shard 1/3)` and
`shard 2/3` flipped from fail to pass, but the remaining 4 failures (all in
this file, all still `bash: -c:`) were exactly the gsd_stall_watch describe
block, which runWatch() serves. runWatch() passes NO extra positional args
at all, so this also rules out the trailing-args theory from the prior
commit: the ~73-line, quote-dense script itself is what does not survive
Windows argv serialization when passed as a single `-c` element, regardless
of how many (if any) further argv elements follow it.

Extracted one shared runBashScript(script, args, opts) helper — write to a
fs.mkdtempSync'd file, run `bash <file> [args...]`, clean up in `finally` —
and routed all three bash-invoking call sites in this file through it
(runShouldRecover, runWatch, and the -mmin freshness test that builds its
own script inline for the `sleep` stub). One transport seam means a fourth
call site in this file cannot silently reintroduce the bug in isolation,
which is exactly what happened here with a second call site.

Corrected extractStallHelpersBash()'s doc comment a second time to state
the mechanism precisely (script content, not argv-element count) and cite
the live evidence (11->4 failures, shards 1 and 2 flipping green) so the
next reader does not have to rediscover it.

Audited every other bash-invoking call site in files this branch touches,
per review request:
- tests/code-review-pipeline-regression.test.cjs (runPostProcessing),
  tests/graphify-visualization.test.cjs (runBlock), and
  tests/drift-detection.test.cjs (two execFileSync('bash', ['-c', ...])
  sites, one of them carrying the same giant runtime-launcher preamble
  text) — all pre-existing, UNCHANGED by this branch (only touched for the
  readFileNormalized() CRLF swap), and already exercised on `next`'s last
  six Windows CI runs per the reviewer's own citation. Left as-is: no
  evidence of failure, and converting untested pre-existing code outside
  #2650's scope on an unverifiable guess would be its own risk.
- tests/git-base-branch.test.cjs (runHandleBranchingStep) and
  tests/quick-branching.test.cjs (runStep) already use the same temp-file
  pattern. No action needed.
- tests/runtime-launcher-parity.test.cjs (runResolver) uses `bash -c` but
  is explicitly `if (process.platform === 'win32') return '';` guarded off
  on Windows entirely, for an unrelated extension-less-PATH-stub reason —
  never reaches Windows argv transport at all. No action needed.
- tests/worktree-cleanup.test.cjs (runGuard) confirmed by the reviewer as
  correct and verified; not touched, per instruction.

Do not touch: the -mmin fix, the drift-ack fragment, the changeset — all
three confirmed correct in prior rounds and left untouched here.

Note: the remote gsd-test runner is Linux-only and cannot confirm this;
only the windows-latest lanes on #3015 can.

* fix(#2650): give runBashScript a default timeout

runShouldRecover() was the only one of the three call sites through
runBashScript() with no timeout — runWatch() and the -mmin test both pass
timeout: 10000 explicitly. Not a regression (this path never had a bound
before), but CONTEXT.md's unbounded-subprocess guidance applies directly,
and runShouldRecover() is driven repeatedly by a fast-check property test:
one pathological input that fails to terminate would hang CI indefinitely
instead of failing.

timeout: 10000 is now the helper's own default, with ...opts spread after
it so the two existing explicit timeout: 10000 call sites are unchanged
and any future caller inherits a bound automatically.

* fix(#2650): build the -mmin freshness test's glob with forward slashes

Windows CI on d6ddda6ea reported the last failure: the -mmin regression
test expected 'active' but got 'waiting' — find matched nothing, the same
silent-degradation shape as the macOS -newermt defect, but this time in the
test's own fixture rather than the shipped bash.

Traced what production actually passes: every gsd_stall_watch call site in
plan-phase.md builds artifact_glob as `"${PHASE_DIR}"'/*-PLAN.md'` —
PHASE_DIR is a POSIX-style .planning/phases/NN-slug value, and the whole
thing runs under Git Bash regardless of host OS, so production's glob is
always forward-slash. The test instead built it with
`path.join(tmp, '*-PLAN.md')`, which on Windows yields a backslash path
(C:\Users\RUNNER~1\...\*-PLAN.md). In bash pathname expansion a backslash
escapes the next character, so that pattern can never match a real path —
find silently returns empty under the existing 2>/dev/null, same shape as
the macOS bug. Confirmed as a test artifact, not a production defect:
production never constructs the glob this way, so no Windows user is
affected.

Fixed by forward-slashing the tmp dir before appending the glob suffix,
matching production's own convention, with a comment recording why (so a
future "simplify this back to path.join" edit doesn't silently reintroduce
the failure). The shipped bash's unquoted $artifact_glob is untouched —
quoting it would break the multi-file glob expansion it exists for.

Note: the remote runner is Linux-only and already passed clean at
d6ddda6ea (0/29,603, both node lanes); only the windows-latest lanes on
#3015 can confirm this fix.

* fix(#2650): forward-slash the three remaining runWatch globs (vacuous-pass CR)

The :353 fix (833c11da9) only converted the -mmin freshness test's glob.
Three sibling tests in the same describe block still built theirs with
path.join(tmp, '*-PLAN.md'), which yields a backslash path on Windows.

Two of those three were silently passing for the wrong reason: the
'-> stalled' and '-> waiting' tests both expect the glob to match nothing,
and on Windows a backslash path matches nothing regardless of whether the
directory is actually empty (bash eats each backslash as an escape before
the pattern is even evaluated). They would have passed identically with
glob expansion completely broken, which is a vacuous pass — not exercising
what they claim to. The third ('-> marker_received') is outcome-independent
of the glob, so it was merely inconsistent rather than wrong.

Converted all three to the same `${tmp.replace(/\\/g, '/')}/*-PLAN.md`
construction already used at the -mmin test, so every glob in the file now
matches production's own forward-slash `"${PHASE_DIR}"'/*-PLAN.md'` shape,
and the two negative tests are meaningful on Windows instead of accidentally
correct. Reworded the trailing comment on the 'stalled' test's glob line:
it now describes the fixture (the tmp dir contains no *-PLAN.md files)
rather than the pattern, since "matches nothing" read as a property of the
glob syntax when it's a property of what's on disk.

No assertion, the sleep stub, runBashScript, or the shipped bash changed.
Smoke-tested all three updated tests manually before committing (not via
node --test): marker_received / stalled / waiting, all correct.

* fix(#2650): fix own regression tests for #2993's plan-phase.md relocation

531101843's merge with origin/next brought in #2993 (unrelated, epic #1671
Phase 6.2), which extracted plan-phase.md's whole "Chunked Planning Mode"
section into gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md,
leaving a <!-- gsd:section --> pointer behind. tests/plan-phase-drift-guard.
test.cjs (#913) was already updated to read the combined surface (host file
+ every steps/*.md) so its label count didn't go blind — my own #2650
regression tests were not, and searched plan-phase.md alone for the two
chunked spawn sites' headings, which no longer exist there. Two tests
failed outright (indexOf returning -1); a third ("standard planner spawn")
was silently weakened to an unbounded slice-to-EOF by the same relocation,
since its own end-boundary heading also moved — passing by accident rather
than by testing what it claimed.

Promoted the drift guard's local readPlanPhaseCombined() to a shared,
exported tests/helpers.cjs readWorkflowCombined(workflowPath) (host file +
sorted steps/*.md, CRLF-normalized at the read boundary) so a second,
divergent implementation is never written — the drift guard now delegates
to it via a same-named local wrapper, unchanged at every existing call site.

Fixed the three affected tests in tests/fix-2650-plan-phase-stall-detection.
test.cjs:
- "standard planner spawn (step 8)": end boundary changed from the now-gone
  "## 8.5. Chunked Planning Mode" heading to "## 9. Handle Planner Return",
  which still exists in plan-phase.md.
- "chunked outline spawn (8.5.1)" / "chunked per-plan spawn (8.5.2)": now
  read gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md
  directly (not the generic multi-file combined blob, whose file-sort
  ordering would put unrelated step files between 8.5.2's slice and any
  downstream anchor) — the same heading-to-heading slicing as before still
  works because the file is small and self-contained.
- Extended the "no unbound $PLANNER_OUTPUT_FILE/$CHECKER_OUTPUT_FILE" check
  to also scan chunked-planning-mode.md, since two of the five spawn sites
  now live there.
- Added a new count-based test asserting exactly 5 (not "at least one")
  `gsd_stall_watch "$TS" "{outputFile}"` invocations across the combined
  surface, mirroring #913's own label-count guard, so every one of the five
  spawns stays provably bounded and a future relocation can't silently drop
  one without a test noticing.

Also added a small positive test that plan-phase.md's <!-- gsd:section -->
pointer to chunked-planning-mode.md exists (#2993 is unrelated to #2650 but
its presence is now load-bearing for where 2 of the 5 spawn sites live).

Audited every other test file in the repo for a stale reference to content
#2993 relocated (searched for the moved headings/prose and for
"chunked-planning-mode"/"CHUNKED_MODE" across all *.test.cjs): only this
file and the drift guard needed changes.
tests/issue-2762-plan-reviews-chunked.test.cjs already reads
chunked-planning-mode.md directly (brought in correct by the same merge).
gen-section-manifest.test.cjs, init.test.cjs, and workflow-fragments.test.cjs
reference "chunked-planning-mode" only as a manifest/section-id fixture
value for #2993 itself, not as a stale pointer to relocated content.

Did not touch: the ported ORCHESTRATOR RULE lines, run_in_background=true,
the glob constructions, runBashScript, the -mmin change, the timeout
default, or the drift-ack fragment (confirmed correct against the stale
local `next` ref two rounds ago and left alone).

---------

Co-authored-by: sim <sim@local>
2026-08-03 10:46:22 -04:00

808 lines
31 KiB
JavaScript

'use strict';
// Tests for graphify.cjs — staleness, mvp-viz, and regressions describe blocks.
// Split from the consolidated 2336-LOC file. Refs #3761.
const { describe, test, beforeEach, afterEach, before, after } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('fs');
const path = require('path');
const os = require('node:os');
const { execFileSync } = require('child_process');
const { createTempProject, createTempGitProject, cleanup } = require('./helpers.cjs');
const {
graphifyStatus,
} = require('../gsd-core/bin/lib/graphify.cjs');
const {
enableGraphify,
writeGraphJson,
gitHead,
commitEmpty,
SAMPLE_NODES_MINIMAL,
} = require('./helpers/graphify.cjs');
// ─── staleness describe ──────────────────────────────────────────────────────
describe('staleness', () => {
// Regression for #3170: graphifyStatus surfaces built_at_commit staleness.
// graphify v0.7+ embeds `built_at_commit` into graph.json at write time.
// Tri-state on commit_stale: null means "we don't know" (pre-v0.7 graph or
// no git), which is semantically distinct from false ("known fresh").
describe('git-aware', () => {
let tmpDir;
let planningDir;
beforeEach(() => {
tmpDir = createTempGitProject();
planningDir = path.join(tmpDir, '.planning');
enableGraphify(planningDir);
});
afterEach(() => cleanup(tmpDir));
test('graph rebuilt at HEAD: commits_behind=0, commit_stale=false', () => {
const head = gitHead(tmpDir);
writeGraphJson(planningDir, { nodes: SAMPLE_NODES_MINIMAL, edges: [], built_at_commit: head });
const result = graphifyStatus(tmpDir);
assert.equal(result.built_at_commit, head.slice(0, 7),
'short hash from graph.built_at_commit');
assert.equal(result.current_commit, head.slice(0, 7),
'short hash of git HEAD');
assert.equal(result.commits_behind, 0,
'zero commits between HEAD and itself');
assert.equal(result.commit_stale, false,
'commit_stale is explicitly false when commits_behind === 0');
});
test('graph 5 commits behind HEAD: commits_behind=5, commit_stale=true', () => {
const built = gitHead(tmpDir);
for (let i = 0; i < 5; i += 1) commitEmpty(tmpDir, `c${i}`);
writeGraphJson(planningDir, { nodes: SAMPLE_NODES_MINIMAL, edges: [], built_at_commit: built });
const result = graphifyStatus(tmpDir);
assert.equal(result.commits_behind, 5);
assert.equal(result.commit_stale, true);
assert.equal(result.built_at_commit, built.slice(0, 7));
assert.notEqual(result.current_commit, built.slice(0, 7),
'current_commit reflects HEAD, not graph build commit');
});
test('built_at_commit absent (pre-v0.7 graph): all four new fields null', () => {
// No built_at_commit on the graph -- GSD must not fabricate one.
writeGraphJson(planningDir, { nodes: SAMPLE_NODES_MINIMAL, edges: [] });
const result = graphifyStatus(tmpDir);
assert.equal(result.built_at_commit, null);
assert.equal(result.commits_behind, null);
assert.equal(result.commit_stale, null,
'tri-state: null means "we do not know", not "fresh"');
// current_commit may still be non-null since we are in a git repo,
// but without a baseline it cannot drive staleness.
assert.notEqual(result.current_commit, undefined,
'current_commit field is always present even when null');
});
test('rebased-away built_at_commit: commits_behind=null, commit_stale=null', () => {
// built_at_commit references a commit that never existed in this repo.
const ghostHash = '0000000000000000000000000000000000000001';
writeGraphJson(planningDir, { nodes: SAMPLE_NODES_MINIMAL, edges: [], built_at_commit: ghostHash });
const result = graphifyStatus(tmpDir);
assert.equal(result.built_at_commit, ghostHash.slice(0, 7),
'echoes the field even if unreachable -- caller can decide what to do');
assert.equal(result.commits_behind, null,
'cannot count commits to an unreachable commit');
assert.equal(result.commit_stale, null,
'unknown distance means unknown staleness');
});
test('malformed built_at_commit (dashed argv): rejected before git invocation', () => {
// Argument-injection fence: a graph.json with a hostile built_at_commit
// must never reach `git` as an argv element. The implementation should
// validate /^[0-9a-f]{4,40}$/i and treat anything else as absent.
const malicious = '--upload-pack=evil';
writeGraphJson(planningDir, { nodes: SAMPLE_NODES_MINIMAL, edges: [], built_at_commit: malicious });
const result = graphifyStatus(tmpDir);
assert.equal(result.built_at_commit, null,
'malformed value is rejected, not echoed');
assert.equal(result.commits_behind, null);
assert.equal(result.commit_stale, null);
});
});
describe('non-git cwd', () => {
let tmpDir;
let planningDir;
beforeEach(() => {
tmpDir = createTempProject();
planningDir = path.join(tmpDir, '.planning');
enableGraphify(planningDir);
});
afterEach(() => cleanup(tmpDir));
test('cwd has no .git: current_commit=null, derived fields=null', () => {
const built = 'abcdef1234567890abcdef1234567890abcdef12';
writeGraphJson(planningDir, { nodes: SAMPLE_NODES_MINIMAL, edges: [], built_at_commit: built });
const result = graphifyStatus(tmpDir);
assert.equal(result.built_at_commit, built.slice(0, 7),
'graph field is echoed even without a local repo');
assert.equal(result.current_commit, null,
'no HEAD without git');
assert.equal(result.commits_behind, null);
assert.equal(result.commit_stale, null);
});
});
describe('back-compat', () => {
let tmpDir;
let planningDir;
beforeEach(() => {
tmpDir = createTempGitProject();
planningDir = path.join(tmpDir, '.planning');
enableGraphify(planningDir);
writeGraphJson(planningDir, {
nodes: SAMPLE_NODES_MINIMAL,
edges: [{ source: 'n1', target: 'n2', label: 'x', confidence: 'EXTRACTED' }],
hyperedges: [],
built_at_commit: gitHead(tmpDir),
});
});
afterEach(() => cleanup(tmpDir));
test('existing fields are unchanged when commit-staleness fields are added', () => {
const result = graphifyStatus(tmpDir);
// Existing contract — must not regress.
assert.equal(result.exists, true);
assert.equal(result.node_count, 2);
assert.equal(result.edge_count, 1);
assert.equal(result.hyperedge_count, 0);
assert.equal(typeof result.last_build, 'string');
assert.equal(typeof result.stale, 'boolean',
'mtime-based stale flag stays as-is for back-compat');
assert.equal(typeof result.age_hours, 'number');
});
test('disabled response is unchanged (commit-staleness fields not added)', () => {
const tmp2 = createTempProject();
try {
const result = graphifyStatus(tmp2);
assert.equal(result.disabled, true,
'disabled path returns the existing shape, no commit fields');
assert.equal(result.built_at_commit, undefined,
'commit-staleness fields are only added on the success path');
} finally {
cleanup(tmp2);
}
});
});
});
// ─── mvp-viz describe ─────────────────────────────────────────────────────────
describe('mvp-viz', () => {
// Contract: commands/gsd/graphify.md documents MVP visual differentiation.
// Per PRD Q5: distinct node color + 'MVP' label suffix.
// Tests parse the markdown skill into structured IR (YAML frontmatter +
// fenced code blocks) and assert on the parsed structures, not raw text.
const CMD = path.join(__dirname, '..', 'commands', 'gsd', 'graphify.md');
/**
* Parse the narrow YAML subset used in this skill's frontmatter:
* key: scalar
* key:
* - item
* - item
*/
function parseSkillFrontmatter(text) {
const lines = text.split(/\r?\n/);
const out = {};
let _activeKey = null;
let activeList = null;
for (const raw of lines) {
const listItem = raw.match(/^\s+-\s+(.+?)\s*$/);
if (listItem && activeList) {
activeList.push(listItem[1]);
continue;
}
const kv = raw.match(/^([A-Za-z][A-Za-z0-9_-]*):\s*(.*)$/);
if (!kv) continue;
const [, key, rawValue] = kv;
const value = rawValue.trim();
if (value === '') {
_activeKey = key;
activeList = [];
out[key] = activeList;
} else {
_activeKey = null;
activeList = null;
out[key] = value;
}
}
return out;
}
/**
* Walk markdown body line-by-line and return every fenced code block as
* { lang, content } records. Tracks fence state explicitly.
*/
function extractFencedBlocks(body) {
const lines = body.split(/\r?\n/);
const blocks = [];
let active = null;
for (const line of lines) {
const open = line.match(/^```(\S*)\s*$/);
if (active === null) {
if (open) active = { lang: open[1] || '', lines: [] };
continue;
}
if (line.trim() === '```') {
blocks.push({ lang: active.lang, content: active.lines.join('\n') });
active = null;
continue;
}
active.lines.push(line);
}
return blocks;
}
function loadSkill() {
// Local rename (`markdown` not `content`) so the no-source-grep lint
// doesn't conflate this readFileSync-bound variable with the
// `b.content.includes(...)` calls below — those operate on parsed
// fenced-block records, not raw file text.
const markdown = fs.readFileSync(CMD, 'utf8');
const lines = markdown.split(/\r?\n/);
const delims = [];
for (let i = 0; i < lines.length; i += 1) {
if (lines[i].trim() === '---') delims.push(i);
if (delims.length === 2) break;
}
assert.equal(delims.length, 2, 'graphify.md must have a closed frontmatter block');
const frontmatterText = lines.slice(delims[0] + 1, delims[1]).join('\n');
const body = lines.slice(delims[1] + 1).join('\n');
return {
frontmatter: parseSkillFrontmatter(frontmatterText),
body,
fencedBlocks: extractFencedBlocks(body),
};
}
// Parse MVP section from graphify.md body as structured IR (not raw grep).
// Extracts: mentionsMvp, colorRuleLine, labelRuleLine, fallbackLine.
function parseMvpVizContract(body) {
const lines = body.split(/\r?\n/);
const lowerLines = lines.map(line => line.toLowerCase());
const mvpLines = lines.filter(line => line.toLowerCase().includes('mvp'));
return {
mentionsMvp: mvpLines.length > 0,
colorRuleLine: mvpLines.find(line => {
const lower = line.toLowerCase();
return lower.includes('color') || lower.includes('fill') || line.includes('#');
}) || '',
labelRuleLine: mvpLines.find(line => {
const lower = line.toLowerCase();
return lower.includes('label') || lower.includes('suffix');
}) || '',
fallbackLine: lowerLines.find(line =>
(line.includes('mode') && (line.includes('null') || line.includes('absent') || line.includes('not mvp'))) ||
(line.includes('standard') && (line.includes('render') || line.includes('fallback')))
) || '',
};
}
test('graphify.md documents distinct color for MVP-mode phases', () => {
const { body } = loadSkill();
const contract = parseMvpVizContract(body);
assert.ok(contract.mentionsMvp, 'must mention MVP in color rule');
assert.ok(contract.colorRuleLine.length > 0, 'must reference a color/fill rule for MVP nodes');
});
test('graphify.md documents MVP label suffix on node text', () => {
const { body } = loadSkill();
const contract = parseMvpVizContract(body);
assert.ok(contract.labelRuleLine.length > 0, 'must add an MVP label/suffix to node text');
});
test('graphify.md specifies fallback when phase mode is null/absent', () => {
const { body } = loadSkill();
const contract = parseMvpVizContract(body);
assert.ok(contract.fallbackLine.length > 0, 'must specify fallback when mode is not mvp');
});
// Counter-test: a non-mvp phase must NOT carry mode:'mvp' in the contract.
// The fallbackLine ensures standard rendering is documented for the non-mvp case.
test('non-mvp phase render path is documented (counter-test)', () => {
const { body } = loadSkill();
const contract = parseMvpVizContract(body);
// The fallback line is required precisely because non-mvp phases exist;
// its presence is the counter-assertion that mvp rendering is NOT applied globally.
assert.ok(
contract.fallbackLine.length > 0,
'fallback documentation confirms mvp rendering is not applied to non-mvp phases',
);
// Additionally: the MVP label should only be a suffix, not a full replacement;
// so the standard label path (no MVP suffix) must be documented.
assert.ok(
contract.mentionsMvp,
'mvp mention is present, meaning mvp is treated as a special case, not the default',
);
});
});
// ─── regressions describe ─────────────────────────────────────────────────────
describe('regressions', () => {
// ── Regression for #3166 ────────────────────────────────────────────────────
// /gsd-graphify build lost artifacts because the skill spawned a Task
// sub-agent that backgrounded `graphify update .`. Sub-agent isolation
// SIGTERM'd the post-extraction phase before graph.json / graph.html /
// GRAPH_REPORT.md were written.
// Fix: skill runs the build inline in a single foreground Bash call.
// Structural fence: skill is parsed into (a) a YAML frontmatter map and
// (b) a list of fenced code blocks. Assertions run against parsed structures,
// never against raw markdown text.
const SKILL_PATH = path.join(__dirname, '..', 'commands', 'gsd', 'graphify.md');
function parseBug3166SkillFrontmatter(text) {
const lines = text.split(/\r?\n/);
const out = {};
let _activeKey = null;
let activeList = null;
for (const raw of lines) {
const listItem = raw.match(/^\s+-\s+(.+?)\s*$/);
if (listItem && activeList) {
activeList.push(listItem[1]);
continue;
}
const kv = raw.match(/^([A-Za-z][A-Za-z0-9_-]*):\s*(.*)$/);
if (!kv) continue;
const [, key, rawValue] = kv;
const value = rawValue.trim();
if (value === '') {
_activeKey = key;
activeList = [];
out[key] = activeList;
} else {
_activeKey = null;
activeList = null;
out[key] = value;
}
}
return out;
}
function extractBug3166FencedBlocks(body) {
const lines = body.split(/\r?\n/);
const blocks = [];
let active = null;
for (const line of lines) {
const open = line.match(/^```(\S*)\s*$/);
if (active === null) {
if (open) active = { lang: open[1] || '', lines: [] };
continue;
}
if (line.trim() === '```') {
blocks.push({ lang: active.lang, content: active.lines.join('\n') });
active = null;
continue;
}
active.lines.push(line);
}
return blocks;
}
function loadBug3166Skill() {
const markdown = fs.readFileSync(SKILL_PATH, 'utf8');
const lines = markdown.split(/\r?\n/);
const delims = [];
for (let i = 0; i < lines.length; i += 1) {
if (lines[i].trim() === '---') delims.push(i);
if (delims.length === 2) break;
}
assert.equal(delims.length, 2, 'graphify.md must have a closed frontmatter block');
const frontmatterText = lines.slice(delims[0] + 1, delims[1]).join('\n');
const body = lines.slice(delims[1] + 1).join('\n');
return {
frontmatter: parseBug3166SkillFrontmatter(frontmatterText),
body,
fencedBlocks: extractBug3166FencedBlocks(body),
};
}
// Regression for #3166
test('graphify.md allowed-tools does not include Task (inline build fence)', () => {
const { frontmatter } = loadBug3166Skill();
assert.ok(Array.isArray(frontmatter['allowed-tools']),
'allowed-tools must be a YAML block list');
assert.ok(frontmatter['allowed-tools'].length > 0,
'allowed-tools must declare at least one tool');
assert.ok(!frontmatter['allowed-tools'].includes('Task'),
'Task must NOT be in allowed-tools — sub-agent isolation truncates ' +
'graphify v0.7+ post-extraction phase (#3166). Build runs inline.');
});
// Regression for #3166
test('graphify.md frontmatter retains Read and Bash (inline build prerequisites)', () => {
const { frontmatter } = loadBug3166Skill();
const tools = frontmatter['allowed-tools'];
assert.ok(tools.includes('Read'), 'Read required for config gate');
assert.ok(tools.includes('Bash'), 'Bash required for inline build chain');
});
// Regression for #3166
test('no fenced code block in graphify.md invokes Task() agent spawn syntax', () => {
const { fencedBlocks } = loadBug3166Skill();
const offending = fencedBlocks.filter(b => b.content.includes('Task('));
assert.deepEqual(offending, [],
'no fenced code block in graphify.md may contain `Task(` invocation ' +
'syntax — sub-agent spawning truncates graphify v0.7+ post-extraction ' +
'phase (#3166). Prose mentioning the word "Task" is fine; only the ' +
'call expression inside a code block is forbidden.');
});
// Regression for #3166
test('a bash code block invokes the inline graphify update . pipeline', () => {
const { fencedBlocks } = loadBug3166Skill();
const bashBlocks = fencedBlocks.filter(b => b.lang === 'bash');
assert.ok(bashBlocks.length > 0, 'skill must contain at least one bash block');
assert.ok(
bashBlocks.some(b => b.content.includes('graphify update .')),
'a bash code block must invoke `graphify update .`'
);
assert.ok(
bashBlocks.some(b => /gsd_run\s+graphify build snapshot/.test(b.content)),
'a bash code block must invoke `gsd_run graphify build snapshot`'
);
});
// ── Regression for #3579 ────────────────────────────────────────────────────
// graphify auto-update hook was dead-on-arrival in 1.50.0-canary.x because:
// Gap 1: scripts/build-hooks.js HOOKS_TO_COPY did not include
// gsd-graphify-update.sh
// Gap 2: hooks/lib/gsd-graphify-rebuild.sh not copied by installer
// Test strategy: run the actual build and assert filesystem outcomes.
const REPO_ROOT_3579 = path.resolve(__dirname, '..');
const HOOKS_DIR_3579 = path.join(REPO_ROOT_3579, 'hooks');
const DIST_DIR_3579 = path.join(HOOKS_DIR_3579, 'dist');
const BUILD_SCRIPT_3579 = path.join(REPO_ROOT_3579, 'scripts', 'build-hooks.js');
const INSTALL_SCRIPT_3579 = path.join(REPO_ROOT_3579, 'bin', 'install.js');
// Regression for #3579: Gap 1 — build-hooks.js packages every top-level hooks/*.sh
describe('#3579 Gap 1: build-hooks.js packages every top-level hooks/*.sh into dist', () => {
before(() => {
execFileSync(process.execPath, [BUILD_SCRIPT_3579], { encoding: 'utf-8', stdio: 'pipe' });
});
test('every top-level hooks/*.sh is emitted to hooks/dist/ by the build', () => {
const topLevelSh = fs
.readdirSync(HOOKS_DIR_3579, { withFileTypes: true })
.filter((e) => e.isFile() && e.name.endsWith('.sh'))
.map((e) => e.name);
assert.ok(topLevelSh.length > 0, 'expected at least one top-level hooks/*.sh in source');
const missing = topLevelSh.filter(
(sh) => !fs.existsSync(path.join(DIST_DIR_3579, sh))
);
assert.deepStrictEqual(
missing,
[],
`every top-level hooks/*.sh must be emitted to hooks/dist/ by scripts/build-hooks.js; missing from dist: ${JSON.stringify(missing)}`
);
});
test('hooks/dist/gsd-graphify-update.sh exists after build', () => {
assert.ok(
fs.existsSync(path.join(DIST_DIR_3579, 'gsd-graphify-update.sh')),
'expected hooks/dist/gsd-graphify-update.sh to exist after build (Gap 1)'
);
});
test('hooks/dist/lib/gsd-graphify-rebuild.sh exists after build', () => {
assert.ok(
fs.existsSync(path.join(DIST_DIR_3579, 'lib', 'gsd-graphify-rebuild.sh')),
'expected hooks/dist/lib/gsd-graphify-rebuild.sh to exist after build (Gap 2)'
);
});
});
// Regression for #3579: installer deploys graphify hook + lib helper to target
describe('#3579: installer deploys graphify hook + lib helper to target', () => {
let tmpDir;
let installStdout;
before(() => {
execFileSync(process.execPath, [BUILD_SCRIPT_3579], { encoding: 'utf-8', stdio: 'pipe' });
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3579-install-'));
installStdout = execFileSync(
process.execPath,
[INSTALL_SCRIPT_3579, '--claude', '--global', '--yes', '--no-sdk'],
{
encoding: 'utf-8',
stdio: 'pipe',
env: { ...process.env, CLAUDE_CONFIG_DIR: tmpDir },
}
);
});
after(() => {
cleanup(tmpDir);
});
test('hooks/gsd-graphify-update.sh present at install target', () => {
const dest = path.join(tmpDir, 'hooks', 'gsd-graphify-update.sh');
assert.ok(fs.existsSync(dest), `expected ${dest} to exist after install`);
});
test('hooks/lib/gsd-graphify-rebuild.sh present at install target', () => {
const dest = path.join(tmpDir, 'hooks', 'lib', 'gsd-graphify-rebuild.sh');
assert.ok(fs.existsSync(dest), `expected ${dest} to exist after install`);
});
test('installer does not warn about missing gsd-graphify-update.sh', () => {
assert.ok(
!installStdout.includes('Missing expected hook: gsd-graphify-update.sh'),
`installer output must not warn about missing graphify hook; got:\n${installStdout}`
);
assert.ok(
!installStdout.includes(
'Skipped graphify auto-update hook — gsd-graphify-update.sh not found'
),
`installer must not skip graphify hook configuration; got:\n${installStdout}`
);
});
});
});
// ────────────────────────────────────────────────────────────────────────
// Folded from tests/bug-622-graphify-optional-graph-html.test.cjs — consolidation epic #1969 (B6 #1975)
// ────────────────────────────────────────────────────────────────────────
{
const { describe: __foldDescribe } = require('node:test');
__foldDescribe("folded:bug-622-graphify-optional-graph-html (consolidation epic #1969 B6 #1975)", () => {
// allow-test-rule: source-text-is-the-product (see #622)
// This test extracts the deployed Step 3 shell block from commands/gsd/graphify.md
// and executes it to prove that a skipped graph.html (due to the graphify HTML viz
// node limit) does not abort the chain (#622). The deployed markdown text IS the
// product surface — the block the runtime executes — so asserting on its execution
// behavior requires reading the source text.
'use strict';
/**
* Regression test for bug #622.
*
* The `/gsd-graphify build` Step 3 shell chain in commands/gsd/graphify.md
* aborted when `graph.html` was intentionally skipped (graph exceeds the HTML
* viz node limit, default 5000). The unconditional `cp graphify-out/graph.html`
* failed with "cannot stat", and the `&&` chain aborted before the
* GRAPH_REPORT.md copy, snapshot, and status steps ran.
*
* Fix: guard the graph.html copy with
* `{ [ -f graphify-out/graph.html ] && cp … || true; }`
* so the chain continues when the file is absent.
*/
const { describe, test, before, after } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('fs');
const path = require('path');
const { spawnSync } = require('child_process');
const { createTempDir, cleanup, readFileNormalized } = require('./helpers.cjs');
// Path to the command doc (relative to repo root)
const GRAPHIFY_MD = path.join(__dirname, '..', 'commands', 'gsd', 'graphify.md');
/**
* Extract the Step 3 fenced bash block from graphify.md.
* The block starts with the line `graphify update .` and ends at the next
* closing ``` fence.
*
* Returns the bash source text (without the fence lines themselves).
*
* readFileNormalized() strips \r\n -> \n before the match below runs — the
* extracted block is later spawned via spawnSync('bash', ...) in runBlock(),
* so an un-normalized read on a Windows checkout would break bash mid-script
* (DEFECT.TEST-SHELL-PIPELINE-NONPORTABLE, #2650).
*/
function extractStep3Block() {
const content = readFileNormalized(GRAPHIFY_MD);
// Capture the full body of the ```bash fence that CONTAINS `graphify update .`
// (including any leading preamble line), without crossing into other fences.
const match = content.match(/```bash\r?\n((?:(?!```)[\s\S])*?graphify update \.(?:(?!```)[\s\S])*?)\r?\n```/);
return match ? match[1].trim() : null;
}
// ─── shared sandbox dirs ──────────────────────────────────────────────────────
let sandbox;
let fakeBin;
let fakeHome;
before(() => {
sandbox = createTempDir('gsd-622-sandbox-');
fakeBin = createTempDir('gsd-622-fakebin-');
fakeHome = createTempDir('gsd-622-fakehome-');
});
after(() => {
cleanup(sandbox);
cleanup(fakeBin);
cleanup(fakeHome);
});
// ─── helpers ─────────────────────────────────────────────────────────────────
/**
* Write a minimal fake `graphify` executable into fakeBin.
* It just exits 0 so the `graphify update .` step succeeds.
*/
function writeFakeGraphify() {
const exe = path.join(fakeBin, 'graphify');
fs.writeFileSync(exe, ['#!/bin/sh', 'exit 0'].join('\n'), { mode: 0o755 });
}
/**
* Write a minimal gsd-tools.cjs stub into fakeHome that exits 0 for any
* invocation (covers the `graphify build snapshot` and `graphify status` steps).
*/
function writeFakeGsdTools() {
const binDir = path.join(fakeHome, '.claude', 'gsd-core', 'bin');
fs.mkdirSync(binDir, { recursive: true });
fs.writeFileSync(
path.join(binDir, 'gsd-tools.cjs'),
['#!/usr/bin/env node', 'process.exit(0);'].join('\n'),
{ mode: 0o755 },
);
}
/**
* Populate the sandbox with the minimal directory structure and output files
* that a real `graphify update .` would produce. `includeHtml` controls
* whether graphify-out/graph.html is created (simulating the node-limit skip
* when false).
*/
function populateSandbox(includeHtml) {
// graphify-out/ — simulates graphify CLI output directory
const outDir = path.join(sandbox, 'graphify-out');
fs.mkdirSync(outDir, { recursive: true });
fs.writeFileSync(path.join(outDir, 'graph.json'), '{}');
fs.writeFileSync(path.join(outDir, 'GRAPH_REPORT.md'), '# report');
if (includeHtml) {
fs.writeFileSync(path.join(outDir, 'graph.html'), '<html/>');
}
// .planning/graphs/ — destination directory
const graphsDir = path.join(sandbox, '.planning', 'graphs');
fs.mkdirSync(graphsDir, { recursive: true });
}
/**
* Execute the extracted Step 3 block in the sandbox.
*/
function runBlock(block) {
return spawnSync('bash', ['-c', block], {
cwd: sandbox,
env: {
...process.env,
PATH: fakeBin + ':' + process.env.PATH,
HOME: fakeHome,
},
encoding: 'utf8',
});
}
// ─── tests ───────────────────────────────────────────────────────────────────
describe('bug #622: graph.html absence must not abort the Step 3 shell chain', () => {
let block;
before(() => {
block = extractStep3Block();
});
test('Step 3 bash block is present in graphify.md (sanity gate)', () => {
assert.ok(block !== null, 'Step 3 bash block starting with "graphify update ." was not found in commands/gsd/graphify.md');
assert.ok(block.length > 0, 'Extracted bash block must not be empty');
});
test('graph.html absent: chain exits 0 and all other artifacts are copied (#622 regression)', (t) => {
// Use t.after for per-test cleanup so sandbox is fresh for each test
t.after(() => {
// Remove and recreate sandbox so the next test starts with an empty dir
cleanup(sandbox);
fs.mkdirSync(sandbox, { recursive: true });
});
writeFakeGraphify();
writeFakeGsdTools();
populateSandbox(false); // no graph.html — simulates node-limit skip
const result = runBlock(block);
// Chain must not abort
assert.equal(result.status, 0, [
'Expected exit 0 but got ' + result.status,
'stderr: ' + result.stderr,
'stdout: ' + result.stdout,
].join('\n'));
// graph.json was copied (step before the guarded line)
assert.ok(
fs.existsSync(path.join(sandbox, '.planning', 'graphs', 'graph.json')),
'.planning/graphs/graph.json must be copied even when graph.html is absent',
);
// GRAPH_REPORT.md was copied (step AFTER the guarded line — key regression assertion)
assert.ok(
fs.existsSync(path.join(sandbox, '.planning', 'graphs', 'GRAPH_REPORT.md')),
'.planning/graphs/GRAPH_REPORT.md must be copied (the chain must not abort at graph.html)',
);
// graph.html must NOT exist in the destination (correctly skipped)
assert.ok(
!fs.existsSync(path.join(sandbox, '.planning', 'graphs', 'graph.html')),
'.planning/graphs/graph.html must NOT be created when source is absent',
);
});
test('graph.html present: chain exits 0 and graph.html is copied (happy path)', (t) => {
t.after(() => {
cleanup(sandbox);
fs.mkdirSync(sandbox, { recursive: true });
});
writeFakeGraphify();
writeFakeGsdTools();
populateSandbox(true); // include graph.html
const result = runBlock(block);
assert.equal(result.status, 0, [
'Expected exit 0 but got ' + result.status,
'stderr: ' + result.stderr,
'stdout: ' + result.stdout,
].join('\n'));
// graph.html must exist in the destination (normal copy)
assert.ok(
fs.existsSync(path.join(sandbox, '.planning', 'graphs', 'graph.html')),
'.planning/graphs/graph.html must be copied when the source file is present',
);
// Other artifacts also copied
assert.ok(
fs.existsSync(path.join(sandbox, '.planning', 'graphs', 'graph.json')),
'.planning/graphs/graph.json must be copied',
);
assert.ok(
fs.existsSync(path.join(sandbox, '.planning', 'graphs', 'GRAPH_REPORT.md')),
'.planning/graphs/GRAPH_REPORT.md must be copied',
);
});
});
});
}