Commit Graph

29 Commits

Author SHA1 Message Date
Tom Boucher
362d0434b2 fix(#3370): state checkpoint gate semantics in executor dispatch prompts (#3478)
* fix(#3370): state checkpoint gate semantics in executor dispatch prompts

* fix(#3370): set changeset pr to 3478

* fix(#3370): keep gate rule in routing fragment under phase-6 ceiling

---------

Co-authored-by: sim <sim@local>
2026-08-14 11:42:22 -04:00
Tom Boucher
5452f1a700 fix(#3324): build-time embed execution context instead of literal @-includes (#3462)
* fix(#3324): build-time embed execution context instead of literal @-includes

* chore(#3324): add changeset

* chore(#3324): set changeset pr reference

* fix(#3324): trim embed note to stay under the 93400 margin ceiling

---------

Co-authored-by: sim <sim@local>
2026-08-14 10:33:09 -04:00
Tom Boucher
69e7afd0c7 chore(#3212): bounded quantifiers over document content — prohibition with teeth — Phase 4 (#3441)
* feat(#3415): ship local/no-unbounded-quantifier, burn down ReDoS class

Phase 4 of epic #3212 (ADR-3212 §5/§7, the final phase). New rule flags
an unbounded */+/{n,} quantifier over a broad character class
([\s\S], dotAll ., or a 1-2-unit negated class like [^\n]/[^)\n] — the
exact #2128-fixed shape) applied to a regex whose match target is
data-flow-traced to readFileSync content.

eslint-rules/lib/readfilesync-trace.cjs extracts the data-flow tracer
shared with no-crlf-fragile-split (Phase 2) rather than a second copy
— no-crlf-fragile-split refactored onto it with zero behavior change,
parity-tested.

Real triage, not 798 mechanical edits: the ADR's census (2026-08-08)
screened every unbounded quantifier in the tree unscoped. Correctly
scoped to readFileSync-derived content (matching Phase 2's own G2/G3
scoping), the rule found 162 real hits across two detection waves — the
second wave (93) surfaced only after a genuine off-by-one bug in this
rule's own first draft was caught while writing its RuleTester tests
and fixed (the bug silently missed every directly-quantified [\s\S]*
with no gap before the quantifier — exactly the class this rule exists
to catch). 3 hits landed in production src/ (commands.cts, milestone.cts,
roadmap.cts) and were each empirically timed against adversarial input
(matching #2128's own measured-not-assumed precedent) — all confirmed
linear-time/benign, left unbounded with a measured-evidence comment
rather than mechanically bounded. The remaining 159 are test-file
fixture parsing (test-author-controlled, fixed-size content, not
adversarial input) — each suppressed with a specific, non-generic
reason. Zero functional behavior changed anywhere in this diff.

tests/no-pending-3212-markers.test.cjs locks the epic's own closing
invariant (ADR §7: "assert zero pending #3212 markers remain") — ground
truth confirmed trivially true today (no phase left any such marker
behind), now regression-locked going forward.

Design: .gsd/phase/chore-3415-prohibition-with-teeth/40-design.md
Test matrix: .gsd/phase/chore-3415-prohibition-with-teeth/50-test-matrix.md

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): correct rule category mislabel, add CI test-scope entry

An orthogonal Standards-axis review found eslint-rules/no-unbounded-quantifier.cjs
mistakenly carried meta.docs.category: 'Portability', copied from a sibling
rule without realizing what that implied: docs/contributing/cross-platform-
portability-rules.md governs an ADR-1703 rule family under a hard "zero
escape hatches" contract (tests/portability-rule-disable-ban.test.cjs's
PROTECTED_RULES bans eslint-disable for those rules entirely). This rule is
not part of that family — it's ADR-3212 (ReDoS/CWE-1333), a different epic —
and its eslint-disable-next-line suppressions (159 of them, added earlier
this same phase after empirical benign-verification) are an intentional,
correct design, not a bypass. Corrected to category: 'Best Practices',
matching the actual precedent (no-adhoc-regex-escape.cjs, Phase 1 of the
same epic, which is also correctly outside PROTECTED_RULES), and the rule's
own docstring now states this explicitly so a future reader doesn't have to
re-derive it.

Also registers a new scripts/ci-test-scope.cjs bucket so editing this rule
or the shared eslint-rules/lib/readfilesync-trace.cjs helper re-runs their
own test suites under targeted CI selection — was previously unregistered
and invisible to that fast-path (this PR's own gsd-test checkpoint runs the
full suite regardless, so this only affects future narrowly-scoped PRs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): bound no-unbounded-quantifier's own scanner (CWE-1333, ironic)

Security review found the rule meant to catch algorithmic-complexity bugs
had one of its own: hasUnboundedBroadQuantifier's negated-class inner
scan walked from each `[^` occurrence to the next `]` (or EOF) with no
bound, while the outer loop only ever advanced by one character — O(n²)
total work on a pattern with many unclosed `[^` runs. Runs unconditionally
inside checkPattern on any `new RegExp('literal string')` argument in any
linted file, before the (cheap) readFileSync data-flow gate — so a single
crafted string literal, no valid regex syntax required, could make
`npm run lint` / CI hang.

Empirically confirmed both the bug and the fix: pre-fix, n=4000/8000/
16000/32000 chars took 30.8/115.6/463.8/1874.3ms (~4x work per 2x n,
quadratic); extrapolated, the 300000-char repro from the finding would
run ~165s. Post-fix (bail the inner scan once units exceeds the rule's
own 1-2-unit scope, rather than continuing to hunt for a closing `]`),
the same 300000-char input runs in 8.7ms via the real rule module,
independently reconfirmed at 18ms via a fresh Linter.verify() call.

New regression row in tests/no-unbounded-quantifier.rule.test.cjs
asserts the RuleTester run on a 50000-char adversarial pattern
completes and returns a defined result — no wall-clock assertion
(CLAUDE.md Clock Seams / local/no-elapsed-assertion).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): triage 3 new sites, re-raise ceiling after upstream batch

next merged 12 more PRs during this PR's review. Two consequences:

- tests/edit-phase.test.cjs (fix #3262, unrelated) added 3 new
  content.match(/<tag>([\s\S]*?)<\/tag>/) reads of this repo's own
  workflow .md content — the same Class A pattern as the ~159 sites
  already triaged elsewhere in this PR. Suppressed with the same
  established reason.
- lint-allow-test-rule-refs' ratchet ceiling needed re-raising again
  (301 -> 303) for the same reason as the two prior bumps: organic
  growth from unrelated, already-reviewed PRs landing concurrently,
  not a defect in this branch's own diff.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 10:02:28 -04:00
Tom Boucher
d30c99bc92 chore(#3421): delete orphan verify-phase workflow, migrate live gates to verifier (#3422)
* chore(#1892): delete orphan verify-phase workflow, migrate live gates to verifier reference

* test(#1892): retarget structural suites from verify-phase.md to verifier-phase-gates.md

* chore(#1892): reword retired-workflow mentions for removed-but-needed lint

* test(#1892): correct stale surface labels in retargeted suites

* docs(#1892): add verifier-phase-gates row to locale inventories

* chore(#3421): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-13 21:22:03 -04:00
Tom Boucher
6e59f97dd5 feat(#1955): flag coincidental reliance in goal-backward verification (#3250)
* test(#1955): failing-first contract for verifier coincidental-reliance advisory

* test(#1955): anchor coincidental-reliance assertions on the frontmatter block

* feat(#1955): flag coincidental reliance in goal-backward verification

* chore(#1955): correct stale workflow tier high-water comment

* fix(#1955): close the verify-phase divergence and state the endogeneity limit

* docs(#1955): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-09 12:07:41 -04:00
Tom Boucher
067a4d1c6c fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls (#3015)
* test(#2650): add failing-first regression for plan-phase stall detection

Regression test for gsd_stall_should_recover / gsd_stall_watch and the
planner.stall_* config keys, none of which exist yet — proves RED before
the fix lands in the next commit.

* fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls

Mirrors the already-shipped executor.stall_* pattern (execute-phase.md, bug
#3212) but with a dispatch change the executor's prose-only surveillance
lacks: the standard planner spawn, chunked-outline planner spawn,
chunked-per-plan planner spawn, plan-checker spawn, and revision-loop
planner respawn now dispatch with run_in_background=true and are followed
by a real, bounded bash poll (gsd_stall_watch) that returns control to the
orchestrator on its own schedule instead of waiting indefinitely on a
subagent that may never return. On stall, the existing accept-plans/retry/
stop recovery menu (9a/11a) is auto-surfaced instead of requiring a manual
interrupt.

New config keys planner.stall_detect_interval_minutes (default 5) /
planner.stall_threshold_minutes (default 10) mirror executor.stall_*.

The helper functions (gsd_stall_should_recover, gsd_stall_watch) live in a
new lazily-loaded gsd-core/workflows/plan-phase/steps/stall-detection-
helpers.md rather than inline, and per-site prose is kept minimal, because
plan-phase.md is frozen under the ADR-857 Phase 6 PRE_PHASE6 gate
(tests/phase6-capstone-conformance.test.cjs) with ~36 bytes of headroom at
baseline; the net effect is plan-phase.md.md ships slightly SMALLER than
before (the old unconditional-wait ORCHESTRATOR RULE sentences are gone at
the five touched sites, superseded by the bounded watcher).

Also fixes a stale doc comment in tests/workflow-size-budget.test.cjs that
still described the per-file workflow-size-baseline.json guard removed by
#2724 (ADR-2719 Phase 4) as if it were still the enforcement mechanism —
discovered while verifying this fix's own byte budget.

Researcher and pattern-mapper spawns are untouched (out of scope per the
issue's Agent Brief).

* fix(#2650): make gsd_stall_watch single-cycle; harden numeric config inputs

Two review findings addressed on top of the prior commit:

1. gsd_stall_watch previously looped internally for the full
   threshold+interval duration inside ONE Bash tool call (up to 15 min at
   defaults) — a single call blocking that long risks the host tool's own
   timeout killing it before it ever prints a result, silently defeating the
   fix. Redesigned to a single sleep-and-check cycle per call, taking an
   explicit dispatch_ts so the orchestrator prose can repeat the (short,
   default 5 min) call until it resolves; the outer threshold is now
   enforced by dispatch_ts accumulating across calls, not by one call's
   duration. Documented the resulting trade-off (up to one interval of
   added latency on the success path) in the changeset and reference doc.

2. PLANNER_STALL_INTERVAL_MINUTES/THRESHOLD_MINUTES are config-controlled
   values that flow into bash arithmetic ($(( ))). A review flagged this as
   command injection; empirically verified against both macOS bash 3.2.57
   and Docker bash:5 that this is NOT actually exploitable (bash hard-errors
   on a `$(cmd)`-shaped arithmetic operand rather than invoking it) — but an
   unvalidated malformed value WOULD abort the stall-watcher itself with
   that bash error, silently defeating the exact hang-recovery this issue
   ships. Added integer validation with safe-default fallback, both at the
   config-resolution point and defensively inside gsd_stall_should_recover.

Also adds the previously-missing integration coverage for gsd_stall_watch's
real execution (grep/find/date plumbing), not just the pure classifier.

* fix(#2650): correct AC2 self-test — helpers doc may name teams-status in prose

The AC2 regression test asserted the stall-detection-helpers.md step file
never contains the substring "teams-status" at all, but the file's own
prose explicitly documents its independence from that guard (containing
the word by design). Narrowed the assertion to what actually matters: no
second `query teams-status` call site and no gating on it, not a blanket
absence of the word.

* test(#2650): regenerate golden install-tree fixtures for the new step file

gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md is an
emitted file (installed for every runtime), so adding it changes the
install tree even though it is invisible to docs/INVENTORY.md and
docs/INVENTORY-MANIFEST.json (both explicitly scope to non-recursive
gsd-core/workflows/*.md — verified against the execute-phase #2930 and
pre-existing plan-phase step-file precedent, which are equally absent from
both inventory artifacts). The golden install tree snapshots the sorted
list of emitted relative paths per runtime, so a file invisible to the
inventory is still visible here. Regenerated via `npm run gen:install-tree`
— one line added per runtime fixture (19 files), no other drift.

* fix(#2650): restore 7 ORCHESTRATOR RULE labels; sync runtime-launcher preamble

Two more consequences of extracting helper bodies out of plan-phase.md,
both caught by verification (0017e1a78, 9 unique failures):

1. tests/plan-phase-drift-guard.test.cjs (#913) requires at least 7
   "ORCHESTRATOR RULE — ALL RUNTIMES" labels in plan-phase.md itself, one
   per agent spawn site. Moving the full explanatory blocks to
   plan-phase/steps/stall-detection-helpers.md carried 5 of the 7 labels
   out with them (only the untouched researcher/pattern-mapper sites kept
   theirs). Restored a short label at each of the 5 stall-watch sites,
   trimmed a few more redundant words ("Per 7.99, " — already established
   by the adjacent step-7.99 pointer) to stay under the frozen
   PRE_PHASE6 cap (94497 bytes, 21 bytes headroom).

2. tests/runtime-launcher-parity.test.cjs (#373) requires exactly one
   canonical gsd_run preamble, byte-equal to
   gsd-core/workflows/_runtime-launcher.snippet.sh, before the first
   gsd_run call in any workflow .md that calls it (recursive scan under
   gsd-core/workflows/, unlike the non-recursive inventory/step-tag-balance
   checks). The new step file's config-get calls use gsd_run without one.
   Fixed via `node scripts/sync-runtime-launcher.cjs`, verified: exactly 1
   preamble occurrence, before the first call, including the .claude/ and
   .codex/ home fallback arms.

Also verified (no fix needed, evidence recorded): the generic
`gsd-core-verbatim` identity rule in tests/helpers/emitted-provenance.cjs
(roots: ['gsd-core'], pattern matching workflows/.+) self-attributes any
new gsd-core/workflows/** path to itself, so the new step file needs no
drift-ack entry — consistent with plan-phase.md's own net shrinkage
requiring none either.

* test(#2650): acknowledge plan-phase.md's +14 byte drift

Restoring the 5 ORCHESTRATOR RULE — ALL RUNTIMES labels (#913) flipped
plan-phase.md from -142 bytes (post-extraction) to +14 bytes net growth
against baseline (94483 -> 94497), which the differential attribution
size ratchet (tests/emitted-attribution.test.cjs) correctly flags as
unacknowledged growth. Added tests/emitted-drift-acks/2650-plan-phase-
stall-detection.json, keyed on the bare filename plan-phase.md per the
existing fragment schema (see tests/emitted-drift-acks/2649-diagnose-
execute-plan-base-check.json), explaining the growth as exactly the 5
restored labels — still verified under the PRE_PHASE6 cap (94497 < 94519)
and satisfying #913's 7-label requirement.

* fix(#2650): bind {outputFile} from the real Agent() return — was dead code

Independent review blocker: PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE were
read by every gsd_stall_watch call but never assigned anywhere in the
diff. With the variable permanently empty, `[ -f "$output_file" ]` was
always false, marker_found could never become true, and marker_received
was unreachable — the marker-based detection path was permanently dead.

Worse for the plan-checker spawn specifically: a checker that PASSES
touches no *-PLAN.md files, so it had no working completion signal at
all without the marker path. A healthy plan-checker finishing cleanly in
two minutes would be declared stalled once planner.stall_threshold_minutes
elapsed and the recovery menu would fire on an already-succeeded agent —
worse than the original unbounded hang.

Fixed by replacing the dead bash variable with the `{outputFile}`
orchestrator-substitution token, the same convention docs-update.md:471
already uses for a real run_in_background=true Agent() return ("Read
tool: file_path: `{outputFile from README agent result}`"). This is a
net BYTE SAVING at each site (`"{outputFile}"` is shorter than
`"$PLANNER_OUTPUT_FILE"`), which funded moving the full binding
explanation — including why plan-checker's *-PLAN.md glob alone is not
a working completion signal — into the lazily-loaded reference file to
stay under the frozen PRE_PHASE6 cap (94496 bytes, 22 headroom; net +13
over baseline, acknowledged in tests/emitted-drift-acks/2650-plan-phase-
stall-detection.json).

Added a regression test asserting plan-phase.md itself binds {outputFile}
at all 5 spawn sites and contains no dangling $PLANNER_OUTPUT_FILE /
$CHECKER_OUTPUT_FILE reference — the previous test suite only exercised
gsd_stall_watch's behavior when handed a valid argument, which is why
the dead production wiring survived two rounds of review. Also fixed
tests/fix-2650-plan-phase-stall-detection.test.cjs:170-195's raw
try/finally to use t.after(), per CONTRIBUTING's test-cleanup convention.

* chore(#2650): backfill changeset PR number to 3015

* fix: normalize CRLF at the read boundary in all .md-bash-extraction tests

Maintainer-authorized scope expansion, folded into this PR rather than
deferred: the Windows CI lane on this PR's own tests/fix-2650-plan-phase-
stall-detection.test.cjs exposed DEFECT.TEST-SHELL-PIPELINE-NONPORTABLE
(CONTEXT.md; recurring since #1700) as a repo-wide latent class, not a
one-off. Ten test files parse a fenced ```bash block out of a workflow
.md file and execute it via spawnSync/execFileSync; a Windows checkout
can yield CRLF line endings despite .gitattributes eol=lf, and bash then
treats the trailing \r on every extracted line as part of the token —
"unexpected EOF while looking for matching `"'" or a bare syntax error,
partway through the script.

Added tests/helpers.cjs:readFileNormalized() — strips \r\n -> \n at the
read boundary, before any fence-slicing or regex runs, so every
downstream operation is correct by construction. Migrated all ten call
sites to it:

Previously broken (fs.readFileSync with no normalization anywhere
between read and spawn):
- tests/worktree-cleanup.test.cjs (extractCwdGuardBash) — also fixes a
  misleading comment claiming the fence regex alone was "CRLF-safe"; it
  protected only the fence delimiters, never the captured body.
- tests/new-milestone-clear-phases.test.cjs (extractFenceBetween,
  extractFenceContaining)
- tests/code-review-pipeline-regression.test.cjs (extractPostProcessingScript)
- tests/drift-detection.test.cjs (readGate/bashBlock, plus the snippet-file
  comparison read in the same test)
- tests/graphify-visualization.test.cjs (extractStep3Block)
- tests/pause-work-improvements.test.cjs (extractCheckBlock)
- tests/plan-review-convergence.test.cjs (extractReviewerFlagsParseBlock
  and the inline post-config-gate resolution-block slices)

Already correct (split(/\r?\n/) then join('\n')), migrated to the shared
helper for consistency rather than a fourth/fifth/sixth copy of the same
fix:
- tests/git-base-branch.test.cjs (extractHandleBranchingBash)
- tests/quick-branching.test.cjs (extractStep25Bash)
- tests/runtime-launcher-parity.test.cjs (extractResolverSnippet)

Verified against a simulated Windows CRLF checkout (not assumed): for
both the worktree-cleanup.test.cjs and new-milestone-clear-phases.test.cjs
extraction shapes, confirmed the pre-fix code produces a real bash syntax
error on CRLF input and the post-fix code does not.

One eslint follow-up: local/no-crlf-fragile-split statically flags any
bare `\n` inside a markdown-fence-shaped regex, regardless of whether the
receiver was already normalized — it cannot see the readFileNormalized()
data-flow. Kept `\r?\n` in extractCwdGuardBash's fence regex (redundant
but harmless on pre-normalized input) rather than fight the rule.

Scope note: this diff is broader than issue #2650's own change (plan-
phase.md stall detection) because the Windows lane surfaced a genuine
repo-wide defect class while verifying that fix, and the maintainer
authorized fixing it here rather than filing it separately and shipping
a known-broken pattern.

Runtime impact: none — this is a test-harness-only defect. The live
orchestrator (Claude Code or another runtime) does not do a byte-exact
extract-and-pipe of .md content into a shell the way these tests do; it
reads the instructions and generates its own bash invocation text, which
does not reproduce a raw CRLF pass-through the same way.

Not touched: tests/plan-review-convergence.test.cjs's separate, tracked
spawnSync ETIMEDOUT flake under bench load (#3005, reproduced on
unmodified next) — unrelated load-sensitivity, not a CRLF symptom.

* fix(#2650): remove stale drift-ack fragment — plan-phase.md is self-explaining

tests/emitted-drift-acks/2650-plan-phase-stall-detection.json acknowledged
plan-phase.md's own emitted-path hash move, but plan-phase.md is directly
edited in this diff. Per the emitted-attribution law (ADR-2719,
tests/emitted-attribution.test.cjs), a workflow's emitted key equals its
own source path (gsd-core-verbatim identity rule), so a direct edit to the
source is self-explaining and auto-attributed — no ack was ever needed.

Verified via the pre-merge lint (scripts/lint-emitted-drift-ack.cjs, run
through npm run lint:ci with a fully cleared eslint cache): it passes clean
with the fragment removed, confirming no contradiction between the lint and
the runtime attribution gate — this was simply an unnecessary fragment.

* fix(#2650): restore plan-phase.md drift-ack — size ratchet demands it against next

tests/emitted-drift-acks/2650-plan-phase-stall-detection.json was deleted in
the previous commit because, against an earlier verification base, it was
inert: it explained a moved emitted hash that a direct edit to plan-phase.md
already self-attributes. Against origin/next@f1af47766a the demand is
different: plan-phase.md is 13 bytes larger than the base copy, which trips
the emitted-attribution size ratchet — a job this same ack also performs.

Recreated in the documented shape, keyed on the bare filename plan-phase.md
(not the full path, and not restating the byte delta per review guidance),
describing the actual change: the {outputFile} binding fix for the dead
PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE variables and the 5 restored
ORCHESTRATOR RULE labels required by #913, both at the stall-watch spawn
sites, with explanatory bodies living in the lazily-loaded
gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md reference.

Confirmed no other fragment (on this branch or on next) claims the bare key
"plan-phase.md" before recreating — scripts/lint-emitted-drift-ack.cjs's
duplicate check is an exact string match, and the only other mention of
plan-phase.md in tests/emitted-drift-acks/ (2658-trae-instruction-file-path.json)
uses the full path as its key, so there is no collision.

* fix(#2650): real cause of Windows CI failure — bash -c argv-transport, not CRLF

The CRLF diagnosis for PR #3015's Windows failure was wrong. Proven wrong,
not assumed: .gitattributes' blanket `* text=auto eol=lf` means a Windows
checkout never receives CRLF for stall-detection-helpers.md, and the
extracted fence's line 64 is byte-identical and correctly balanced on every
platform. The real cause: runShouldRecover() passed a 70+ line, quote-dense
script as ONE argv element to `spawnSync('bash', ['-c', script, arg0, ...])`
PLUS four more positional args. Windows has no execve — Node serializes
that whole argv into a single CreateProcess command-line string, and Git
Bash's MSYS layer re-splits and unescapes it with its own rules. The
boundary between the script and the trailing args was not stable across
that round trip (live evidence: one failure's stderr was prefixed
`gsd_stall_should_recover_test:` — arg0 arrived — another `/usr/bin/bash:`
— arg0 did not).

Fixed by writing the script to a temp file and running `bash <file> <args>`
instead — the four values are now normal, quote-free positional args, and
the script itself never enters argv transport at all. Mirrors
tests/quick-branching.test.cjs's extractStep25Bash/runStep, which already
uses this exact shape and is green on Windows on `next`.
tests/worktree-cleanup.test.cjs's extractCwdGuardBash/runGuard stays on
`bash -c` but never appends extra positional args beyond the script itself,
so it never hits the same boundary — checked both siblings per review, not
assumed.

Corrected the now-actively-misleading CRLF comment in
extractStallHelpersBash(), and corrected the changeset's claim that the
repo-wide CRLF-normalization fix (folded into this branch, maintainer-
authorized) explains this PR's own Windows failure — it doesn't, though it
remains defensible on its own merits as general test-portability hardening.

Separately, while auditing the shipped (non-test) gsd_stall_watch for
Windows portability per review request, found and fixed a second, real
user-facing defect: the artifact-freshness check used GNU find's
`-newermt "@<epoch>"` shorthand, which the BSD find(1) actually shipped on
macOS does NOT understand ("Can't parse date/time: @<epoch>", verified live
against /usr/bin/find on both a stale and a genuinely fresh file). With the
adjacent `2>/dev/null`, that failed silently and permanently degraded
artifact_fresh to false on every macOS run — a plan-checker or planner
actively writing plan files could still be reported "stalled." Replaced
with `find $glob -mmin -N` ("modified less than N minutes ago"), which
needs no date-string parsing and is supported identically by GNU find and
BSD find; verified live that the old shape fails and the new shape passes
against the same real fresh file. Added a real-execution regression test
(gsd_stall_watch with `sleep` stubbed to a no-op so the test doesn't
actually wait, but the real `find ... -mmin` line still runs) proving the
fix, replacing the prior "not integration-tested" note for that path.

Note: the remote gsd-test runner is Linux-only, so it cannot itself confirm
the Windows fix — only the actual windows-latest CI lane can.

* fix(#2650): route the third bash -c call site through the same temp-file seam

runWatch() and a `-mmin` regression test still passed their script via
`bash -c <script>` after the previous commit only converted
runShouldRecover() — live Windows CI on 4b86cc57f confirmed the mechanism:
failures went 11 -> 4, and `full test (windows-latest, 22, shard 1/3)` and
`shard 2/3` flipped from fail to pass, but the remaining 4 failures (all in
this file, all still `bash: -c:`) were exactly the gsd_stall_watch describe
block, which runWatch() serves. runWatch() passes NO extra positional args
at all, so this also rules out the trailing-args theory from the prior
commit: the ~73-line, quote-dense script itself is what does not survive
Windows argv serialization when passed as a single `-c` element, regardless
of how many (if any) further argv elements follow it.

Extracted one shared runBashScript(script, args, opts) helper — write to a
fs.mkdtempSync'd file, run `bash <file> [args...]`, clean up in `finally` —
and routed all three bash-invoking call sites in this file through it
(runShouldRecover, runWatch, and the -mmin freshness test that builds its
own script inline for the `sleep` stub). One transport seam means a fourth
call site in this file cannot silently reintroduce the bug in isolation,
which is exactly what happened here with a second call site.

Corrected extractStallHelpersBash()'s doc comment a second time to state
the mechanism precisely (script content, not argv-element count) and cite
the live evidence (11->4 failures, shards 1 and 2 flipping green) so the
next reader does not have to rediscover it.

Audited every other bash-invoking call site in files this branch touches,
per review request:
- tests/code-review-pipeline-regression.test.cjs (runPostProcessing),
  tests/graphify-visualization.test.cjs (runBlock), and
  tests/drift-detection.test.cjs (two execFileSync('bash', ['-c', ...])
  sites, one of them carrying the same giant runtime-launcher preamble
  text) — all pre-existing, UNCHANGED by this branch (only touched for the
  readFileNormalized() CRLF swap), and already exercised on `next`'s last
  six Windows CI runs per the reviewer's own citation. Left as-is: no
  evidence of failure, and converting untested pre-existing code outside
  #2650's scope on an unverifiable guess would be its own risk.
- tests/git-base-branch.test.cjs (runHandleBranchingStep) and
  tests/quick-branching.test.cjs (runStep) already use the same temp-file
  pattern. No action needed.
- tests/runtime-launcher-parity.test.cjs (runResolver) uses `bash -c` but
  is explicitly `if (process.platform === 'win32') return '';` guarded off
  on Windows entirely, for an unrelated extension-less-PATH-stub reason —
  never reaches Windows argv transport at all. No action needed.
- tests/worktree-cleanup.test.cjs (runGuard) confirmed by the reviewer as
  correct and verified; not touched, per instruction.

Do not touch: the -mmin fix, the drift-ack fragment, the changeset — all
three confirmed correct in prior rounds and left untouched here.

Note: the remote gsd-test runner is Linux-only and cannot confirm this;
only the windows-latest lanes on #3015 can.

* fix(#2650): give runBashScript a default timeout

runShouldRecover() was the only one of the three call sites through
runBashScript() with no timeout — runWatch() and the -mmin test both pass
timeout: 10000 explicitly. Not a regression (this path never had a bound
before), but CONTEXT.md's unbounded-subprocess guidance applies directly,
and runShouldRecover() is driven repeatedly by a fast-check property test:
one pathological input that fails to terminate would hang CI indefinitely
instead of failing.

timeout: 10000 is now the helper's own default, with ...opts spread after
it so the two existing explicit timeout: 10000 call sites are unchanged
and any future caller inherits a bound automatically.

* fix(#2650): build the -mmin freshness test's glob with forward slashes

Windows CI on d6ddda6ea reported the last failure: the -mmin regression
test expected 'active' but got 'waiting' — find matched nothing, the same
silent-degradation shape as the macOS -newermt defect, but this time in the
test's own fixture rather than the shipped bash.

Traced what production actually passes: every gsd_stall_watch call site in
plan-phase.md builds artifact_glob as `"${PHASE_DIR}"'/*-PLAN.md'` —
PHASE_DIR is a POSIX-style .planning/phases/NN-slug value, and the whole
thing runs under Git Bash regardless of host OS, so production's glob is
always forward-slash. The test instead built it with
`path.join(tmp, '*-PLAN.md')`, which on Windows yields a backslash path
(C:\Users\RUNNER~1\...\*-PLAN.md). In bash pathname expansion a backslash
escapes the next character, so that pattern can never match a real path —
find silently returns empty under the existing 2>/dev/null, same shape as
the macOS bug. Confirmed as a test artifact, not a production defect:
production never constructs the glob this way, so no Windows user is
affected.

Fixed by forward-slashing the tmp dir before appending the glob suffix,
matching production's own convention, with a comment recording why (so a
future "simplify this back to path.join" edit doesn't silently reintroduce
the failure). The shipped bash's unquoted $artifact_glob is untouched —
quoting it would break the multi-file glob expansion it exists for.

Note: the remote runner is Linux-only and already passed clean at
d6ddda6ea (0/29,603, both node lanes); only the windows-latest lanes on
#3015 can confirm this fix.

* fix(#2650): forward-slash the three remaining runWatch globs (vacuous-pass CR)

The :353 fix (833c11da9) only converted the -mmin freshness test's glob.
Three sibling tests in the same describe block still built theirs with
path.join(tmp, '*-PLAN.md'), which yields a backslash path on Windows.

Two of those three were silently passing for the wrong reason: the
'-> stalled' and '-> waiting' tests both expect the glob to match nothing,
and on Windows a backslash path matches nothing regardless of whether the
directory is actually empty (bash eats each backslash as an escape before
the pattern is even evaluated). They would have passed identically with
glob expansion completely broken, which is a vacuous pass — not exercising
what they claim to. The third ('-> marker_received') is outcome-independent
of the glob, so it was merely inconsistent rather than wrong.

Converted all three to the same `${tmp.replace(/\\/g, '/')}/*-PLAN.md`
construction already used at the -mmin test, so every glob in the file now
matches production's own forward-slash `"${PHASE_DIR}"'/*-PLAN.md'` shape,
and the two negative tests are meaningful on Windows instead of accidentally
correct. Reworded the trailing comment on the 'stalled' test's glob line:
it now describes the fixture (the tmp dir contains no *-PLAN.md files)
rather than the pattern, since "matches nothing" read as a property of the
glob syntax when it's a property of what's on disk.

No assertion, the sleep stub, runBashScript, or the shipped bash changed.
Smoke-tested all three updated tests manually before committing (not via
node --test): marker_received / stalled / waiting, all correct.

* fix(#2650): fix own regression tests for #2993's plan-phase.md relocation

531101843's merge with origin/next brought in #2993 (unrelated, epic #1671
Phase 6.2), which extracted plan-phase.md's whole "Chunked Planning Mode"
section into gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md,
leaving a <!-- gsd:section --> pointer behind. tests/plan-phase-drift-guard.
test.cjs (#913) was already updated to read the combined surface (host file
+ every steps/*.md) so its label count didn't go blind — my own #2650
regression tests were not, and searched plan-phase.md alone for the two
chunked spawn sites' headings, which no longer exist there. Two tests
failed outright (indexOf returning -1); a third ("standard planner spawn")
was silently weakened to an unbounded slice-to-EOF by the same relocation,
since its own end-boundary heading also moved — passing by accident rather
than by testing what it claimed.

Promoted the drift guard's local readPlanPhaseCombined() to a shared,
exported tests/helpers.cjs readWorkflowCombined(workflowPath) (host file +
sorted steps/*.md, CRLF-normalized at the read boundary) so a second,
divergent implementation is never written — the drift guard now delegates
to it via a same-named local wrapper, unchanged at every existing call site.

Fixed the three affected tests in tests/fix-2650-plan-phase-stall-detection.
test.cjs:
- "standard planner spawn (step 8)": end boundary changed from the now-gone
  "## 8.5. Chunked Planning Mode" heading to "## 9. Handle Planner Return",
  which still exists in plan-phase.md.
- "chunked outline spawn (8.5.1)" / "chunked per-plan spawn (8.5.2)": now
  read gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md
  directly (not the generic multi-file combined blob, whose file-sort
  ordering would put unrelated step files between 8.5.2's slice and any
  downstream anchor) — the same heading-to-heading slicing as before still
  works because the file is small and self-contained.
- Extended the "no unbound $PLANNER_OUTPUT_FILE/$CHECKER_OUTPUT_FILE" check
  to also scan chunked-planning-mode.md, since two of the five spawn sites
  now live there.
- Added a new count-based test asserting exactly 5 (not "at least one")
  `gsd_stall_watch "$TS" "{outputFile}"` invocations across the combined
  surface, mirroring #913's own label-count guard, so every one of the five
  spawns stays provably bounded and a future relocation can't silently drop
  one without a test noticing.

Also added a small positive test that plan-phase.md's <!-- gsd:section -->
pointer to chunked-planning-mode.md exists (#2993 is unrelated to #2650 but
its presence is now load-bearing for where 2 of the 5 spawn sites live).

Audited every other test file in the repo for a stale reference to content
#2993 relocated (searched for the moved headings/prose and for
"chunked-planning-mode"/"CHUNKED_MODE" across all *.test.cjs): only this
file and the drift guard needed changes.
tests/issue-2762-plan-reviews-chunked.test.cjs already reads
chunked-planning-mode.md directly (brought in correct by the same merge).
gen-section-manifest.test.cjs, init.test.cjs, and workflow-fragments.test.cjs
reference "chunked-planning-mode" only as a manifest/section-id fixture
value for #2993 itself, not as a stale pointer to relocated content.

Did not touch: the ported ORCHESTRATOR RULE lines, run_in_background=true,
the glob constructions, runBashScript, the -mmin change, the timeout
default, or the drift-ack fragment (confirmed correct against the stale
local `next` ref two rounds ago and left alone).

---------

Co-authored-by: sim <sim@local>
2026-08-03 10:46:22 -04:00
Tom Boucher
1c1af70a4b refactor(#2724): delete the committed golden fixtures and size baselines (#2767)
* test(#2724): delete golden-install-parity fixtures, test, and generator

Removes the 19 committed path->hash manifests, the two per-file size
baselines, tests/golden-install-parity.test.cjs, and
scripts/gen-golden-install-parity-zcode.cjs. These were pure functions
of the source tree (ADR-2719); the differential attribution check
(tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs)
is now the sole gate for emitted-artifact propagation.

tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs
are unchanged (ADR-2719 section 7 exception).

Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's
existence guard, .gitattributes, package.json scripts, the emitted-provenance
totality guard's IO, the differential check's baseline acquisition, CI
wiring to publish/restore the baseline artifact, and docs.

* refactor(#2724): make the differential attribution check self-sufficient

Three fixes required to delete the golden fixtures without breaking CI:

- scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs
  from the three rules that named it. #2759's missingRuleTestFiles guard
  hard-throws at module load if a rule names a test file absent from
  disk, which would break the changes job on every PR the moment the
  fixture-deletion commit landed.

- tests/helpers/emitted-provenance.cjs: loadManifests() read the
  committed golden fixture directory. With that directory deleted at
  every future ref, this would throw at module load forever, taking
  the Phase 2 totality guard down with it. Rebuilt from real installer
  spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest),
  the same shape emitted-runtime.cjs's currentManifests() already uses.

- tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs:
  the real-tree test's baseline acquisition swaps from
  baselineManifestsAtRef(base) (git show at a ref that no longer carries
  fixtures) to resolveBaseline()'s documented precedence: env, then the
  on-disk cache, then an in-job build. The build fallback
  (buildBaselineAtRef, new) checks out base into a throwaway git
  worktree and runs the new scripts/gen-emitted-baseline.cjs there --
  no npm ci needed, since bin/install.js and the test helper shells are
  Node-builtins-only. That script also publishes the baseline artifact
  from CI's push-to-next job (wired in a follow-up commit).

* refactor(#2724): retire the merge-driver bridge and per-file size baselines

The Phase 1 bridge (#2721) is retired now that the artifacts it guarded
are deleted: scripts/git-merge-regen-driver.cjs, its test, and the
'setup:merge-driver' npm script are removed, and the .gitattributes
merge=gsd-regen/linguist-generated block for the three deleted-path
globs is dropped. tests/fixtures/install-tree/*.json keeps its normal
merge behavior, unchanged (ADR-2719 section 7).

scripts/update-size-baseline.cjs and its test are removed: their sole
purpose was regenerating tests/workflow-size-baseline.json and
tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm
script and its step in 'regen:derived' go with it. The per-file
baseline describe blocks in tests/workflow-size-budget.test.cjs and
tests/agent-size-budget.test.cjs are removed for the same reason; the
independent loose-tier hard caps are untouched. The differential
attribution check's size ratchet (tests/emitted-diff.cjs, already
shipped in #2723) is the replacement anti-creep mechanism.

'npm run gen:golden' is replaced by 'npm run gen:install-tree', which
keeps regenerating tests/fixtures/install-tree/*.json (the one artifact
family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's
error messages point at the new command name.

tests/golden-parity-single-source.test.cjs's anti-divergence guard
(#2266) is retargeted from the two deleted golden-parity consumers to
their two replacements (tests/helpers/emitted-runtime.cjs and
tests/helpers/emitted-provenance.cjs), which import buildParityManifest
the same way — the divergence risk the guard exists for is unchanged.

Also wires CI: a new publish-emitted-baseline job runs
scripts/gen-emitted-baseline.cjs after a push to next and caches the
result keyed on the sha; the test and test-full jobs restore that cache
on pull_request events, keyed on the PR's base sha, and export
GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree
test to pick up.

* docs(#2724): flip ADR-2719 to Accepted and update contributor docs

Status: Proposed -> Accepted. Regenerated docs/adr/README.md index.

CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET.
EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET.
AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry)
no longer point at the deleted golden-install-parity fixtures, size
baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver /
git-merge-regen-driver.cjs bridge. Editing shipped content now
requires zero manual fixture regeneration, documented against the
differential attribution check instead of the deleted commands.

* docs(#2724): add changeset for removed golden-parity commands

* fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref

check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs
token in CONTEXT.md resolves to a real file. The RULESET.
EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks,
which the checker reads as a live reference, not historical prose.

* test(#2724): retarget ci-test-scope tests off the deleted golden test

tests/ci-test-scope.test.cjs asserted specific RULES entries select
tests/golden-install-parity.test.cjs, and that every rule selecting it
also selects both emitted gates. Both premises broke when the golden
test was deleted (#2724): the deleted filename never re-appears in
targeted_tests, and there was no longer a third file for the gates to
travel alongside. Retargeted the two selection describe blocks to
assert tests/emitted-provenance.test.cjs directly (the drift guard the
golden gate's rules were retargeted to), and simplified the third block
to assert the two emitted gates always travel together, without
reference to the golden filename.

* docs(#2724): repoint two contributor how-to guides at the differential check

Both guides told contributors to regenerate a baseline against
tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed
at the differential attribution check (tests/emitted-attribution.test.cjs,
ADR-2719), which needs no manual regeneration step.

* fix(#2724): repair phase6-capstone-conformance's deleted-baseline read

An independent orthogonal review caught a real regression this branch
introduced into a test file the branch's diff never touched:
tests/phase6-capstone-conformance.test.cjs read
tests/workflow-size-baseline.json (deleted earlier in this branch) with
no fallback, so the whole suite would throw ENOENT the moment this
branch landed. The test's actual intent — prove the host-loop workflow
files are real, tracked, non-empty docs — is preserved by asserting the
live byte count via the same shared counter (scripts/workflow-size.cjs)
the size guards already use, instead of a committed snapshot.

Also, from the same review: a stale doc comment in
scripts/workflow-size.cjs still named the deleted
scripts/update-size-baseline.cjs as a consumer, and
buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left
two fs.rmSync calls unguarded against masking the primary result/error,
inconsistent with the try/catch already wrapping the git cleanup beside
them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef
explaining why it (and its siblings) are kept despite having no
production caller post-cutover — they still answer real questions
about refs that predate the cutover.

* fix(#2724): repair three real regressions found by remote verification

1. tests/emitted-provenance.test.cjs's two hostile-input tests
   (non-object manifest, unreadable fixture) drove loadManifests(tmp)
   and monkeypatched fs.readFileSync, both premised on the deleted
   fixture-directory read this branch already replaced with real
   installer spawns -- the negative assertions silently stopped firing.
   loadManifests() now accepts injected {families, install, build,
   clean} (defaulting to production values), giving the tests a real
   seam to drive a bad build result and a build failure through the
   ACTUAL loader instead of a reimplementation, and added coverage that
   clean() still runs on both paths.

2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE'
   steps hardcoded shell: bash, which is wrong on windows-latest (native
   pwsh) and on test-full's macos-latest legs (native zsh per that job's
   own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning
   .test.cjs) caught it. Replaced the inline bash script with
   scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a
   bare 'node <path>' command line has no shell-specific syntax, so it
   runs correctly under bash, zsh, and pwsh without a shell override.

tests/phase6-capstone-conformance.test.cjs's deleted-baseline read
(caught by the same remote run, at a commit prior to this one) was
already fixed in d0c3b1242 and is not touched here; verified still
passing after these changes.

* fix(#2724): revive ADR-1610's new-file size cap inside the differential

An isolated review caught a real regression: deleting
tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP
(ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor)
with no successor. tests/helpers/emitted-diff.cjs's size ratchet
already 'continue's past any file absent from sizeBaseline -- exactly
the files this cap exists to bound -- so a brand-new workflow file
sized 32,769-40,960 bytes passed CI clean and shipped, then risked
silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted
and never referenced anywhere in this branch.

Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own
size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline,
name) signal the growth check already computes -- 'new' is exactly
'present in sizeCurrent, absent from sizeBaseline'. Not ack-able,
matching the tier hard caps it sits beside: the fix is extraction, not
an acknowledgment entry. Documented, disclosed narrowing: the pure
differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering
(tests/workflow-size-budget.test.cjs's classification), so a
legitimately large new file must extract rather than tier in, one
release earlier than an existing file would need to. ADR-1610 itself is
left unamended -- this restores its decision rather than re-litigating
it.

Also fixes a stale comment plus a redundant real 19-installer-spawn
assertion left over from the pre-injection-seam version of
tests/emitted-provenance.test.cjs's build-failure test, and annotates
3 of 4 stale golden-fixture citations in
docs/reference/host-integration-capability-matrix.md as superseded
(the 4th is an accurate historical PR narrative, left alone).

* fix(#2724): repair three red CI defects on the golden-fixture cutover

Windows-only provenance false attribution (defect A): the `hooks-built`
provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are
Windows-only installer output (ensureCodexHooksJsonSessionStart /
ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping
the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so
the self-attribution resolved to a path that exists on no platform. Only
windows-latest ever emits the key, so this only failed there. Fixed by
special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated
rule) — a dedicated rule would match zero paths, and therefore report as a
dead rule, on every non-Windows lane of the same totality guard. `sources`
already supported per-match functions; `transforms` is extended to support
the same shape so the attribution can vary by match within one rule.

Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef`
ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but
that script is new in this PR and therefore absent at any base ref that
predates it — every call failed closed with "Cannot find module". Fixed by
running the PR checkout's own generator against the worktree via a new `--dir`
parameter, decoupling "which copy of the script runs" from "which tree it
measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override,
threaded down to `runMinimalInstall`'s new `installScript` override). This is
not just a bootstrap fix: a differential needs ONE measurement schema applied
to both sides, or the two stop being comparable the moment that schema
evolves — running each side's own copy would silently reintroduce that risk.
Verified locally end-to-end against real origin/next: resolves a valid
{version, sha, manifests, sizes} artifact with the correct sha and no leaked
worktree.

Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let
docs-lint evaluate the fragment for the first time; it already passes
(docs/TESTING-SUITES.md and friends already document the removed scripts).

Also fixed while in this file: an eslint no-unused-vars warning surfaced by
the changed lint run (unused `cleanup` import in
tests/emitted-provenance.test.cjs).

Added regression coverage for both A and B: a cross-platform spot-check that
drives the real hooks-built rule against `.cmd` keys directly (not through a
real Windows install), and a real-tree test that drives buildBaselineAtRef
against a base ref verified (via git cat-file) to lack the generator, both
skipping honestly rather than false-passing when their precondition does not
hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test

Two isolated-review findings on PR #2767:

- `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the
  wrapped `hooks/<name>.js` script, asserting a byte-provenance link that
  does not exist — traced against buildCodexHookWindowsShimIR
  (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that
  same file) flows into the .cmd bytes, never its content. Point `sources`
  at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention
  used elsewhere in the table. Since `sources` is checked before
  `transforms` in the differential, the wrong mapping silently excused any
  .cmd byte movement caused by editing the wrapped .js file.

- The `buildBaselineAtRef` regression test skipped unless a resolvable base
  ref still lacked scripts/gen-emitted-baseline.cjs — true only until this
  PR merges, after which every base ref carries the file and the test skips
  forever with zero ongoing coverage. Rebuilt hermetically: synthesize the
  missing-generator condition in-place via git plumbing (a throwaway commit,
  child of HEAD, with just that one file removed from a scratch index),
  never touching the real working tree, HEAD, or index, and never depending
  on ambient history or remotes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path

The runner container mounts the repo at a path owned by a different uid than
the process running the suite, so git's dubious-ownership protection refuses
every git operation there. GitHub Actions never hits this because
actions/checkout registers the workspace as safe automatically; this
runner's container does not.

buildBaselineAtRef is the production build-fallback the sole remaining
emitted gate depends on (resolveBaseline's in-job-build leg), not just a
test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs)
that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's
worktree add/remove/prune, and the hermetic regression test added in the
prior commit — funnels through, plus gen-emitted-baseline.cjs's own
rev-parse (now reusing that same wrapper instead of a second execFileSync,
so the fix has one source of truth). Each call declares -c
safe.directory=<the exact directory it already operates on>, never the *
wildcard.

Audited every other helper on this surface (emitted-diff.cjs,
emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them
shell out to git at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:41:43 -04:00
Behruz Nassre Esfahani
23254ca5a7 fix(#1936): reconstruct OpenCode review from JSON events (#1992)
* fix(#1936): reconstruct OpenCode review from JSON events; diagnosable empty-output stub

On a large review prompt, OpenCode's default `build` agent runs a few read
tool calls then ends its turn with zero output tokens (reason:"stop",
output:0), so `opencode run --format default` emits empty stdout. The reviewer
block redirected stderr to /dev/null and wrote a generic "failed or returned
empty output" stub — so the phase silently lost its second independent reviewer
with no diagnostic and no timeout.

Rewrite the OpenCode reviewer block to invoke `--format json` as the primary
call and reconstruct the review from the assistant `text` parts (jq). Capture
stderr to a `.err` sidecar (mirrors the Codex block). When the agent emits no
text, surface the stop reason, output-token count, and stderr so the failure is
diagnosable. Gate the stub on the extracted CONTENT, not the output file size —
an empty jq extraction still prints a lone newline that a `[ -s file ]` check
would treat as populated. Document the wall-clock timeout as a Bash-tool param
(macOS lacks GNU timeout; opencode has no native timeout flag).

review.md was already at the DEFAULT size-tier ceiling (40956/40960), so the
fix cannot fit without reclassifying it into the LARGE tier (it is a
multi-reviewer orchestration file that outgrew "focused single-purpose"; 43.4 KB
sits well under the LARGE high-water mark). Recapture the 16 golden-install
fixtures — the diff is exactly one review.md hash per runtime. Regression block
folded into review-default-reviewers-workflow.test.cjs (new bug-NNNN test files
are not accepted).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1936): add changeset

* test(#1936): property-test the OpenCode review jq reconstruction

Address the re-review's one actionable finding: the jq JSON-event → text
reconstruction had no fast-check property test.

Add tests/opencode-review-reconstruction.property.test.cjs. It extracts the two
shipped jq programs (OPENCODE_REVIEW, OPENCODE_DIAG) verbatim from
gsd-core/workflows/review.md and runs the real jq — not a reimplementation — so
the shipped logic is what gets tested. Properties: the reconstructed review
equals the newline-join of every assistant text part (order preserved); a stream
with no text part reconstructs to empty (drives the #1936 stub); null/absent text
parts are dropped, never rendered as "null". Plus example-based coverage of the
diagnostic edges the reviewer cited: missing .tokens.output and no step_finish
degrade to "?"; non-JSON stdout makes jq fail rather than masquerade as a review.

Verified the invariant has teeth (a comma-join jq fails the property).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test when jq is absent

The property test shells out to `jq`, which GitHub's windows-latest runners do
not ship (macOS/Linux runners do). `execFileSync('jq')` therefore ENOENT-failed
the whole file on `test (windows-latest, *)`. Probe `jq --version` at load and
skip the suite when jq is not on PATH — the reconstruction logic is
platform-independent, so the assertions still run in full on every jq-present
runner (mirrors how golden-install-parity skips on win32).

Verified: jq present → 7 pass; jq removed from PATH → 7 skipped, 0 fail.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test on Windows, not just when jq is absent

The prior guard skipped only when `jq` was absent from PATH — but the
windows-latest runners DO ship jq, so the suite still ran there and failed with
`jq: parse error: Invalid numeric literal` (confirmed from the CI job log). Root
cause is Node's child_process argument quoting mangling the jq program (it embeds
double quotes) on Windows, not the shipped review.md logic — the macOS/Linux legs
pass. Gate the suite on `process.platform === 'win32'` (still also skipping when
jq is absent), mirroring golden-install-parity's win32 skip. Logic is
platform-independent and fully asserted on every macOS/Linux CI leg.

Verified: macOS → 7 pass; simulated win32 → skips.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 14:07:07 -04:00
Tom Boucher
3a3b2135c2 chore(#1073): purge phantom pre-migration issue refs from source, tests, docs (#1471)
#2551/#3182/#2361 are pre-migration get-shit-done-redux issue numbers with no
equivalent in open-gsd/gsd-core; they mislead triage and manufacture phantom
blockers. Repoint to real successors (#717 byte-budget rework, #720) or rewrite
as prose referencing the discuss-phase/modes progressive-disclosure split.
Correct co-located 'line budget'/'<500 lines' framing to the byte-based reality
(#717). Add a CI guard (tests/no-phantom-issue-refs.test.cjs) that fails if a
phantom ref is reintroduced. SSH-key patterns (id_ed25519) left untouched.
No user-facing runtime behavior change.

Closes #1073
2026-06-20 13:36:57 -04:00
Rezolv
b055ca14e2 test(#1074): swap workflow size enforcement to baseline + loose hard caps (PR 2/3) (#1096)
Completes the #1074 migration for workflows. The per-file baseline (PR 1) is
now the primary anti-creep guard, so the tier-max tighten-only ceilings are
retired here.

- Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests) and
  the GRACE constant — per-file baseline already guards every file by name,
  strictly stronger than max(tier).
- Convert the per-file tier test into 'SIZE: workflow tier hard caps': absolute
  red lines (XL 96 KiB / LARGE 60 KiB / DEFAULT 40 KiB) that mean 'extract, do
  not raise', each with real headroom above its high-water file.
- Add a 32 KiB (Codex project_doc_max_bytes) cap for net-new workflow files not
  yet in the baseline and not explicitly tiered.
- Rewrite the header doc comment for the new two-guard model; drop the
  assertTightCeiling import (now unused here; still exported + used by the agent
  test until PR 3).

Addresses the three PR-2 items from the #1089 review:
- Finish the enumeration consolidation (Minor #1): the tier hard-cap and
  new-file guards now read both their file list and byte sizes from
  measureWorkflows()/listWorkflowStems(), removing the inline readdirSync +
  per-file byteCount split-brain. Enumeration and measurement share one source.
- Rewrite CONTEXT.md RULESET.WORKFLOW_SIZE_BUDGET (Minor #2): stale caps
  (XL<=90000/LARGE<=54000/DEFAULT<=38000) replaced with the baseline-first
  model, current hard caps, the new-file anchor, and the size:baseline
  remediation.
- Ship docs (contract-change requirement): a Diataxis how-to + reference for
  the size guard in docs/TESTING-SUITES.md, beside the sibling regression-name
  ratchet (the 'file grew, CI red -> npm run size:baseline, commit the one-line
  diff, justify or extract lazily' workflow). CONTRIBUTING.md's workflow-tree
  note is updated to the baseline model and points at the new guide.

Negative proof: each guard fails independently when violated (new-file cap at
33 KB; hard cap with baseline current; baseline on any per-file growth).

Refs #1074. Part 2 of 3.
2026-06-12 09:30:44 -04:00
Rezolv
74d7bc8239 test(#1074): add additive per-file workflow size baseline guard (PR 1/3) (#1089)
* test(#1074): add additive per-file workflow size baseline guard (PR 1/3)

Introduces a committed per-file size baseline scheme alongside (not replacing)
the existing tier anti-creep tests. Green by construction — the baseline
records current sizes, so both schemes pass side by side during migration.

- scripts/lib/allowlist-ratchet.cjs: add assertFileBaseline (third pure helper,
  same injected-fail style) — per-file growth/shrink/add/remove diff vs baseline.
- scripts/workflow-size.cjs: single source of truth for LF-normalized byte
  counting (#683) + workflow enumeration, shared by the guard and the generator
  so they can never measure differently. Lives in scripts/ root (NOT scripts/lib/)
  because it is dev/CI-only tooling — scripts/lib/ is bundled into the installed
  runtime, scripts/ root is not, so this keeps it out of the shipped payload.
- scripts/update-size-baseline.cjs + npm run size:baseline: regenerate the
  snapshot (sorted keys, trailing newline, idempotent).
- tests/workflow-size-baseline.json: generated snapshot (88 workflows).
- tests/workflow-size-budget.test.cjs: import the shared counter (drops the
  duplicated local byteCount) and add the per-file baseline describe block.
- Tests for the helper, the shared module, and the generator (incl. round-trip
  and fault-injection cases).

Refs #1074. Part 1 of 3; PR 2 swaps enforcement, PR 3 covers the agent test.

* test(#1074): regenerate workflow baseline after Update-branch merge with next

The 'Update branch' merge (652a916b) pulled in next's update.md change (#1090)
without regenerating the snapshot, leaving the per-file baseline stale by one
file. Re-ran `npm run size:baseline` so the committed baseline matches the
merged workflow files.

Refs #1074.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-11 23:59:56 -04:00
Tom Boucher
7c07fce70f fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes (#1084)
* fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes

On runtimes that execute each fenced bash block in a separate shell process
(e.g. Claude Code — documented behavior: each Bash command is a separate
process; inline shell functions and exported vars do not persist between
calls), the once-per-file gsd_run() function was undefined in every block
after the preamble block, and the call was swallowed by
`2>/dev/null || echo "{}"` into silent empty state.

Fix (budget-neutral session-level resolution):
- Ship gsd-core/bin/gsd_run, a POSIX sh wrapper that symlink-resolves its own
  location and execs the co-located gsd-tools.cjs. Exposed on PATH via the npm
  `bin` field (global installs) and shipped to local installs via the recursive
  gsd-core/ copy.
- The per-file launcher preamble now appends `export PATH='<bindir>':"$PATH"`
  to the file named by $CLAUDE_ENV_FILE (Claude Code's documented
  env-persistence mechanism) so later fresh-shell blocks resolve gsd_run from
  PATH. Guarded as a strict no-op when CLAUDE_ENV_FILE is unset; the inline
  gsd_run() definition remains the fallback for all other runtimes. The
  single-quoted dir neutralizes shell metacharacters at source time.
- Propagated via scripts/sync-runtime-launcher.cjs to all launcher-using files.
- XL workflow byte budget 93000 -> 93200 (the ~130B clause pushes plan-phase.md
  to 93135; legitimate content growth, ratchet-up per #717).

Regression tests (I)/(J) in runtime-launcher-parity.test.cjs cover wrapper
delegation and end-to-end PATH persistence (sourcing the env file with a
space-bearing install path).

Known limitation: an install path containing a literal single-quote yields a
malformed env-file line and falls back to the status quo (no regression);
rare on sanitized home directories.

Closes #381

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#381): add changeset for gsd_run fresh-shell reachability fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#381): scope test (J) bare-PATH execution to POSIX (Windows Git Bash exec bit)

Windows Git Bash (msys2) does not honor Node's chmod exec bit for
PATH-executing extension-less scripts, so the bare `gsd_run` command lookup
failed there even though the env-file PATH persistence was correct. The
env-file content assertions (the fix's actual cross-platform logic) still run
on every platform; only the final source-and-execute sub-step is gated to
non-win32. Global installs on Windows are covered by npm's generated bin shim.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 21:42:52 -04:00
Tom Boucher
b866b95296 fix(#921,#922): orchestrators must not fork; plan-phase Agent gate is attempt-based (#926)
`context: fork` strips the `Agent` tool from a subagent's environment.
Spawning orchestrators (`/gsd-autonomous`, `/gsd-execute-phase`,
`/gsd-plan-phase`) depend on `Agent` to dispatch sub-agents; running
them forked silently disables the core capability they exist to provide
(#921). Remove `context: fork` from all three command frontmatter files.
`effort: xhigh` (introduced by #769) is preserved.

The `<runtime_compatibility>` Agent-availability guard added by #913 was
checking whether `Agent` was present *before* attempting the call. On
runtimes where the tool list is dynamically resolved this produced
false-negative aborts in sessions that have the tool (#922). Replace
the introspection-based pattern with an attempt-based gate: always
attempt the `Agent()` call; stop only if a real tool-unavailable error
is returned. This preserves #853's backgrounded-session close-off and
#913's intent of preventing inline role-collapse, while eliminating
false negatives.

Tests updated: enh-769-context-fork-effort.install.test.cjs asserts the
three orchestrators lack `context: fork` and that the converter still
passes the field through for non-orchestrator commands; plan-phase-drift-
guard.test.cjs adds four assertions for the attempt-based gate language;
workflow-size-budget unchanged (budgets not exceeded).

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 08:42:25 -04:00
Tom Boucher
5bf77a527f fix(#913): guard top-level Claude Code plan-phase against role collapse (#915)
Three-part fix for the top-level inline collapse bug:

1. plan-phase.md: add <runtime_compatibility> block after
   </available_agent_types> that makes the Agent-availability
   requirement explicit; workflow fails-closed (stops with a clear
   log) in genuinely Agent-less contexts.

2. plan-phase.md: rename 7 "ORCHESTRATOR RULE — CODEX RUNTIME"
   labels to "ALL RUNTIMES" so the spawn guard applies universally
   (not just when Codex is detected).

3. execute-phase.md: scope the existing "Other runtimes" inline-
   fallback prose to non-Claude contexts, preserving the #853
   backgrounded-agent behaviour for Claude Code background agents.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 23:32:34 -04:00
Tom Boucher
a90c654745 fix(#891): probe non-Claude runtime homes in gsd-tools launcher shim detection (#911)
- Updated `gsd-core/workflows/_runtime-launcher.snippet.sh` with 15 new
  `elif` arms covering Hermes, Cursor, Codex, Gemini, Copilot, Windsurf,
  Augment, Trae, Qwen, CodeBuddy, Cline, Grok, Antigravity, OpenCode, and
  Kilo (respecting each runtime's env-var override with a `$HOME`-relative
  default).
- Re-ran `scripts/sync-runtime-launcher.cjs` to propagate the expanded
  snippet into all `gsd-core/workflows/*.md` files (~70 files).
- Manually applied the same snippet update to `commands/gsd/import.md`
  (1 occurrence) and `commands/gsd/graphify.md` (5 occurrences) — these
  are not covered by the sync script.
- Updated `tests/workflow-size-budget.test.cjs` budgets (XL/LARGE/DEFAULT
  + discuss-phase target) to account for the ~3 KB snippet expansion.
- Added regression test `tests/bug-891-non-claude-runtime-home-fallback.test.cjs`
  (6 tests: structural probe presence, ordering, behavioral HERMES_HOME
  env-var + default-path stubs, resolution order, and workflow propagation).
- Added `.changeset/891-launcher-non-claude-runtime-homes.md` (Fixed).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 22:51:40 -04:00
Tom Boucher
cf8bd3cd5e fix(#683): auto-degrade phase execution to sequential on worktree base mismatch (#749)
* fix(#683): auto-degrade phase execution to sequential on worktree base mismatch

Claude Code forks worktree-isolated executors off the repository default
branch (origin/HEAD), not the orchestrator's HEAD. Running /gsd-execute-phase
on a branch diverged from the default (unmerged milestone/feature branch) left
every executor without the phase's plan files and tripped the
worktree-branch-check guard with `exit 42` — 100% reproducible, all OSes.

- New module src/worktree-base-ref.cts: HEAD-vs-fork-base drift detection
  (origin/HEAD with symbolic-ref fallback) and no-clobber worktree.baseRef
  management, exposed as `worktree base-check` / `worktree set-baseref`.
- execute-phase.md: pre-dispatch, for Claude Code with worktrees enabled,
  auto-degrades the run to sequential on the main tree when a base mismatch
  is detected, recommending worktree.baseRef:"head". The exit-42 guard stays
  as a backstop.
- Installer: fresh local Claude installs set worktree.baseRef:"head" in
  .claude/settings.local.json (no-clobber, respecting an explicit shared
  settings.json value); upgrades print an opt-in notice pointing at
  `gsd-tools worktree set-baseref`.
- Docs: how-to guide, CLI/config reference, planning-config cross-ref.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): auto-apply worktree.baseRef on upgrade; gate fresh+upgrade on use_worktrees

Per maintainer direction: on a local Claude Code UPGRADE, set
worktree.baseRef:"head" automatically (no opt-in notice) when the project's
workflow.use_worktrees is enabled, instead of merely printing a remediation
notice. For consistency the FRESH path is now gated the same way: both paths
compute worktrees-enabled once (bounded walk-up read of .planning/config.json,
default enabled unless workflow.use_worktrees === false) and apply the
no-clobber baseRef only when enabled — never overwriting an explicit value in
settings.local.json or a shared settings.json. gsd-tools worktree set-baseref
remains for manual use. Docs + changeset updated; tests hardened (file-exists
assertions, fresh+disabled case, upgrade idempotency).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): measure workflow byte-budget on LF, fixing Windows-only CI failure

The workflow-size-budget test failed only on Windows: git checks out the .md
files as CRLF (no eol=lf in .gitattributes) and byteCount used
fs.statSync().size (raw on-disk bytes), counting an extra \r per line. That
inflated execute-phase.md — the XL high-water-mark file pinned near its ceiling
by the tighten-only ratchet — from 88492 LF bytes to ~90245 on Windows, over
the 90000 XL ceiling, while passing on the LF-checkout Mac/Linux runners.

The ceilings are explicitly "calibrated against raw `wc -c`" on an LF checkout,
so the measurement should be LF-based on every platform. byteCount now reads the
file and counts Buffer.byteLength after stripping CR, making the budget
platform-independent (a no-op on LF checkouts; verified statSync === normalized
for all 88 workflow files). No ceilings changed. Added a regression test
asserting CRLF and LF content of the same file count identically.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): make worktree-base-ref test path mocks Windows-safe (path.join)

tests/worktree-base-ref.test.cjs keyed its injected readFile/writeFile mocks
(and a few expected `file` values) with forward-slash template literals like
`${claudeDir}/settings.local.json`. The module composes those paths with
path.join(), which emits backslashes on Windows, so the mock keys never matched
the module's lookup → readFile returned null → resolveEffectiveBaseRef /
cmdWorktreeBaseCheck / cmdWorktreeSetBaseRef (and the JSONC variants) failed on
the Windows full-test runner only (they passed on Mac/Linux, and the install
tests passed because they use the real filesystem). The module is correct;
only the test fixtures hardcoded '/'.

All mock keys and path assertions now use path.join(base, ...) mirroring the
module, so they match on every platform (no-op on POSIX). 19 path references
across 16 lines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 23:40:24 -04:00
Tom Boucher
1bea220d58 refactor(#720): lazy-load MVP-only reference bodies on non-MVP runs (#746)
* refactor(#720): lazy-load MVP-only reference bodies (eager @-import → gated Read)

Convert eager @-imports of MVP-only reference bodies into lazy "Read" instructions
gated on MVP_MODE / WALKING_SKELETON / MVP+TDD, so non-MVP planning/execution runs
no longer pull MVP guidance into context. Covers both the workflow files and the
planner/executor agent definitions (the dominant context-cost path):

- workflows/plan-phase.md: planner-mvp-mode.md + skeleton-template.md (L146/936/937/941)
- workflows/execute-phase.md: execute-mvp-tdd.md halt-report ref, now gated on gate-trip (L191)
- agents/gsd-planner.md: planner-mvp-mode.md, user-story-template.md, skeleton-template.md
- agents/gsd-executor.md: execute-mvp-tdd.md

The dedicated always-MVP mvp-phase workflow keeps its eager imports (intentional).
Behaviour is unchanged; non-MVP runs simply carry less loaded context. Adds a
regression guard mirroring the discuss-phase lazy-load test, and documents the
conformance in docs/ARCHITECTURE.md.

Refs #720

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#720): add changeset fragment (pr #746)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 18:23:17 -04:00
Tom Boucher
31edeb1b54 feat(#717): re-base workflow size budget on bytes + document quality rationale (#719)
* feat(#717): re-base workflow size budget on bytes + document quality rationale

Re-base tests/workflow-size-budget.test.cjs from line counts to byte
counts (matches Codex's 32,768-byte project_doc_max_bytes cap; deterministic,
no tokenizer). Tier ceilings: XL=90000, LARGE=54000, DEFAULT=38000, GRACE=3000;
discuss-phase target re-expressed as <30 KB. The #597 tighten-only ratchet and
per-file budget semantics are preserved unchanged — only the unit swaps.

byteCount() uses fs.statSync().size to match `wc -c` (includes trailing
newline), deliberately not lineCount()'s newline-stripping.

Document the context-rot / attention-budget QUALITY rationale (independent of
prompt caching) in the test JSDoc and docs/ARCHITECTURE.md, plus the
Goodhart caveat: the byte budget measures one file, so the real goal is
bounded *loaded* context — eager @-imports game the proxy; legitimate
extraction is lazy. Update CONTEXT.md RULESET.WORKFLOW_SIZE_BUDGET to bytes and
remove a stale duplicate ruleset entry that still said "1800 lines".

Defers the #3182 MVP-mode split (tracked separately): MVP is a cross-cutting
concern woven through plan-phase/execute-phase, not a discrete extractable mode.

Closes #717

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#717): add changeset fragment for byte-budget re-base

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 20:19:08 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
a28dcec981 chore(#597): replace count-based ratchet guards with AST lint + named-set allowlists (#603)
The windows-test-parity ratchet greps test source for fs.rmSync-without-
maxRetries (and six other Windows-portability anti-patterns), failing when an
integer offender COUNT exceeds a frozen baseline (rmSync: 95). A count ratchet
is a Goodhart metric: fixing one offender and adding another keeps the count
constant, so a new defect slips through green. Replace it — and every other
count ratchet in the repo — with a layered, masking-proof design.

Behavioral seam test
- tests/helpers-cleanup.test.cjs proves helpers.cleanup() carries the Windows
  EBUSY retry budget. cleanup() delegates retries to Node's fs.rmSync via
  maxRetries (it owns no loop), so the test asserts the option contract
  (recursive/force/maxRetries>0/retryDelay>0) + real-FS removal + the cwd-guard,
  rather than a loop that does not exist. The EBUSY risk is now tested ONCE at
  the helper, not approximated textually at every call site.

Write-time ESLint rule (AST-accurate, replaces the grep)
- eslint-rules/no-raw-rmsync-in-tests.cjs (error in tests/**/*.test.cjs) bans
  raw fs.rmSync, steering to cleanup(). Catches member, computed (fs['rmSync']),
  destructured and aliased forms; escape hatch is inline
  `// eslint-disable-next-line local/no-raw-rmsync-in-tests -- <reason>` only.
- Migrated 336 raw fs.rmSync teardown calls across ~116 test files to cleanup().
  ~18 genuinely load-bearing sites (mid-test SUT/fault-injection removals,
  error-swallowing or name-colliding local teardown helpers) keep the raw call
  with an inline eslint-disable + reason.

Shared anti-ratchet primitive
- scripts/lib/allowlist-ratchet.cjs:
  - assertWithinAllowlist: fails on NOVEL ids (new offender introduced) AND on
    STALE ids (a known offender was fixed but not pruned) — identity, not count,
    and a ratchet DOWN toward zero.
  - assertTightCeiling: a size/length budget whose ceiling must stay within a
    grace band of the high-water mark, so budgets may only tighten, never creep.

Ratchets converted onto the primitive
- windows-test-parity-guard.test.cjs: rmSync rule deleted (now ESLint-enforced);
  the remaining six patterns moved from integer baselines to named-set
  allowlists with ratchet-down.
- scripts/lint-test-file-count.{cjs,allowlist.json}: per-module integer counts →
  named filename sets (closes the swap-a-file-keep-the-count blind spot); a
  module dropping under cap now FAILS to force pruning its allowlist entry.
- enh-2790 skill-count `<= 63` → named skill allowlist (ratchets toward ~58).

Size budgets hardened (tighten-only)
- agent-size / workflow-size / feat-3039 help-tiered: ceilings lowered to the
  current high-water mark and an assertTightCeiling anti-creep check added per
  tier. Fixed external-contract limits (description ≤100 chars, agent ≤100 KB)
  are intentionally left as-is — they are not grandfathered creeping budgets.

No user-facing behavior change (tests + tooling only); no USER_FACING_PREFIXES
touched, so no changeset fragment is required.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 22:43:49 -04:00
Tom Boucher
79002a00cb chore(#518): rename npm package + bin to @opengsd/gsd-core (#519)
* chore: rename npm package + bin to @opengsd/gsd-core (functional)

- package.json: name @opengsd/get-shit-done-redux → @opengsd/gsd-core,
  bin key get-shit-done-redux → gsd-core, repository/homepage/bugs URLs
- package-lock.json: regenerated (npm install --package-lock-only)
- tests/**, scripts/**, bin/**, .github/**, agents/**, commands/**,
  get-shit-done/bin/**, get-shit-done/workflows/**:
  applied the 4-rule replacement (scoped npm ref, GitHub repo path,
  bin/clone invocations) per #505 single-source refactor

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: sweep live references to @opengsd/gsd-core

Update all live documentation (README.md + translations, docs/**,
CONTRIBUTING.md, VERSIONING.md, SECURITY.md, CONTEXT.md,
docs/CANARY.md) to reflect the renamed package and repository.

Rules applied:
- @opengsd/get-shit-done-redux → @opengsd/gsd-core (scoped npm name)
- open-gsd/get-shit-done-redux → open-gsd/gsd-core (GitHub repo)
- GSD-redux/get-shit-done-redux → open-gsd/gsd-core (stale badge org)
- bare bin/clone refs → gsd-core

CHANGELOG.md, docs/adr/**, docs/RELEASE-*.md, docs/research/**,
and .changeset/** are preserved byte-identical.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: add negative lookbehind to slash-command regex in bug-2954 test

The extractSlashReferences regex matched /gsd-core inside npm package
URLs (@opengsd/gsd-core), producing a false /gsd:core command reference.
Adding a negative lookbehind (?<![a-z]) excludes matches preceded by a
letter, so only standalone /gsd-<cmd> and /gsd:<cmd> tokens are found.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#518): add changeset for package rename

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#518): update package-identity expectations to the renamed coordinates

The rebase regenerated the seam to @opengsd/gsd-core (bin gsd-core, repo
open-gsd/gsd-core). The #498 seam tests assert deriveIdentity against the REAL
package.json, so their expected literals must follow the rename. The drift-lint
unit test is left as-is — its SEAM is a self-consistent fixture and its
stale-literal detection cases would shift if altered; the live-repo scan in it
already passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-30 17:25:02 -04:00
Tom Boucher
b8c33647d8 refactor(tests): retire output-grep & source-grep via typed surfaces (finish #2974) (#462)
* refactor(#455): implement typed surfaces to retire grep tests

Production surfaces added:
- hooks/managed-hooks-registry.cjs: new CJS module exporting MANAGED_HOOKS
  as a typed array; gsd-check-update-worker.js now requires it instead of
  declaring an inline array
- bin/install.js: elevate inline gsdHooks to module-level GSD_UNINSTALL_HOOKS,
  export it alongside runtimeMap/allRuntimes (already exported)
- scripts/build-hooks.js: export HOOKS_TO_COPY; guard build() behind
  require.main===module so tests can require the file without triggering a build
- get-shit-done/bin/lib/init.cjs: add --json mode to agent-skills command,
  emitting typed IR { agent_type, block, skills_count } for test assertions
- get-shit-done/bin/gsd-tools.cjs: wire --json flag for agent-skills dispatch

Category-B source-grep migrations:
- tests/managed-hooks.test.cjs: require MANAGED_HOOKS from registry, drop fs.readFileSync+regex
- tests/orphaned-hooks.test.cjs: require MANAGED_HOOKS+HOOKS_TO_COPY as typed exports
- tests/hooks-opt-in.test.cjs: replace gsdHooks regex-parse with GSD_UNINSTALL_HOOKS import
- tests/install-minimal-hooks.test.cjs: replace gsdHooks regex-parse with GSD_UNINSTALL_HOOKS
- tests/copilot-install.test.cjs: replace src.includes() checks with typed
  assertions on runtimeMap, allRuntimes, parseRuntimeInput, buildRuntimePromptText
- tests/agent-skills.test.cjs: migrate to --json typed IR assertions

pending-migration-to-typed-ir token cleared (87 of 87 files):
- 78 files already had source-text-is-the-product; removed duplicate token
- 5 files already used typed assertions; reclassified or annotated
- 4 files required individual reclassification to source-text-is-the-product
  or architectural-invariant

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): update workflow-guard test to typed GSD_UNINSTALL_HOOKS import; isolate HOME in runtime-launcher (D) test

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): guard install.js main() behind require.main===module so the typed export is require-safe

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#455): document --json typed surfaces for agent-skills, progress, validate context

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#455): add changeset fragment for new --json surfaces

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): complete grep migration for files flagged by lint-tests

The branch commit 4e630d99 stripped `allow-test-rule: pending-migration-to-typed-ir`
from ~80 test files without replacing their assertions or adding the correct
exemption annotation. The files were NOT source-grep tests — they read .md
workflow/agent/command/reference files (source-text-is-the-product) or hook
source files for structural invariants (structural-regression-guard). No
assertion logic was changed; only the correct allow-test-rule annotation was
added to each file per CONTRIBUTING.md exception matrix.

73 files: `source-text-is-the-product` — workflow/agent/command/reference .md
7 files:  `structural-regression-guard` — hook .js / bin/install.js structural checks

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 11:39:52 -04:00
Tom Boucher
2a915c1b82 chore: migrate references from gsd-build to open-gsd/get-shit-done-redux (#120) (#121)
Security-motivated migration of all stale repository and npm-scope references.

Three categories of changes (58 files, 174 substitutions):

1. gsd-build → open-gsd (security-critical):
   - .github/workflows/release-sdk.yml — npm token comment, tarball filename pattern
   - .github/workflows/hotfix.yml — same
   - .changeset/fix-3406-detect-stale-sdk-shadow.md — @gsd-build/sdk → @open-gsd/sdk
   - .changeset/sharp-quails-leap.md — same
   - get-shit-done/workflows/update.md — CHANGELOG raw GitHub URL

2. GSD-redux org slug → open-gsd (canonical rename):
   - package.json + sdk/package.json — repository/homepage/bugs metadata
   - All README.*.md — live badge and link sections
   - CONTRIBUTING.md, CONTEXT.md, QUICK-WINS-CONFIRMED-BUGS.md
   - .coderabbit.yaml, .release-monitor.sh, scripts/sync-rulesets.sh
   - docs/** — all live agent/ADR/user-facing documentation
   - tests/** — repo slug assertions and test fixtures
   - scripts/changeset/cli.cjs + github-release-notes.cjs
   - .github/ISSUE_TEMPLATE/*, .github/pull_request_template.md
   - bin/install.js, get-shit-done/bin/lib/model-catalog.cjs
   - sdk/HANDOVER-*.md, sdk/src/*.test.ts

3. CLAUDE.md (gitignored local file — not in this commit):
   Updated separately outside git: --repo gsd-build/get-shit-done →
   --repo open-gsd/get-shit-done-redux with security warning.

Intentionally unchanged: CHANGELOG.md, docs/RELEASE-*.md,
.changeset/README.md, .changeset/build-hooks-atomic-write.md,
README.md migration table (historical fork record),
tests/changeset-serialize.test.cjs line 78 (serialization fixture).

The gsd-build/get-shit-done repo is compromised (rug-pull documented in
README.md). Do not push to or interact with that repo.

Closes #120
2026-05-22 12:28:16 -04:00
Tom Boucher
dff176bfd2 chore: rebrand to GSD-redux/get-shit-done-redux
Mirror of code, issues, and PRs from the upstream gsd-build/get-shit-done,
which appears compromised or abandoned (maintainer unreachable since
2026-04-01; $GSD token linked to rug-pull).

- Adds rebrand notice block at top of English README
- Removes $GSD token badge and @gsd_foundation X badge (keeps Discord)
- Renames npm packages: get-shit-done-cc -> get-shit-done-redux,
  @gsd-build/sdk -> @gsd-redux/sdk
- Updates all repo URLs across docs, workflows, package.json, bin/
- Updates ci@gsd-build -> ci@gsd-redux in workflow git identities
- Leaves CHANGELOG and .changeset/* alone (historical, time-stamped)
2026-05-22 08:27:07 -04:00
Tom Boucher
ca2644a71a fix(worktree): unlock-retry on locked cleanup + startup orphan sweep (#3707) (#3719)
* fix(worktree): unlock-retry on locked cleanup + startup orphan sweep (#3707)

Two root causes fixed:

1. **In-session cleanup blocked**: `executeWorktreeWaveCleanupPlan` now attempts
   `git worktree unlock <path>` then retries `git worktree remove --force` when the
   initial single-force remove fails on a locked worktree. Previously every cleanup
   after a successful merge was silently blocked.

2. **Cross-session orphan accumulation**: new `reapOrphanWorktrees` helper sweeps
   `.git/worktrees/*/locked` at startup. It reaps entries where the pid is dead,
   the branch tip is an ancestor of the default branch (ancestry guard prevents data
   loss on squash-merge repos), and the lock mtime is older than 5 minutes (race
   guard). Wired into `quick.md` and `execute-phase.md` startup blocks guarded by
   `USE_WORKTREES != false`.

SDK: adds `worktree.reap-orphans` query command (routes through gsd-tools.cjs).
Tests: 11 real-fs tests covering unlock-retry, dead-pid reap, live-pid skip,
unmerged skip, fresh-mtime skip, idempotent double-call, and structural wiring.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(changeset): add Fixed fragment for PR #3707 (worktree orphan cleanup)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(worktree): fix test portability on Windows + macOS for bug-3707 reap tests

- worktreeMeta helper: replace /\/\.git$/ with /[/\\]\.git$/ so the
  gitdir path suffix is stripped on both Windows (backslash) and Unix.
- worktreeMeta helper: normalize CRLF→LF before splitting porcelain
  blocks, fixing block parsing when git emits CRLF on Windows.
- reapOrphanWorktrees: replace single 'main' rev-parse with a
  [defaultBranch, 'main', 'master'] candidate loop so test fixtures
  without a remote origin (where branch may be 'master') don't bail
  early. Intentionally excludes 'HEAD' to prevent false reaping when
  HEAD is detached or on a feature branch (Codex adversarial finding).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(worktree): CI green — macOS symlink path, Windows test helper, pid portability, EPERM liveness

Four fixes to get macOS + Windows CI from red to green:

1. **macOS symlink mismatch** (worktree-safety.cjs): `reapOrphanWorktrees` now
   builds a canonical→listed path map from `git worktree list --porcelain` using
   `fs.realpathSync.native`. Uses the listed path (as git knows it) for
   `git worktree unlock/remove`, not the gitdir-derived path.  Fixes the
   `/var/folders` vs `/private/var/folders` discrepancy on GitHub macOS runners
   where `git worktree unlock <realpath>` was silently failing because git's
   list stored the unresolved symlink path.

2. **Windows path separator in test helper** (test file): `worktreeMeta`
   `.replace(/\/\.git$/, '')` → `.replace(/[/\\]\.git$/, '')`. On Windows,
   git writes backslash separators in the gitdir file; the Unix-only regex was
   causing `Cannot find .git/worktrees/<name>` for all Suite 2 tests.

3. **Non-portable PID in tests** (test file): All `'999999'` dead-PID literals
   replaced with `deadPid()` helper that spawns a real short-lived child, captures
   its PID, and returns it after exit. Eliminates flakiness on Linux systems where
   `pid_max` can reach 4194304, making 999999 a live PID.

4. **EPERM fail-closed in isPidAlive** (worktree-safety.cjs): `catch { return false }`
   → checks `err.code === 'EPERM'` and returns `true` (alive). On Windows and
   cross-user scenarios, `process.kill(pid, 0)` throws EPERM for live but
   inaccessible processes; treating that as dead would reap a live worktree.

Adversarial review via codex confirmed:
- Squash-merge repos: fail-closed (CONCERN, not BUG — by design, not data-loss)
- canonicalToListed map: SAFE (fail-closed on realpathSync error)
- Concurrent reapers: SAFE (both prune; second gets skipped: remove_failed)
- Startup blocking: CONCERN (no global cap, 10s/call × N worktrees) — tracked,
  not fixed here (requires separate perf work)
- gsd-sdk missing: SAFE (quick.md checks and fails fast with guidance)

All 27 local tests + Docker (holodeck) green.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(worktree): address codex adversarial findings — fail-closed default branch + CRLF map

Two fixes from codex adversarial review of PR 3718:

1. **Default branch resolution (data-loss risk)**: `reapOrphanWorktrees` now
   uses `refs/remotes/origin/<branch>` exclusively when a remote is configured.
   If `origin/HEAD` is absent but a remote exists, we bail out (fail-closed)
   rather than falling back to a local `main`/`master` that may not be the
   real integration branch.  The `main`/`master` fallback is only used when
   there is provably no remote (local-only test fixtures).

2. **CRLF normalization in canonical-path mapper**: The `worktree list
   --porcelain` output was split on '\n\n' without normalizing CRLF first.
   On Windows, git emits CRLF, which caused block-splitting to fail and
   left the canonicalToListed map only partially populated, weakening the
   symlink/path-mismatch fix introduced earlier.

3. **Windows 8.3 short-path fix (test helper)**: Both `beforeEach` blocks
   now call `resolvedTmpDir()` which pre-resolves `os.tmpdir()` via
   `fs.realpathSync.native` so temp paths avoid RUNNER~1-style short names
   that git stores in long form, causing worktreeMeta path comparisons to
   fail on Windows CI.

All 11 real-fs + 16 unit tests green locally.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(worktree): adversarial findings + macOS CI path-mismatch fix

## Root cause (macOS CI fail)
`reapOrphanWorktrees` stored `worktreePath` (gitdir-derived, real path via
git's symlink resolution, e.g. `/private/var/folders/…`) in results, while
the test's `wtDir` used the unresolved symlink form (`/var/folders/…`).  After
reaping, `canonicalPath(wtDir)` can no longer call `realpathSync.native`
(directory gone), so it falls back to `path.resolve` — which returns the
symlink form — causing the `result.find()` comparison to miss.

## Fixes applied

### Source — worktree-safety.cjs
1. **Finding 1 (fail-closed PID check)**: Non-parseable lock content (e.g.
   `"Locked by claude-code agent-xxx"`) is now treated as ALIVE with reason
   `lock_owner_unknown`, not as dead.  Previously it fell through as dead.
2. **Finding 1b (EPERM safe)**: `isPidAlive` call wrapped in try/catch; any
   thrown error (EPERM = process exists but cross-user on Windows) → ALIVE.
3. **Finding 2 (startup warning)**: `cmdWorktreeReapOrphans` now writes a
   one-line stderr warning when ≥1 entry is skipped or when reaper throws,
   while keeping exit-zero so workflows don't break.
4. **Finding 3 (default-branch discovery)**: Local-only fallback now tries
   `init.defaultBranch` config and HEAD symref before `main`/`master`, so
   repos configured with `trunk`, `dev`, etc. get correct orphan detection.
5. **macOS path fix**: Result entry for reaped worktrees now uses `gitKnownPath`
   (from `git worktree list`) instead of `worktreePath` (from gitdir file),
   ensuring the caller always sees the path git uses for the worktree.

### Test — bug-3707-locked-worktree-cleanup.test.cjs
6. **macOS CI fix**: Pre-compute `wtDirCanonical = canonicalPath(wtDir)` before
   calling `reapOrphanWorktrees` so the comparison works after removal.
7. **Gap 1**: New test — Claude Code lock format (`"Locked by claude-code …"`)
   must not be reaped; asserts `status=skipped, reason=lock_owner_unknown`.
8. **Gap 2**: New test — `isPidAlive` throwing EPERM → must not reap.
9. **Gap 3**: New test — repo with `init.defaultBranch=trunk`; merged worktree
   must be reaped (verifies trunk is discovered as the integration branch).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(test): raise waitForStoppedAt timeout 2 s → 5 s for Windows/Node22 CI load

Subprocess write latency exceeds 2 s on loaded windows-latest/Node22 runners
(test duration was 6181 ms); 5 s gives sufficient headroom without changing
any production behaviour.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-20 15:22:12 -04:00
Andreas Brauchli
6a2bf05de7 feat(#3039): tier /gsd-help output (--brief, default, --full, <topic>) (#3040)
Replace the single 747-line /gsd-help reference with a progressive-disclosure
dispatcher (#2551 pattern). Newcomers get a one-page tour; returning users get
a 10-line refresher with --brief; the complete reference stays available behind
--full; /gsd-help <topic> emits one section; and /gsd-help --brief <topic>
is a compact scoped lookup (signature + one-line summary).

- workflows/help.md becomes a small dispatcher routing on $ARGUMENTS
- workflows/help/modes/{brief,default,full,topic}.md hold the tier bodies
- commands/gsd/help.md passes $ARGUMENTS through, advertises composable form
- docs/COMMANDS.md documents the new flags and topic form
- existing tests that read help.md repointed at help/modes/full.md
  (bug-2836, bug-2950, bug-2954, cursor-reviewer, execute-phase-wave)
- new feat-3039-help-tiered test enforces structure, size budgets,
  dispatcher routing, shim arg passthrough, topic→section coverage,
  orphan-heading detection, conflict-resolution rules, routing preamble,
  and compact-scope rule

Trek-e review fixes (PR #3040):
- topic.md output rules split into 5a/5b/5c — explicit handling for single
  sections, multi-section "plus" joins, and bold-line sub-block anchors;
  each rule also takes scope (full vs compact) into account
- explicit resolved-routing preamble line emitted by topic.md before content
  ("**Topic:** `<alias>` → `<heading>` *(scope: full | compact)*") so the
  user sees which alias matched at which scope (review finding #3)
- composable `--brief <topic>` invokes topic.md in compact scope: heading
  + first `**/gsd:*`** signature line + one-line summary. Dispatcher and
  topic.md cooperate via $ARGUMENTS pass-through (review finding #4)
- full.md capped at LARGE-tier budget (FULL_BUDGET = 1500); the non-recursive
  workflow-size-budget test does not reach modes/ subdirs
- structural <progressive_disclosure> table parse (5-row assertion) replaces
  substring-soup regex matching — 5 rows = 4 base tiers + composable scope
- forward /gsd:* sub-block token coverage + reverse orphan-heading allowlist
  catch alias-table drift in both directions
- four conflict-resolution tests guard dispatcher promises (--brief+--full
  without topic → --full; --brief <topic> → compact; --full <topic> or bare
  → full; dispatcher retains --brief when delegating to topic.md)
- hardcoded topic lists removed from docs/COMMANDS.md and full.md (drift
  surfaces reduced from 5 to 2)
- topic.md alias bloat trimmed (~75 → ~25 rows); cleanup/update split into
  distinct sub-block rows under ### Utility Commands
- comment-rot ("~750 lines") removed from default.md and full.md
- dispatcher size guard tightened from < 100 to <= 40 lines
- commands/gsd/help.md <process> block trimmed to one line
- MD040 fence languages added to all plain code blocks across mode files

Main-merge conflict resolution:
- workflows/help.md kept as dispatcher (body lives in help/modes/full.md)
- /gsd-<cmd> → /gsd:<cmd> rename from #3452 reapplied to the mode files
  (full.md, default.md, brief.md, topic.md) — the six namespace routers
  (/gsd-context, /gsd-ideate, /gsd-manage, /gsd-project, /gsd-quality,
  /gsd-workflow) and wildcards (/gsd-*) preserved in hyphen form per
  main's convention
- bug-2950 test combines branch's path repointing with main's namespaced
  replacement strings
2026-05-15 22:02:55 -04:00
Tom Boucher
2d32ad82be fix(plan-phase): remove agent: directive that caused OpenCode subagent dispatch (#3156) (#3206)
* feat(roadmap): parse **Mode:** field on phase sections

Adds a 'mode' field to roadmap.get-phase and roadmap.analyze outputs.
Recognizes '**Mode:** mvp' lines in phase sections; lowercased + trimmed.
Forward-compat: unrecognized values preserved verbatim, no enum check.

Foundation for --mvp flag in plan-phase (PRD: vertical-mvp-slice).

* feat(plan-phase): parse --mvp flag and resolve MVP_MODE

Resolution order: CLI flag → ROADMAP **Mode:** field → workflow.mvp_mode
config → false. Walking Skeleton gate fires for new-project Phase 1.
Wires MVP_MODE + WALKING_SKELETON into gsd-planner subagent prompt.

Per PRD vertical-mvp-slice Phase 1 (Q1, Q2, Q4).

* docs(planner): add vertical-slice planning reference

New reference loaded by gsd-planner when MVP_MODE=true. Defines slice
ordering, Walking Skeleton rules, and anti-patterns. Referenced from
plan-phase workflow MVP_MODE wiring.

* docs(planner): add SKELETON.md template

Template emitted by gsd-planner under WALKING_SKELETON=true. Captures
architectural decisions and out-of-scope list for new-project Phase 1.

* chore(inventory): register new planner references

Added planner-mvp-mode.md and skeleton-template.md to INVENTORY.md and
INVENTORY-MANIFEST.json. References now: 53.

* feat(gsd-planner): add MVP Mode Detection section

Mode-switched branch in the existing planner agent (per Q4: single agent).
Vertical-slice decomposition rules, Walking Skeleton handling, and
TDD-mode compatibility. Heavy guidance lives in references/planner-mvp-mode.md.

* test(plan-phase): add --mvp resolution-chain integration cases

Validates roadmap.get-phase --pick mode and confirms workflow.mvp_mode
default is unset in fresh projects.

* docs(changelog): announce --mvp vertical-slice planning (#2826)

* feat(mvp-phase): add /gsd mvp-phase slash command

Standalone command for vertical MVP planning. Frontmatter only;
heavyweight workflow at get-shit-done/workflows/mvp-phase.md follows
in next commit. Mirrors discuss-phase/edit-phase command shape.

* docs(planner): add user-story-template reference

Defines the canonical 'As a / I want to / So that' format and the
ROADMAP.md / PLAN.md emit rules. Used by mvp-phase workflow and
gsd-planner agent under MVP_MODE.

* docs(planner): add SPIDR splitting reference

Defines size signals, the five SPIDR axes (Spike/Paths/Interfaces/Data/Rules),
the interactive workflow, and anti-patterns. Per PRD Q3 decision: full
interactive flow, not lightweight check. Used by mvp-phase workflow.

* fix(mvp-phase): trim description to fit 100-char budget

* feat(mvp-phase): add mvp-phase workflow

Standalone workflow: phase validation -> user story prompts (As a / I want to /
So that) -> SPIDR splitting check -> ROADMAP write (Mode + Goal) -> delegation
to plan-phase. Per PRD Phase 2 (Q3 full SPIDR; Phase-2-A/B/C/D decisions).

Plan-phase auto-detects MVP via Phase 1's resolution chain, so no flags
are needed when delegating.

* feat(gsd-planner): emit user-story header in PLAN.md under MVP mode

Extends the MVP Mode Detection section (added in Phase 1) so the planner
sources the user story from ROADMAP **Goal:** and emits the bolded
**As a** / **I want to** / **so that** form as the first content under
the phase header in PLAN.md. References user-story-template.md.

* test(mvp-phase): integration smoke test for ROADMAP mutation

Validates roadmap.get-phase output after a workflow-spec'd ROADMAP write:
mode=mvp and goal=full user story. Catches schema drift between workflow
emit and parser expectation. Includes a long-story case (>120 chars) to
confirm SPIDR-rejected stories still parse correctly.

* chore(inventory): register mvp-phase command + 2 new references

Adds /gsd mvp-phase to commands list, mvp-phase workflow to workflows list,
and user-story-template.md + spidr-splitting.md to references. References
count: 53 -> 55.

* docs(changelog): announce /gsd mvp-phase command (#2826)

* fix(mvp-phase): add TEXT_MODE plain-text fallback for non-Claude runtimes (#2012)

* docs(executor): add MVP+TDD gate reference

Defines the runtime gate semantics for execute-phase when both
MVP_MODE and TDD_MODE are true: pre-task verification of failing-test
commit, end-of-phase review escalation from advisory to blocking,
behavior-adding task definition. Loaded conditionally by
execute-phase workflow and gsd-executor agent.

* feat(execute-phase): MVP+TDD runtime gate + blocking review

Resolves MVP_MODE in Step 1 (CLI flag -> roadmap mode -> config -> false).
Adds per-task gate that halts before behavior-adding tasks run if no
failing-test commit exists for the plan. Escalates end-of-phase TDD
review from advisory to blocking when both MVP_MODE and TDD_MODE active.

Also updates INVENTORY-MANIFEST.json to register execute-mvp-tdd.md
(added by Task 1) so manifest-sync tests pass.

Per PRD vertical-mvp-slice Phase 3a (decisions Phase-3-A, Phase-3-Split).

* feat(gsd-executor): add MVP+TDD Gate section

Mirrors the planner's MVP Mode Detection pattern from Phase 1.
Instructs halt-and-report when the runtime gate trips, references
execute-mvp-tdd.md for full semantics. No agent changes outside the
new section.

* test(execute-phase): add MVP+TDD resolution-chain integration cases

Validates roadmap.get-phase --pick mode and confirms workflow.mvp_mode
default is unset in fresh projects. Mirrors the Phase 1 plan-phase
resolution-chain integration test.

* chore(inventory): register execute-mvp-tdd reference

Bumps References count 55 -> 56. Registers execute-mvp-tdd.md.
Adds "init" to PROSE_ALLOWLIST in registry integration test so
bare `gsd-sdk query init` prose examples in plan docs don't
trigger the unregistered-handler guard (real commands are all
init.<subcommand>).

* docs(changelog): announce MVP+TDD runtime gate in execute-phase (#2826)

* docs(verifier): add verify-mvp-mode reference

Defines UAT framing under MVP mode: user-flow walk-through first,
technical checks deferred, coverage check as goal-backward narrowing
to the user story's outcome clause. Loaded conditionally by
verify-work workflow and gsd-verifier agent.

* feat(verify-work): MVP-mode UAT framing — user flow first

Resolves MVP_MODE from phase mode field. Under MVP mode, generates UAT
in three ordered sections: user-flow walk-through (derived from user
story), technical checks (deferred), coverage check (goal-backward).
Falls back to standard UAT generation when mode is null/absent.
User-story-format guard refuses to verify a mode:mvp phase with a
non-user-story goal.

Also updates docs/INVENTORY.md (56 references) and
docs/INVENTORY-MANIFEST.json to register verify-mvp-mode.md added
in Task 1.

Per PRD vertical-mvp-slice Phase 3b (decisions Phase-3-B,
Phase-3-Verify-Structure).

* feat(gsd-verifier): add MVP Mode Verification section

Narrows goal-backward verification to the user-story [outcome] clause
when phase mode is mvp. References verify-mvp-mode.md. Preserves
existing goal-backward methodology for non-MVP phases. User-story-format
guard refuses to verify a mode:mvp phase with a non-user-story goal.

* docs(changelog): announce MVP-mode UAT framing in verify-work (#2826)

* feat(new-project): add Vertical MVP vs Horizontal Layers mode prompt

Asks user at project init how to structure the project. Vertical MVP
emits **Mode:** mvp on every initial roadmap phase (per-phase mode
preserved per PRD Q1). Horizontal Layers falls back to standard
template — no behavioral change for existing flows.

Per PRD vertical-mvp-slice Phase 4 (decision Phase-4-Persistence).

* feat(progress): add MVP-mode user-flow display

When phase has **Mode:** mvp, progress renders user-flow status from
PLAN.md task names alongside standard task progress. Tasks that aren't
user-flow-shaped (technical-sounding) are filtered out of the user-flow
sub-block. Falls back to standard display when mode is null/absent.

Per PRD vertical-mvp-slice Phase 4 (decision Phase-4-Progress).

* feat(stats): add MVP phase count summary

Reads roadmap.analyze (which surfaces mode per phase from Phase 1) and
emits 'Phases: N total | M MVP | K standard' summary line. Suppressed
when MVP_COUNT == 0 to avoid clutter on non-MVP projects.

Per PRD vertical-mvp-slice Phase 4.

* feat(graphify): add MVP-mode visual differentiation

MVP-mode phases render with #22c55e fill color AND ' (MVP)' label
suffix — two-channel signaling for color-blind and grayscale renders.
Standard phases unchanged.

Per PRD vertical-mvp-slice Phase 4 (PRD Q5: distinct visual treatment).

* docs(changelog): announce Phase 4 discovery & progress (#2826)

* chore(release): bump dev to 1.50.0-canary.0 for first 1.50.0 canary

Sets the base version that .github/workflows/canary.yml derives the canary
tag from (strips suffix → base 1.50.0 → next available v1.50.0-canary.N).

This kicks off the 1.50.0 release train, opened by the MVP/TDD/UAT vertical
slice landed across PRs #2867, #2874, #2878, #2880, #2883.

* docs: add CANARY stream README + v1.50.0-canary.1 release notes

- docs/CANARY.md — explains the dev→@canary stream policy, install/rollback
  paths, and when (not) to install canary builds
- docs/RELEASE-v1.50.0-canary.1.md — release notes for the first 1.50.0
  canary cut: vertical MVP/TDD/UAT slice (#2867 + #2874 + #2878 + #2880 +
  #2883), opening the 1.50.0 train under PRD #2826
- docs/README.md — index entry + quick link for the canary stream

* fix(ci/canary): publish gate checks dev branch, not main

Four publish-step `if:` conditions in .github/workflows/canary.yml were
checking `github.ref == 'refs/heads/main'`. Those steps (Tag and push,
Publish to npm, Publish SDK to npm, Verify publish) therefore always
skipped on every workflow_dispatch invocation since canary runs from dev,
never main.

The workflow's own header comment is unambiguous: `dev → @canary`. The
gate was a copy-paste from release.yml (which correctly targets main for
the @next/@latest streams) that was never corrected for the canary stream.

This is why the 1.50.0-canary.1 publish hadn't materialized despite three
green workflow runs. With the gate corrected, the next dispatch will
actually publish.

* ci(release-sdk): make release-sdk.yml dispatchable from the dev branch

The workflow lives on main only, so the GitHub Actions "Use workflow
from" dropdown doesn't list dev — meaning dev → @dev publishes can't be
triggered from the dev branch directly. Add the file to dev so an
operator can dispatch it with branch=dev and tag=dev.

Per project release-stream policy: dev branch publishes canary (@dev).
This is the stream that needs the file most, since main never publishes
@dev itself (main does @next / @latest).

File is byte-identical to main's release-sdk.yml — straight propagation,
no behavioral change. Tracking issues #2925, #2929.

* docs(mvp): canary-prep concept cleanup — CONTEXT.md, mvp-concepts index, --prd interaction (#3176)

* chore(mvp): concept cleanup + cross-ref index for v1.50.0-canary.2 prep

- CONTEXT.md gains 7 MVP domain terms (MVP Mode, User Story, Walking
  Skeleton, Vertical Slice, Behavior-Adding Task, MVP+TDD Gate, SPIDR
  Splitting) so the project glossary matches the shipped surface.
- New get-shit-done/references/mvp-concepts.md indexes the six MVP
  reference files and concept-to-file map so agents and contributors
  can find the right canonical doc without grepping.
- plan-phase.md Walking Skeleton block now documents that --mvp and
  --prd compose orthogonally on Phase 1; no precedence needed.
- INVENTORY/INVENTORY-MANIFEST refreshed for the new reference (58 -> 59).

No behavior change. Canary-prep cleanup ahead of v1.50.0-canary.2.

Surfaced for follow-up (not in this PR):
- MVP_MODE resolution shell block duplicated across plan-phase,
  execute-phase, verify-work workflows (needs a shared workflow-include
  mechanism; structural change).
- Behavior-Adding Task predicate is prose-only; no shared utility.
- User Story regex hardcoded in verify-work; would benefit from a
  central definition consumed by the verifier and the mvp-phase command.

* chore(changeset): set PR number for mvp concept cleanup

* feat(mvp): centralize resolution surfaces + fix SDK roadmap mode parity (#3178)

Three new SDK query verbs replace the architectural duplication surfaced by
the v1.50.0-canary.2 review against dev tip 12c4e565:

  phase.mvp-mode <N> [--cli-flag]
    Single canonical precedence resolver (CLI flag -> ROADMAP **Mode:** mvp
    -> workflow.mvp_mode config -> false). Replaces 4-8 lines of bash that
    were duplicated across plan-phase.md, execute-phase.md, verify-work.md,
    and progress.md. Returns {active, source, roadmap_mode, config_mvp_mode,
    cli_flag_present}.

  task.is-behavior-adding <plan-file> | --task-content <xml>
    Behavior-Adding Task predicate (tdd="true" + <behavior> block + non-test
    source files in <files>). Replaces prose-only specification in
    references/execute-mvp-tdd.md; gsd-executor agent now invokes the verb
    instead of re-inlining the three checks. Returns {is_behavior_adding,
    checks, reason}.

  user-story.validate <text> | --story <text>
    Owns the canonical User Story regex /^As a .+, I want to .+, so that .+\.$/
    previously hardcoded in verify-work.md prose. Consumed by gsd-verifier
    (phase-goal guard) and /gsd-mvp-phase (interactive-prompt validation).
    Returns {valid, slots: {role, capability, outcome}, errors[]}.

Bug fix bundled: sdk/src/query/roadmap.ts searchPhaseInContent now extracts
the mode field from **Mode:**, restoring parity with roadmap.cjs:120-123.
Without this, roadmap.get-phase --pick mode returned null on the native
dispatch path even when the phase had **Mode:** mvp set, causing MVP_MODE
to silently fall through to the config/false branch in every consuming
workflow. The original PRs Phase 1 (#2885) shipped the CJS parser but the
SDK port omitted the field; this fix brings them back to parity.

Workflows + agents updated to call the verbs:
  - plan-phase.md, execute-phase.md, verify-work.md, progress.md call
    phase.mvp-mode (one line replaces the duplicated bash chains).
  - execute-phase.md MVP+TDD gate calls task.is-behavior-adding.
  - verify-work.md goal guard calls user-story.validate.
  - mvp-phase.md interactive prompt validates via user-story.validate.
  - gsd-executor agent references task.is-behavior-adding instead of prose.
  - gsd-verifier agent references user-story.validate instead of inlined regex.

Tests: 24 new vitest tests in sdk/src/query/mvp.test.ts cover all three
verbs + the regression. Two existing contract tests (progress, verify)
updated to assert on the new verb shape. All 60 existing MVP contract
tests pass; golden integration suite (38 + 42 tests) passes.

Closes #3177

* fix(canary.2): unblock release gates for v1.50.0-canary.2

Run 25451329660 (Release SDK Bundle on dev, 2026-05-06T17:41) failed at the
test-suite step with 3 deterministic content/structure gate failures, all
attributable to the MVP umbrella integration in #3178 and the docs sweep
in #3180.

Failure 1: /gsd-mvp-phase undocumented in workflows/help.md
  - tests/bug-2954-help-md-slash-command-stubs.test.cjs requires every
    shipped commands/gsd/<X>.md to have a /gsd-<X> mention in help.md
  - PR #3180 updated docs/COMMANDS.md but missed help.md (which the AI
    agents load in-product)
  - Fix: add a /gsd-mvp-phase entry to help.md right before /gsd-plan-phase

Failures 2 + 3: execute-phase.md (1727) and plan-phase.md (1714) over XL budget (1700)
  - PR #3178 added MVP-mode verb calls (phase.mvp-mode, task.is-behavior-adding,
    user-story.validate) to both workflow files, pushing them past 1700 lines
  - Fix: bump XL_BUDGET 1700 -> 1800 with inline comment pointing at the
    structural follow-up (extract MVP bodies to <workflow>/modes/mvp.md per
    the discuss-phase/modes/ precedent)
  - The structural extract is the right long-term fix but is bigger than
    canary unblock scope; will land in a follow-up after canary cycles

Local verification:
  $ node --test tests/bug-2954-help-md-slash-command-stubs.test.cjs                 tests/workflow-size-budget.test.cjs
  tests 111  pass 111  fail 0

After this lands, re-trigger Release SDK Bundle on dev for v1.50.0-canary.2.

* chore(changeset): set PR number for canary.2 unblock

* fix(codex): generate-claude-md writes to AGENTS.md on Codex runtime

When config.runtime === 'codex' or GSD_RUNTIME=codex, override the
output target to AGENTS.md regardless of claude_md_path, so Codex
projects no longer have GSD sections written to CLAUDE.md by mistake.

Fixes both the CJS (gsd-tools) and SDK (profile-output.ts) paths.
Explicit --output flags are still honoured in both paths.

Closes #3163

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(plan-phase): remove agent: directive that caused OpenCode subagent dispatch

On OpenCode, any command with `agent: <name>` in its frontmatter is
auto-dispatched to a subagent context where the Agent tool is unavailable.
plan-phase.md and mvp-phase.md both carried `agent: gsd-planner`, causing
them to run inside gsd-planner's subagent context with no ability to spawn
researcher/planner/checker subagents — the orchestrator fell back to inline
execution for all three phases.

Fix: remove `agent: gsd-planner` from both command files so they run in the
main agent context. Also replace the stale `Task` tool in allowed-tools with
`Agent` (the correct dispatcher tool name post-#3168 rename).

Adds a structural regression test that parses YAML frontmatter of every
commands/gsd/*.md file and asserts no command carries an `agent:` directive.

Closes #3156

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(mvp): address CodeRabbit workflow and contract findings

* fix(execute-phase): use registered state.update query command

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-06 21:51:38 -04:00
Tom Boucher
918f987a19 feat(#2982): extend no-source-grep lint to catch var-binding readFileSync.includes() (#2985)
* feat(#2982): extend no-source-grep lint to catch var-binding readFileSync.includes()

The base lint (scripts/lint-no-source-grep.cjs) only catches
readFileSync(...).<text-method>() chained directly. The much more
common var-binding form escapes it:

  const src = fs.readFileSync(p, 'utf8');
  // 50 lines later
  if (src.includes('foo')) {}        // ← still grep, lint missed it

Scan of the test suite found ~141 files using this pattern.

Implementation built TDD per #2982 with structured-IR assertions:

  scripts/lint-no-source-grep-extras.cjs
    - detectVarBindingViolations(src) — pure detector, two passes:
      pass 1 collects vars bound from readFileSync, pass 2 finds any
      <var>.<includes|startsWith|endsWith|match|search>( on those vars.
    - detectWrappedAssertOkMatch(src) — flags
      assert.ok(<expr>.match(...)) which escapes the assert.match rule.
    - VIOLATION enum exposes stable codes for tests to assert on.

  scripts/lint-no-source-grep.cjs
    - Wires the new detectors into the existing per-file check; one
      additional violation row per file with the first 3 sample tokens.

  tests/bug-2982-lint-var-binding.test.cjs
    - 13 tests, all assertions on typed VIOLATION enum / structured
      records. Covers all 5 text-match methods, multi-var, no-bind,
      string literal (must NOT trigger), wrapped assert.ok(.match),
      and assert.match (must NOT double-flag).

Migration backlog (#2974 expanded scope):

  - 42 files annotated `// allow-test-rule: source-text-is-the-product`
    (legitimate — they read .md/.json/.yml files whose deployed text
    IS the product)
  - 3 files annotated `// allow-test-rule: pending-migration-to-typed-ir [#2974]`
    (read .cjs/.js source — clear migration debt)
  - 95 files annotated `pending-migration-to-typed-ir [#2974]` with
    `Per-file review may reclassify as source-text-is-the-product
    during migration` (mixed — manual review under #2974)

After this lands the lint reports 0 violations on main; new
violations in PRs surface immediately.

Closes #2982
Refs #2974

* test(#2982): fix truncated test name per CR

The label ended with a bare '(' from a copy-paste mishap. Now reads
'does NOT flag .matchAll(...) — matchAll is not match, so
assert.ok(.matchAll(...)) is not flagged'.

* chore(#2982): add changeset fragment for PR #2985

* chore(#2982): add changeset fragment for PR #2985
2026-05-01 19:50:10 -04:00
Tom Boucher
41dc475c46 refactor(workflows): extract discuss-phase modes/templates/advisor for progressive disclosure (closes #2551) (#2607)
* refactor(workflows): extract discuss-phase modes/templates/advisor for progressive disclosure (closes #2551)

Splits 1,347-line workflows/discuss-phase.md into a 495-line dispatcher plus
per-mode files in workflows/discuss-phase/modes/ and templates in
workflows/discuss-phase/templates/. Mirrors the progressive-disclosure
pattern that #2361 enforced for agents.

- Per-mode files: power, all, auto, chain, text, batch, analyze, default, advisor
- Templates lazy-loaded at the step that produces the artifact (CONTEXT.md
  template at write_context, DISCUSSION-LOG.md template at git_commit,
  checkpoint.json schema when checkpointing)
- Advisor mode gated behind `[ -f $HOME/.claude/get-shit-done/USER-PROFILE.md ]`
  — inverse of #2174's --advisor flag (don't pay the cost when unused)
- scout_codebase phase-type→map selection table extracted to
  references/scout-codebase.md
- New tests/workflow-size-budget.test.cjs enforces tiered budgets across
  all workflows/*.md (XL=1700 / LARGE=1500 / DEFAULT=1000) plus the
  explicit <500 ceiling for discuss-phase.md per #2551
- Existing tests updated to read from the new file locations after the
  split (functional equivalence preserved — content moved, not removed)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(#2607): align modes/auto.md check_existing with parent (Update it, not Skip)

CodeRabbit flagged drift between the parent step (which auto-selects "Update
it") and modes/auto.md (which documented "Skip"). The pre-refactor file had
both — line 182 said "Skip" in the overview, line 250 said "Update it" in the
actual step. The step is authoritative. Fix the new mode file to match.

Refs: PR #2607 review comment 3127783430

* test(#2607): harden discuss-phase regression tests after #2551 split

CodeRabbit identified four test smells where the split weakened coverage:

- workflow-size-budget: assertion was unreachable (entered if-block on match,
  then asserted occurrences === 0 — always failed). Now unconditional.
- bug-2549-2550-2552: bounded-read assertion checked concatenated source, so
  src.includes('3') was satisfied by unrelated content in scout-codebase.md
  (e.g., "3-5 most relevant files"). Now reads parent only with a stricter
  regex. Also asserts SCOUT_REF exists.
- chain-flag-plan-phase: filter(existsSync) silently skipped a missing
  modes/chain.md. Now fails loudly via explicit asserts.
- discuss-checkpoint: same silent-filter pattern across three sources. Now
  asserts each required path before reading.

Refs: PR #2607 review comments 3127783457, 3127783452, plus nitpicks for
chain-flag-plan-phase.test.cjs:21-24 and discuss-checkpoint.test.cjs:22-27

* docs(#2607): fix INVENTORY count, context.md placeholders, scout grep portability

- INVENTORY.md: subdirectory note said "50 top-level references" but the
  section header now says 51. Updated to 51.
- templates/context.md: footer hardcoded XX-name instead of declared
  placeholders [X]/[Name], which would leak sample text into generated
  CONTEXT.md files. Now uses the declared placeholders.
- references/scout-codebase.md: no-maps fallback used grep -rl with
  "\\|" alternation (GNU grep only — silent on BSD/macOS grep). Switched
  to grep -rlE with extended regex for portability.

Refs: PR #2607 review comments 3127783404, 3127783448, plus nitpick for
scout-codebase.md:32-40

* docs(#2607): label fenced examples + clarify overlay/advisor precedence

- analyze.md / text.md / default.md: add language tags (markdown/text) to
  fenced example blocks to silence markdownlint MD040 warnings flagged by
  CodeRabbit (one fence in analyze.md, two in text.md, five in default.md).
- discuss-phase.md: document overlay stacking rules in discuss_areas — fixed
  outer→inner order --analyze → --batch → --text, with a pointer to each
  overlay file for mode-specific precedence.
- advisor.md: add tie-breaker rules for NON_TECHNICAL_OWNER signals — explicit
  technical_background overrides inferred signals; otherwise OR-aggregate;
  contradictory explanation_depth values resolve by most-recent-wins.

Refs: PR #2607 review comments 3127783415, 3127783437, plus nitpicks for
default.md:24, discuss-phase.md:345-365, and advisor.md:51-56

* fix(#2607): extract codebase_drift_gate body to keep execute-phase under XL budget

PR #2605 added 80 lines to execute-phase.md (1622 -> 1702), pushing it over
the XL_BUDGET=1700 line cap enforced by tests/workflow-size-budget.test.cjs
(introduced by this PR). Per the test's own remediation hint and #2551's
progressive-disclosure pattern, extract the codebase_drift_gate step body to
get-shit-done/workflows/execute-phase/steps/codebase-drift-gate.md and leave
a brief pointer in the workflow. execute-phase.md is now 1633 lines.

Budget is NOT relaxed; the offending workflow is tightened.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 21:57:24 -04:00