Commit Graph

529 Commits

Author SHA1 Message Date
Tom Boucher
4d65c248e5 fix(#4641): make test-conformance the sole Windows selector and narrow the tier to 28.5% (#4643)
* test(#4641): failing-first tests for the tier ceiling and a single Windows selector

Tests only, committed ahead of the implementation so the RED run is real.

- tests/platform-conformance-tier.test.cjs: tier-size ceiling asserted as a
  ratio against a live denominator (Windows 33%, macOS 25%); per-helper negative
  cases proving seam calls and path-call-plus-slash-literal are not platform
  signals; positive pins that genuine platform content, seam-bypassing spawns,
  chmod and symlink still classify in; macOS signal set and generated list
  unchanged.
- tests/ci-full-lane-sharding.test.cjs: the test job has zero windows-latest
  rows and test-conformance still has 3 windows + 1 macOS.
- tests/ci-test-scope.test.cjs: windows_tests is absent rather than empty, a
  non-tier test file no longer forces full_matrix, a RULE-pulled windows-hint
  test does, and resolveSelection rejects the retired windows scope.

Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): delete the second Windows selector and narrow the conformance tier

Epic #4589's goal — the OS-agnostic bulk on Linux, a small explicitly-scoped
conformance tier on real Windows/macOS — was not met. Measured on PR #4640
(run 34618834118): 7 non-Linux jobs, a 546/930 (58.7%) "tier", and 5 of 7
changed test files running on a real Windows runner twice.

Two selectors, only one in the epic's scope. The test job's three scope:windows
shards predate the epic (#494, sharded #3057) and gate on product_changed, not
full_matrix, so they fire on every product PR whatever Phase 3's classifier
decides. They are deleted; test-conformance becomes the sole Windows selector,
as it already was for macOS. Non-Linux jobs 7 -> 4.

Gating the lane instead was rejected as provably redundant: for a test file
reachesConformanceTierOrSeam is literally CONFORMANCE_TIER_FILES.includes(file),
and that same predicate sets full_matrix, which turns test-conformance on. Every
file a gated lane would run is already covered in the same run. The lane's one
non-redundant residue -- RULE-pulled tests matched by the isWindowsHint filename
heuristic -- is ported into reachesConformanceTierOrSeam so it sets full_matrix
instead of feeding a parallel lane.

Two detectors matched the repo's own test idiom rather than any platform signal
and carried 226 of the tier's sole-signal membership against 41 for the other
eight: process-seam-subprocess (335 files, 118 unique) matches the
tests/helpers.cjs entry points nearly every CLI test uses, and going through the
seam is the opposite of a platform signal since shell-command-projection takes
platform as an injected parameter; hardcoded-path-vs-path-call (328, 108) needs
only a path call anywhere plus a slash literal anywhere, and that class is
already enforced by ADR-1703's Linux-runnable ESLint rules. Both are removed.
Tier 546 -> 254 (27.3%). src/ reachability is unchanged at 28 files, measured.

Adds the size gate Phase 2 never had, as a ratio against a live denominator so
it cannot stop binding as the suite grows.

292 files leave real-OS Windows execution. The drop-out set was audited: 14 have
a platform-suggestive filename and all 14 are static source-text analyses or
seam-mediated CLI tests. raw-child-process was investigated as a suspected false
negative and left unchanged -- relaxing it adds 13 files, all false positives.

macOS is untouched: MACOS_CATEGORIES is a separate array and the regenerated
macos-conformance-tier.generated.cjs is byte-identical at 196 files.

Fixes #4641
Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): register the new ADR path in the docs-guard exempt baseline

tests/ci-test-scope.test.cjs references docs/adr/4641-windows-selector-consolidation.md
in a comment justifying the retired windows scope; lint-docs-guard-registration
tracks that reference set, so the baseline needs the new path. Verified the
exemption still holds: the path is prose, not a filesystem read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): make the escalation tier-backed and drop every hardcoded count

Three follow-ups from measuring the first pass rather than trusting it.

The windows-hint escalation now requires tier membership as well as the
filename hint. Setting full_matrix runs test-conformance, which runs only the
tier; escalating on a test that is NOT in the tier costs four jobs and still
never runs that test on Windows. Measured over the 16 RULES entries the
narrowed predicate fires on exactly the same rules today, so this is
correct-by-construction rather than a behavior change. The broader variant --
escalate on any tier member a rule pulls in, ignoring the hint -- was measured
at 14/16 rules and rejected as over-broad.

Removes the hardcoded counts. A hardcoded macOS tier length of 196 broke as
soon as the rebase pulled in one new test file from #4253, which is the whole
argument against them: the ceilings are ratios against a live denominator, the
committed lists are pinned by comparison against a fresh classification of the
live tree, and the three named probe files now assert on their SIGNAL rather
than on membership in a literal list -- asserting by filename is the exact
error this PR fixes in the classifier.

Regenerates both lists against the rebased tree. Same-tree figures are now
547 -> 255 of 931 eligible (58.8% -> 27.4%), 292 entries removed and none
added; macOS is unchanged at 197 with a zero-line diff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): restore real-shell-spawn coverage and repair assertions the narrowing broke

An isolated adversarial review found a real false negative. Removing the
blanket process-seam-subprocess detector also removed the only coverage for
tests that spawn a REAL shell: tests/helpers/process-seam.cjs's runHook
spawns options.interpreter via real spawnSync, so
runHook('-c', [script], { interpreter: 'bash' }) runs a real bash binary
executing a shell script extracted from workflow markdown. The seam argument
holds for src/shell-command-projection.cts, which takes platform as an
injected parameter; it does NOT hold for the test helpers, which spawn real
binaries. Conflating the two is what made the blanket detector look purely
noisy -- it was 99% noise wrapping a real signal.

Adds a narrow shell-interpreter-spawn category keyed on a real interpreter
option. Measured 2026-09-11: 33 files match, 9 were outside the tier and are
added back, taking it 255 -> 264 of 931 (27.4% -> 28.4%), still under the 33%
ceiling. All 9 confirmed by reading the matching source line, zero comment or
fixture matches. runGit-alone and non-node-spawnSeam alternatives were measured
and rejected -- each adds 9 files but misses the counterexample entirely.

Fixes a real bug the suite caught: jobs.test is ubuntu-only now that its
scope:windows rows are gone, so it must wire GSD_STRICT_LIVE_CONFIG_GUARD
strictly rather than carrying the Windows report-only carve-out. The carve-out
now lives solely on test-conformance, whose matrix does include windows.

Repairs seven pre-existing assertions the category removal invalidated,
preserving each case's purpose rather than deleting coverage, and converts the
last hardcoded tier bounds to live-derived ratios -- including the macOS
sanity range that was still a magic [100, 350].

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): keep the confinement test on a real OS via a documented allowlist

A security review found tests/external-descriptor-confinement.test.cjs had
dropped out of the Windows tier. It must stay in, and no content signal can
express why: it exercises isPathConfined (src/external-descriptor-trust.cts),
which uses the AMBIENT path module -- path.resolve(root, target) and path.sep
-- with no injection. Its win32 semantics (drive letters, UNC, separator) are
only reachable by actually running on Windows, and it is a security-relevant
write-confinement gate. A content classifier cannot see 'this module reads the
ambient path module', so no regex belongs here.

Adds ALWAYS_REAL_OS, a Map of path -> recorded reason, unioned into the Windows
tier only. A Map rather than a list so an entry without a reason is impossible
by construction, and tests assert every entry names a file that exists on disk
so a stale entry fails loudly instead of rotting. This is the centrally-
enumerated single source of truth epic #4589 Phase 2 asked for and ADR-1703's
portability-vocab.cjs already models -- deliberately not a heuristic.

Windows tier 264 -> 265 of 931 (28.5%), still under the 33% ceiling. macOS is
untouched and byte-identical: the win32 concern does not apply to a POSIX
runner, and a test asserts the allowlist does not leak into that tier.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): inject the path impl into isPathConfined and correct the ADR count

Two review findings, both fixed rather than dispositioned.

A security review found tests/external-descriptor-confinement.test.cjs had left
real-OS execution. The allowlist pinned it back, but that only restored
INCIDENTAL coverage: isPathConfined used the ambient path module, and its test
carried POSIX-only literals, so a win32 confinement escape was unverified on
every platform including Windows. isPathConfined now takes an optional third
parameter carrying the path implementation, defaulting to the ambient module.
Blast radius is CRITICAL -- 53 affected symbols across 19 files -- so the change
is purely additive and every existing two-argument caller is byte-identical.

Tests now inject path.win32 and path.posix, covering a different drive letter,
a cross-drive absolute, backslash and forward-slash traversal, UNC, and the
startsWith prefix-boundary bug (.gsdEVIL against root .gsd) on both separators.
Proved load-bearing: dropping the + p.sep from the prefix check fails exactly
the two boundary cases and nothing else. Callers' suites 149/149.

The spec review caught an off-by-one: the ADR narrated a 264-file tier while the
committed list holds 265. The ADR now records the full chain 547 -> 255 -> 264
-> 265 (28.5%).

Also corrects a stale comment in scripts/docs-guard-registry.cjs that narrated
classify() as zeroing windows_tests, a key this change removes -- kept as
historical narration but labelled as such.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#131): make the unwritable-HOME test actually test something

Found by sweeping for the root-bypass class after fixing commit-files-deletion.
This one is the silent variant, and it was broken twice over.

First, the condition: the test made a fake HOME unwritable with chmod 0o500.
The gsd-test Docker bench runs as root, root bypasses mode bits, so HOME stayed
writable and the hostile condition never existed. Replaced with a HOME whose
PARENT is a regular file, so every write under it fails ENOTDIR at the VFS
layer for every uid -- no permission check is involved at all.

Second, and more fundamental: the probe was npm --version, which on npm 11.19.0
performs zero filesystem I/O against HOME. Proven rather than assumed --
neutralizing runNpm()'s isolation turned the sibling test red while this one
stayed green, so its assertion could never detect the regression it guards, on
any uid, with or without the condition fix. npm config get cache was tried next
and proved vacuous the same way (it only string-resolves the path). The probe is
now npm cache verify, which really does mkdir _cacache under HOME.

Re-proved load-bearing after the change: with isolation neutralized the test now
fails with ENOTDIR on <blocker>/home/.npm/_cacache. tests/helpers.cjs was
restored and verified diff-clean; suite 13/13.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the net drop-out figure in ADR-4641

The Consequences section still said 292 files leave real-OS Windows execution.
That was the count before the narrow shell-interpreter-spawn replacement
restored 9 and ALWAYS_REAL_OS pinned 1. Net is 282. Also names both real-binary
categories rather than only raw-child-process, and clarifies that the 14-file
filename audit was against the 292 initially dropped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the rejected concentration ceiling and its measurement

Applying Goodhart's own question to the new ceiling -- how would you make this
metric look good without improving what it represents -- surfaces a real
weakness: a ratio can be satisfied by inflating the denominator, so adding
OS-agnostic tests loosens it without narrowing the tier.

The obvious companion gate was a sole-signal concentration ceiling, since the
original defect was one detector carrying half the tier. Measured and rejected:
peak concentration post-fix is raw-child-process at 53/265 = 20.0%, against the
historic offenders at 21.6% and 19.8%. Any threshold above 20% misses the
original defect; any threshold below it fails on a legitimate category. The
discriminator is whether a signal is platform-meaningful, which no threshold
encodes. Weakness disclosed rather than covered by a gate that does not bind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4641): add the changeset fragment for the confinement-check change

changeset-lint failed on PR #4643: the PR touches user-facing paths and carried
no fragment. The earlier no-changeset call matched #4604's CI-only precedent and
was correct then; it was not revisited once the PR grew a src/ change, which is
my miss.

The fragment describes the real user-visible improvement: the external-descriptor
write-confinement check's Windows semantics are now verified deterministically
rather than only when the suite happened to run on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the tier count in TESTING-SUITES.md

Said the tier narrowed from 546 to 254. The final committed list is 265 of 931
eligible (58.8% -> 28.5%) after the shell-interpreter-spawn replacement restored
9 files and ALWAYS_REAL_OS pinned 1. Same error class the spec review caught in
the ADR, in a live reference page rather than a dated record, so it states the
current truth rather than carrying an amendment note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the measured aggregate from real CI job lists

Epic #4589's closeout asserted its reduction from a static count; #4641's
acceptance criterion asks for a figure read off a real run. Recorded here:
test.yml job count 21 -> 15 and non-Linux 7 -> 4, comparing PR #4640's run
against this PR's own. Against the true pre-epic baseline of 9, that is 9 -> 4.

Also states the caveat that a PR's total CHECK count is not a clean before/after
comparison, since many gates are path-scoped and this change touches a broader
path set -- the like-for-like figure is the test.yml job count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): compare job totals the same way on both sides

The measured-aggregate table put #4640's COMPLETED run total (21) against this
run's count at matrix-expansion time (15). Those are not the same measurement:
the completed total includes the post-test Coverage gate and baseline-publisher
jobs. Counted identically, it is 21 -> 17. The load-bearing figure, non-Linux
jobs 7 -> 4, was correct and is unchanged.

Called out in the table rather than silently corrected -- comparing two
differently-derived numbers is exactly the error class this ADR is about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record measured conformance wall-clock and date the stale counterfactual

Adds the per-job durations from both runs. The honest read is that this is a
correctness win more than a speed one: file count fell 52% but wall-clock only
9-29%, because what was removed were the cheap static tests and what remains is
concentrated in expensive spawn-heavy work. Stated explicitly so nobody expects
a future narrowing to buy time proportional to file count.

The load-bearing figure is windows shard 3/3: 40m24s against a 45-minute cap on
the 547-file tier -- 90% of the cliff #869 and #3057 were both filed about --
pulled back to 31m27s. macOS moved the wrong way (17m48s -> 21m02s) while its
tier was UNCHANGED at 197 files, which fixes that as runner variance and is
noted as a caution against reading a single duration as signal.

Also dates the symlink-keyword counterfactual, which cited a 254-file tier from
before the replacement category and allowlist took it to its final 265.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): re-measure against the rebased tree and disclose the allowlist's zero

next gained #4644 mid-flight, so every absolute count shifted. Re-measured on
the tree this actually ships against (932 eligible): 548 -> 257 by detector
removal, 257 -> 266 once shell-interpreter-spawn restores 9. Net 282 removed,
9 restored. macOS 198, unchanged by this PR.

The percentages did not move across three rebases (58.8% -> 28.5%), which is
the whole argument for expressing the ceilings as ratios rather than counts --
noted in the ADR since it is now evidence rather than assertion.

Also discloses that ALWAYS_REAL_OS now contributes ZERO files: this PR's own
win32 test cases introduced the literal win32 into the pinned file, so it
classifies in on content via win32-darwin-literal. The entry stays and the
reason is written down, because the file's real-OS need is a property of the
code under test (isPathConfined reads the ambient path module), not of the
test's text -- the text that currently saves it is incidental and could be
refactored away silently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 17:00:11 -04:00
Tom Boucher
db4d8a9bae fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic (#4644)
* fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic

$((10#${PHASE_NUMBER})) is a hard bash/zsh syntax error when PHASE_NUMBER is
decimal (01.1, from an inserted phase) or N-segment (23.1.2) — neither is
valid shell-arithmetic syntax at all, and the failed expansion aborts the
rest of the snippet in a non-interactive shell. safe_resume_gate runs
unconditionally before trusting STATE.md or dispatching any executor, so
execute-phase failed at its own gate before the first executor on any
decimal phase, regardless of workflow.tdd_mode. Regression from #4194.

Fixes all 4 sites: safe_resume_gate and the TDD gate in
workflows/execute-phase.md, the completion-signal spot-check fallback in
workflows/execute-phase/steps/completion-reconciliation.md, and the
executor gate validation example in references/tdd.md. Each now zero-strips
only the leading integer segment into a *_INT variable (via %%.* / #
parameter expansion — always valid shell syntax regardless of what follows)
and keeps the remainder as an escaped-dot string for the anchored commit-
scope regex, exactly as issue #4619 verified in both bash and zsh. A plain
integer phase (12, 01) computes byte-identically to before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4619): pin the decimal/N-segment fix and characterize the pre-fix bug

Behavioral coverage via real bash execution: the old $((10#01.1)) form
throws (characterizes the bug, matching the issue's own reproduction); the
new form resolves 01.1 -> 1\.1 and 23.1.2 -> 23\.1\.2, unchanged for plain
integers (12 -> 12, 01 -> 1); the resulting anchored ERE matches
feat(01.1-03):/test(1.1-3): and correctly rejects feat(01-03):,
feat(01.2-03):, feat(011-03):, feat(12-03): for a decimal phase — mirroring
issue #4619's own verified table exactly. Updates
safe-resume-gate-anchoring.test.cjs's 4 existing source-text assertions
(one per site) to the new fixed text.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): refine the shell-arith drift detector to distinguish safe from unsafe arithmetic

With #4619's fix in place, the guard's original "ban $((10#... outright,
match any occurrence" was too blunt: it flagged a comment merely mentioning
the pattern in prose, the now-safe $((10#$PHASE_INT)) arithmetic on an
already-%%.*-stripped integer, and the always-safe plan-id arithmetic
(plan ids are plain integers, never decimal). Refines the detector to skip
full-line comments and to only flag a captured variable/placeholder name
that contains "phase" and does NOT end in _INT/_int — the naming convention
the #4619 fix establishes at all four sites for "already reduced to a safe
integer." A plan-id variable was never phase-number arithmetic in the first
place and is excluded on the same basis.

This closes epic #4634's D6 ("lint-phase-id-drift... passes with no new
exemptions") and D7 ("a decimal and N-segment phase id survive an
end-to-end execute-phase selection without error") for real — the guard now
reports zero violations across all five .cts/.md rules.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: regenerate conformance-tier manifests for the new test file

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4619): cover the plain-padded-integer near-miss matrix too

Review found the anchored-ERE near-miss coverage only exercised the
decimal case (PHASE_NUMBER=01.1); issue #4619's own worked table also
verifies the plain padded-integer case (01 -> PHASE_N=1) against its own
near-miss set (matches 01-03, rejects 01.1-03/011-03/12-03). Adds the
missing assertion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4619): add Fixed changeset

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4619): correct JS backslash-escaping in safe-resume-gate anchoring test

The test's string-literal assertions for the PHASE_FRAC//./\\.} pattern wrote
only 2 backslash characters in JS source, which single-quoted-string parsing
collapses to 1 real backslash at runtime -- but the workflow/reference files
actually contain 2 raw backslash bytes at that position (needed so bash's
${var//pattern/replacement} produces the correct single-backslash output).
Write 4 backslash characters in the JS source at all 4 occurrences so the
runtime string matches the files' real bytes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4619): refresh the committed compact-content benchmark baseline

The new PHASE_INT/PHASE_FRAC arithmetic lines added to
gsd-core/workflows/execute-phase.md shifted its committed compaction-ratio
baseline. Regenerate via `node scripts/benchmark-compact-content.cjs --write`.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4619): note the safe_resume_gate arithmetic growth in the test header

The emitted-attribution gate flags execute-phase.md growing 91253 -> 91846
bytes (593 bytes). The growth is the fix: the safe_resume_gate and TDD RED
block now derive PHASE_INT/PHASE_FRAC before computing PHASE_N, so a
decimal/N-segment phase number (e.g. 01.1, 2.3.1) zero-strips its leading
integer segment via base-10 arithmetic instead of forcing the whole value
through $((10#...)) and hitting a hard shell syntax error on the first dot.

A blank line previously separated the Emitted-Drift-Ack-Growth trailer from
the Co-Authored-By trailer below it, which splits git's trailer-block
detection: only the last contiguous non-blank run of Key: Value lines at the
end of a commit message is recognized as trailers, so the growth ack was
silently read as ordinary body text and the differential-attribution gate
failed with the growth unacknowledged. Joining the two trailers into one
contiguous block fixes it.

Emitted-Drift-Ack-Growth: execute-phase.md — adds PHASE_INT/PHASE_FRAC derivation to the safe_resume_gate and TDD RED commit-scope grep so a decimal/N-segment phase number zero-strips its leading integer segment via base-10 arithmetic instead of failing on a non-numeric value (#4619)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4208): replace chmod-based restore-failure injection with a root-proof git shim

`tests/commit-files-deletion.test.cjs`'s two restore-failure tests simulated
an unwritable index via a `post-index-change` hook running `chmod a-w` on
the git dir. That relies on the OS enforcing the *owner's own* permission
bits against itself, which uid 0 (a routine identity inside this repo's
Docker-based gsd-test benches) does not: every DAC check short-circuits true
for root, so the write the chmod meant to block silently succeeds, the
restore comes back clean, and the disclosure/rollback behavior under test
never actually gets exercised.

This is CLAUDE.md's own named anti-pattern for I/O-failure injection
("Cross-platform test IO-failure injection" — chmod tricks fail under root
Docker/CI). It is confirmed as the actual root cause here, not a production
defect: `src/commands.cts`'s `restoreRemovedEntries`/rollback-disclosure
logic (added by #4253, merged just before this run) was hand-traced and
manually reproduced end to end on an unprivileged workstation against a
freshly built `gsd-core/bin/lib/commands.cjs`, and it already produces
exactly the `staging_failed` + "could not be restored" / "could NOT be
restored during rollback" results both tests assert. The other
`post-index-change`-based tests in this file (a `sleep` to force a timeout;
a real `update-index` to flip a restored entry's mode) are unaffected
because neither depends on a permission check — consistent with only the
two chmod-based tests failing on the real remote run.

Replaces the chmod fixture with a fake `git` placed ahead of the real one on
PATH that fails only `update-index --add --cacheinfo` — the one call the
restore makes — unconditionally, regardless of privilege level. Every other
git invocation execs straight through to the real binary, so the rest of
each scenario (`rm --cached`, the restore's own `ls-files` verification,
etc.) is exercised exactly as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4619): backfill changeset pr number to 4644

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4619): feed the bash fixture script via stdin, not argv, to fix Windows CI

Passing the script as a `-c "<script>"` argv element made it subject to
Windows' CreateProcess command-line argument encoding, which silently
dropped the escaped-dot backslashes before bash ever saw them (observed on
PR #4644's windows-latest CI shard: `1\.1` came back as `1.1`). Feeding the
same script via stdin instead removes argv entirely from the transport, so
there is nothing for Windows to re-encode. POSIX behavior is unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 15:47:14 -04:00
Tom Boucher
5e0a7b1b56 fix(#4433,#4569,#4126): consolidate the phase-identity seam at name-validity, allocation, and branch-slug (#4640)
* fix(#4433): apply the name-validity guard symmetrically to every milestone-name capture

extractMilestoneHeadingName already refused a punctuation-only captured name
(#4134), but its two sibling capture sites in getMilestoneInfo — the
STATE.md-anchored 🚧-bullet match and the no-STATE.md in-progress 🚧-bullet
fallback — skipped straight to a bare truthiness check, so a malformed bullet
whose only content past the version was punctuation passed through as a real
milestone name.

Extracts the existing inline /[\p{L}\p{N}]/u check into a single shared
hasNameableContent predicate and applies it at all three capture sites, so
the guard is one owner rather than a copy that happened to land at only one
of them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4433): pin the name-validity guard at all three milestone-name capture sites

Failing-first coverage for the hasNameableContent extraction: a
punctuation-only 🚧-bullet name must not surface as a real milestone name,
either on the STATE.md-anchored path or the no-STATE.md in-progress
fallback, while a real name (including a digits-only one) still resolves
COMPLETE exactly as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4569): consolidate decimal-phase-number allocation into one function

cmdPhaseInsert allocated its next decimal sub-phase number by scanning only
on-disk phases/ directories and ### Phase N.M: headings, never the roadmap
summary checklist — so a decimal that existed only as a checklist bullet
(no heading yet, no on-disk directory yet) was invisible, and phase insert
could silently reallocate an already-used number. It also always nested one
level deeper under afterPhase, with no way to request a sibling.

cmdPhaseNextDecimal had its own separate, near-identical two-source scan
(missing the checklist source too) — the exact "duplicate implementations
kept in sync instead of deleted" pattern this issue exists to close.

Extracts scanExistingDecimalPhaseNumbers (directories + headings + checklist
bullets, in one place) and migrates both cmdPhaseInsert and
cmdPhaseNextDecimal onto it — deleting cmdPhaseNextDecimal's own copy rather
than patching it in parallel. Adds an allocation: 'nested' | 'sibling'
argument to cmdPhaseInsert (default 'nested', matching every existing
caller's behavior); a top-level phase with no existing decimal segment falls
back to nested since there is no sibling level to join. No CLI flag wires
'sibling' yet — that is a separate, disclosed follow-up.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4569): pin decimal-allocation coverage across phase insert and next-decimal

Failing-first coverage for scanExistingDecimalPhaseNumbers: a checklist-only
decimal must not be reallocated by phase insert; a decimal present in
heading, checklist, and on-disk directory simultaneously must count once;
an unrelated phase family's checklist bullet must not cross-pollute; and
phase next-decimal (migrated onto the same shared helper) must see a
checklist-only decimal too, closing the same gap in a second command.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): extend the phase-id drift guard for name-validity and shell arithmetic

The epic's ratchet requirement: lint-phase-id-drift.cjs must cover the two
new predicates this PR introduces, and must also scan shell inside
gsd-core/workflows/**/*.md and gsd-core/references/**/*.md for
integer-coercing phase-number arithmetic ($((10#...)) and friends), which
neither the canonical TypeScript module nor a source-only lint can reach.

Adds findNameValidityDrift (bans re-deriving /[\p{L}\p{N}]/u outside
hasNameableContent's owner file) and findShellPhaseArithDrift +
scanMarkdownShellArith (bans $((10#...)) in workflow/reference markdown,
sanctioned via <!-- phase-id-owner: --> on the preceding line). scanRepo
keeps its existing, narrower contract (src/**/*.cts only) so the
already-passing "the live repo is clean" test is untouched; a new scanAll
merges both for the CLI's full report.

Running the guard directly against this tree correctly reports the 7
pre-existing #4619 shell sites (workflows/execute-phase.md x4,
workflows/execute-phase/steps/completion-reconciliation.md x2,
references/tdd.md x1) as violations — demonstrating the ratchet works, not
fixing them. #4619 is a live regression tracked and fixed separately; this
PR does not touch those markdown files. A characterization test pins the
current count of 7 so a future change to that number is investigated rather
than silently absorbed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4569): wire --sibling through phase insert's CLI so the argument is reachable

cmdPhaseInsert's allocation parameter had no CLI path to 'sibling' — shipped,
untested, unreachable code (code-review finding: a guaranteed surviving
mutant). Adds --sibling to phase insert's argument parsing, threads it
through, and documents the flag in docs/CLI-TOOLS.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4569): exercise --sibling end-to-end through the real CLI

Confirms --sibling joins afterPhase's parent decimal level rather than
nesting, and falls back to nested when afterPhase has no existing decimal
segment (no sibling level to join).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4634): demonstrate the two new drift detectors end-to-end via a planted violation

The epic asks for the guard to be "demonstrated by watching it go red" on a
reintroduced copy. The two new detectors (name-validity, shell-arith) had
only unit-level fixture tests; mirrors the existing bracket-rule's
planted-violation-in-a-temp-tree test for both, proving they're actually
wired into scanRepo/scanMarkdownShellArith end-to-end, not just correct in
isolation.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): consolidate the drift guard's own owner-sanction-check logic

Standards review flagged the "walk to nearest preceding non-blank line,
check for a phase-id-owner comment" logic as duplicated across all four
detector functions in a PR whose whole point is eliminating exactly that
pattern. Extracts isSanctionedByPrecedingComment, shared by all four;
behavior-preserving (verified: identical output before/after, same 7 known
#4619 violations, zero token/bracket/name-validity).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): add Fixed changeset for the name-validity guard and allocation consolidation

pr:0 placeholder — backfilled once the real PR number exists.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4126): consolidate branch-name slug substitution into one shared renderer

cmdCommit (commands.cts) and cmdInitExecutePhase (init.cts) each
independently implemented branch-name template substitution, and both
substituted the literal string 'phase' when phase_slug was empty or
undeliverable — producing a non-identifying branch name (gsd/phase-08-phase)
that contradicted the honestly-reported phase_slug: null in the same
payload. Same structural defect as the other three gaps in this epic: two
consumers reimplementing one concept independently instead of sharing an
owner.

Adds renderPhaseBranchName (src/phase-id.cts) as the sole owner: a real slug
substitutes normally; an empty/undeliverable one drops the {slug} token plus
one adjacent separator (collapsing/trimming the result) rather than
substituting a placeholder word, for the shipped default template and any
user-configured shape alike. Both call sites now delegate to it; the old
inline duplicates are deleted, not kept in sync. {project} substitution
stays a separate step in init.cts, unchanged, since it is a config-level
field with its own fallback contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4126): pin renderPhaseBranchName and both migrated call sites

Property-based coverage for the shared renderer's degrade-path invariant
(output, when non-null, never contains {slug} and never starts/ends with a
separator), plus example coverage for real-slug substitution, empty/null/
non-string slug, token position at either edge, a doubled-separator
template, and the only-{slug} -> null case. One regression test each in
commands.test.cjs and init.test.cjs confirms a phase with no derivable slug
no longer produces a branch name ending in the literal '-phase'.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: route scanExistingDecimalPhaseNumbers through the canonical enumeration owner

Caught by an actual gsd-test run, not a hypothesis: the new decimal-scan
helper (fix(#4569)) enumerated phases/ directories via a raw
fs.readdirSync, which the pre-existing phase-enumeration drift guard
(#3185/#3882) correctly flags as an unsanctioned re-derivation outside its
canonical owner (listAllPhaseDirs / isSentinelPhaseId). Ironic given this
epic's own thesis, and exactly why the guard exists: consolidating one seam
can reintroduce drift in an adjacent one if the new code doesn't route
through what's already there. Migrates the enumeration to listAllPhaseDirs;
identical decimal-detection output for every existing case.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): extend the drift guard for branch-slug fallback; fix a real regex bug

Adds the fourth detector the epic's ratchet section names ("both
branch-name sites"): bans a `.replace('{slug}', ... || 'phase')` call
outright, sanctioned via renderPhaseBranchName or a dedicated comment.
Wired into scanRepo (no per-file exemption — this is a banned anti-pattern
everywhere, not a grammar with one legitimate owner). Now that #4126's fix
(prior commit) has landed, scanRepo reports zero violations across all four
.cts-scanning rules, restoring the simple "the live repo is clean" assertion
instead of a pinned-known-count characterization.

Also fixes a real bug an actual gsd-test run caught: findNameValidityDrift's
regex didn't tolerate the doubled-backslash template-string form its own
test claimed to cover (0 !== 1) — widened to \{1,2} matching
TOKEN_DRIFT_RE's existing tolerance for the same two forms.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4126): document the {slug} degrade behavior; update changeset for the full seam

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: detectPhaseNumberFromFiles wrongly rejected bare, slug-less phase directories

Caught by an actual gsd-test run on the #4126 regression test, not a
hypothesis: a bare phase directory with no slug remainder (e.g.
.planning/phases/01/) has extractPhaseToken correctly return "01" — which is
simply identical to the directory name in that case, not its no-match
fallback. A stale `token !== phaseDir` check treated that equality as "no
numeric token found" and rejected it regardless, leaving phaseNum null and
silently skipping cmdCommit's phase-branching block entirely (the commit
proceeded on whatever branch was already checked out instead of the
phase branch).

phaseTokenShape.test(normalized) already excludes every genuine non-phase
case on its own: extractPhaseToken's real no-match fallback only fires for a
dirName that doesn't start with a digit or short letter+digit prefix, and
normalizePhaseName's leading-\d+ requirement rejects those regardless. The
equality check was redundant for real rejections and actively wrong for
bare-numeric directories.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: backfill changeset PR number to 4640

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 12:42:04 -04:00
0xdhx
4cc2a466b5 fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec (#4253)
* fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec

`cmdCommit`'s `--files` list can stage an addition but never a deletion:
the #2014 guard skips a missing explicit entry because the filesystem
cannot tell "moved away" from "not written yet". A caller that moves a
file therefore had two forms, both wrong — a directory entry records the
move but also commits every unrelated file in that directory (a
concurrent session's in-flight todo, in the unattended execute-phase
sweep), and a file entry leaves the old path's deletion dangling with the
todo tracked at both paths.

`--files-removed <paths>` is the caller-declared delete intent. Each entry
names a file, or a directory whose tracked-but-absent files are the
removals; those paths are staged with `git rm --cached` and join the
commit pathspec. `--files` keeps its skip-if-missing contract untouched.
A file entry still present on disk fails the commit closed with the
existing staging-failure rollback; a never-tracked path is a no-op.
`--files-removed` alone is a declared scope, not the unscoped .planning/
sweep.

The dispatcher previously folded every non-flag token after `--files`
into that list, so a second list flag could not exist; each list now
runs from its flag to the next `--` token.

The execute-phase todo sweep names the moved todos on both sides from
CLOSED[@], and cleanup's archive commit moves .planning/phases/ and
.planning/quick/ under --files-removed.

Fixes #4208

Emitted-Drift-Ack-Growth: cleanup.md — the archive commit moves phases/ and quick/ under --files-removed; the growth is one paragraph stating why those two directories must not be --files entries

* chore(#4208): set changeset fragment pr to 4253

* fix(#4208): fit execute-phase.md under the ADR-857 ceiling and re-point the #2415 guard

Three CI failures, all consequences of this PR's own change.

1. gsd-core/workflows/execute-phase.md was 93,577 bytes against the
   ADR-857 Phase 6 margin gate's <= 93,400 (hard ceiling 93,600). The
   three-line rationale comment plus the four-line array-building block
   added 318 bytes to a file that had only 141 of headroom on next.

   Move the rationale to docs/CLI-TOOLS.md -- which this PR already
   extends with the --files-removed contract, and which is where the
   ADR-857 gate wants call-site detail to live rather than in the host
   workflow -- and fold the array build onto one line. 93,577 -> 93,372.

2/3. tests/close-phase-todos-stage-deletion.test.cjs pinned the #2415
   guarantee to its old MECHANISM: it regex-matched the literal
   .planning/todos/{completed,pending}/ directory pathspecs in the
   commit --files list. This PR deliberately replaced those with named
   files (a directory entry also committed an unrelated todo a
   concurrent session dropped in mid-close), so the guard failed on a
   change it should have accepted.

   Re-point it at the new mechanism without weakening it: assert the
   ADDED array reaches --files, the REMOVED array reaches
   --files-removed, STATE.md is still committed, and -- newly -- that
   the two arrays are built from $COMPLETED_DIR and $PENDING_DIR
   respectively. Verified by negative control: deleting
   --files-removed "${REMOVED[@]}" from the workflow still fails the
   test, so the #2415 regression remains caught.

Note for the merge queue: #4233 also grows execute-phase.md (+114). The
two are additive -- different regions, no textual conflict -- so with
both landed the file reaches ~93,486, over the 93,400 margin though
under the 93,600 hard ceiling. Whichever merges second will need to
reclaim ~86 bytes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0183892Y3fxxirte4WNmBKbv

* fix(#4208): reclaim execute-phase.md bytes so the PR is net-neutral under the ADR-857 margin

Rebasing onto next surfaced the byte-gate collision flagged earlier on
this PR: #4284 grew execute-phase.md by 95 bytes (93,259 -> 93,354),
so this PR's +113 landed at 93,467 against the <= 93,400 margin in
tests/claude-orchestration.test.cjs.

Compact the close_phase_todos step this PR already edits -- drop the
PHASE_NUM indirection, fold the normaliser and the match guard, print
the closed list with one printf, shorten the step's prose -- without
touching the mechanism the #2415 guard pins (ADDED/REMOVED arrays, the
plain mv). 93,467 -> 93,349: 5 bytes under the base, so the PR no
longer spends any of next's 46 bytes of headroom.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): classify absent index entries before staging a removal; restore removed entries exactly on rollback

Review of #4253 found three Majors with one root cause: the removal
side judged presence by fs.lstatSync alone, where the addition side
already reads `git ls-files -v` state. Absence from the worktree is not
removal:

- a submodule gitlink (mode 160000) whose directory was deleted by hand
  lists like a file and was `rm --cached` with no .gitmodules cleanup;
- a skip-worktree path is never materialised by a cone-mode sparse
  checkout, so a directory entry over a sparse-excluded tree dropped
  that whole tree from the index;
- an assume-unchanged path's worktree state is not something git
  itself consults;
- an intent-to-add entry (`git add -N`) renders as a plain cached entry
  on the empty blob, yet nothing tracked exists to remove and no
  rollback can restore the flag.

The index listing now carries each entry's `ls-files -v -s` tag, mode
and stage. Only a plain cached (H), stage-0, non-gitlink entry is a
removal candidate; every other state is left alone under a directory
entry (exactly like a present file) and fails closed when named
directly, with the state in the error. "Named directly" is decided on
RESOLVED paths, not strings -- realpath of the longest existing prefix
with the absent tail re-appended: an absolute path, `./x`, `--cwd`, or a
symlinked spelling of the tree (macOS `/var` ->
`/private/var`, where `process.cwd()` is the real path and the caller's
absolute path is not -- CI on this round's first push) all resolve to the
same entry, where a string compare against git's cwd-relative output
silently took the directory polarity (pre-push review, driven; the
symlink case is driven with an aliased fixture directory). The enumeration's domain is what
`ls-files -v -s` can emit for an index entry, stated at the classifier.

The third Major -- on an unborn HEAD a successful `rm --cached` was
never rolled back when a later entry failed -- is fixed differently
from the review's suggestion. Pushing the path into stagedPaths would
put it on the commit pathspec, which a root commit refuses ("pathspec
did not match", driven), and `git reset -- <path>` cannot restore an
entry with no HEAD anyway. Instead every index entry this call removes
is recorded (mode, blob) before the `rm` and put back with
`update-index --cacheinfo` on rollback. That also restores a
caller-pre-staged blob at a removed path exactly, where a reset would
have silently replaced it with HEAD's version. The rollback is
best-effort, as the addition-side reset already was, and the docs say
so.

Eight tests: gitlink under a directory entry, named directly, and named
by absolute path; skip-worktree both forms; intent-to-add both forms;
assume-unchanged named; unborn-HEAD partial failure restores the
removal; pre-staged blob survives the rollback.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): drop the empty fenced block left dangling in cleanup.md's commit step

Review nit on #4253: inserting the --files-removed rationale between the
original bash block and its closing fence left an empty ```bash``` pair
before </step>. Harmless at runtime, a formatting artifact of this PR's
own diff; removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): a boolean flag inside a commit path list no longer ends the list

Review minor on #4253: collectList stopped at the next `--` token, so a
positional wedged between a boolean flag and the next list flag
(`--files a --amend b --files-removed c`) was claimed by neither list
and silently dropped -- a regression in shape against the old
slice-to-end parse, which filtered `--` tokens and kept `b`. No current
call site interleaves that way, but the gap was real.

A list now runs to the next LIST flag (`--files` / `--files-removed`)
and skips boolean flags on the way, and a REPEATED list flag merges
its runs (`--files a --files b` -> [a, b]) as the slice-to-end parse
did -- a first cut stopped at the repeat and dropped `b`, the same
silent-drop shape one level over (pre-post comment audit). The only
change #4208 makes to parsing is that a second list flag can exist.
Tests: STATE.md wedged between --no-verify and --files-removed lands
in the commit; both runs of a repeated --files reach it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* test(#4208): drive the reappearance window with a post-index-change hook

Review nit on #4253: the defensive re-check for a file recreated between
the absence test and `git rm --cached` -- the concurrent-session race
this PR's own changeset names -- had no test. git fires
post-index-change the moment `rm --cached` writes the index, so a hook
that copies the file back exactly then exercises the window
deterministically. The call reports staging_failed / "reappeared on
disk", commits nothing, and the rollback restores the removed entry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): restore a staged removal when the call records nothing

A `git rm --cached` that succeeds mutates the index whether or not a commit
follows. Only the staging-failure rollback put those entries back, so a call
that reached `nothing_to_commit` reported no state change while the removal sat
staged -- riding along on the caller's next commit.

The review named the unborn-HEAD, removal-only shape. Keying on `headExists`
would have fixed half of it: the guard also fires with a real HEAD when the
removed path is index-only (added, never committed), because `diff HEAD` reads
clean with the path absent on both sides. Both shapes now restore, at both
`nothing_to_commit` exits. The failure exits are deliberately left alone --
they report a failure rather than no-change, and the addition side leaves its
own staged paths there too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* refactor(#4208): lift declared-removal staging out of the cmdCommit hotspot

`cmdCommit` was a critical-risk hotspot before this flag existed, and #4208 had
inlined another ~270 lines into it. `stageDeclaredRemovals(cwd, removedDeclared)`
now owns the index-state classification, path canonicalisation and entry
recording, returning the pathspec entries and the recorded removals its caller
merges.

Pure motion: no branch, message or probe changed. Only the two accumulators
became local names, and `restoreRemovedEntries` stays with the caller because
the exits that restore are the caller's. cmdCommit 888 -> 625 lines here; the
extracted helper is 277.

(Figures corrected after publication: an earlier version of this message said
854 -> 591 and claimed the result was below cmdCommit's pre-#4208 shape. Both
were wrong -- the count came from a faulty brace scanner, and `next`'s cmdCommit
is 581, so this is above it, not below.)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): property-test the two-list commit parser

RULESET.TESTS.property-based-testing asks a parser for at least one property
test asserting a domain invariant; `collectList` had only hand-picked examples,
one per shape a review round had already broken.

Hoisted it to module scope as `collectListFlagValues` and exported it in the
file's existing exported-for-tests convention -- a parser reachable only by
spawning the CLI can be tested one example at a time and no faster.

Three properties over generated argv: every positional lands in exactly the run
open at it whatever the flag order or count; no positional after the first list
flag is dropped or double-claimed; and with `--files-removed` absent the parse
equals the pre-#4208 slice-to-end parse. Controlled against two mutants -- a run
ending at any `--` token, and a repeated list flag that does not merge -- each
of which the properties catch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): pin cleanup.md's archive commit to --files-removed

execute-phase.md's rewrite is pinned by the #2415 guard in this file;
cleanup.md's equivalent was not, so reverting its routing would have been
caught by nothing -- the mechanism's unit tests never read this file and pass
either way.

Asserts the two archived directories are under --files-removed and NOT under
--files (where a directory entry sweeps in a concurrent session's in-flight
writes), and that the destinations and STATE.md stay on the additive half.
Controlled by restoring the pre-#4208 sweep, which fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): pin that a symlink to a directory is one tracked path

Review of #4253 read the `lstatSync(...).isDirectory()` test as a
symlink-following defect. Driving it says the opposite: git tracks the link as
a single blob (mode 120000) and does not traverse it, so the tracked paths
"under" it live at the real directory and were never named by the caller.
Following the link would stage those -- the directory sweep #4208 exists to
remove -- while the named entry still sat present on disk.

Pinned rather than changed, with the premise driven in the test body. Swapping
`lstatSync` for `statSync` -- the prescription as written -- fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4208): refresh the compact-content baseline for this PR's execute-phase edit

The base range added `tests/benchmark-compact-content.test.cjs` and a committed
token baseline over the compacted workflows. This PR edits
`gsd-core/workflows/execute-phase.md`, so the baseline drifts by +12 tokens on
that entry and on the aggregate.

Refreshed with `node scripts/benchmark-compact-content.cjs --write`; the diff is
those two entries and nothing else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): report a removal the call could not put back

Round review of this round found the restore itself unchecked: the helper
ignored `update-index`'s exit code, so a FAILED restore still reported
`nothing_to_commit` -- the same false "no state changed" the restore exists to
prevent, surviving one level down on the restore-failure path.

It now returns a boolean. The two no-change exits report `staging_failed`
naming the paths left staged; the staging-failure rollback still ignores it,
deliberately, because it is already reporting a failure and an unwritable index
is usually the failure being reported.

Driven with a post-index-change hook that makes the git dir unwritable the
moment `rm --cached` lands, so the restore cannot take its lock. Reverting both
guards fails the test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): disclose a removal the rollback could not restore

Round review refuted the reasoning behind leaving the rollback path's restore
unchecked. The claim was that this exit is already reporting a failure, so the
restore's result adds nothing. The counterexample is the ordinary case: the
reported failure is usually a DIFFERENT cause -- a contradictory declaration, a
reappeared path -- so a caller reading `failures` sees only that cause and
learns nothing about the removal still sitting in its index.

The rollback now appends a disclosure entry per un-restored removal, naming the
path. The reason and `file` still report the failure that caused the rollback;
the disclosure is additive.

Also moves the restore-failure test's chmod into a `finally`: `t.after` runs
AFTER the parent `afterEach`, so a throw before it left the fixture undeletable.

Both driven; reverting the disclosure fails the new test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): decide index state by observation, never by an exit code

The restore added two commits earlier keyed both its record decision and its
success verdict on git's exit code. An exit code answers "did the command
succeed", never "did the index change" -- execGit collapses a spawn timeout to
a non-zero exit, and a killed git can already have written the index. Round
review drove four failures from that one assumption, in both directions:

  - a failed `rm` still contributed an entry, so the rollback disclosed a
    removal that was never staged (stale index.lock);
  - a timed-out `rm` whose write DID land contributed none, so a real mutation
    was neither restored nor disclosed;
  - a timed-out `update-index` whose write landed reported failure, publishing
    a "could NOT be restored" disclosure that was false;
  - and the read-back that replaced it omitted `-z`, so core.quotePath rendered
    `café.md` as `"caf\303\251.md"` and an exactly-restored entry read as not
    restored -- the same quoting defect this PR already fixed for `preStaged`.

Everything now observes the index. A failed `rm` re-reads `ls-files -z` for the
path: gone means this call owns the removal and records it; still there means
nothing was staged; a probe that cannot answer becomes its own failure entry
rather than an assumption. The restore verifies the same way, comparing the
WHOLE entry (mode, blob, stage), because `--cacheinfo` restores all three and a
path-only test accepts an entry that came back as something else.

The verdict is three-valued -- `restored` / `not-restored` / `unverified` --
and the unverified wording says the restore could not be VERIFIED rather than
that it failed. The rm's own failure is pushed ahead of any probe diagnostic so
a timed-out removal keeps `timed_out: true` and its own message as the reported
cause.

Five regression cases, each negative-controlled against the shape it pins.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): treat a declared removal path as a path, not a pathspec

An index path handed back to git is parsed as a PATHSPEC, and the removal side
handed several back. Three driven harms, all of them the sweep-in this flag
exists to remove, arriving through the operand rather than through a directory
entry:

  - a tracked file literally named `.planning/*.md` made `rm --cached` GLOB: it
    removed `peer.md` and `stays.md` too, only the declared entry was recorded,
    so the rollback restored one of three and the other two rode out as staged
    deletions the result disclosed nowhere;
  - the same name reached `git commit -- <paths>`, which globbed and committed
    an undeclared `M peer.md` alongside the declared removal;
  - and the intent-to-add probe (`diff --cached` over the path) matched a
    STAGED PEER instead of itself, so an `add -N` entry was misclassified as
    ordinary content, removed, and restored by `--cacheinfo` -- which cannot
    restore the intent flag. It came back as a real staged addition.

Every operand on this path is now `:(literal)`: the `rm`, both index probes,
the intent-to-add probe, the restore read-back, the entry-level `ls-files` /
`ls-tree`, and -- for the REMOVAL-derived entries only -- the downstream
`ls-files` / dry-run / `diff HEAD` / `commit` pathspec. `--files` entries keep
whatever pathspec behaviour they have today; that is not this change's to
alter. `:(literal)` still resolves a directory to its descendants (driven), so
the directory form is unchanged.

Closes what an earlier cut of this commit declared as a residual: a filename
beginning with `:` is now removable end to end, because the commit pathspec no
longer reinterprets it.

Also fixes a MINOR from the same review: cleanup.md's contract test checked the
destinations' position relative to `--files-removed` but never that `--files`
was present at all, so deleting the flag still passed.

Un-literalising the seven sites fails three of the new tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): scope the rollback to the caller's own name space

Round review drove a rollback that destroyed the caller's own staged work. Two
causes, one of them pre-existing:

  - `git diff --cached` prints REPO-relative paths whatever the cwd, while
    `stagedPaths` holds the caller's cwd-relative names. In a project nested
    inside its repo (`<repo>/sub/.planning/...`) the two name spaces never
    intersect, so `preStaged` matched NOTHING, every path landed in `toUnstage`,
    and the reset unstaged a caller-staged deletion and modification that this
    call had never touched. `--relative` makes the two sets comparable, and is a
    no-op when the project IS the repo root. This governs the `--files` side too
    and predates this flag.
  - the rollback's `reset` was the last place a removal-derived name reached git
    as a bare pathspec; it takes `asPathspec` like every other site.

Driven on a nested fixture; dropping `--relative` fails the new test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): gate six fixtures that Windows cannot construct

CI's `test (windows-latest, 24, shard 2/3)` went red on this round. Two
primitives the new fixtures rely on do not exist on Windows, both driven on a
real Windows host rather than inferred:

  - a filename containing `*` or `:` cannot be created at all (`IOException` /
    `FileNotFoundException`), which is four of the pathspec fixtures;
  - `chmod` cannot make a directory unwritable — a write into a ReadOnly
    directory succeeds — so the two restore-failure fixtures cannot drive the
    failure they exist to drive.

Each is skipped on win32 with its measured reason, in the repo's existing
`{ skip: process.platform === 'win32' ? '<reason>' : false }` form. The
behaviours they pin are platform-independent; only the fixtures are not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): build git's index-syntax path with forward slashes

The remaining Windows red was mine, not the platform's: `git rev-parse :<path>`
takes a forward-slash path, and `path.join` yields backslashes there, so git
rejected it as an ambiguous argument. The hook in the same test already used
the slash form.

Not gated — the behaviour it pins is portable; only the argument was not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4208): refresh the compact-content baseline against the rebased base

`next` moved the `new-project` split and the aggregate under this PR's
execute-phase entry; regenerated with `scripts/benchmark-compact-content.cjs
--write` so the only leaves differing from the base's copy are the
execute-phase split and the aggregate it feeds.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh

* chore(#4208): regenerate the macOS conformance tier for this PR's fixtures

`next` gained the macOS-specific conformance tier (#4593) after this branch
was cut. Its classifier (`scripts/gen-platform-conformance-tier.cjs --target
macos`) now selects `tests/commit-files-deletion.test.cjs` on the
`chmod-mode-bit` and `symlink-keyword` signals the PR's fixtures carry (the
chmod-driven failed-restore cases and the symlink-to-directory case).
Regenerated with `--target macos --write`; the platform tier was already in
sync. The file was modified, not added, which is why the added-files check
did not surface it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-11 12:26:55 -04:00
Tom Boucher
523be34133 fix(#4282): register PATTERNS.md as a canonical .planning/ artifact (#4618)
* test(#4282): prove PATTERNS.md is unrecognized by the artifact registry

Regression test only, no fix yet: CANONICAL_EXACT in src/artifacts.cts was
never updated when workflows/graduation.md started writing .planning/
PATTERNS.md, same omission class as the already-fixed #3224 (WINDOWS.md).

Expected RED on this commit (src/artifacts.cts is unchanged).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4282): register PATTERNS.md as a canonical .planning/ artifact

CANONICAL_EXACT in src/artifacts.cts was never updated when
workflows/graduation.md started writing .planning/PATTERNS.md for the
`patterns` graduation-target category -- same omission class as the
already-fixed #3224 (WINDOWS.md). validate.health's W019 falsely flagged it
as unrecognized on every repo that has run the graduation scan.

Also backfilled 5 other pre-existing stale rows in
gsd-core/templates/README.md's artifact table (WINDOWS.md, STATE-ARCHIVE.md,
milestone.lock, state.json, skill-manifest.json) that were already in the
source registry but missing from the docs table -- found while fixing this
exact drift class, cheap to close alongside it.

RED proven on f3dd791fb8cda18196803e7144ce20e506d6490b (test-only commit,
gsd-test outcome:failed, exactly the new PATTERNS.md test failing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4282): fix stale function name in skill-manifest.json comment

Review finding: both the source comment and the new docs row said
"routeSkillManifest" -- no such symbol exists (verified via Memtrace); the
actual function is cmdSkillManifest (src/init.cts). Copied verbatim from a
pre-existing comment, not introduced by this PR, but cheap to fix alongside.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4282): add changeset fragment

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4282): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: isolate lint-vendored-deps-manifest.test.cjs's fixRow tests from the real vendor file

Genuine, pre-existing defect found and fixed per this repo's no-defer policy
(discovered while investigating a real CI failure during this PR's own
merge attempt, user-directed investigation -- not deferred to a separate
issue since it was actively blocking work and root-caused with concrete
evidence, not speculation).

Root cause: fixRow(row) (scripts/lint-vendored-deps.cjs) unconditionally
does fs.copyFileSync(upstreamCjs, vendoredCjs) as its first line. All three
tests in the #4573 describe block called fixRow(row) with the REAL js-yaml
row, so all three wrote to the real, shared gsd-core/bin/lib/vendor/
js-yaml.cjs -- a file other test files' require() calls can read at any
moment, since node --test runs files concurrently in this repo.
fs.copyFileSync's write is not atomic against a concurrent reader on every
filesystem; a concurrent require() elsewhere caught the file mid-overwrite
and read a truncated file, crashing an entirely unrelated test
(m9-statelock-write-error-orphan.test.cjs) with a SyntaxError.

Confirmed via two real CI log fetches, not assumed: the exact same shard
grouping (same 308 files) ran clean ~90 minutes earlier during PR #4615's
own final merge CI, with the identical #3660 reap-fix code already present
-- ruling out a deterministic connection to that change and confirming a
genuine, non-deterministic timing race in this pre-existing test design.

Fix: all three tests now redirect row.vendoredCjs to a private os.tmpdir()
path via a cloned row object before calling fixRow, so the real vendored
file is never touched. upstreamCjs stays pointed at the real node_modules
copy (read-only, safe to share).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: also isolate fixRow's package.json pin-rewrite from the real file

Review finding (major) on the previous race-condition fix: fixRow's pin
-rewrite path still hardcoded path.join(ROOT, 'package.json'), so the third
#4573 test still wrote the real, shared package.json -- read at module
top-level by dozens of other test files, the same concurrent-file race
class already fixed for the vendored .cjs copy.

Adds an optional pkgRoot parameter (defaults to the real ROOT) threaded
through readPinState/checkRow/fixRow -- fully backward-compatible, every
existing call site (the CLI --fix path, any other caller) is unaffected
since the default is unchanged. The pin-rewrite test now builds an isolated
temp root (its own package.json + node_modules/js-yaml/package.json) and
passes it explicitly, so the real package.json is never touched either.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: use helpers.cleanup instead of raw fs.rmSync in test cleanup

CI caught it: local/no-raw-rmsync-in-tests flagged the three t.after temp-dir
cleanup calls added for the fixRow isolation fix. helpers.cleanup() carries
the Windows-EBUSY retry budget (maxRetries/retryDelay) that raw fs.rmSync
lacks.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 22:46:18 -04:00
Tom Boucher
1e47560e34 feat(#4593): add a macOS-specific conformance tier, final phase of epic #4589 (#4607)
test-conformance's macos-latest leg (Phase 2, #4591) has been running the
same 546-file, Windows-oriented conformance-tier list as windows-latest --
built from signals like windows-shell-token/windows-env-var that have
nothing to do with macOS. Issue #4593 asked for macOS coverage sized to
its own evidence-backed surface (zsh dispatch, case-sensitivity, darwin-
specific behavior) instead.

Issue #4593 was filed before Phase 5 (#4603) existed and referenced
updating test-full's macOS legs -- that job is gone. Corrected the issue's
body before any code was touched: the "shrink from full replay" half of
the original ask was already done by Phase 5; what remained was narrowing
the still-Windows-oriented tier macOS was inheriting.

Two design assumptions were measured and rejected before accepting a
design (documented in docs/adr/4593-macos-conformance-tier-architecture.md):
- Reusing the general tier's signals minus its 3 Windows-specific
  categories barely narrows anything (546 -> 424, 78% retained) -- most
  files match multiple signals and only need one to survive exclusion.
- A standalone CRLF/autocrlf signal, despite the issue naming
  "CRLF-checkout behavior": even narrowed to /\bCRLF\b|autocrlf/i it hit
  143/930 files. Root cause: CRLF is primarily a Windows checkout concern
  in this codebase (ADR-1703 files it under DEFECT.WINDOWS-TEST-
  PORTABILITY), so the signal was really re-selecting Windows-relevant
  files already covered by the general tier, not narrowing macOS
  specifically.

Built 5 new, genuinely macOS-specific signals instead: darwin-literal
(darwin alone, not the general tier's win32-OR-darwin), zsh-dispatch,
case-sensitivity, plus chmod-mode-bit and symlink-keyword reused verbatim
from the general tier (genuinely Unix-relevant, not Windows-motivated).
Measured against the real tree: 196 of 930 eligible unit-suite files
(21%), versus the general tier's 546 (59%) -- a real, evidence-backed
narrowing.

scripts/gen-platform-conformance-tier.cjs gains classifyMacosContent/
classifyMacosTree/renderMacosGeneratedFile and a --target windows
(default, unchanged)/--target macos CLI flag, so the same generator
produces two independent, gated outputs rather than needing a second
script. New committed output: scripts/lib/macos-conformance-tier.
generated.cjs. .github/workflows/test.yml's test-conformance job: only
the macos-latest leg's file-list source changes; windows-latest is
byte-for-byte untouched. New shipped-file ripples handled proactively
(19 install-tree fixtures regenerated, bin/install.js registered).

An isolated code-review pass found one real defect: the ADR's per-
category count table had drifted by 1 (zsh-dispatch, case-sensitivity)
because the new test file's own fixture strings joined the tree it
classifies after the table was authored -- fixed, with the union total
(196, what CI actually gates on) confirmed unaffected. An isolated
security-review pass found no qualifying findings.

The ADR also records an explicit requirement for any future widening
proposal: check whether the motivating regression is already covered by
Phase 1's no-rendered-text-length-assert lint rule (#4590) before
re-proposing full macOS/Linux parity, since that is exactly what #4421's
root cause was (a rendered-text-length assertion, not a real behavioral
divergence).

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 14:37:15 -04:00
Tom Boucher
0b928fe28c feat(#4592): replace the blanket test-file full_matrix rule with reachability (#4602)
scripts/ci-test-scope.cjs's classify() previously set full_matrix=true for
ANY changed tests/**/*.test.cjs file, unconditionally (restored by #4421
after #962's narrowing let a real macOS-only regression, PR #4384, land
undetected). This replaces that blanket rule with a reachability check
against real data instead of a path prefix:

- A changed test file forces full_matrix only when it is present in Phase
  2's committed CONFORMANCE_TIER_FILES list (scripts/lib/platform-
  conformance-tier.generated.cjs) -- direct membership, not a graph walk.
- A changed src/ file forces full_matrix when its own content carries a
  genuine platform-conditional signal, reusing gen-platform-conformance-
  tier.cjs's classifyContent with a narrowed, source-code-safe signal
  subset (excludes two categories -- hardcoded-path-vs-path-call and
  symlink-keyword -- empirically found to flag 100/235 src/ files when
  applied verbatim, versus 28/235 with the narrow subset, all verified to
  carry genuine platform branches). New export: NOISY_FOR_SOURCE_REACHABILITY.
- A change to the classification mechanism's own definition files
  (gen-platform-conformance-tier.cjs, the generated tier list, or
  suite-detection.cjs) always forces full_matrix -- the mechanism being
  changed cannot presume its own new output is safe.
- Any computation error (a require/read failure, a malformed module) fails
  safe to full_matrix=true, per the issue's explicit requirement.

The existing RULES array entries with their own fullMatrix:true (workflow
automation, installer/package layout, hooks, environment/dependency gates,
test harness) are deliberately left untouched -- they are curated,
narrowly-scoped triggers for "this diff changes the CI/installer/hooks
mechanism itself," a different and still-valid reason than "product code
might reach a platform branch." Disclosed in .gsd/phase/.../40-design.md
as a scope decision, since the issue's "Done when" wording read broader
than its "Proposed work" bullets.

Two design assumptions were caught and corrected before any code was
written (rubber-duck pass, documented in 40-design.md): (1) reusing Phase
2's classifyContent verbatim against src/ was far too noisy; (2) a single
hardcoded seam file (src/shell-command-projection.cts only, per CLAUDE.md's
"single platform seam" framing) would have silently missed genuine,
independent platform branches in src/runtime-hooks-surface.cts,
src/capability-lock.cts, src/capability-ledger.cts, and src/surface.cts --
reintroducing the #4421 failure shape inside src/ instead of tests/.

An isolated code-review pass found and fixed one real defect (a dead,
untested branch that would have survived Stryker mutation testing) and one
design-doc completeness gap (2 of 8 "narrow" signal categories were left
implicitly rather than explicitly audited). An isolated security-review
pass found no qualifying findings.

tests/ci-test-scope.test.cjs gains the full #4592 boundary-case matrix
(.gsd/phase/.../50-test-matrix.md), including a named #4421 regression case
proving tests/state-todos-render.test.cjs still forces full_matrix, now for
the documented reason instead of the removed blanket rule. Two pre-existing
tests were corrected: one used a nonexistent fixture path (src/semver.cts
-> src/semver-compare.cts, a real file); one (A3) asserted the exact old
blanket-rule behavior this issue removes, updated to the new, verified-
correct expectation.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 12:46:21 -04:00
Tom Boucher
181c4c8659 chore(#4603): retire the test-full CI job (#4604)
* chore(#4603): retire the test-full CI job

Phase 2 (#4591) added test-conformance but left test-full (the pre-existing
full-suite Windows/macOS replay) running unchanged, gated on the same
full_matrix flag, downgraded only from a hard gate to a non-blocking
::warning:: -- framed as "a non-gating safety net for one release cycle."
No phase or issue ever retired it. Result: every full_matrix=true PR ran
10 OS-specific jobs (test-full's 6 + test-conformance's 4, purely
additive) instead of the original 6 -- the epic's own goal (reduce
runner-minutes) was measurably regressing, not improving, for the
majority of PRs.

This phase was missing from the original 4-phase epic decomposition; the
epic (#4589) has been amended to add it as Phase 5 (see its comment
thread), and this issue was filed as the tracked sub-issue.

Deletes the test-full job from .github/workflows/test.yml entirely, along
with every reference to it: required-tests' needs/FULL_TEST_RESULT
warning branch, ci-timeout-report.cjs's JOB_RULES entry,
ci-test-job-timeout-budget.test.cjs's LANE_COSTS/staticLanes/testFullRule
entries, ci-test-scope.test.cjs's test-full-specific tests (preserving
three unrelated tests that were nested in the same describe block, moved
under a renamed describe rather than deleted), and docs mentions.
test-conformance is now the sole gating signal for real-OS coverage.

Two separate defects found and fixed while auditing every test-full
reference:
- tests/ci-pr-mergeability.test.cjs's GATED['test.yml'] safety-critical
  array (jobs that must needs: the mergeability preflight) had test-full
  but was missing test-conformance entirely -- Phase 2 never added it.
  Verified the real workflow wiring was already correct (test-conformance
  does have needs: [changes, preflight]); this was a test-coverage gap,
  not a live defect. Fixed by swapping the array entry.
- docs/TESTING-SUITES.md's "## CI matrix" section was substantially stale
  independent of this phase (predating even #2952's coverage-gate split).
  Rewritten against the real, current job topology, verified directly
  against test.yml rather than trusted from memory.

An isolated code-review pass found and fixed two minor inaccuracies in the
rewritten docs table (two jobs' "Gated on" column didn't match their real
if: condition exactly). An isolated security-review pass found no
qualifying findings -- every compute-provisioning job already carries
needs: preflight directly, unaffected by this deletion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(ci): isolate 7 more heavy test files from chunk-weight packing

`next`'s own push-triggered Tests run failed: `conformance test
(windows-latest, 24, shard 2/3)` chunk 3/6 was killed after 600019ms.
Root cause: state.test.cjs (weight 21.35, measured) was packed alongside
companions by run-tests.cjs's LPT chunk packer, the same failure mode
that previously hit codex-config.test.cjs (weight 17.87) twice and got a
dedicated fix (ISOLATED_HEAVY_FILES, #4497) -- but state.test.cjs was
never added to that set.

This is a direct, unintended consequence of epic #4589 Phase 2: the new
platform-conformance-tier job packs only ~546 files per shard (vs. the
~950-file full suite the packer used to balance against), so the same
absolute-weight outlier now represents a larger share of a smaller, more
homogeneous pool -- the LPT packer has fewer light files to pad around
it with. This was a real, foreseeable side effect of shrinking the
packing pool that nobody checked for when Phase 2 shipped.

A first attempt at this fix hand-picked 4 candidates by eyeballing a
truncated weight list and missed 3 heavier ones -- caught by an isolated
code-review pass (blocker: emitted-attribution.test.cjs at 66.2% of the
Windows chunk budget, install-minimal-hooks.test.cjs at 61.1%,
install.test.cjs at 47.1%, all above codex-config.test.cjs's own
44.7% -- the ratio that already proved dangerous twice). Corrected by
systematically computing weight/budget for every unit-suite file and
isolating everything at or above that same ratio: 7 files total, plus
the pre-existing codex-config.test.cjs (8 total).

Added a durable regression test (tests/run-tests-harness.test.cjs) that
re-derives this exact computation from the live tests/test-timings.json
on every run, so a future heavy file crossing this threshold fails the
test instead of silently reintroducing this failure -- not just a
one-time manual sweep.

Verified end-to-end: simulated the real 3-way windows shard split of the
actual conformance-tier file list with the real packing functions. Max
packable-chunk weight across all 3 shards is now 27.04 / 24.10 / 23.91
(shard 2 is the exact shard that failed on next), comfortably under the
40 budget -- versus 40+ and a 600s kill before this fix.

A second isolated code-review + security-review pass on the corrected
diff found nothing further.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 12:46:04 -04:00
Tom Boucher
bcd99696d3 chore(#4591): add platform-conformance-tier classifier + gate CI on it (#4598) 2026-09-10 08:36:13 -04:00
Tom Boucher
615b74ff45 fix(#4460): correct two stale changesets left by an admin-merge race (#4572)
* fix: address orthogonal-review findings on the new work in this PR

Isolated code-review + security-review of everything added to this PR
since its original review (hono override, check-env.cjs rewrite/revert,
new lib file, its test, installer enumeration). Security review: clean,
no findings. Code review found:

- BLOCKER: .changeset/silly-hens-relax.md described a hono override
  this PR no longer actually makes -- PR #4560 landed the identical fix
  on next first, and this branch's own hono commit became a genuine
  no-op the moment it was rebased onto that updated next (git diff
  origin/next -- package.json package-lock.json is empty). Deleted the
  orphaned changeset; next already carries #4560's equivalent one
  (.changeset/zesty-seals-click.md).
- HIGH: .changeset/tame-hens-jump.md's body still described the
  execNpm-routing approach that was tried and reverted -- stale text
  from before that revert, would have shipped a release note for code
  that isn't actually in the diff. Rewritten to describe what actually
  shipped (self-contained spawnSync, 15s timeout, accurate ENOENT vs.
  timeout vs. non-zero-exit diagnosis).
- LOW: no comment explaining why the spawnSync call has no try/catch
  (safe -- its documented contract routes failures through the returned
  result, never a throw -- but worth stating given this file's whole
  purpose is graceful degradation). Added one.
- nit: exitCode 0 + empty stdout fell through to "npm binary not found
  on PATH", misdescribing a real npm binary that simply printed
  nothing. Gave it its own message; updated the corresponding test.

Manually re-verified describeNpmVersionCheckFailure's branches and the
real check:env success path before re-running gsd-test, since this
repo blocks local node --test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: rest of the orthogonal-review fixes (previous commit only caught the deletion)

Tooling mistake in the previous commit: a git add with the already-staged
deleted changeset mixed into the same pathspec list errored out and
silently skipped staging the other four files, so only the changeset
deletion actually committed. This commit carries the rest of that same
change: tame-hens-jump.md's rewritten body, check-env.cjs's no-try/catch
comment, npm-version-check-diagnosis.cjs's exitCode-0-empty-stdout fix,
and the corresponding test update.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4460): fix changeset pr field to point at this PR, not the original

.changeset/tame-hens-jump.md's pr field still said 4552 (the PR its
original text was authored under), but this PR (#4572) is what's
actually landing the corrected body -- changeset-lint's own
DEFECT.CHANGESET-PR-FIELD-DRIFT check caught it: "pr: 4552, expected
pr: 4572".

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 16:26:25 -04:00
Tom Boucher
37b965c0d1 enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code (#4553)
* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code

ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills
(src/init.cts) now selects between a canonical agents/<name>.md and a
token-minimized agents/<name>.compact.md sibling based on
workflow.compact_content, resolved in code (a real function call with a real
exit code) rather than a prose config-get gate — the same precedent stream 1's
spine/detail split established for a load-bearing seam, applied here because
this seam already runs through TypeScript instead of an eager @-include.

A missing compact sibling falls back to the canonical persona and discloses
the fallback in the served payload itself (a leading HTML-comment provenance
line), so the Done-when contract — compact when on, canonical when off, never
silent or empty — holds even for an agent nobody has compacted yet.

Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md),
each an independent, complete rewrite (not an extraction — nothing is "moved"
the way spine/detail moves text) that preserves frontmatter, every @-include,
every output-format contract, and every guardrail verbatim while cutting
restatement and verbose framing. Verified mechanically: every pair registers
(a canonical sibling exists), every compact file is strictly smaller, and the
full @-include set matches canonical's — including which references are
standalone eager-load lines versus inline prose mentions, since demoting one
to inline changes what the host actually substitutes.

Traced the install path before writing any code (.gsd/phase/.../40-design.md):
stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no
stem filtering under the default full profile, so the new .compact.md files
install for free with zero installer changes — matching issue #4407's stated
scope. A tiered agent profile that doesn't stage a compact sibling degrades
through the same fallback-with-provenance path already required for an
unauthored one, so no installer change is needed there either.

Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export
(deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are
reached by a generic code construction rather than a literal path in prose,
and checkReachability's markdown-search shape has nothing to find there).
Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new
"#4407 compact payload selection" describe block spawns gsd_run agent-skills
against real compact/canonical fixture pairs and asserts on the served
payload, which can only pass if the seam genuinely wires through.

Fixed a pre-existing test whose agents/*.md glob incidentally matched the new
.compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift
guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's
roster, both real, unrelated-to-content defects the new files' mere existence
surfaced.

Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files
under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token
benchmark baseline (npm run benchmark:compact-content-variants --write).

Closes #4407.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): apply orthogonal review findings from the compact-payload seam

Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath)
to collapse the duplicated read-and-empty-check shape between the compact
and canonical branches in cmdAgentSkills, and updated the adjacent comment
enumerating flat JSON extras to name agent_payload_variant alongside
source/degraded (added by the prior commit, comment left stale).

Security review and the Spec axis found no defects requiring a code change;
their non-blocking observations (a pre-existing, unmodified path-construction
pattern; the reasoned, documented substitution of a behavioral test for the
literal reachability check) are recorded in
.gsd/phase/enhance-4407-agent-skill-seam/60-review.json.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents

Root-caused via a real gsd-test run (93 failures) rather than guessing which
tests glob agents/ naively. Two classes of defect, both genuine:

1. Identity-roster confusion (11 files/areas): many tests and one production
   script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f
   => f.endsWith('.md'))`, which incidentally matched the new .compact.md
   variant siblings too — a compact file is a rendering of an EXISTING agent
   identity, not a new one. Fixed at the shared root
   (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests
   already consolidated on) and at each independent glob that didn't use it:
   agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix
   before checking XL/LARGE membership, so a compact file inherits its
   canonical sibling's tier instead of silently falling through to DEFAULT),
   agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual
   script, not just its test), codex-config.test.cjs (confirmed directly
   against generateCodexAgentToml that a compact role's derived sandbox_mode
   is byte-identical to its canonical sibling's before excluding it — not
   assumed), and copilot-install.test.cjs (two counts that legitimately DO
   need both files — an installed-file count and a full-conversion smoke test
   — fixed to expect 70, not stay pinned to 35).

   no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix:
   two compact files reproduce descriptive prose already allowlisted at their
   canonical file's line number; added matching entries at the compact files'
   own line numbers rather than excluding them from the scan (a genuine bare
   gsd-tools command-position bug in a compact file would be as real a defect
   as in canonical).

2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree
   run): six agents' compact renditions (gsd-debugger, gsd-executor,
   gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed
   the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction —
   confirmed structural, not a compaction-quality gap: each is dominated by
   content this phase's own rules require verbatim (the ~2.6 KB gsd_run
   bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in
   every agent that calls gsd_run, output-format contracts, guardrails).
   ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing
   spot in cmdAgentSkills's single-file synchronous read. Removed these 6
   compact files rather than ship an over-cap file or invent a multi-part
   read mechanism out of scope for this phase; recorded by name with the
   reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and
   50-test-matrix.md, per #4407's own "or explicitly recorded as not worth
   covering" allowance. Their canonical personas are served correctly today
   via the fallback-with-disclosed-provenance path this phase's own Done-when
   #2 already requires — 29 of 35 agents now have a compact variant.

Also fixes an unrelated, genuinely pre-existing defect this gsd-test run
surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own
DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in
this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md
| wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed:
two meaning-preserving trims in the <success_criteria> block (a repeated
parenthetical replaced with a same-exception reference; one redundant
qualifier dropped) bring it to 40,940 bytes.

Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant
benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6
now-orphaned roster rows removed alongside them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): make .compact.md-aware roster checks resilient to partial coverage

Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent
has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception),
breaking once 6 stems legitimately have none.

- tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was
  picking up the "### Compact Payload Variants" subsection's rows as
  phantom/uncounted entries in the primary/advanced/inventory-only
  classification this test validates — a compact row documents an existing
  agent's alternate rendition and never gets its own AGENTS.md heading, so it
  was never meant to participate in that classification. Excluded at the
  parser, not per-assertion.
- tests/copilot-install.test.cjs: the derived expected-file-list generator
  assumed every listAgentFiles() stem has a .compact.md source sibling;
  checks disk per stem now instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4407): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 12:38:59 -04:00
Tom Boucher
6bd5c22f68 ci: auto-refresh vendored js-yaml/re2js on Dependabot PRs (#4573) (#4576)
lint-vendored-deps.cjs gates gsd-core/bin/lib/vendor/{js-yaml.cjs,re2js.cjs}
for byte-freshness against node_modules and requires package.json's pin to
literally match the installed version. Dependabot regularly opens
lockfile-only PRs bumping these packages within the existing semver range,
which it can never satisfy (it has no awareness of the vendor copy or the
pin), so every such PR sits permanently red until a human manually runs the
refresh and pushes a fixup commit. PR #4565 was the latest instance.

Adds `--fix` to lint-vendored-deps.cjs (fixRow): mechanically re-copies the
upstream build artifact over the vendored .cjs/.d.cts twins and bumps the
manifest pin, preserving its range-operator style. It never touches a
hand-authored twin's declared type surface (js-yaml.d.cts) — a remaining
finding there means a real upstream API break, and --fix leaves it failing
rather than mask it.

Adds .github/workflows/dependabot-vendor-refresh.yml: on a same-repo
dependabot[bot] PR touching package.json/package-lock.json (same
defense-in-depth identity check as dependabot-auto-merge.yml), runs
`--fix` and, only on a clean result, commits and pushes the refresh back
to the PR branch via GSD_BOT_PR_TOKEN (falling back to GITHUB_TOKEN, same
pattern as auto-backmerge.yml) so the push re-triggers `synchronize` and
the real lint-vendored-deps check in test.yml genuinely re-passes. A
non-clean --fix result (real incompatibility) makes no commit, leaving
the actual failure visible for a human — this fixes the check's own
complaint, it does not bypass or weaken the check.

Closes #4573

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:46:12 -04:00
sim
8731a90be6 fix: revert execNpm import in check-env.cjs, keep the diagnosis improvement
The execNpm-routing redesign (previous commit) broke every real CI job:
check-env.cjs runs as its own standalone "Environment check" step BEFORE
`npm ci` / `npm run build:lib` -- a deliberate pre-flight, run before
there is even a node_modules to build with. Its require of
../gsd-core/bin/lib/shell-command-projection.cjs (a tsc-compiled artifact
that plain does not exist at that point in the pipeline) crashed with
MODULE_NOT_FOUND on every platform, immediately, confirmed via the real
CI log. My own local gsd-test run never caught this because it doesn't
replicate that exact pre-build step ordering.

Reverted the cross-module require entirely; check-env.cjs is back to a
self-contained spawnSync(npmCmd, ...) call, no requires reaching into
gsd-core/bin/lib. Kept the two things actually worth keeping from that
detour:
  - the 15_000ms timeout (matches execNpm's own default elsewhere in this
    repo -- not invented, an existing precedent -- vs. the original 10s
    that failed twice under real Windows CI contention);
  - computing `timedOut` via `error.code === 'ETIMEDOUT'` inline (the
    same canonical, cross-platform-correct predicate that seam uses),
    rather than the earlier signal === 'SIGTERM' check, which that seam's
    own docstring documents as platform-fragile with a Windows-specific
    false-negative risk.

Manually verified check-env.cjs runs correctly with gsd-core/bin/lib
temporarily removed entirely (simulating the real pre-npm-ci CI
ordering) before re-running gsd-test, since this repo blocks local
node --test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 08:07:34 -04:00
sim
1cd17cc3d6 fix: route check-env.cjs's npm-version check through the canonical execNpm seam
Per /research + /diagnose direction: the timeout fix landed earlier this
session correctly diagnosed the failure (a real spawnSync timeout under
Windows CI contention, not npm being absent) but it recurred on the very
next push -- same chunk, same ~51-file load. Rather than raise the
hand-rolled 10s timeout myself (CLAUDE.md's own rule: fix the cost, not
the tolerance, and never touch a timeout without explicit instruction),
investigated the repo's own precedent first.

Found: this repo already has a canonical OS-shell-projection seam for
exactly this (src/shell-command-projection.cts's execNpm), already used
by dozens of other scripts/*.cjs files (require('../gsd-core/bin/lib/...')
is an extremely well-established pattern), with:
  - the same npm.cmd/shell:true Windows handling check-env.cjs was
    hand-rolling, but centralized;
  - a 15s default timeout (vs. check-env.cjs's 10s) -- not invented here,
    an EXISTING value already governing npm subprocess calls elsewhere;
  - isSpawnTimeout / result.timedOut, the canonical cross-platform timeout
    predicate (error.code === 'ETIMEDOUT'), whose own docstring explicitly
    warns that checking signal === 'SIGTERM' (what my first fix did) is
    "platform-fragile" with a Windows-specific false-negative risk -- the
    exact platform this bug lives on.

check-env.cjs's npm-version check now calls execNpm(['--version']) instead
of hand-rolling spawnSync + npmCmd + shell:true, and
describeNpmVersionCheckFailure now operates on execNpm's SpawnResultOutput
shape (using timedOut, not signal) rather than a raw spawnSync result.
This is a genuine architectural fix, not just a bigger number: it removes
a duplicate, slightly-divergent re-implementation of an existing seam and
inherits whatever that seam's timeout/handling becomes in the future.

Manually verified end-to-end (npm run check:env against the real
environment) and re-verified describeNpmVersionCheckFailure's branches
directly against execNpm's actual return shape before wiring the test
file, since this repo blocks local node --test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 08:07:34 -04:00
sim
0aeafc6425 fix: report the real cause when check-env.cjs's npm-version check fails
Discovered blocking this PR's own Windows CI (unrelated to this PR's
actual diff, fixed inline per this repo's no-defer policy): PR #4552's
"full test (windows-latest, 24, shard 2/3)" job failed
tests/check-env.test.cjs's npm-version subtest with "npm binary not
found on PATH" under chunk 5/9's heavy load (51 concurrent files,
5+ minutes). Root-caused via the CI log: the check's spawnSync call
used a 10s timeout, and every failure mode -- ENOENT, a signal-killed
timeout, a non-zero exit, a thrown spawn error -- collapsed into that
one message (only `res.status === 0 && res.stdout` was checked), so a
genuine npm.cmd cold-start timeout under contention was indistinguishable
from npm actually being absent.

Extracted the reason-selection into describeNpmVersionCheckFailure, a
pure function in the new scripts/lib/npm-version-check-diagnosis.cjs
(kept out of check-env.cjs itself, which runs its CLI unconditionally
on require with no `require.main === module` guard, so the pure logic
can be unit-tested without triggering a real environment check).
Reports ENOENT, signal-kill, non-zero-exit, and thrown-error cases
distinctly. Does NOT raise the 10s timeout itself -- a slow subprocess
under contention is a cost to reduce, not a tolerance to widen.

Also fixed a stale tsconfig.build.tsbuildinfo incremental-build cache
discovered while verifying this change: npm run build:lib was silently
omitting gsd-core/bin/lib/markdown-table.cjs (a real, needed compiled
module -- src/state-document.cts requires it), which only surfaced via
npm run lint:generated-sync's gen-health-docs check failing with
Cannot find module. Deleting the cache and rebuilding fresh restored
it; docs/INVENTORY-MANIFEST.json needed no net change once the build
was genuinely complete.

Manually verified describeNpmVersionCheckFailure's five branches
directly (ENOENT, signal-kill, non-zero exit, thrown error, defensive
default) before wiring the test file, since this repo blocks local
node --test. Re-ran node scripts/check-env.cjs directly to confirm the
real success path is unaffected.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 08:07:34 -04:00
Tom Boucher
a27cb6b2fa enhance(#4139): Phase 6 — the lazily-read remainder and the artifact templates (#4540)
* enhance(#4406): the lazily-read remainder and the artifact templates

ADR-4139 Decision 3, Phase 6 of the #4139 Compact Content epic. Covers stream 1b
(gsd-core/workflows/<name>/{modes,steps,templates}/*.md) and stream 4
(gsd-core/templates/**) with a variant-swap mechanism, confirmed with the user:
two independent, complete files per covered path (canonical + .compact.md
sibling), with the gate picking which one gets Read at the call site. This is
a different shape from Phase 5's spine+detail partition, and is safe here
specifically because these files are already reached only by a runtime Read —
a missed Read already means zero overlay content today, with or without
workflow.compact_content, so selecting between two independently-complete
files introduces no new failure mode (documented in
gsd-core/references/compact-content-gate.md's new "Streams 1b and 4" section).

Disposition, after inspecting every candidate rather than trusting a byte-size
threshold (same rigor Phase 5 applied to review.md):

- Stream 1b: 1 of 78 files compacted (help/modes/full.md, a user-facing
  reference doc emitted verbatim, not orchestrator instruction). The other 9
  size-threshold candidates are dominated by fail-closed guards, exact CLI
  invocations, or output-format contracts (AskUserQuestion blocks) — recorded
  not-worth-compacting, same reasoning as Phase 5's review.md.
- Stream 4: a ground-truth reachability audit replaced the initial size-only
  candidate list. Two files (summary.md, user-setup.md) got compact variants;
  a third (spec.md) was drafted, then dropped after discovering its only two
  call sites are eager @-includes, not a runtime Read — stream-1 material
  hiding under gsd-core/templates/, not stream-4's actual mechanism. summary.md
  itself has 3 eager call sites and only 1 genuine runtime-Read call site
  (execute-plan.md); only that one was wired, so the compact variant's savings
  apply to the sequential single-plan execution path only.
- Discovered while auditing reachability: 12 gsd-core/templates/** files with
  zero references anywhere in workflow/agent/command prose, compiled source,
  or tests — dead scaffolding predating this phase. Deleted in this same PR
  per this repo's no-defer policy, after re-verifying against a computed
  path.join(...) pattern (not just a plain-string search) that nearly caused
  two genuinely load-bearing templates (user-profile.md, dev-preferences.md)
  to be misclassified as dead.

New checker (tests/helpers/compact-content-variant.cjs): registration,
reachability, protected-content-preserved, size-smaller — replacing Phase
3/5's disjointness/completeness checks, which assume a partition rather than
two deliberately-overlapping documents. The reachability check's own
"unprefixed match" guard had a real bug (rejected the repo's own
`~/.claude/gsd-core/...` convention), caught by running it against the
already-wired help/modes/full.compact.md pair rather than only synthetic
fixtures — fixed to anchor on the nearest `gsd-core` path segment instead.

Template consumer parity (tests/compact-content-template-variant-parity.test.cjs):
proves each compact variant's `## File Template` fenced block — the actual
output-format contract a generated SUMMARY.md/USER-SETUP.md is parsed
against — is byte-identical to the canonical file, then runs the one real
deterministic consumer (gsd-core/bin/lib/coverage.cjs's classifyContent,
backing `gsd-tools uat classify-coverage`) against content built from that
shared contract.

Added a sibling benchmark script (scripts/benchmark-compact-content-variants.cjs)
rather than extending the existing spine/detail one — different data shape,
and the existing script's own contract deliberately isolates it from a
test-only helper's shape changing.

Emitted-drift acknowledgement: not needed. Every changed/added path in this
diff is hand-authored and present in the diff itself, so diffEmitted's
attribution loop resolves `via` to the path's own source before reaching the
ack-lookup branch (same reasoning Phase 5 verified for its own diff).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* enhance(#4406): address code-review findings on the variant-swap gate

- docs/CONFIGURATION.md and gsd-core/references/planning-config.md's
  workflow.compact_content entries described only the spine+detail mechanism
  (Phase 5) and were missing this phase's variant-swap mechanism and its
  benchmark:compact-content-variants script entirely — required since this
  PR's changeset is type Added (CLAUDE.md's "Missing Docs for Changesets"
  rule). Both now describe both mechanisms and which call sites are wired.
- Added the missing RED^-1/no-op fixture for checkProtectedContentPreserved:
  a canonical file with zero <!-- gsd:protected --> blocks must be a
  no-op, not a violation — the only branch of that function the existing
  fixtures didn't exercise.
- Collapsed findCompactFiles/findMarkdownFiles in
  tests/helpers/compact-content-variant.cjs into one findFilesWithSuffix
  helper — the two were identical recursive walks differing only in the
  extension predicate (minor Duplicated-Code finding).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4406): restore copilot-instructions.md, a false-positive dead-template classification

gsd-test caught this, not static analysis: 10 real failures in
tests/copilot-install.test.cjs, tests/installer-migration-install.integration.test.cjs,
and tests/repo-layout.test.cjs — all downstream of bin/install.js's Copilot install
path, which does
fs.readFileSync(path.join(targetDir, 'gsd-core', 'templates', 'copilot-instructions.md'))
after copying gsd-core/templates/** into the target project, then merges it into both
.github/copilot-instructions.md and (local installs) AGENTS.md. The reachability audit
that flagged this file as dead checked src/*.cts and gsd-core/bin/*.cjs but never the
repo-root bin/install.js — a separately maintained installer bundle outside the
src/-to-gsd-core/bin/lib/ compiled-output convention. The fs.existsSync guard around
that read degrades to a silent skip rather than a crash when the template is missing,
which is why this surfaced only once the real E2E install test ran, not from any
static check.

Re-verified the remaining 11 deleted filenames against bin/install.js specifically
(plain substring and quoted-filename search) before trusting that list — all 11 have
zero hits there, confirmed dead by the same standard this one file failed.

Regenerated the installer emitted-tree goldens (tests/fixtures/install-tree/*.json) to
reflect the restored file, and corrected the "Removed" changeset (jolly-lynx-sprint.md)
and the phase design doc from 12 to 11 deleted files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: execute-plan.md — call-site wiring for the summary.md and user-setup.md .compact.md variants
Emitted-Drift-Ack-Growth: help.md — call-site wiring for full.compact.md, same variant-resolution rule

* docs(#4406): backfill changeset PR numbers

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4406): resolve removed-but-needed lint findings on the dead-template deletion

CI's own full-test matrix (not gsd-test's matrix, which does not run this
check) caught 4 more false-positive dead-template classifications via
tests/removed-but-needed-lint.test.cjs / scripts/lint-removed-but-needed.cjs
— a literal, word-boundary basename check across .github/workflows/,
gsd-core/, and docs/ (excluding docs/adr/** and docs/research/**) for every
file a PR deletes. It has no semantic awareness, so a deleted template's
basename colliding with something else entirely still fires:

- claude-md.md: gsd-core/templates/README.md had a stale table row claiming
  /gsd-profile reads this template to generate CLAUDE.md. Verified false (no
  code reads it anywhere, same search that already covered bin/install.js) —
  fixed the row to *(inline)*, matching every other command-generated
  artifact in that table. File stays deleted.
- codebase/testing.md: collided with docs/guides/testing.md, an illustrative
  example row in docs-update.md's sample output table (an unrelated real
  generated-docs path). Swapped the example topic to "contributing" — the
  row is illustrative, any topic works. File stays deleted.
- codebase/architecture.md, codebase/stack.md: collided with docs/reference/
  planning-artifacts.md's directory listing of a user's own generated
  .planning/codebase/architecture.md and stack.md output — the same
  semantic mismatch already investigated and dismissed as unrelated earlier
  in this phase's audit, now caught by a gate instead of judgment. That
  listing repeats across 5 locale copies of the doc.
- continue-here.md: collided with the real .continue-here.md pause-work
  artifact, referenced across 15+ locale and workflow files.

For the last two, the lint's own error message offers "restore the file or
update every consumer in the same commit." Rewording 15+ files across
languages I cannot verify translation quality for, to shave 2 already-tiny
templates that were merely presumed dead, is disproportionate to this PR's
actual scope — restored codebase/architecture.md, codebase/stack.md, and
continue-here.md instead, and corrected docs/ARCHITECTURE.md's Templates
section accordingly.

Final confirmed-dead set: claude-md.md, codebase/concerns.md,
codebase/conventions.md, codebase/integrations.md, codebase/structure.md,
codebase/testing.md, debug-subagent-prompt.md, discovery.md — 8 files, down
from the original 12. Verified locally: GSD_REMOVED_BUT_NEEDED_BASE=next
node scripts/lint-removed-but-needed.cjs now passes clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: docs-update.md — swapped an illustrative example-table topic (testing -> contributing) to avoid a removed-but-needed basename collision with the deleted codebase/testing.md template; net +10 bytes

* fix(#4406): split codex-config.test.cjs to fix a genuine Windows CI timeout

Root cause of the `full test (windows-latest, 24, shard 2/3)` failure the
user asked to be actually fixed, not just re-run past: PR #4497 (landed
2026-09-07, one day before this PR's CI run) isolated
tests/codex-config.test.cjs into its own dedicated chunk because its
measured weight (17.87, ~45% of the post-cut Windows budget) made it unsafe
to share a chunk with any other file. That isolation was necessary but not
sufficient — even alone, with zero companion-file contention, the file's
real Windows execution time sits right at the 600s per-chunk ceiling. Two
independent CI runs on two unrelated PRs (this one and #4154) were both
killed within ~1.4s of the identical 600000ms mark — not random contention,
a deterministic near-miss the isolation fix couldn't address because it
never reduced the file's own cost, only removed the risk of a companion
file's cost stacking on top of it (which the PR #4497 comment explicitly
anticipated: "if a future profiling pass genuinely speeds up
codex-config.test.cjs itself, this isolation can be revisited").

The file itself explains why it's this heavy: 11,262 lines / 433 tests / 79
describe blocks, accumulated over dozens of bug-fix PRs (#2695, #2760,
#3245, #3285, #3346, #3426, #3427, #3562, #3566, #3582, #3808, and more),
several of which are explicitly documented as "folded" in from separate
files that were never actually split back out ("Verified non-duplicate
against both the pre-existing target and the other three folded sources").

Split into 4 files by top-level AST statement boundaries (never a naive
column-0 regex — an early attempt at that overcounted 79 apparent
"describe(" matches when only 21 are genuinely top-level; the rest are
nested inside a handful of large folded-in blocks, which a regex can't tell
apart from real top-level statements). Verified lossless twice: the split
script asserts byte-for-byte reconstruction of every source character, and
independently, total test()/describe() call counts match exactly between
the original file and the sum across all 4 new files (433/79 both sides).
Each new file carries the complete original shared header (imports/helpers)
for safety; per-file unused-import warnings from that duplication are
resolved via ESLint-precise alias renames (`{ foo: _foo }`, the standard
form for an intentionally-unused destructured binding — never a bare `{
_foo }`, which would destructure a different, nonexistent property).

No change needed to scripts/run-tests.cjs's ISOLATED_HEAVY_FILES or its
pinned test in tests/run-tests-harness.test.cjs: the file that keeps the
original name (tests/codex-config.test.cjs) is now only ~28% of the
original's size and safely isolated in its own chunk as before; the other
three new files re-enter normal weight-balanced packing, none individually
close to disproportionate. Confirmed no other file hardcodes the hardcoded
filename anywhere that would silently stop these tests from running (the
CI test-selection scripts determine scope algorithmically, not by literal
filename).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 14:31:04 -04:00
Tom Boucher
147c89a9b8 fix(#4456): forward --ws to every downstream new-milestone.md call (#4545)
* fix(#4456): forward --ws to every downstream new-milestone.md call

new-milestone.md's Step 1 parses --ws <name> into GSD_WS, but each
workflow step's bash fence is a separate shell invocation — GSD_WS set in
Step 1 never survived to Steps 5, 6, or 7. Four call sites never
forwarded it: init.new-milestone (both calls), state.milestone-switch,
and both phases.clear branches. Under GSD_WORKSTREAM env or a stored
session pointer differing from the explicitly requested --ws, every
downstream operation silently operated on the wrong workstream (or root)
instead of the one the caller asked for.

Confirmed --ws is a universally-parsed CLI flag (gsd-core/bin/gsd-tools.cjs:
4867, resolveActiveWorkstream) — stripped from argv and written into
process.env.GSD_WORKSTREAM for the rest of that process, so appending it to
ANY gsd_run query call works uniformly. Fixed by persisting GSD_WS to
.planning/.gsd-ws-arg right after Step 1 parses it (mirroring the
established .gsd-outgoing-milestone round-trip idiom this same file
already uses for the identical cross-fence problem), reading it back in
each later step, and appending it unquoted (matching the ${GSD_WS}
splicing convention documented in workstream-flag.md). Cleaned up after
its last use in Step 7.

Bundled, in-scope fixes found while implementing the above (per this
repo's no-defer policy):

- Step 6's phase-archive `git add .planning/milestones/ .planning/phases/`
  hardcoded literal ROOT paths — both directories are workstream-scoped
  (matching cmdMilestoneComplete's established #1911 precedent), so under
  a workstream this staged nothing real. Added phases_dir/archive_dir
  fields to cmdInitNewMilestone and resolved through them instead.
- Step 6's milestone-start commit hardcoded .planning/STATE.md — also
  workstream-scoped, so it would commit the wrong (or a stale) file under
  a workstream. Resolved through init.new-milestone's existing state_path
  field instead; PROJECT.md correctly stays a literal-shaped-but-resolved
  root path (shared, per the #4455 follow-up already merged).
- cmdInitNewMilestone's config_path field: config.json is ALSO a shared
  file (marked `# Shared` in workstream-flag.md's directory diagram, same
  as PROJECT.md) but was resolved via the workstream-aware planningDir —
  fixed alongside cmdInitNewProject's identical instance of the same bug
  (found via grep, matching the precedent from the #4455 follow-up of
  fixing every occurrence of an identically-evidenced bug uniformly).

Verified: direct CLI invocation confirms phases_dir/archive_dir/state_path
resolve into the workstream while project_path/config_path stay root under
GSD_WORKSTREAM=alpha. Manual bash-fence execution of every modified fence
(Step 1 parse+persist, Step 5 forwarding, both phases.clear branches, the
git add fence, Step 7's forward+cleanup, the commit fence) confirms correct
behavior in both flat and --ws modes, including two flags composing
together (--archive-version + --ws; --reset-phase-numbers + --ws).

Rewrote the pre-existing "step 6: commit stages PROJECT.md" test, which
asserted the literal (buggy) --files string verbatim — it now asserts the
resolved paths via a JSON-returning stub, and gained isolated per-test tmp
dirs (the prior version ran with no explicit cwd, at real risk of writing
a stray .gsd-ws-arg into this repo's own .planning/ once Step 1's fence
started performing a real write).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4456): Steps 9/10 also commit workstream-scoped files via literal root paths

A fresh isolated code-review pass on the first version of this fix found
the identical bug in two more places, missed in the initial sweep:

- Step 9's requirements commit (`gsd_run query commit ... --files
  .planning/REQUIREMENTS.md`) and Step 10's roadmap commit (`--files
  .planning/ROADMAP.md .planning/STATE.md .planning/REQUIREMENTS.md`)
  both hardcoded literal ROOT paths for files that are workstream-scoped.
- Worse: `.planning/.gsd-ws-arg` was being deleted at the end of Step 7,
  but Steps 9 and 10 run AFTER Step 7 and still needed to re-read it —
  the round-trip mechanism this fix builds was already gone before its
  two remaining consumers ran.

Fixed by moving the `.gsd-ws-arg` cleanup to Step 10 (its true last
consumer, after the roadmap commit) and adding the same
fetch-then-_gsd_field-extract pattern already used in Step 6 to Steps 9
and 10, resolving `requirements_path`/`roadmap_path`/`state_path` through
`init.new-milestone $GSD_WS_ARG` instead of literal paths.

Also fixed (MEDIUM, same review pass): Step 1's `.gsd-ws-arg` write had
no `2>/dev/null || true`, unlike every other round-trip write in this
same file — brought into line with the established idiom.

Verified: reproduced the pre-fix bug directly (Step 9/10 fences echoing
the literal root paths regardless of --ws), confirmed both fences now
resolve the workstream-scoped paths correctly, and confirmed the
round-trip file survives Step 7 and is only removed after Step 10.
Updated the Step 7 test that previously asserted premature cleanup
(inverted to assert the file survives); added new coverage for Steps 9
and 10 in both flat and --ws modes.

Two remaining LOW/pre-existing findings from the same review pass,
deliberately left as-is: `phase_archive_path` (src/init.cts, untouched by
this diff) resolves via the same root-only `getLatestCompletedMilestone`
this fix's earlier commit already declined to touch, for the same
genuine-product-intent-ambiguity reason (workstream-scoped vs
project-pooled "latest completed milestone" is not resolvable from the
code alone). `.planning/research/` staying root-scoped in the #222
self-heal prose is consistent with the existing (unchanged) `research_dir`
field, not a new inconsistency.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4456): revert wrong config_path change; fix isolated-cwd test env

gsd-test caught two real regressions from this fix's earlier commits:

1. config_path is NOT shared like PROJECT.md. The prior commit's
   grep-and-replace ("fix six more functions with the identical bug")
   also touched cmdInitExecutePhase's config_path (a fourth call site
   beyond the two I'd manually checked) — but tests/init.test.cjs's
   pre-existing, ADR-0006-governed "init handlers honor GSD_WORKSTREAM"
   coverage explicitly asserts config_path IS workstream-scoped for
   execute-phase/plan-phase/phase-op/milestone-op. workstream-flag.md's
   "# Shared" marking for config.json is stale (the same class of
   staleness already found for milestones/ during the #4455 follow-up);
   ADR-0006 plus its real, passing tests is the authoritative source.
   Reverted config_path to the plain workstream-aware planningDir(cwd)
   in all four functions it was wrongly changed in.

2. Isolating cwd to a tmpDir (needed once Step 1's fence started
   performing a real .gsd-ws-arg write) broke the runtime-launcher
   preamble's own gsd-tools.cjs discovery — no git repo at an isolated
   tmpDir, no global gsd_run on the CI bench's PATH. Fixed by passing
   RUNTIME_DIR explicitly in every isolated-cwd test's env, matching
   the preamble's own documented override precedence.

Verified: direct CLI invocation confirms execute-phase's config_path is
workstream-scoped again under GSD_WORKSTREAM=wsx; the RUNTIME_DIR fix
confirmed against a stripped PATH (no global gsd_run), matching the
bench condition that surfaced the original failure.

Emitted-Drift-Ack-Growth: new-milestone.md — #4456 forwards --ws to every downstream gsd_run call across 7 fences (Steps 1/5/6x3/7/9/10), adding a persisted round-trip file plus resolved-path fetches that replace several hardcoded literal paths
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4456): backfill changeset PR number and correct final scope

pr: 0 -> pr: 4545, and removed the changeset's claim that config.json
is a shared file -- that was the change this same PR later reverted
after gsd-test caught it contradicting ADR-0006's established,
workstream-scoped config_path contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4456): baseline the 10 new SC2086 findings from --ws forwarding

The lint-tests CI job failed with a hard exit 1. Diagnosis (not assumed):
the log's two `fatal: ambiguous argument 'origin/next...HEAD'` git errors
(lines 244/248) are a red herring — both belong to
lint-removed-but-needed.cjs, which prints its own "could not resolve
origin/next, skipping" message and exits gracefully, exactly like the
already-handled two-dot-form error from lint-fix-has-regression-tests
earlier in the same log. Neither contributes to the actual failure.

The real cause is lint-workflow-shellcheck: this fix's new fences append
$GSD_WS_ARG unquoted to gsd_run calls (deliberately, so it splits into 0
or 2 argv tokens — the same idiom gsd-core/workflows/verify-work.md
already uses for ${GSD_WS} and already has baselined). ShellCheck
correctly flags each as SC2086, and lint-workflow-shellcheck.cjs's
baseline is a deliberate ratchet (#4109) requiring new findings to be
explicitly accepted, not auto-passed. new-milestone.md previously had
zero baselined SC2086 findings, so all 10 new (correct, intentional)
occurrences were reported as new and failed the gate.

Added 10 {file, code, message} entries to
scripts/lint-workflow-shellcheck-baseline.json for
gsd-core/workflows/new-milestone.md's SC2086 findings, matching the
established, already-accepted precedent for the identical pattern in
verify-work.md. No source or workflow file changed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 09:23:20 -04:00
Tom Boucher
e03921c7d8 enhance(#4405): split the rest of the eager-window workflows worth splitting (#4536) 2026-09-08 00:17:22 -04:00
Dennis Alexis Valin Dittrich
18c899def5 enhance(#4209): optional external source reviewer lanes for /gsd:code-review (#4323)
* test(01-01): define reviewer-support trait contract

Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02):
validator rejects non-boolean values with an exact field path, accepts
missing/true/false, and the real code-review capability.json steps
must declare supportsReviewerLanes: true. Add loop-resolver projection
coverage proving the trait reaches activeHooks verbatim for a
provider-neutral synthetic step (not code-review-specific), and that
omitted/false values stay inert (no key on the active hook).

All 8 new assertions fail today: the validator has no such field, and
loop-resolver has nothing to project. RED before GREEN.

* feat(01-01): declare reviewer-capable steps

Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional
boolean opt-in trait, step-scoped (not capability-wide). Only a
literal true validates and projects; false/omitted stay inert (no
key on the projected active hook), and every non-boolean type fails
capability-validator.cjs with an exact field-path error.

Opt both existing code-review steps (execute:post, execute:wave:post)
into the trait in capabilities/code-review/capability.json. Project
the validated field through src/loop-resolver.cts into activeHooks
so a provider-neutral generic interpreter can read it without any
code-review-specific knowledge. Document the field in
docs/reference/capability-manifest.md and regenerate
gsd-core/bin/lib/capability-registry.cjs via the generator (never
hand-edited).

Makes all 8 RED assertions from the prior commit pass.

* test(01-02): define shared reviewer dispatch

- Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes:
  inert when the supportsReviewerLanes trait is off or nothing is selected,
  exactly-once plan/invoke per selected lane, duplicate-alias dedup, the
  bounded metadata-only source-review prompt (repo root, paths+baseSha,
  depth, four fixed prohibitions), and capability-neutral reuse via a
  second synthetic step context.
- RED: module under test (src/reviewer-step-dispatch.cts) does not exist
  yet, so require() fails and every assertion is unreached.

* feat(01-02): dispatch reviewers for opted-in steps

- Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps),
  ONE interpreter for a step's supportsReviewerLanes trait. Reuses
  resolveReviewerSelection for selection and resolveLanePlan for planning
  (both already-existing, pure building blocks); invocation is the one
  required, caller-injected seam (deps.invoke) since runLane needs
  OS-aware spawn plumbing this module does not own.
- trait !== true, or a selection resolving to zero lanes, dispatches
  nothing (zero plan/invoke calls). Each selected lane is planned and
  invoked exactly once, in the selector's deduped/sorted order.
- buildSourceReviewPrompt assembles a metadata-only bounded prompt
  (repo root, canonical paths + base SHA, depth, four fixed
  prohibitions) — never file contents — written once per dispatch and
  shared across every invoked lane.
- GREEN: tests/reviewer-step-dispatch.test.cjs now passes.

* test(01-02): define reviewer dispatch failures

- Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed
  matrix: an explicitly requested lane the selector could not resolve
  still lets the OTHER resolved lane run, but the aggregate result must
  never read as a clean success (and 'every explicit lane unavailable'
  must be distinguishable from the plain no-flags-passed inert case);
  request-level validation (path traversal, absolute paths outside
  repoRoot, empty/non-string paths, missing depth/base SHA) halts the
  whole dispatch before any lane is planned or invoked; a per-lane
  prompt-budget overflow hard-fails only that lane before invoke while
  its sibling still runs.
- RED: src/reviewer-step-dispatch.cts does not yet implement any of
  these guards, so 9 of the new assertions fail against the current
  (Task 1) implementation.

* fix(01-02): fail closed in reviewer dispatch

- src/reviewer-step-dispatch.cts: add the fail-closed guards the prior
  commit deliberately left out. An explicitly requested lane the
  selector could not resolve no longer lets the aggregate read as a
  clean success — lanes that DID resolve still run and keep their
  results (never narrow the requested set), but selection.errors now
  flips the aggregate ok to false, and 'every explicit lane
  unavailable' is now distinguishable (SELECTION_FAILED) from the
  plain no-flags-passed inert case (NO_LANES_SELECTED).
- Add request-level validation (validatePaths, depth/baseSha presence)
  that halts the WHOLE dispatch before any lane is planned or invoked:
  path traversal, absolute paths outside repoRoot, empty/non-string
  paths, and missing provenance are all rejected up front.
- Add per-lane prompt-budget enforcement (resolveBudget, mirroring
  gsd-tools.cjs's budgetFor convention including budget 0 = unbounded):
  a lane whose resolved budget the prompt exceeds hard-fails before
  invoke runs for it, without cancelling a sibling lane already
  planned.
- Document the supportsReviewerLanes trait and its dispatch-step
  interpreter in gsd-core/references/loop-hook-dispatch.md.
- GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass;
  no regressions in the review-lane/reviewer-selection/prompt-budget
  suites (356 passing).

* test(01-03): define optional source reviewer flow

RED: assert code-review.md dispatches roster-derived reviewer-lane flags
through a single review-lane dispatch-step call (DISP-01..05), that the
no-flag path stays byte-for-behavior unchanged (COMP-01), and that
external evidence reaching the internal reviewer prompt is marked
unverified (CONS-02). Also covers the CLI contract directly: no-op with
no explicit selection, and fail-closed on an explicit unknown lane
(SAFE-07) via real gsd-tools.cjs subprocess calls.

* feat(01-03): route optional source reviewers

GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches
canonical reviewer-lane flags against the merged first-party + installed
roster (never a hand-maintained list) and, only when at least one is
present, calls the shared reviewer-step interpreter exactly once with the
already-resolved repo root, file scope, depth, and base SHA. Its evidence
paths are appended to the internal reviewer prompt via
${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane
flag leaves the internal-only dispatch byte-for-behavior unchanged
(COMP-01).

Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane
dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI
route `dispatchReviewerLanes` wires through, but never implemented the
gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add
it to the existing review-lane router, reusing the same effort-aware plan
building and runner deps `plan`/`invoke` already use (factored into
buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard
the CLI's own `detected` set on whether an explicit flag was passed:
resolveReviewerSelection's no-explicit-selection fallback is "select every
detected reviewer" (the correct default for /gsd:review), and passing it
an unconditionally non-empty detected set would silently invoke the whole
roster on every no-flag code review, violating COMP-01.

* test(01-03): define external finding consolidation

RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as
untrusted input — independently re-verifies every claim against the actual
current source, resists a prompt-injection attempt embedded in evidence
text, and folds a verified claim into the existing Narrative Findings
section with no second REVIEW.md schema (CONS-01..03). Also assert
code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* feat(01-03): consolidate external review evidence

GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence>
as untrusted data, independently re-verifies every cited claim against the
actual current source before it can appear in REVIEW.md, and explicitly
resists prompt injection embedded in evidence text (never a command, no
matter what it claims to be). A verified claim folds into the existing
Narrative Findings section with (external: {slug}) provenance — one
REVIEW.md schema only, no separate external-findings section.
code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* fix(01-02): gitignore the reviewer-step-dispatch build artifact

01-02 added src/reviewer-step-dispatch.cts but never added its
npm run build:lib output to .gitignore, unlike every sibling
gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked
noise in git status.

* docs(01-04): publish user and command contract for reviewer-lane source review

- Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md
  and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on
  failure, findings independently consolidated into the single REVIEW.md
- Add the same contract to the docs/features/code-review-pipeline.md
  fragment and regenerate docs/FEATURES.md from it
- Preserve /gsd-review as the plan-review command; cross-reference it
  rather than duplicating the reviewer roster
- Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md
  drift owned by source already shipped in Plans 01-01/01-03 but never
  regenerated (npm run regen:derived had not been run in this worktree)

* docs(01-04): align architecture and agent ownership docs for reviewer-lane trait

- ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes)
  through the shared dispatchReviewerLanes interpreter to the existing
  review-lane plan/invoke machinery, ending at gsd-code-reviewer as the
  sole REVIEW.md consolidator
- AGENTS.md: document gsd-code-reviewer's full-context verification scope
  and its treatment of external reviewer evidence as unverified input
- No new diagram, abstraction, or config key; docs/CONFIGURATION.md is
  unchanged since the feature adds no setting or default

* fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact

Same gap as the earlier .gitignore fix: 01-02 added
src/reviewer-step-dispatch.cts but never added its generated
gsd-core/bin/lib/reviewer-step-dispatch.cjs output to
eslint.config.mjs's ignore list like every sibling generated file,
so tsc's emitted __importDefault CommonJS-interop var tripped
no-var.

* fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md

01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists
cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster
row in docs/INVENTORY.md — required by design, since a role sentence
cannot be generated — was never added.

* fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes

refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook
strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added
supportsReviewerLanes: true to that step and this fixture was not updated.

* chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md

Both files grew as a direct, intended consequence of wiring optional
reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes
step and the untrusted-evidence consolidation contract) — not
incidental drift.

Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209)
Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209)

* test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes

From internal code review: dispatched must be false when zero lanes
actually reached plan(), and a throwing plan()/invoke() for one lane
must not discard results already collected for a sibling lane —
matching the fail-closed pattern gsd-tools.cjs already uses for the
same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794).

Refs: gsd-core-dks.16, gsd-core-dks.17

* fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review

- WR-01: dispatched now tracks whether any lane actually reached
  plan(), not results.length — an unresolvable selected slug no
  longer reports dispatched:true.
- WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a
  throw for one lane can never discard results already collected
  for a sibling lane, matching the same guard gsd-tools.cjs already
  has around the identical resolveLanePlan call.
- IN-01: documents the intentional budget===0-is-unbounded
  convention (#2797) the caller already relies on.
- IN-02: review-lane dispatch-step no longer blocks indefinitely on
  an un-piped interactive TTY; fails closed to empty paths instead.

Refs: gsd-core-dks.16, gsd-core-dks.17

* docs(01-05): add changeset fragment for PR #17

* fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract

agents/gsd-code-reviewer.md's untrusted-evidence section and its
pinning regression test both quote injection phrases as the exact
attack they defend against/detect — same
DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing
allowlist entries, not an actual injection vector.

* test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip

From CodeRabbit review: WR-02's earlier fix only wrapped plan() —
writePromptFile()/deps.invoke() still ran unguarded, so a throw
there still aborted every later selected lane. Also covers the
dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit
lanes silently not running when no prior review and no phase-start
commit exist).

* fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved

Previously an explicit reviewer-lane request with no prior review and
no resolvable phase-start commit reached dispatch-step with an empty
--base-sha, which fails closed via missing_provenance — correct, but
silent about why explicitly requested lanes didn't run. Now skip
dispatch entirely in that case with a stderr warning naming the
actual cause.

* fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan()

WR-02's original fix only guarded plan() — a throw from
writePromptFile() or deps.invoke() still aborted the whole dispatch,
discarding results already collected for lanes processed earlier in
the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it.

* fix(01-05): WR-02b mock must throw only on the first writePromptFile() call

The committed mock threw unconditionally, so codex's retry also threw and
failed for the same reason as claude's — the test could not distinguish
'sibling still runs' from 'sibling also breaks'. Gate the throw to the
first call, matching WR-02/WR-02c's single-failure intent.

* fix(#4209): close review findings from adversarial + critical-code-reviewer pass

Two independent reviews (agy adversarial review, Opus critical-code-reviewer +
ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch
wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and
fixed here:

- dispatch-step's reducer silently swallowed whole-dispatch rejections
  (invalid paths, missing provenance, etc); it now checks parsed.ok/reason.
- spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the
  LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review;
  now shares the single compute_file_scope derivation.
- the external reviewer prompt had no actual review request or citation
  requirement, only prohibitions; added both.
- removed the supportsReviewerLanes trait plumbing (capability registry,
  validator, loop-resolver, docs, tests) — it was never consulted by the
  real dispatch path, which gates on explicit CLI flags instead.
- flag-resolution require() was a fragile cwd-relative literal that failed
  silently on non-vendored installs; now resolves via GSD_TOOLS's own
  directory and warns instead of swallowing failure.
- reducer didn't unwrap the @file: overflow protocol for large payloads.
- deduplicated resolveBudget/budgetFor into one resolveLaneBudget.
- lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a
  second dispatch can't overwrite prior evidence.
- validatePaths rejects control characters, closing a markdown-injection
  vector into the external prompt via crafted filenames.
- reworded the one line that tripped prompt-injection-scan.sh instead of
  allowlisting the whole production prompt file.
- fixed a stale docstring range and a dispatched-field ordering bug.
- added 3 integration tests executing the actual reducer against synthetic
  dispatch-step JSON, replacing markdown-substring-only assertions.

771/771 tests pass across every touched suite; tsc --noEmit clean.

* fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait

The maintainer's approval on issue #4209 explicitly redirected implementation
shape: reviewer-lane dispatch must be a reusable capability/step-dispatch
trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call
itself. My previous commit (e2558326) deleted that trait entirely after
finding it declared-but-never-consulted, which was backwards — the fix was to
wire it, not remove it.

Restores the trait (capability.json, generated registry, validator,
loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes
now resolves its own active hook via `gsd_run loop render-hooks` for the
configured workflow.code_review_point and only proceeds to CLI-flag matching
when supportsReviewerLanes reads true. Explicit flags no longer bypass the
trait; a matching flag with the trait false resolves zero slugs (proven by a
new integration test executing the real fence with both trait states).

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the
dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer
redirect requires the capability layer, not the workflow, own the opt-in
decision).

* fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point

Both an agy adversarial review and an Opus critical-code-reviewer pass
independently found the same gap in my previous commit (9b2c3773d): the trait
check I wired into code-review.md only protected code-review's OWN
invocation — gsd-tools.cjs's dispatch-step handler still hardcoded
`trait: true` unconditionally, so a second capability declaring
supportsReviewerLanes would get zero enforcement from the shared CLI unless
it correctly re-implemented the ~15-line render-hooks scrape itself. That is
exactly the "each workflow.md hand-wiring the call" the maintainer's redirect
said to eliminate.

Moves the trait check into dispatch-step itself: given --cap-id/--point, it
self-invokes `loop render-hooks <point>` (relocating the one subprocess
code-review.md used to spawn for this, not adding a new one) and derives the
real trait from that capId's active hook, rather than trusting a
caller-passed boolean. code-review.md now only passes
--cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or
gates on the trait itself — the ~20-line scrape it previously carried is
gone. Any other capability opts into the identical enforcement by declaring
the trait and passing the same two flags.

Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input
variable (they proved a bash branch honors a variable, not that the variable
reflects the real capability manifest) with three integration tests that
invoke the real dispatch-step CLI against the real first-party capability
registry: the real code-review trait resolves true, an unknown --cap-id
resolves false (trait_not_enabled, fail-closed), and omitting
--cap-id/--point entirely resolves false (no context means no opt-in).

Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check
(agy-F1 was incomplete), and delete the promptWritten per-lane coupling
flag — the prompt write is idempotent, so writing it once per lane instead
of gating on "did any lane write it yet" removes a latent bug where a
deps.plan override that ever varies promptPath per lane would silently skip
writing for a later lane.

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count
drops (the trait scrape moved into dispatch-step), but the file still grew
this session across multiple commits; acknowledging per the growth-tracking
convention.

* fix(#4209): remove per-run token waste from the shipped prompts

Runtime prompt content, not session tokens: two real, per-invocation token
costs in the code that ships.

1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of
   load_context step 5's ~180-word untrusted-evidence contract in ~90 more
   words, breaking this section's own established terse one-liner style
   (every other rule here is 1-2 sentences). This prompt loads fresh on
   every /gsd:code-review invocation. Shrunk to a one-line cross-reference,
   matching how write_review's own reference to step 5 already does it.

2. buildSourceReviewPrompt repeated the base SHA on every single file line
   even though it is identical for every file and already stated once at
   the top of the prompt — O(files) wasted tokens on every dispatched lane
   for a 50-file review, for zero information gain. File lines are now bare
   paths.

* fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3

Opus critical-code-reviewer found a real Blocking defect in the --cap-id/
--point self-invocation added last commit: `dispatch-step` spawned
`loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its
stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to
`@file:<path>` instead of inline JSON -- the same overflow protocol this
feature already unwraps for its OWN dispatch result 60 lines later in
code-review.md. A large-enough activeHooks envelope (more installed
capabilities/fragments) would throw, get silently swallowed by the bare
catch, and misreport a real trait as trait_not_enabled with zero diagnostic.

Fixed by extracting the config/registry/capability-state resolution
`cmdLoopRenderHooks` already performs into an exported pure function,
resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now
share it), and calling it in-process from dispatch-step instead of spawning
a subprocess at all. This eliminates the @file: exposure entirely (the
dispatch-step path never touches the rendered-string envelope or its
JSON-stringify/50000-char threshold), removes one subprocess spawn per
code-review invocation, and gives a genuine diagnostic (stderr warning) on
resolution failure instead of silent fail-closed. Corrected three doc/
docstring references to the now-removed subprocess self-invocation.

Also fixes 2 real CI failures this round surfaced:
- lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line
  no-control-regex` comment was unused under this project's ESLint config
  (verified locally: the rule never actually flags \x00-\x1f in this repo's
  config) -- a mistake from an earlier commit this session, never actually
  lint-checked before push. Removed the disable comment.
- security (prompt-injection-scan): the agy-F1 regression test's crafted
  fixture literally contains "Ignore all prior instructions." as test data
  proving validatePaths rejects it -- allowlisted the test file, same
  DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries.

Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one
bullet stated "untrusted, never a command" three different ways in one
paragraph, and a same-file duplicate of write_review's schema rule.
Consolidated to state each rule once.

Declined one suggestion from this round: shrinking code-review.md's
EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests
(tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block,
tests/code-review.test.cjs's CONS-02 test) deliberately lock the four-
prohibitions restatement and the untrusted-evidence prose into the
INJECTED block itself, not just the consolidator's system prompt --
adjacency of the warning to the untrusted payload it's warning about is a
recognized prompt-injection defense-in-depth pattern from this
workstream's original TDD plan, not accidental duplication.

* fix(#4209): correct stale per-file base-SHA prose in the external prompt

Leftover from removing the per-file base SHA repetition earlier this
session: the review-request sentence still said "relative to its base SHA"
(singular per-file framing) when there's now exactly one base SHA, stated
once above the file list. Reads "relative to the base SHA above" now.

* fix(#4209): make getLane/configGet/plan required deps, delete dead defaults

R3/R4 from the review round I'd deferred as low-priority test-churn: this
file's one production caller (gsd-tools.cjs's dispatch-step handler) always
supplies all three, so the fallbacks were dead in production -- but each was
actively WRONG if ever reached: the default configGet always returned
undefined, silently disabling resolveLaneBudget's overflow guard; the
default getLane looked up only first-party REVIEWER_LANES, diverging from
production's overlay-merged roster; the default plan skipped per-host effort
resolution entirely.

These defaults were introduced by this PR's own earlier work (this file did
not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from
elsewhere, so there's no external caller depending on the lenient contract.

Turned out free to fix: making the three deps required and deleting
defaultGetLane/defaultPlan needed zero test changes -- every existing test
that actually reaches the per-lane loop already supplies getLane/plan
explicitly, and configGet's only real dependent (the budget-overflow tests)
already supplies it too. 788/788 tests pass unchanged, tsc/lint clean.

* fix(#4209): define depth semantics for the external reviewer lane

Verified this was a real bug, not a match to existing convention as I'd
claimed when declining the suggestion earlier this session: the internal
gsd-code-reviewer agent's own system prompt carries a full <depth_levels>
block defining what quick/standard/deep mean and do (agents/gsd-code-
reviewer.md:68-99). The external reviewer lane has no access to that
persona at all -- it only ever sees buildSourceReviewPrompt's bounded text,
which sent the bare depth label with zero definition to a third-party CLI
with no other source of truth for what "standard" means.

Added depthMeaning(), condensed from the internal reviewer's own
<depth_levels> definitions so the two stay consistent, and interpolated it
into the review-request sentence. 150/150 tests pass, tsc/lint clean.

* fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation

CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the
roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a
SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide
whether to dispatch at all. This file's own documented rule (its
depth-resolution guard, stated explicitly a few hundred lines earlier) is
that a guard and the extraction it protects must run as one shell
control-flow decision, because markdown-fenced blocks do not share shell
state -- this step violated its own file's rule for the entire feature's
gating condition.

Merged the roster-resolution fence and the dispatch-decision fence into one
continuous bash block, removing the intervening prose that split them.
Fixed the stderr-based failure detection in the same edit (RQ-01: checking
whether stderr is non-empty misfires on any benign Node warning; now checks
the actual exit status of the roster-resolution command).

Verified by extracting the merged fence and executing it standalone, driving
both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a
real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty,
SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests
pass, tsc/lint clean.

* fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write

Batch of Required/Suggestion fixes from the Opus critical-code-reviewer +
writing-for-agents pass:

- CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch
  blocks, commented-out code) and deep (error propagation, state mutation
  consistency, circular dependencies) relative to the real <depth_levels>
  block, and had zero test coverage. Restored full accuracy and added tests
  that read the real agents/gsd-code-reviewer.md file directly, so drift
  between the two can't recur silently. Unrecognised depth now normalizes to
  standard's definition, matching that agent's own documented rule, instead
  of rendering an undefined bare label.

- RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt
  `paths` does, but weren't checked for control characters like paths were
  (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and
  applied it to all four fields at the same provenance-check boundary.
  runDir previously had zero validation at all.

- S1: deleted the dead `identity` parameter on `invoke` -- the one production
  caller already ignores it, no test read it by name.

- S2: hoisted the shared prompt write above the per-lane loop -- promptPath
  is derived from runDir alone (constant across lanes by construction), so
  writing it once is both correct and cheaper than the per-lane write R1
  introduced earlier this session. Discovered and fixed a real regression
  from the naive version of this hoist: an unguarded throw would have
  escaped dispatchReviewerLanes as an uncaught exception instead of a clean
  per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason,
  matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with
  a dedicated regression test.

- S3: moved `planned = true` past the budget-overflow gate, so `dispatched`
  only reports true once a lane has cleared BOTH plan and budget checks.

- S5: relayed gsd-code-reviewer.md's own "performance issues are out of
  scope unless also correctness issues" policy into the external-lane
  prompt, which previously had no such guidance and could return findings
  the internal reviewer's own contract excludes.

- RQ-05 (partial): shrunk this file's own header docstring's restatement of
  the trait-reuse architecture to a pointer at
  gsd-core/references/loop-hook-dispatch.md, the canonical home.

234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion

RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the
SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke`
already share. code-review.md's ~18-line inline `node -e` reimplementing
`loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in
gsd-tools.cjs) is now a single call to this subcommand -- the exact
violation code-review-flags.cjs's own header warns against ("this is the
canonical flag-parsing surface -- do not replicate inline bash parsing").

RQ-03: an empty --cap-id XOR --point now warns distinctly from the
legitimate no-context opt-out (both absent) -- a caller that named a
capability without its point was silently indistinguishable from a correct
opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only
ever fires when the config-get COMMAND ITSELF fails (config-get already
resolves the manifest's own schema default in the normal case), but that
failure was previously silent.

RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait
resolved inside dispatch-step" explanation was restated in full in 5
places across this session's own review cycles. Consolidated to ONE
canonical statement in gsd-core/references/loop-hook-dispatch.md; the other
4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md,
code-review.md's step-opening comment) now point at it instead.

W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two
inert cases when capability-validator.cjs already rejects non-boolean at
load -- restated as the two cases that actually reach this code. Removed a
"do not hand-roll trait resolution" prohibition whose target no longer
exists once the positive description precedes it.

W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing
block means proceed as normal") -- an absent optional block already means
proceed as normal without being told.

W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the
made-up compound "byte-for-behavior [un]changed" with the token this
session's own docs already coined for this concept (inert) and the word
that means what byte-for-behavior was reaching for (unchanged).

W-10: dispatch_reviewer_lanes had no completion criterion -- added one
sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set,
either populated or empty). This exact sentence would have caught the
cross-fence bug fixed two commits ago at authoring time.

Declined from this round, with reasoning: W-02/W-03 (trim the
untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) --
two tests deliberately lock this as intentional adjacency-based
prompt-injection defense-in-depth, not accidental duplication (see this
branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site
`trap ... EXIT`) -- would fire at the end of the CREATING fence, before
spawn_reviewer's agent ever reads the evidence files, given this file's own
documented fenced-block execution model; the existing named cross-reference
between creation and cleanup already satisfies the co-location concern
without introducing that regression.

853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex

Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed
for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get
fallback lived in an earlier, separate fence from the fence that consumes it
via --point, split only by prose (not a guard, per this step's own documented
rule). Merged into the single continuous fence and added a structural test
asserting exactly one bash fence in the step.

The new end-to-end regression test for this used --codex, which drives the
fence's real `review-lane dispatch-step` call and, with the codex binary
present on PATH, spawns the real external CLI — which then blocks on
interactive auth with no stdin (BL-01). Stubbed gsd_run for
`review-lane dispatch-step` only (captures argv instead of executing),
keeping the real config-get/explicit-from-argv calls the test is actually
about.

* fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment

Round-5 review (Opus) warning-tier findings:

- WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but
  a control-character injection attempt" — a caller distinguishing a config
  problem from a security event couldn't tell them apart. Split into
  MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid).
- WR-05: validatePaths' containment check was lexical only (path.resolve),
  so a symlink whose own path sits inside repoRoot could still point outside
  it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can
  legitimately name a file already deleted in a stale worktree), realpathing
  repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't
  false-positive-reject its own real children.
- WR-08: a comment in the per-lane loop still said a throwing writePromptFile()
  was caught there — stale since the prompt write was hoisted above the loop
  in an earlier round.

WR-03 (validate depth against the quick/standard/deep enum) was considered
and declined: this dispatcher is deliberately capability-neutral (see the
existing "synthetic step context" test, which passes a non-code-review depth
label on purpose to prove no code-review-specific special-casing exists).
WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and
WR-07 (reason omitted on the aggregate return) were verified against source
and are not bugs — see review notes.

* docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap

Round-5 review (Opus, BL-03) flagged that an early exit between
dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A
trap-based cleanup was considered and rejected: if a step genuinely runs as
a separate process, a trap set at creation time would fire at the end of
that SAME fence, deleting the directory before spawn_reviewer/commit_review
ever read it — worse than the leak it would fix.

review.md's own gather_context/cleanup pair for the identical resource class
(a run-scoped reviewer temp dir) already makes and documents this exact
trade-off: cleanup runs only on a documented success path, and a leftover
$TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording
that precedent here so this isn't re-raised as a live gap in a future review.

* fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path

reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes
paths: ['docs/spec.md'] as a synthetic, never-read path proving the
dispatcher has no code-review-specific special-casing. lint-docs-guard-
registration correctly flagged this as an unregistered docs/ path reference —
add the docs-guard-exempt marker and its pinned baseline entry, the same
pattern every other synthetic docs/ literal in this test suite already uses.

* fix(#4209): backfill changeset pr: field with the real upstream PR number

changeset-lint's fail_pr_field_drift caught the fragment still pointing at
the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this
branch is now also open against.

* docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam

trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every
decision in ADR-2782 (D1-D9) and every prior dated amendment governs the
`role: "reviewer"` capability body and its one consumer, /gsd:review. This
PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary
feature capability's `steps[]` entry, projected through loop-resolver.cts
and resolved in-process via resolveActiveHooksForPoint - is a different
capability axis (steps/gates/contributions) that the ADR's own scope note
explicitly places out of reach. Per docs/contributor-standards.md's
"Amending an accepted ADR", an in-place dated section is the established,
lighter-weight path for an addition that stays within the ADR's existing
decisions - used twice already in this same file - so this appends a third
dated entry documenting the new seam, its consumer, and why it reuses the
existing D1-D9-governed plan/invoke machinery rather than adding a second
one. No decision is reversed; no new Amends/Amended-by pair is needed since
the steps/gates/contributions axis already carries reciprocal links to
ADR-857 and ADR-894.

* fix(#4209): close two test-quality gaps trek-e's review found

Minor 1: validatePaths (a path-shape parser guarding the prompt-
injection/path-traversal trust boundary) had only example-based coverage,
violating ADR-456's rule that parsers/budget limits carry at least one
fast-check property test. Adds three: safe-segment paths are never
rejected, a single leading "../" always escapes the one-segment repoRoot,
and a control character anywhere is always rejected - one property per
rejection reason validatePaths owns.

Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only
ever exercised far below budget or at budget:0 (unbounded), never at the
exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds
three exact-boundary tests using the real estimateTokens/
buildSourceReviewPrompt the module calls internally, so the resolved
token count is exact rather than approximated: budget == estimate (must
pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must
pass).

Also extracts okPlan()'s fixture timeoutMs into a named constant -
local/no-adhoc-timeout-literal (#4446) landed on next after this branch
was authored and flagged the pre-existing literal on rebase; it is fixture
data for a synthetic plan object dispatchReviewerLanes never waits on, a
distinct class from tests/helpers/timeouts.cjs's real subprocess norms.

* fix(#4209): update docs-guard-registration baseline for the new ADR citation

reviewer-step-dispatch.test.cjs's new fast-check property tests cite
docs/adr/456-test-rigor-architecture.md in a justifying comment (never a
real read). lint-docs-guard-registration fingerprints every docs/ path
string an exempted test file mentions and fails on drift so a human
re-confirms the exemption still holds - re-confirmed, and the baseline is
updated to match.

* fix(#4209): point changeset pr: field at the fork PR for CI validation

changeset-lint's fail_pr_field_drift check compares the fragment's pr:
field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH),
not a fixed target. Rehearsing this branch on fork PR
davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's
pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but
fails here. Backfill to 4323 happens again, as the last commit, immediately
before the approved push to open-gsd#4323 - never leaving pr: 17 on the
branch that ships upstream.

* fix(#4209): reject promptChannel:none lanes from source-review dispatch

CodeRabbit found a real scope mismatch: coderabbit's lane declares
promptChannel: 'none' and reviews the working tree on its own terms,
fed nothing (review.md:367). Silently dispatching it through
dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope
buildSourceReviewPrompt promises and let the lane review whatever it
independently sees fit, violating this interpreter's own scoped,
metadata-only contract. Reject before plan()/invoke(), same as an
unresolved slug.

* fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file

CodeRabbit found the whole-file match on workflowContent would still
pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts
of this 1000+-line workflow, proving nothing about the actual evidence
block's contract. Line-filtered via splitLines (not a bare-\n regex
spanning readFileSync content) so this stays CRLF-portable and passes
local/no-unbounded-quantifier and local/no-crlf-fragile-split.

* fix(#4209): guard DISPATCH_JSON substitution and capture its stderr

CodeRabbit found the dispatch-step command substitution unguarded: a
non-zero exit could leave DISPATCH_JSON empty (or halt the step under
errexit with no warning), and the downstream reducer would only ever
report the generic unparseable_dispatch_output reason, discarding the
command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/
EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface
it in a warning on failure, and fall back to a parseable dispatch_
command_failed JSON stub so the reducer's existing reason-reporting
path still fires.

* docs(#4209): fix byte-for-behavior wording and missing colon, regenerate

CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the
established repo term for output-identical unchanged behavior) and a
missing colon after the bold "Optional external reviewer lanes (#4209)"
lead-in in docs/features/code-review-pipeline.md. Fixed in the two
hand-authored sources (commands/gsd/code-review.md, docs/features/
code-review-pipeline.md) and regenerated the two derived projections
(skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/
FEATURES.md via gen-features.cjs) so they stay in sync.

* fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI)

The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds
double-quoted JSON keys inside a single-quoted shell literal. That
extra quote density, inside an already quote-heavy ~8KB driver string,
passed bash -n and the full local suite on Linux but broke Windows
Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end
to end (#4209 round 5)` failed on two Windows CI shards with `bash -c:
unexpected EOF while looking for matching '''` — a Windows argv-to-
command-line re-quoting edge case, reproducible on rerun, not a flake.
Root-caused via gh api job logs plus a byte-identical local
reconstruction of the test's own driver script.

Fix: drop the fabricated stub. The downstream node -e reducer already
falls back to reason `unparseable_dispatch_output` on any JSON.parse
failure, so an empty/partial DISPATCH_JSON on command failure is still
handled correctly, with zero new quoting risk.

* revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI)

Two materially different mechanisms for the same CodeRabbit Nitpick
("Trivial | Quick win") both broke Windows Git-Bash reproducibly:
a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ...
matching '''") and, after removing that, a plain `head -1 "$VAR"`
inside a nested command substitution ("unexpected EOF ... matching
'"'"). Both passed bash -n and the full local suite on Linux every
time; both failed the SAME test deterministically on Windows CI. Two
attempts at the same class of fix (nested-quote construction near
this exact step) is the retry limit - reverting to the original,
already-shipped, Windows-verified unguarded form rather than
continuing to guess at a third quoting mechanism for a Trivial-
severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for
anyone attempting this again: the fix belongs outside this specific
markdown-fence-driver test harness (e.g., a real .sh helper script)
if it's worth doing at all.

* fix(#4209): backfill changeset pr: field to the real upstream PR before push

Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy
changeset-lint's PR-number check while rehearsing there; this is the
last commit before the approved push to the real upstream PR
(open-gsd/gsd-core#4323), so the field points at that PR number again.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 22:52:33 -04:00
Michel Moreira
733bec3ad1 enhance(#4261): report size-cap headroom on every run, with a reserved margin (#4418)
* enhance(#4261): report size-cap headroom on every run, with a reserved margin

The tier caps are red lines and none of them moves here. What was missing is
everything below the red line: a passing run said nothing, so a contributor
at 99.6% of a cap and one at 60% got identical feedback, and the density that
produces merge-time collisions was invisible to the people creating it.

Two levels, matching the shape execute-phase.md already carries by hand (a
hard ceiling plus a lower margin "so minor future edits don't re-trip the
gate") and which was until now the only capped file with one:

  1. a headroom census printed every run, green included, sorted
     least-headroom-first, and appended to the GitHub job summary
  2. a 95% reserved margin that names the files inside it and REPORTS
     rather than fails

The margin deliberately does not fail. A cap breach is a red line; a file at
96% is not broken, it is a file whose next contributor should extract before
adding. Failing there would create a second red line and force exactly the
+N bumps the policy forbids. Neither level asserts a count, so this adds no
snapshot to regenerate — the per-file size baseline was deleted by #2724 for
conflicting on 7 of 7 PRs that touched it.

Also deletes rather than refreshes the per-tier high-water comments in both
guard files. They were measured once and then diverged from the tree: the
LARGE line still claimed "gsd-executor 42,342 -> ~6.8 KB headroom" while the
real high-water sat at 99.6% of that cap, so the comment documenting the
margin was itself why nobody noticed the margin was gone.

As measured on next by the new census, the pressure has grown since the
issue was filed: gsd-plan-checker.md has 9 bytes of headroom, gsd-verifier.md
21, and plan-phase.md 14.

* chore(#4261): add changeset for the size-cap headroom census

* test(#4261): exercise reserved-margin boundaries

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 22:04:48 -04:00
Tom Boucher
93e141a006 enhance(#4139): Phase 4 — measure the window instead of asserting it (#4502)
* enhance(#4404): add offline token benchmark for compact-content splits

ADR-4139 Decision 2 requires the finite-attention justification for
workflow.compact_content to be measured, not asserted. `npm run
benchmark:compact-content` computes, per registered spine/detail split
discovered under gsd-core/workflows/, the token count with the split
active (spine alone) vs inactive (spine + all detail parts read back
in), using gpt-tokenizer (pinned exact devDependency — Anthropic
publishes no tokenizer for Claude 3+, so every output surface labels
this a PROXY-TOKENIZER comparison: the on/off delta is exact under one
tokenizer applied identically to both sides, the absolute counts are
not Claude's real ones).

Reporting-only by design and verified so: --check diffs the live
recompute against a committed baseline (tests/fixtures/compact-content-benchmark-baseline.json)
and prints drift, but never exits non-zero for a drifted or missing
baseline — the only thing allowed to fail this script is a genuine I/O
error reading a source .md file it's measuring. Not wired into lint:ci
or pretest.

Discovery is deliberately reimplemented rather than importing
tests/helpers/compact-content-split.cjs (Phase 3, #4403), keeping a
scripts/ reporting tool from depending on a test-only module.

tests/fixtures/deny-network.cjs preloads via NODE_OPTIONS=--require to
prove the benchmark makes no network call, monkeypatching http/https/
net/dns/fetch to throw rather than relying on sandboxing.

docs/CONFIGURATION.md documents the new benchmark against the
workflow.compact_content key to satisfy this repo's docs-required gate
for an Added-type changeset.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4404): address orthogonal review findings on the token benchmark

Standards axis found two hard violations against documented rules:
- CLAUDE.md's Generative Fix Divergence rule requires a parity assertion
  for shared discovery logic maintained in two places. Added a test
  comparing benchmark-compact-content.cjs's own discoverRegisteredSplits
  against tests/helpers/compact-content-split.cjs's version on the real
  repo tree, so the two can never silently drift apart.
- The changeset body closed its bold span with a period and continued
  as a second sentence, instead of the canonical
  `**<phrase>** — <explanation>.` shape CONTRIBUTING.md documents.

Spec axis found the "network disabled + identical output across two
runs" Done-when criterion was verified as two separate properties
(determinism tested without network denial, offline survival tested as
a single run) rather than as one combined property. Added a test that
runs the benchmark twice under the deny-network preload and asserts
byte-identical stdout.

Security axis found tests/fixtures/deny-network.cjs didn't patch
dns.promises (a separate binding from the callback dns API), tls.connect,
or http2.connect — inert today since nothing in the benchmark calls
them, but a silent gap in what the preload's own header claims to
guarantee. Patched all three.

CLAUDE.md's Property-Based Testing rule also requires a fast-check test
for budget-limit arithmetic; added one for computeAggregate's off/on
summation (true sum over N splits, never NaN/Infinity, never exceeds
100% when off >= on for every split).

Standards axis's remaining two findings (a Data Clumps observation on
the {offTokens, onTokens, reductionPct} triple, and mild duplication in
formatDriftReport's three line-formatters) are left as judgement calls:
introducing a named type for a 3-field local tuple, or a formatter
abstraction for three short lines, would be exactly the premature
abstraction CLAUDE.md's engineering guidance warns against for a script
this size.

All changes verified directly (parity logic, the fast-check property,
and the three newly-denied network surfaces actually throwing under the
preload) via node -e before committing; full npm run lint:ci passes
with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4404): backfill changeset pr number to 4502

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 19:19:03 -04:00
Tom Boucher
a0f8f956c4 enhance(#4139): Phase 3 — partition rules + the five checks (#4497)
* enhance(#4139): Phase 3 — partition rules + the five checks

ADR-4139 Decision 5, epic #4139 Phase 3. Issue #4403's own "Proposed behavior"
section lists four checks; the ADR's Decision 5 and its own phase table ("partition
rules + the five checks") list five — the same four plus "boundary moves are
declared, ongoing". Same issue-vs-ADR drift Phase 2 hit on the detail.md vs
detail/*.md layout: the ADR is the locked, reviewed document, so it wins. This PR
implements all five.

docs/PARTITION-RULES.md (new) is the partition-rules document: the partition rule
itself, the protected-content list and <!-- gsd:protected --> sentinel syntax
(relocated unchanged from gsd-core/references/compact-content-protected-content.md,
now deleted — it was never referenced by any runtime workflow Read, only by the
predecessor test as documentation, so nothing at runtime regresses, and removing it
from gsd-core/references/ also drops it from all 19 installed-project shipped-content
trees for a file nothing ever read), and the five checks explained for a human
reader. Referenced from a new CONTRIBUTING.md subsection under "Editing shipped
content".

tests/helpers/compact-content-split.cjs (new) is the shared mechanics: split
discovery (any gsd-core/workflows/<name>/detail/*.md paired with <name>.md — no
registry file, a pair is registered by existing on disk), line normalization
(carries forward Phase 2's bare-label-line isTrivial fix and the canonical
gsd_run-launcher-preamble exclusion), sentinel extraction, and a
Boundary-Move-Declared commit-trailer reader that is a direct structural port of
tests/helpers/emitted-runtime.cjs's Emitted-Drift-Ack-Hash/-Growth trailer reader
(ADR-3942) — same merge-base range, same fail-closed throw on an uncomputable range,
same dedupe/conflict rules.

tests/compact-content-partition-guard.test.cjs (new) is the actual guard, superseding
tests/plan-phase-compact-split.test.cjs (deleted — its per-pair checks are now the
general guard's job for plan-phase specifically). Checks 2 (disjointness) and 3
(registration + size cap) run unconditionally against every registered split. Checks
1 (completeness, fires once per split on the PR that introduces a new detail/ path),
4 (protected content — no trailer can ever excuse this one, unlike check 5) and 5
(boundary moves declared) are PR-diff-scoped against the resolved base ref and skip
cleanly when there's nothing to compare (a fresh clone, no PR in flight) — a
deliberate asymmetry from check 5's trailer reader, which must throw rather than
silently pass when ITS range is uncomputable, since that function is answering "did
this PR declare its moves" rather than "is there even a diff to look at". Each of the
five checks carries a RED (deliberately broken fixture) / GREEN (fixed) test pair,
built against synthetic temp files or real throwaway git repos, per this repo's rule
that a guard nobody has seen go red is not yet a guard. Building the real fixtures
caught and fixed one real bug before it shipped: check 4's line-presence test was
using the trivial-line-filtered normalizer, so a byte-identical spine falsely
reported its own protected code-fence line as "deleted" — fixed with a
non-filtering membership check.

Extending docs/INVENTORY.md's "Workflow Sub-Files" table for `detail` surfaced a
pre-existing, unrelated gap in the SAME area: gsd-core/workflows/<name>/templates/*.md
is a fourth workflow sub-file kind that already existed on disk and was already known
to lint-response-language-coverage.cjs's FRAGMENT_DIRS, but was invisible to
gen-inventory-manifest.cjs and undocumented in that table. Fixed alongside it, same
pattern, same PR, rather than deferred.

Also, mechanically required by the new fourth sub-file kind:
- scripts/lint-response-language-coverage.cjs: `detail` added to FRAGMENT_DIRS
  alongside modes/steps/templates — a detail/<part>.md inherits its parent's
  response_language coverage through the same per-file proof, not a parallel one.
- tests/workflow-size-budget.test.cjs: explicit regression test locking that
  detail/ files are governed solely by the hard, non-waivable NEW_FILE_CAP
  (tests/helpers/emitted-diff.cjs) and never by the XL/LARGE/DEFAULT spine tiers —
  true by construction (measureWorkflows/listWorkflowStems don't recurse), made
  explicit per the issue's own Done-when item rather than left true-by-omission.
- scripts/gen-inventory-manifest.cjs: `workflow_detail` and `workflow_templates`
  NESTED_FAMILIES entries; docs/INVENTORY-MANIFEST.json regenerated
  (plan-phase/detail/elaboration.md, discuss-phase/templates/*.md now tracked);
  docs/INVENTORY.md's table updated to four kinds.

Verified: `npm run lint:ci` clean with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4403): review findings + a real gsd-test failure in the new guard

Two orthogonal review passes (Standards + Spec, isolated sub-agents) plus a
separate security review ran against the prior commit. Fixed everything each
surfaced:

- Security (Low, path-traversal existence oracle): checkRegistration's
  dangling-reference check extracted detail-path-shaped substrings from spine
  PROSE via a regex that permits `.`/`/` freely, then joined them onto repoRoot
  and probed fs.existsSync with no containment check — a spine file containing
  `../../../etc/detail/passwd.md`-shaped text could make the guard test file
  existence outside the repo. Added a path.relative-based containment check
  before the fs.existsSync call; anything that resolves outside repoRoot is now
  reported as a dangling reference directly, never probed on disk.
- Standards (Boundary Coverage): the size-cap fixtures covered NEW_FILE_CAP and
  NEW_FILE_CAP-1 but not NEW_FILE_CAP+1 — added the third boundary-point case
  CLAUDE.md's TEST RULES require (limit-1/limit/limit+1).
- Standards (Property-Based Testing): extractProtectedBlocks (a sentinel
  parser) and the new parseBoundaryMoveTrailerValues (a declare/dedupe/conflict
  parser, bijective-shaped) had no fast-check property test. Added three: a
  render/parse bijectivity property for the trailer parser (mirroring the exact
  ADR-3942 sibling test's alphabet/idiom), a dedupe-is-idempotent property for
  the same parser, and a well-formed-sentinel-round-trips property for
  extractProtectedBlocks.

Then dispatched gsd-test on the resulting commit. It found a real bug the
reviews couldn't have caught (none of them can run inside gsd-test's sandbox):
checks 4/5's real-repo assertion failed against plan-phase's own split,
reporting DISK_PLANS/#3218-comment lines as "undeclared boundary moves" —
content Phase 2 (#4402) legitimately moved into detail/elaboration.md months
before this PR's Boundary-Move-Declared mechanism existed to require a
trailer for it. Root cause: `resolveBase()`'s own doc comment already documents
that no `origin/*` remote-tracking ref exists inside the gsd-test sandbox
container, and its fallback candidate (a bare `next` branch) can resolve to a
point in history that predates an already-merged, already-reviewed split —
making that split look "newly introduced" from the sandbox's vantage point.
Check 1 (completeness) already scopes itself correctly to only genuinely-new
detail paths (git diff status 'A'); checks 4 and 5 did not share that scoping,
so a stale base made them re-litigate a settled split retroactively. Fixed by
having checks 4/5 skip any split name check 1 already counted as newly-split —
their own premise ("did an EXISTING split shed/undeclare something") does not
apply to a split that is, from the resolved base's vantage point, brand new;
that is check 1's domain alone. Verified locally (25/25 tests pass via a
direct `node -e` require, since `node --test` is blocked in this repo) and via
re-reasoning through the exact real-repo scenario the gsd-test failure showed.

Also regenerated all 19 tests/fixtures/install-tree/*.json goldens — the
prior commit's deletion of gsd-core/references/compact-content-protected-content.md
was never reflected there, which is what golden-install-tree.test.cjs's other
19 failures in the same gsd-test run were.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4403): backfill changeset pr number to 4497

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4403): isolate codex-config.test.cjs into its own chunk, root-causing the Windows CI failure

PR #4497's "full test (windows-latest, 24, shard 2/3)" job failed: run-tests
killed chunk 3/8 at the 600s per-chunk backstop, with codex-config.test.cjs
(weight 17.87, by far the chunk's dominant cost) packed alongside 39 other
files. Traced, not assumed:

- scripts/run-tests.cjs's own timeout-headroom comment for the OUTER
  per-shard timeout documents that "adding one test file reshuffled 115 of
  268 unit files between shards" — shard/chunk composition is architecturally
  known to be unstable to single-file additions, which is exactly what this
  PR's own new tests/compact-content-partition-guard.test.cjs is.
- A second comment, dated 2026-09-06 (one day before this PR, PR #4428's own
  CI), already documents the SAME chunk hitting the SAME 600s backstop with
  the SAME file (codex-config.test.cjs, "a genuinely MEASURED weight of
  17.87 — not a stale-table miss") dominating it — the fix then was cutting
  the Windows per-chunk budget from 60 to 40. That cut clearly was not
  enough: two documented incidents in two days, at two different budget
  settings, both centered on one file that alone consumes ~45% of even the
  reduced Windows budget.
- tests/test-timings.json's own header confirms its source data
  (test-events-linux-node22/24.jsonl) is Linux-only, and run-tests.cjs's own
  chunk-timeout diagnostic already prints "real Windows cost runs ~2.2x the
  recorded figure" — the packer's weight-balancing is working off data that
  is both stale (table last regenerated 2026-08-07) and known to
  underestimate the platform where the failure occurs.

Given codex-config.test.cjs is disproportionately heavy AND every companion
sharing its chunk is decided by a packing algorithm already documented as
reshuffling unpredictably on any new file, tuning the shared budget a third
time only moves the marginal line to wherever the next new file happens to
land — it does not remove the gamble. Isolating codex-config.test.cjs into
its own dedicated single-file chunk, unconditionally and on every platform,
removes it at the source: the file never enters the pool packChunks balances,
so no other file's packing changes, and no future single-file addition
(mine or anyone else's) can silently reintroduce this exact failure by
landing in its chunk.

Extracted as a small pure function, partitionIsolatedFiles (mirroring this
file's existing pattern of pulling packing/analysis logic out of main() for
in-process unit coverage — see computeSweepProtectSet, analyzeChunkEvents),
with 6 new tests in tests/run-tests-harness.test.cjs covering basename
matching across path separators, near-miss non-matches, the empty-list case,
and the isolated-set contents.

Root cause is now closed rather than papered over with a retry: this failure
is a property of one specific heavy file's chunk placement, not something
that recurs randomly. If codex-config.test.cjs itself is ever genuinely sped
up, this isolation can be revisited — this is a packing-side mitigation for
a known file's cost, not a claim the cost is irreducible.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 16:44:53 -04:00
Michel Moreira
ae40529d31 chore(#4394): lint allowed-tools parity — Bash without Grep (#4431)
gen-plugin-skills.cjs --check already guarantees skills/*/SKILL.md matches
what commands/gsd/*.md generates, so the two trees cannot silently diverge
FROM EACH OTHER. Nothing guarded the shape #3085 actually found: a command
shipping Bash without Grep purely by omission, identical in both trees and
therefore invisible to a parity check that only compares them to each other.
That drift ran until 29 of 71 skills lacked a tool most of their siblings
declared, and a manual audit — not a gate — is what surfaced it.

Detection only: the lint never edits a command's allowed-tools.

Two failure classes, not one. Violations are the rule itself. Stale
exemptions are the other half: an entry whose command is gone, or which no
longer declares Bash without Grep, fails just as loudly. An exemption list
that can only grow becomes a list of things nobody re-examined, and a
pre-forgiven command silently absorbs the next omission.

That check earned its keep immediately. #4394 named eight exemptions from the
#3085 review — the six ns-* dispatchers, help, and surface — and seven of
them do not declare Bash at all, so the rule never reaches them. Listing them
would have pre-forgiven seven commands for a condition none of them has. Only
surface needs an entry.

The tests drive synthetic fixtures rather than the live corpus: asserting
"the real tree is clean" would say nothing about whether the rule can detect
anything, which is the exact failure mode this lint exists to close. The one
live-corpus arm asserts no stale exemptions — a property of this script's own
list — and deliberately does not pin a violation count, which would make it a
baseline every fix has to update.

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 09:44:38 -04:00
Brenden Smerbeck
e54d3aa159 enhance(#4401): register workflow.compact_content as a validated config key (#4441)
* feat(#4401): register workflow.compact_content as a validated config key

- Add compact_content: false to the nested workflow object in
  gsd-core/bin/shared/config-defaults.manifest.json
- Add 'workflow.compact_content': false to SCHEMA_DEFAULTS in src/config.cts
  so an absent key resolves to false via config-get --raw
- validKeys entry in config-schema.manifest.json already present

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(#4401): behavioral and boundary tests for workflow.compact_content

- 19 behavioral tests covering config-set/config-get round trip, invalid-shape
  rejection (banana, 42, empty string), the corrected null-unset semantics
  (#2046), absent-key resolution against config-defaults.manifest.json,
  config-new-project wiring, and doc-row shape assertions
- Drops the install-tree fixture-parity block (and its docstring item) that
  asserted gsd-core/references/compact-content-gate.md and
  gsd-core/workflows/compact/map-codebase.md fixture entries — those paths
  belong to #4402 and do not exist on this filtered branch

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(#4401): document workflow.compact_content in both config references

- One 4-cell row in docs/CONFIGURATION.md (workflow.* run)
- One 5-cell row under Workflow Fields in gsd-core/references/planning-config.md
- Both cross-reference ADR-4139

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(#4401): add changeset

- Added-type fragment, pr: 4401 (issue number; backfill to the real PR number
  is a required follow-up once the PR is opened, per D-08 and CHANGESET-PR-
  FIELD-DRIFT)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(#4401): backfill changeset pr field to #4441

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(#4401): derive workflow.compact_content default from CONFIG_DEFAULTS

SCHEMA_DEFAULTS['workflow.compact_content'] hardcoded the literal false
instead of deriving it from CONFIG_DEFAULTS the way 3 of its 8 sibling
entries do (smart_zone_tokens, pr_strict, inline_plan_threshold), leaving
a single-source-of-truth drift risk: a future manifest-only edit to the
default could silently diverge from this literal, only caught later by
the D-03 test if it ever happened to manifest.

Adds compact_content to CONFIG_DEFAULTS in src/config-loader.cts and
derives SCHEMA_DEFAULTS from it in src/config.cts, matching the majority
sibling pattern. Found during maintainer review (review-open-prs) of
this PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4401): map compact_content in config-field-docs NAMESPACE_MAP

The previous commit added compact_content to CONFIG_DEFAULTS in
src/config-loader.cts but missed the matching entry in
tests/config-field-docs.test.cjs's NAMESPACE_MAP, which maps flat
CONFIG_DEFAULTS keys to their namespaced doc form before checking
gsd-core/references/planning-config.md for a match. Without it, the
test looked for a bare `compact_content` doc reference instead of the
actual `workflow.compact_content` row, and failed:
"CONFIG_DEFAULTS keys missing from planning-config.md: compact_content".

Found by actually running gsd-test against the branch rather than
trusting the plausible-looking fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4401): register compact-content-4139 test in the docs-guard lane

tests/compact-content-4139.test.cjs's D-06 tests read docs/CONFIGURATION.md
directly (fs.readFileSync) to assert the workflow.compact_content doc row's
shape, which makes it a doc-reading test file under the #3753 docs-guard
lane. It was never added to scripts/docs-guard-registry.cjs's
DOCS_GUARD_TESTS map and carries no docs-guard-exempt marker, so
tests/ci-docs-guard-registry.test.cjs's registration lint correctly failed:
"compact-content-4139.test.cjs reads a docs/ path but is not registered in
the docs-guard lane and carries no docs-guard-exempt marker".

Registers it with ['docs/CONFIGURATION.md'] (the only real docs/-prefixed
path it reads; gsd-core/references/planning-config.md is outside this
registry's docs/ scope, matching the sibling config-field-docs.test.cjs
entry's existing convention).

Found by actually running gsd-test against the branch — this gap predates
the maintainer's config-loader.cts fix and was already present in the
original PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: sim <sim@local>
2026-09-06 19:52:59 -04:00
Tom Boucher
1c0acb2359 feat(#4422): block merging into next/main while the base branch's Tests run is red (#4428)
* feat(#4422): block merging into next/main while the base branch's Tests run is red

Adds a next-health job to test.yml that checks the base branch's own last
push-triggered Tests run via the GitHub API and fails the existing "Required
tests" required check when it's red, with a maintainer-applied "fix-next"
label as the explicit escape hatch for the fix-forward PR itself. No
branch-protection config change needed — it rides the already-required
check. The job is deliberately not gated behind preflight, same reasoning
as the changes job: a compute-free API read has nothing to save by waiting.

Documents the fix-next label in CONTRIBUTING.md and adds a property test
locking the CLEAN/RED/INDETERMINATE classification's iff-relationship.

This closes the second half of the 2026-09-06 RCA: three unrelated PRs
merged on top of an already-broken next before anyone noticed it was red.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: close two zero-margin CI timing gaps found while verifying #4422

Discovered while watching this branch's own CI, root-caused via /diagnose
rather than dismissed as Windows flakiness:

1. tests/gsd-check-update-worker-atomic-cache.test.cjs's outer timeout
   (15000ms) exactly matched the inner npm-view timeout the worker wraps
   (NPM_VIEW_TIMEOUT_MS, gsd-core/bin/check-latest-version.cjs). A slow
   registry response raced two SIGKILLs at the same instant, killing the
   worker before it could catch its own timeout and degrade gracefully.
   Windows's shell-wrapped npm subprocess made the race lose more often
   there, but the zero margin was platform-agnostic. Fixed by giving the
   test real headroom (+10s) beyond the named constant it wraps, plus an
   invariant test so the two values can't silently collide again.

2. scripts/run-tests.cjs's per-chunk weight budget (MAX_FILES_PER_CHUNK)
   let a Windows full-matrix chunk that was well under budget by the
   Linux/macOS-calibrated weight table (~32/60 units) still exceed the
   600s wall-clock backstop — codex-config.test.cjs's genuinely-measured
   weight (17.87) doesn't transfer 1:1 to Windows's slower install/
   subprocess overhead. Windows now gets its own lower cap (40 vs 60).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 17:29:54 -04:00
Tom Boucher
38e4ce5f62 fix(#4186): anchored status vocabulary, record-session arg guard, recount pin (#4381)
* fix(#4186): anchored status vocabulary, record-session arg guard, recount pin

Three defects from #4186:

1. normalizeStateStatus ran a first-match-wins SUBSTRING chain over the
   free-prose body Status field, so prose merely mentioning a status word
   was silently rewritten to a credible wrong token (a .planning/ path in
   Italian prose -> status: planning; verifica -> verifying; completezza ->
   completed). Recognition is now an ANCHORED whole-field match against a
   declared vocabulary (STATUS_EXACT_TOKENS + STATUS_ANCHORED_PATTERNS,
   state-document.cts) — case/whitespace-tolerant, branch-order artifacts
   preserved (Planning complete -> planning; Phase complete — ready for
   verification -> verifying). The recorded lenient fallback (#3873 row 26)
   stands: unrecognized prose passes through verbatim. Read-side consumers
   (W011, statusline) ride the same function.

2. The progress recount skew (stray *-SUMMARY.md inflating
   completed_plans) is already dead on next via #1988/PR #2016
   (countMatchedSummaries pairs summaries to plans) — verified live and
   pinned with regression rows composed against the #4129/#4359 ratchet.

3. state record-session with no args executed and wrote STATE.md; it now
   errors like state update (stopped-at or resume-file required), handler-
   side so SDK callers are covered too. Four tests pinning the bare-call
   write are updated to the new contract.

* fix(#4186): update status pins to the anchored vocabulary contract

Bench round 1 follow-ups:

- Legacy bare 'Milestone complete' kept as reader-side vocabulary
  (ADR-2207 removed the writers, not recognition of legacy files).
- state.test pins updated: 'Paused at Plan 3' and round-trip
  'Executing Plan 5' were pins of the substring guessing itself —
  the round-trip now uses the real handler form 'Executing Phase 5'.
- record-session no-op/no-fields tests repurposed to the usage-error
  contract (CLI + SDK-level ExitError), byte-unchanged assertions kept.
- statusline tests repinned: vocabulary values collapse to keywords;
  narratives render the documented first-word fallback instead of a
  guessed token. Hook doc comment updated to match.
- docs-guard exempt baseline: state.test.cjs now cites docs/CLI-TOOLS.md.
- docs/CLI-TOOLS.md: record-session signature notes the required flag.

* fix(#4186): repair a dangling sentence in the schema docstring

* test(#4186): bound the completed_plans scan regex (#2128 class)

* chore(#4186): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-06 17:08:24 -04:00
Tom Boucher
2920bbc022 fix(#4421): rescind #494's macOS full-matrix skip on changed test files (#4427) 2026-09-06 15:42:23 -04:00
Tom Boucher
b7406b293f enhance(#2618): render pending todos as one bounded bullet per todo (#4384) 2026-09-06 08:06:39 -04:00
Tom Boucher
0be5bf865a enhance(#3783): audit-uat summary segments current-milestone vs archived debt (#4336)
* test(#3783): add failing coverage for audit-uat summary segmentation

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3783): segment audit-uat summary into current_milestone and archived buckets

Additive: current_milestone/archived are new; total_items, total_files, parse_gap_files, by_phase, and by_category are unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3783): add changeset fragment for audit-uat summary segmentation

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3783): allowlist the new audit-uat-summary-segmentation test file

lint-test-file-count.cjs baselines the "audit" module (keyed off bin/lib/audit.cjs)
at 6 pre-existing files; this adds the new dedicated suite as a 7th, matching the
module's existing one-file-per-feature-slice precedent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#3783): fix phase/file number mismatch in the mixed-milestone fixture

The active phase fixture used dir "02-current" with file "01-UAT.md" — a
cross-phase stray per phase-id.cts's isPhaseArtifact/scopeToPhase (#3511),
so the file was silently excluded from the scan and current_milestone read
{files:0, items:0} instead of {files:1, items:1}. Confirmed by direct CLI
run against a hand-built fixture before recommitting. Renamed the file to
02-UAT.md to match its directory's phase number, matching every other
fixture in this suite.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3783): backfill changeset PR number to 4336

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 19:03:14 -04:00
Cody Anderson
77e2472ca0 enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration

Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on
Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets
(the .env.example/.sample/.template/.dist templates stay readable).
Read checks file_path; Grep checks an explicit path and judges the glob
per brace alternative; Bash runs a two-pass token scan (quotes, comments,
redirects with fd digits, separators, $( )/backtick/<( ) recursion,
heredoc bodies never scanned as commands, nested bash -c/eval rescans,
git <ref>:<path> shapes) with a closed non-reading exemption set for
existence checks. Fail-open crash policy; 1 MiB commands are denied as
command-too-large; more than 64 glob alternatives as glob-too-complex.

Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt
for approval whenever any Read() deny rule exists, even in auto mode. A
hook denial is not a permission rule and never arms that check. The
installer-written deny rules are retired in the follow-up commit.

Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks
HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking
guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell),
shell-command-projection managed sets, installer-migration-report,
OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch),
docs tables in five locales, ADR-766 always-on list, regen:derived
fixtures, and a new table-driven unit suite.

* test(#4221): pin the secret-read guard in existing hook gates

Register gsd-secret-read-guard.js in every existing hook gate: the
hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest
REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity
EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal-
hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS,
kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and
typed-payload floors, the OpenCode adapter (grep mapping, include ->
glob, three dispatch tests) and a Kimi TOML matcher assertion.

* fix(#4221): retire installer Read() deny rules (legacy filter)

Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS
and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets)
strings. mergeClaudePermissions now only filters them out of an existing
permissions.deny: an absent deny key stays absent, a malformed one is
still repaired to [], and an array emptied by the filter is deleted so
no `"deny": []` residue is left. Uninstall filters the same legacy list
and, symmetric with the Antigravity branch, drops an emptied allow or
deny key and an emptied permissions object.

Unlike the #2278 allow-side migration there is no surviving current
deny list, so the constant is renamed rather than mirrored. Removal is
byte-exact: a hand-written identical rule is indistinguishable from the
installer's and is removed too (the manifest never recorded permission
strings). USER-GUIDE and CONTEXT.md updated.

* test(#4221): flip install-regressions deny-rule assertions to the retired shape

The fresh-merge, non-destructive merge, idempotency, end-to-end install,
reinstall and uninstall assertions now expect no Read(.env*) deny rules
and no permissions.deny key on a fresh install; the deny:null repair case
is kept. A new describe block covers the legacy filter: retired strings
removed with a user entry kept, partial sets, near-miss strings
untouched, idempotency, GSD-only deny array deleted, a pre-existing
empty deny preserved, and uninstall symmetry for allow/deny/permissions.

* chore(#4221): add changeset fragment for PR #4236

* fix(#4221): case-fold names; scan shell stdin and xargs pipes

Review round 1 (trek-e):

- Blocker: secret-name matching is now case-insensitive in the Read,
  Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a
  case-insensitive filesystem are recognized as the same secret file.
- Major: a shell interpreter's script is now scanned wherever it comes
  from. The tokenizer keeps heredoc bodies as per-segment tokens and
  records separator operators; pass 2 groups by segment id and resolves
  bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined
  `-lc`) scans the script operand, a file operand is checked as a file
  (a `<( )` operand's echo/printf output is reconstructed), otherwise
  stdin is the script and heredocs, here-strings and a piped echo/printf
  source are scanned. `eval` joins all its operands; `source`/`.` handle
  process substitution. Data heredocs (`cat <<EOF`, the commit-message
  shape) stay unscanned.
- Major: `… | xargs <cmd>` checks the upstream segment's operands as
  file names when the sub-command reads (`echo .env | xargs cat`,
  `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the
  inference; a shell sub-command's `-c` script is scanned.

Header, USER-GUIDE bullet and changeset updated; documented gaps now
include piped scripts from non-echo sources and `exec`/`timeout`
wrappers. 60 new suite cases pin the block and allow shapes.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 04:00:08 -04:00
JusticeWay
0fca71eaae enhance(#2529): cover every workflow with response-language directives + CI lint (#2558)
* enhance(#2529): cover every workflow with response-language directives + CI lint

Every workflow now carries response-language coverage in one of three forms,
and a CI lint keeps it that way.

- 43 workflows load the new shared reference,
  `gsd-core/references/response-language-directive.md`, by eager `@`-import.
- Lazy-loaded modes/steps/templates, which cannot rely on an eager import,
  carry an exact inline directive; 35 such paths are pinned by exact path.
- Fragments dispatched by a covered parent inherit coverage, proven per file
  rather than granted per directory.

The 45 workflows whose directive covered only "questions, prompts, and
explanations" now name inter-tool narration, which is the defect #2529
reports: the running commentary between tool calls stayed English while the
answers around it were translated.

`scripts/lint-response-language-coverage.cjs` enforces it and fails closed on
three independent discovery failures (unreadable catalog, empty catalog,
unfollowed symlink). It resolves which reference a workflow imports and applies
the same four-predicate test to that file, so a weakened shared reference
uncovers its importers instead of passing silently, reported once as a systemic
failure rather than 43 times. The walk follows symlinked subtrees with a
realpath cycle bound. `lint:ci` invokes it by name.

REQ-LANG-03 and REQ-LANG-04 state the contract in docs/FEATURES.md;
REQ-LANG-04 names the two forms that satisfy it ("narration", "between tool
calls") rather than enumerating class members an author cannot use verbatim,
and a test pins that text to what the matcher accepts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): register the coverage test in the docs-guard lane

`107eb8c1` (#3787) landed the docs-guard lane on `next` while this PR was
open: a test that reads a `docs/` path must be named in
`scripts/docs-guard-registry.cjs` or carry a `docs-guard-exempt` marker,
so the guards that read a doc run on the PR that changes it.

`tests/response-language-coverage.test.cjs` reads `docs/FEATURES.md` -- it
extracts every form REQ-LANG-04 offers an author and runs each through the
matcher that enforces it. Registration, not exemption, is the correct side
of that gate: a reword of the requirement with no code change is precisely
the diff this test exists to catch, and it is the diff the lane would
otherwise skip.

Registered narrowly (`['docs/FEATURES.md']`) rather than with the `'*'`
sentinel, so an unrelated docs change does not pull this test into the lane.

Verified: lint-docs-guard-registration 0 violations, tests/ci-docs-guard-registry.test.cjs
51/51, lint:ci exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): consolidate this PR's emitted-growth acks into its own fragment

This PR ripples emitted bytes across 85 workflow paths. Until now each ripple
was acknowledged by appending to whichever live fragment owned that path,
because two ack sources may never name the same path.

`a84f7563` (#3078) swept all 45 fully-spent fragments off `next`. Forty-two of
the paths this PR grows were owned by swept fragments, so those keys are now
unowned and this PR's own fragment declares them directly -- one path, one
source, and no dependence on a fragment that no longer exists. Each adopted
entry keeps its measurement and records where it came from.

Two paths are handled differently, because the sweep did not free them:

- `review.md` is now owned by `3034-parallel-reviewer-lanes.json`, which
  landed on `next` after the sweep. Its entry is live, so the old route still
  applies: this PR's note is appended to that entry rather than declared a
  second time.
- `plan-review-convergence.md` keeps the arrangement made in round 24.

Result: 3 fragments in the directory, 85 keys in this PR's own,
0 cross-source duplicates. `lint-emitted-drift-ack` exit 0,
`tests/emitted-attribution.test.cjs` green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): move REQ-LANG-03/04 into the feature fragment that now generates them

`36375513` (#3845) made docs/FEATURES.md a generated projection of
docs/features/*.md, marked "do not edit by hand". This PR wrote REQ-LANG-03
and REQ-LANG-04 straight into the generated file, so the rebase left the
requirement present in the projection and absent from its source -- the next
regeneration would have deleted both, and `tests/features-index-gate.test.cjs`
was already red on the mismatch.

Both requirements now live in docs/features/response-language-config.md
alongside REQ-LANG-01 and -02. Regenerating produces a docs/FEATURES.md that is
byte-identical to the committed one, so the text this PR shipped is unchanged --
only its source of truth moved to where #3840 put it.

The docs-guard registration is widened to name the fragment as well as the
projection. The requirement's source is the fragment now, and an edit there
that skips regeneration would otherwise reach this guard through neither path.

Verified: features-index-gate 68/68, lint-docs-guard-registration 0 violations,
ci-docs-guard-registry + response-language-coverage 142/142.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): hand the plan-phase ack back to its new live owner

`c933184b` (#3825) landed `3172-stated-failing-direction.json` on `next` after
fragment had adopted that path when the sweep left it unowned, so the merged
tree named it from two sources -- a hard failure in
`scripts/lint-emitted-drift-ack.cjs`.

The path has a live owner again, so the append route applies: this PR's note
joins that entry, carrying its own measurement, and the key is dropped from
this PR's fragment (84 keys left, the others untouched). The provenance
sentence written for the swept-fragment case is removed rather than reused --
this path was never orphaned, so that account of it would be false.

Same shape as `review.md` and `plan-review-convergence.md`: ownership is a
property of the merged tree, and a fragment landing upstream after a push can
reclaim a key no local check would have flagged.

Verified: lint-emitted-drift-ack exit 0, lint:ci exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): state byte figures that are true against the tree

The reference claimed `execute-phase.md` has "2 bytes of headroom under the
ceiling named below". That was true when the sentence was written -- the file
sat at 93398 against the 93400 comfort assert -- and upstream has since shrunk
it to 91493 against a 93600 hard ceiling, so the figure now understates the
headroom by three orders of magnitude. The rationale the sentence supports does
not depend on the number, so the number is gone rather than refreshed: a
restated figure would go stale again on the next upstream edit, and nothing
parses it.

Audited every other numeric claim this PR ships the same way, mechanically
against the merge base: all 82 FILE-delta claims in the ack fragment match the
real per-file delta exactly, and the 1,629-byte reference and 63-byte import
line check out. One class was imprecise: the 41 notes for workflows whose
inline directive was rewritten in place quoted the conversion counterfactual as
"+1,692 bytes more loaded context", which is the reference form's whole weight,
not the increase over the inline directive those files already carry. Each now
names both quantities and the net (+1,605 / +1,609 / +1,584).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): one rule for pinned vs inherited coverage, and the docs to pick it

Review measured that 14 of the 35 pinned fragments would pass by inheritance
anyway, and that the PR asserted both readings at once: inheritance is real
coverage (so those 14 pins are noise) or it is not (so 30 inheriting fragments
are green-but-uncovered). Only one can be true.

Inheritance is real: the predicate proves it per file -- the parent must
dispatch this exact path from a read/execute context AND be covered itself --
so the parent's directive is in the loaded context by the time the fragment is
read. The 14 pins are therefore removed along with the directive lines they
pinned, and those files inherit like the 30 structurally identical ones. The
rule is now stated where the set is declared, and enforced from the other side
by a test: no member of the pinned set may be one that would have inherited.
That is what decides the form for the next fragment.

- pinned set 35 -> 21; 14 workflow files revert to their base content
- `findViolations` no longer returns early on a pinned path: a file that becomes
  eagerly loaded and takes the shared reference is strictly better off, and the
  gate must not red that. The reference form is admitted because its own wording
  is validated in turn; an arbitrary reworded inline line still fails.
- the reference-directive cache is keyed by size and mtime, not by path alone,
  so a rewritten reference re-asked in one process no longer returns the stale
  verdict
- `carriesInlineDirective` names its negation blindness: four independent hits
  read vocabulary, not polarity
- the real-tree scan asserts each source produced files instead of `> 152`, a
  constant that read as the workflow count and would have passed a scan that
  lost one of its two directories
- the pinned-set size assertion goes the same way: the size follows from the
  rule, so the rule is what the suite asserts

Docs, for the gate that now governs every future workflow:
- `docs/contributing/response-language-coverage.md` -- why the narration class
  is the discriminator, the four coverage forms, the decision order that picks
  one, the pinned line, and what each failure message means
- a row in CONTRIBUTING.md's CI checks table, matching the docs-guard row
- `docs/CONFIGURATION.md` points at it from the `response_language` entry

Also: the changeset said 45 reworded workflows; it is 44 (42 @-reference + 21
pinned + 44 rewritten = 107 touched). That text ships to CHANGELOG.md.

`3707-parse-gap-reporting.json` landed on `next` reclaiming `audit-uat.md` and
`progress.md`; both handed back by the append route, leaving 82 keys here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): correct the reference-taker count, 43 -> 42

The ack notes said the import line is byte-identical "in each of the 43
workflows that take the reference" and that the alternative would be "43 inline
copies". The shared reference has 42 importers; the 43rd file in review's table
is `execute-phase.md`, which imports the OTHER reference. Corrected in all 41
notes that carry the sentence, across this PR's fragment and the two it appends
to.

Found by re-running the numeric audit from the previous round after the rebase,
which also re-verified all 84 FILE-delta claims against the new base -- all
exact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): migrate the emitted-drift ack from a fragment to commit trailers

ADR-3942 (#3954) landed while this PR was open: the acknowledgment is now a commit
trailer and tests/emitted-drift-acks/ no longer exists. The fragment is deleted and
each key it declared becomes one trailer, reasons unchanged.

The four keys this PR had handed to 3034-*, 3172-* and 3707-* under the one-source
rule come home here. That rule was the whole reason for the hand-backs, and the
trailer model has no shared namespace to collide in -- five of this PR's rounds were
spent on exactly those collisions.

Emitted-Drift-Ack-Growth: add-backlog.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: add-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: add-tests.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: add-todo.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: ai-integration-phase.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by ddf85287 (fix(#3357), #3513), and no fragment on `next` declares this path now. The ack therefore returns to this PR's own fragment, which is the only live source for it — the change to the path is this PR's.
Emitted-Drift-Ack-Growth: analyze-dependencies.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: audit-fix.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 19: this PR declared the path in its own fragment, and `3602-workflow-subagent-model-resolution.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3602-workflow-subagent-model-resolution.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: audit-milestone.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed from `2962-zsh-nomatch-for-glob-portability.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: audit-uat.md — A live/archived split was added and then reverted on this branch (see `$comment`): the split's extra rule in `initialize`, the narrowed Unparsed-table filter, and the separate 'Unparsed UAT Files in Archived Milestones' informational section are all removed, so the file settles at origin/next 5582 -> 7124 bytes (+1542, final). #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: autonomous.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 16: `3210-autonomous-precondition-gate.json` landed on `next` in 8fc88f66 (fix(#3210), #3528) and declares this path today. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3210-autonomous-precondition-gate.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: check-todos.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: cleanup.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 18: this PR declared the path in its own fragment, and `2142-quick-task-archival.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `2142-quick-task-archival.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: code-review-fix.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 13: `3190-code-review-fix-auto-rewrite-review.json` landed on `next` in 1d5d7795 (fix(#3190), #3434) and declares this path too. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3190-code-review-fix-auto-rewrite-review.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: code-review.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 19: the fragment that carried this sentence (`3503-diff-base-scope-anchor.json`) was retired on `next` by 2fca0e17 (enhance(#2554), #3695), and `2554-code-review-depth-overrides.json` declares this path today. One path takes exactly one ack source, so the sentence follows the path to its live owner. Re-homed from `2554-code-review-depth-overrides.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: complete-milestone.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 15: the fragment carrying it (`3458-audit-open-acknowledge-wiring.json`) was retired on `next` and the path is declared by `3409-unreachable-guard-arms.json` today. One path takes exactly one ack source, so the sentence follows the path to its live owner. Re-homed from `3409-unreachable-guard-arms.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: debug.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 14: the fragment that carried this sentence (`3149-init-debug-entry-point.json`) was retired on `next` by 26f8015c (fix(#3448), #3476), and `3448-debug-autoresume-next-action.json` declares the path today. One path takes exactly one ack source, so the sentence follows the path to its live owner rather than being dropped or re-armed under a retired number. Re-homed from `3448-debug-autoresume-next-action.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: diagnose-issues.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: an inline copy in every workflow would be that many places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 19: this PR declared the path in its own fragment, and `3602-workflow-subagent-model-resolution.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3602-workflow-subagent-model-resolution.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: discuss-phase.md — #2529 round 36: this workflow's inline directive was rewritten in round 10 to name inter-tool narration, but in the compressed form, and that rewrite came to −1 byte against `next` — so it declared no growth and this key was absent from this PR's ack set until now. Round 36 replaces the compressed clause with the same enumeration the other rewordings carry — narration between tool calls, status updates, progress notes, findings, questions, prompts, and explanations — because `discuss-phase.md` started from the identical upstream sentence as `verify-work.md` and `new-milestone.md` and those two took the full list, so the shorthand was an inconsistency rather than a decision. +87 bytes against `next`, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. The directive stays INLINE rather than becoming an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have cost 1,692 bytes of loaded context against the 87 this sentence costs. `commands/gsd/discuss-phase.md` dispatches this workflow lazily (`Read and execute ...`) rather than `@`-importing it, so the 87 bytes land in the installed file and are read once the workflow is dispatched, not on every command invocation.
Emitted-Drift-Ack-Growth: discuss-phase-assumptions.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 15: `3409-unreachable-guard-arms.json` landed on `next` in #3558 and declares this path too. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3409-unreachable-guard-arms.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: discuss-phase-power.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: do.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: docs-update.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +83 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +83 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 83 bytes this inline directive costs, a net +1,609. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,609 bytes more loaded context per invocation. Re-homed in round 19: this PR declared the path in its own fragment, and `3602-workflow-subagent-model-resolution.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3602-workflow-subagent-model-resolution.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: edit-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 13: `3262-editphase-milestone-scope-guard.json` landed on `next` in fd4715f8 (fix(#3262), #3446) and declares this path too. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3262-editphase-milestone-scope-guard.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: eval-review.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by ddf85287 (fix(#3357), #3513), and no fragment on `next` declares this path now. The ack therefore returns to this PR's own fragment, which is the only live source for it — the change to the path is this PR's.
Emitted-Drift-Ack-Growth: execute-plan.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 14: the fragment that carried this sentence (`2652-quick-diagnose-dispatch-isolation.json`) was retired on `next` by 362d0434 (fix(#3370), #3478), and `3370-execute-phase-gate-conflation.json` declares the path today. One path takes exactly one ack source, so the sentence follows the path to its live owner rather than being dropped or re-armed under a retired number. Re-homed from `3370-execute-phase-gate-conflation.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: explore.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed from `2229-explore-claim-disposition.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: extract-learnings.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: fast.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 18: this PR declared the path in its own fragment, and `3585-planning-commit-guard.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3585-planning-commit-guard.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: forensics.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: graduation.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: health.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 13: the fragment that carried this sentence (`2573-state-head-freshness.json`) was retired on `next` by 7ddcc198 (fix(#3309)), and `3309-health-docs-generated.json` declares the path today. One path takes exactly one ack source, so the sentence follows the path to its live owner rather than being dropped or re-armed under a retired number. Re-homed from `3309-health-docs-generated.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: help.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: import.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 18: this PR declared the path in its own fragment, and `3576-references-canonical-cites.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3576-references-canonical-cites.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: inbox.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: ingest-docs.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed from `2658-trae-instruction-file-path.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: insert-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: list-phase-assumptions.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: list-seeds.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: list-workspaces.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 …

* fix(#2529): read the catalog-relative dispatch spelling, and the plural of "output"

Two false positives in the coverage lint, both surfaced by this round's work
rather than by a red gate finding them for us.

#3552 landed `execute-phase/steps/protected-branch.md` on next while this PR was
open, dispatched from execute-phase.md's `"none"` arm with the path written
RELATIVE to the catalog. `namesFragmentAsEntryPoint` only ever looked for the
`gsd-core/workflows/`-rooted spelling, so it read a live dispatch as no dispatch
and the new fragment as uncovered. It now accepts both spellings and matches the
relative one on a path boundary, so `vendor/<path>` cannot vouch for `<path>`.

Recognizing that spelling makes one pin redundant: execute-phase.md dispatches
executor-isolation-dispatch.md the same way, so the fragment inherits and its
own copy of the sentence comes back out. That is the rule round 29 encoded,
enforced by the test that measures it rather than by hand.

`output` was the one term in USER_OUTPUT_RE without an `s?`, so "translate all
outputs, including narration between tool calls" read as uncovered. The new
property tests caught it on their first run.

Those properties pin the rule the hand-written cases are instances of: four
signals on ONE line accept, dropping any one rejects, spreading them across
lines rejects. The vocabulary is written out in the test rather than read back
from the script's regexes, per CONTRIBUTING.md "Fixture provenance (#2371)" -- a
generator seeded from the matcher can only re-derive what the matcher already
believes, and that independence is what caught the plural. Both new properties
are mutation-verified: dropping the narration predicate reds the necessity
property, and collapsing the document to a single line reds the cross-line one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#2529): spell the narration enumeration out in the last two shorthand directives

discuss-phase.md and plan-phase.md were the only two of this PR's 44
rewordings that abbreviated the inserted clause to "narration between tool
calls included" instead of naming the output classes the way the rest of them
do. Both forms satisfy the lint's four predicates, so nothing was broken --
but the point of #2529 is that an author reading one workflow should not have
to infer what the neighbouring one means by "included".

Both abbreviations were size decisions rather than wording ones, and both
reasons have since expired because next shrank the files. discuss-phase.md
sat 25 bytes under the 32,000-byte #717 dispatcher budget and now has 1,825;
plan-phase.md sat 87 bytes under the 94,519-byte ADR-857 capstone ratchet
against a +108 clause and now has 3,180. Neither budget is raised here and no
unrelated prose is trimmed; workflow-size-budget and
phase6-capstone-conformance both pass.

discuss-phase.md started from the identical upstream sentence as verify-work.md
and new-milestone.md ("All user-facing questions, prompts, and explanations in
this workflow"), and those two received the full enumeration; it now matches
them exactly. plan-phase.md keeps its own scope word ("orchestrator output") and
its subagent pass-through instruction, both upstream's, and only trades the
shorthand for the enumeration.

The shorthand now appears nowhere in the catalog. The two remaining variants
(plan-review-convergence.md, spec-phase.md) keep upstream's own verb and scope
and end on "report prose", which is what those workflows actually emit --
rewriting those would change a directive's strength, not its wording.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): match the workflow extension case-insensitively in coverage discovery

`findMarkdownFilesRecursive` filtered on `entry.name.endsWith('.md')`, so
`SETTINGS.MD` — the same file to Windows and macOS, a different one to Linux —
was skipped on the only platform whose verdict gates the merge. The direction of
that failure is the problem: a workflow the walk declines to see is a workflow
this lint certifies by omission, which is the same vacuous pass `main()` already
refuses when discovery returns nothing at all.

The filter is now an allowlist keyed on the lowercased `path.extname`.
`.mdx` stays out on purpose: admitting an extension states what a workflow IS,
and that claim has a second half — `inheritsParentCoverage` resolves a
fragment's parent as `<workflow>.md`. An `.mdx` entry belongs here next to the
parent resolution it would have to move with, not ahead of it.

Two tests: an uppercase-extension file is discovered AND lands as a violation
rather than an exemption, and every admitted extension is spelled so the
lowercasing match can reach it (an uppercase or dotless entry would be dead
configuration that reads like coverage).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#2529): cover the quick-batch workflow that #3676 landed uncovered

`next` gained `quick-batch.md` and nine step fragments in 2f64e6230 (#3676,
PR #4212) with no response-language directive, so the merge result reds this
PR's own lint with 10 violations. The lint is doing exactly what it exists to
do; the coverage is what has to move.

`quick-batch.md` takes the shared @-reference on line 1, the same as the other
42 top-level workflows, and eight of the nine fragments then inherit through
its `read and execute` stubs. The ninth does not:
`quick-batch/steps/plan-checker-loop.md` is dispatched by a SIBLING fragment
(`planner-wave.md:134`) and named in the parent only inside a parenthetical
with no dispatch verb, which is the shape round 29's rule already covers for
`execute-phase/steps/regression-gate-run.md` and
`plan-phase/steps/prd-express-path.md`. It carries the pinned inline directive
and joins `EXACT_INLINE_DIRECTIVE_WORKFLOWS`; the comment above that set now
names four such fragments instead of three. Coverage: 163 workflows.

`FULL_BUDGET` in tests/skill-frontmatter-contract.test.cjs moves 844 -> 846.
The same commit grew `help/modes/full.md` from 834 to 844 lines, landing it
exactly on the ceiling with zero slack, and the two lines this PR adds there
are its pinned directive and the blank separating it. That is a coverage
contract every workflow carries, not the content creep the budget guards.
The #597 ratchet rule holds: actualMax 846, slack 0, well inside LARGE_GRACE.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: quick-batch.md — #2529: the workflow arrived on `next` in 2f64e6230 (#3676, PR #4212) with no response-language directive, so this PR's lint reds on the merge result; covering it is the PR's whole contract, not an optional extra. It gains the shared directive as a single eager `@`-reference line, the identical form the other 42 top-level workflows take. FILE delta: +62 bytes, as the gate measures it. LOADED-CONTEXT delta: +1,691 bytes — the import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates read the FILE and not the transitive inline, so they see 62 of those 1,691 bytes; the remaining 1,629 are declared here because no gate reads them. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Nine `quick-batch/steps/*` fragments are covered without a byte of their own — eight inherit through the parent's dispatch stubs, and the ninth takes the pinned inline sentence, which the emitted surface does not measure.

* fix(#2529): scope row 48 by what a diff says, not by which paths it names

`tests/gsd-quick-batch-quick-regression.test.cjs` treats any branch touching a
`quick-batch` path as #3676 phase work, then forbids it from editing ordinary
`quick.md`. This PR covers EVERY workflow with the shared response-language
directive — quick-batch.md and its fragments included — so the scope check
turned true, and the row read this PR's one-line directive on `quick.md` as a
phase violation.

That is the false positive the row's own #3730 note already scoped away from,
arriving by the other door: not an unrelated branch that misses the surface,
but a catalog-wide sweep that touches all of it. A path now counts as phase
work only when its diff says something other than the coverage contract, and
the two accepted directive forms are read from
`scripts/lint-response-language-coverage.cjs` rather than restated, so a
reworded contract cannot leave the carve-out matching prose the lint no longer
recognizes. A file the branch ADDED still counts — every line is new, which is
what a real #3676-phase branch looks like.

The invariant is unweakened in the direction that matters: a phase branch that
edits `commands/gsd/quick.md`, `gsd-core/workflows/quick.md` or anything under
`quick/steps/` for any reason other than the directive still fails the row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-04 21:12:12 -04:00
Tom Boucher
2f4f7538e9 fix(#4264): wire both TDD dispatch backends to phase.tdd-applicable (#4284) 2026-09-04 16:08:42 -04:00
Tom Boucher
75ee7b0214 enhance(#4273): add phase.tdd-applicable single-owner predicate (#4277)
* enhance(#4273): add phase.tdd-applicable single-owner predicate

One query verb computes TDD-applicability for a plan (CLI flag, plan
type: tdd frontmatter, a task's tdd="true" attribute, or the
workflow.tdd_mode config default), mirroring phase.mvp-mode's
precedence-cascade shape. Foundation for epic #4272 Phase 2, which
wires both dispatch backends to consume it instead of restating the
predicate independently.

Also fixes workflow.tdd_mode, workflow.research, and
workflow.nyquist_validation, which never reached
cmdInitExecutePhase/cmdInitPlanPhase/cmdInitDebug/cmdInitNewMilestone
because loadConfig() never populates config.workflow — a dead
accessor found while wiring this verb's own config read, fixed inline
per the no-defer rule rather than left alongside it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4273): document phase.tdd-applicable's FEATURES.md entry

Add a docs/features/ fragment for the new phase.tdd-applicable query
verb and regenerate docs/FEATURES.md. docs/COMMANDS.md is left
untouched: it documents /gsd-* slash commands only, and the sibling
verb phase.tdd-applicable mirrors (phase.mvp-mode) has no formal CLI
reference entry anywhere in docs/ either -- only inline prose mentions
in docs/reference/workflow-fragments.md -- so there is no COMMANDS.md
precedent to extend.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): use PHASE_NOT_FOUND reason code, remove try/finally from tests

Two orthogonal code reviews flagged a mistyped error reason and a CONTRIBUTING.md-banned try/finally pattern in the phase.tdd-applicable change; both are corrected here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): stop whitelisting capability-owned config keys centrally

workflow.tdd_mode, workflow.research, and workflow.nyquist_validation are
each already owned by their own first-party capability's federated config
schema (the tdd/research/nyquist capabilities declare them under their own
capability.json `config`), resolved via isCapabilityConfigKey. Adding them
to gsd-core/bin/shared/config-schema.manifest.json's central validKeys, as
the prior commit in this branch did (mirroring workflow.mvp_mode, which
genuinely is central-only), declares the same key in two places at once.
That collision breaks capability-loader.cts's loadRegistry composition:
gsd-test caught this as 84-85 unrelated failures across
capability-cli/capability-command-dispatch/capability-lifecycle test files,
every one showing "unknown capability: <id>" for a freshly-installed
third-party capability that should have resolved fine.

Verified directly (not asserted): reverting only this file, keeping the
config-loader.cts tdd_mode/research/nyquist_validation flattening and the
init.cts call-site fixes from the prior commit, and re-running the exact
capability install + capability set repro from
tests/capability-cli.test.cjs's "issue-2322" test locally reproduces the
failure with the whitelist entries present and clears it without them.
loadConfig() still surfaces all three flattened values correctly with no
central whitelist entry (confirmed directly against the compiled module) —
the whitelist additions were never required for the #4273 fix to work; they
were an incorrect over-application of the mvp_mode precedent to keys that
aren't central.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): use getNested for tdd_mode (no legacy top-level fallback), allowlist new test file

Both fixes address defects found by a gsd-test bench run: tdd_mode routed through get() invented an undocumented top-level alias that silently outranked the canonical workflow.tdd_mode key, and the new phase-tdd-applicable test file was missing from the file-count allowlist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4273): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 14:14:56 -04:00
Tom Boucher
18e5cfff8a fix(#4250, #4260): distinguish a timed-out npm audit from a JSON parse failure, retry with backoff (#4251)
* fix(#4250): distinguish a timed-out npm audit from a JSON parse failure

npm-audit-baseline.cjs's runPackageLockAudit, and the near-identical
auditProductionVulns helper in npm-integrity-gate.test.cjs, both grabbed
e.stdout whenever an npm audit child process exited non-zero -- without
checking whether the process was actually killed by its 180s timeout.
A timeout-killed process's stdout is truncated mid-write, not complete
JSON, so JSON.parse threw a misleading "Unexpected end of JSON input"
instead of naming npm's registry timeout as the real cause.

Root-caused live during a CI investigation: npm's own status page
reported degraded service, and the registry's bulk-advisories endpoint
was returning 503/hanging, causing npm audit to sit until the timeout
fired.

Adds a shared isTimeoutKill(error) predicate (checks execFileSync's
documented killed/signal fields) and checks it first in both catch
blocks, throwing a clear, actionable error before ever reaching
JSON.parse. The pre-existing "non-zero exit with complete JSON"
recovery path is unchanged and still covered by regression tests.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4250): add changeset for npm-audit timeout fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4250): share the timeout-kill error message and cover auditProductionVulns

Two independent review passes (standards + spec) on the first commit found
real gaps: the timeout-kill error message was duplicated verbatim between
runPackageLockAudit and the near-identical auditProductionVulns helper in
tests/npm-integrity-gate.test.cjs (this repo's own Generative Fix Divergence
anti-pattern -- shared logic across parallel surfaces with no parity check),
and auditProductionVulns picked up the same production fix with zero test
coverage of its own.

Extracts buildTimeoutKillError(cwd), used by both callers so the message
cannot independently drift. Gives auditProductionVulns the same injectable
execFileSyncImpl seam runPackageLockAudit already had, and adds the matching
regression tests (timeout-kill throws the clear error; the pre-existing
non-zero-exit-with-complete-JSON path still recovers correctly).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4250): backfill changeset PR number to #4251

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* diag(#4250): surface captured stderr in the timeout-kill error

The killed child process's stderr is buffered in-memory by execFileSync
and attached to the thrown error, but nothing surfaced it -- the timeout
message named the timeout but discarded the one piece of data that could
show WHY npm was still running when it fired (DNS stall, TLS handshake
stall, a registry-side retry loop, all look identical without it).

buildTimeoutKillError now takes the killed error and includes its stderr
(or an explicit 'no stderr was captured' note) in the message. This is a
diagnostic improvement for the next CI occurrence, not a behavior fix --
local reproduction has directly ruled out npm version (installed the
exact CI-bundled 11.17.0 and ran it against this repo: 0.49s, clean),
general npm registry reachability (0.4-1.4s locally, repeatedly), and
npm ci speed (2m, succeeded) as explanations for the 180s CI hangs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4260): bounded retry with backoff for npm audit calls, finish the extraction

The audit backend has real, independent latency variance from the rest of
the npm registry -- measured (see #4260): a bulk-advisories POST took
43.41s vs 0.20s for a plain registry fetch on the same host, and the same
endpoint returned no response at all (000) twice in the same window,
while status.npmjs.org reported fully operational throughout. Against
that, runPackageLockAudit and its near-duplicate auditProductionVulns
each made exactly one attempt with no retry -- any single bad moment
failed a REQUIRED CI gate on a transport hiccup, not a real advisory.

Replaces the single 180s attempt with runNpmAuditWithRetry: up to 3
attempts at 60s each (comfortably above the worst measured working
latency) with exponential backoff between them. Only a confirmed
timeout-kill is retried; a genuine non-timeout failure still fails
immediately, and exhausting all attempts still fails the gate -- per
#4260's own caveat, silently disarming a required security check on a
transport error is worse than occasionally re-running CI.

Also finishes the extraction #4260 flagged as stopped halfway:
auditProductionVulns (tests/npm-integrity-gate.test.cjs) duplicated
runPackageLockAudit's entire candidate loop, recovery branch, and timeout
classification, differing only in npm args and precondition check. It is
now a thin wrapper delegating to the newly-exported runInstalledTreeAudit,
which shares runNpmAuditWithRetry with runPackageLockAudit -- one
implementation instead of two that could independently drift.

buildTimeoutKillError now reports attempt count and still surfaces
captured stderr from the last kill.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4260): update changeset for retry/backoff scope

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4260): budget for two sequential retry-audit calls, close coverage gaps

Two review passes on the retry/backoff commit found real gaps:

- TEST_TIMEOUT_MS budgeted only one retry-audit call's worst case (210s),
  but checkTreeAgainstBaseline makes two sequential calls (HEAD tree via
  auditProductionVulns, baseline tree via runPackageLockAudit) -- combined
  worst case is ~372s. If both genuinely exhausted retries, node:test's
  own timeout would fire first and mask buildTimeoutKillError's clear
  message, undercutting #4250's own fix in that edge case. Recomputed
  using the same backoff formula the production code uses, so it can't
  independently drift.

- buildTimeoutKillError's default-attempts(1) singular-phrasing branch had
  zero direct test coverage (nothing calls it with a single attempt
  anymore) -- a real mutation-testing risk. Added direct tests for both
  phrasing branches plus the no-error-object case.

- runInstalledTreeAudit's null-guard skip paths (missing package.json,
  missing node_modules) had no tests, unlike runPackageLockAudit's
  matching paths. Added for parity.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 10:02:17 -04:00
Tom Boucher
97ce61dee2 fix(#3990): state the RED/GREEN/REFACTOR cycle once, embed tdd.md conditionally (#4228)
* test(#3990): the RED/GREEN/REFACTOR cycle is stated once, embedded conditionally

* fix(#3990): state the cycle once — pointers in consumers, conditional tdd.md embeds

Emitted-Drift-Ack-Growth: execute-phase.md — #3990 conditions the tdd.md embed on the dispatch being TDD

* chore(#3990): changeset for the single-statement TDD cycle

* chore(#3990): backfill changeset pr number

* fix(#4228): linear cycle check — the lazy-span regex pinned a Windows core for the whole job cap

* test(#3990): allowlist pin tracks the rebased line

* fix(tests): npm-integrity gate names an empty audit output explicitly — empty stdout crashed the parse as a bare SyntaxError

* fix: name an empty npm-audit stdout explicitly — it crashed the parse as a bare SyntaxError

Observed on CI (several branches, all lanes): spawnSync npm ETIMEDOUT with
empty stdout; the empty string survived the recovery path and surfaced as
'SyntaxError: Unexpected end of JSON input', hiding the captured error. The
recovery path now requires non-empty stdout, and an empty result throws with
the captured stdout/stderr/message so the actual error is on the record.
Root cause of the ETIMEDOUT itself is NOT diagnosed here — this change only
stops masking it.

---------

Co-authored-by: sim <sim@local>
2026-09-03 21:41:14 -04:00
Tom Boucher
1fe85cd43e chore(#4244): ESLint rules for the #4220 Windows dirname-walk / TMPDIR-triad bug class (#4246)
* fix(#4244): repoint TEMP/TMP alongside TMPDIR and fix the sweepProtectSet fixed-point walk

Repo-wide sweep (ahead of adding lint rules for these exact bug classes)
found both incident patterns still live and unfixed on `next`:

- scripts/run-tests.cjs's sweepProtectSet walk stopped on
  `cur !== runTempRoot && cur.length > 1` — a POSIX-only sentinel.
  win32 dirname('D:\') is a fixed point (length 3, never satisfies
  `> 1`... wait, it does satisfy length>1), so a selected file living
  outside runTempRoot (the common case) spins the walk forever on
  Windows. Extracted a pure, exported computeSweepProtectSet helper
  that terminates on dirname(cur) === cur instead, with in-process
  RuleTester-style coverage for both win32 and posix paths.

- tests/run-tests-temp-root.test.cjs's own #4020 regression test set
  only TMPDIR on its runNode(...) child env. Node's os.tmpdir() never
  reads TMPDIR on Windows (only TEMP, then TMP), so the redirect
  silently no-oped there — masked because Windows CI died in the
  dirname-walk hang above before ever reaching this test.

- tests/config-schema.property.test.cjs's fallow config-set test had
  the same TMPDIR-only pattern, direct process.env assignment this
  time, restored in its own finally block.

Origin: #4220 and its shared root cause #4020.

* feat(#4244): require-full-tmpdir-triad and no-unbounded-dirname-walk ESLint rules

Two custom local ESLint rules catch the #4220 / #4020 Windows CI hang bug
class at author time, joining the ADR-1703 DEFECT.WINDOWS-TEST-PORTABILITY
catalog. Neither eslint-plugin-unicorn nor eslint-plugin-n has a rule for
either shape.

- local/require-full-tmpdir-triad: flags a TMPDIR environment override
  (direct process.env.TMPDIR assignment, or a TMPDIR property in a
  spawn-like call's env: object literal) not accompanied by TEMP and TMP
  in the same scope. Node's os.tmpdir() never reads TMPDIR on Windows.
  Registered on tests/**/*.cjs, matching the require-userprofile-with-home
  precedent.

- local/no-unbounded-dirname-walk: flags a while/do-while loop reassigning
  from dirname() with no fixed-point termination guard
  (dirname(cur) !== cur, or path.parse(cur).root). path.dirname() is a
  no-op at the platform root, but the value differs by platform
  (win32 'D:\' is length 3, posix '/' is length 1), so a POSIX-shaped
  length/equality bound never fires on Windows. Registered on BOTH
  tests/**/*.cjs and scripts/**/*.cjs — the real #4020 bug lived in
  scripts/run-tests.cjs, not tests/.

Both rules join the zero-escape-hatch discipline already established for
this catalog (no bespoke comment marker; PROTECTED_RULES in
tests/portability-rule-disable-ban.test.cjs independently bans
eslint-disable of either). ADR-1703 and its two companion contributing
docs get an amendment documenting the mechanism, code examples, and the
repo-wide sweep (three live instances found and fixed in the prior
commit; no others found). CI test-scope selection updated so an edit to
either rule or to scripts/run-tests.cjs re-runs the right suites.

* fix(#4244): no-unbounded-dirname-walk must analyze a single-condition loop test too

checkWhile bailed out early unless node.test was a LogicalExpression,
so a single-condition loop -- while (cur !== root) { cur = dirname(cur); } --
was silently skipped and never reported. That is the EXACT minimal
shape of the original #4020/#4220 bug, and it is literally the shape
used by this rule's own shipped RuleTester fixtures (the "equality-only
bound" invalid cases), which were failing (0 errors reported, 1
expected) until this fix -- confirmed by running RuleTester directly
against both fixtures, not just via a passing test-runner exit code.

The conjunct-collection helper already handled a non-LogicalExpression
test correctly (it pushes a single node as the sole conjunct); only the
early-return gate needed to stop requiring a compound && / || test.

Verified: RuleTester run directly against both previously-broken
fixtures plus two new sanity cases (a guarded single-condition loop
stays valid; an unrelated single-condition loop stays silent), and a
fresh `npx eslint .` across the whole repo remains clean (no other
single-condition dirname-walk shape exists in the tree).

* fix(#4244): require-full-tmpdir-triad must recognize a destructured child_process call

isSpawnLikeCallee only recognized a MemberExpression callee
(child_process.spawnSync(...)) or a bare identifier in
ENV_LOCAL_HELPER_NAMES (runNode). A destructured import called bare --
const { spawnSync } = require('child_process'); spawnSync(...) -- has an
Identifier callee named "spawnSync", which matched neither branch, so
the whole env-literal check was skipped. gsd-test caught this: both
"invalid: child_process.spawnSync with TMPDIR-only env" cases in
tests/require-full-tmpdir-triad.rule.test.cjs were failing (0 errors
reported, 1 expected).

Widened the bare-identifier branch to also match any of the known
ENV_CHILD_PROCESS_METHODS names, matched by name only -- the same
lightweight convention this repo's other eslint-rules/*.cjs use (e.g.
no-hardcoded-tmp.cjs's isFsMethodCall), not full import data-flow
tracing.

Verified: RuleTester run directly against all 11 cases in
tests/require-full-tmpdir-triad.rule.test.cjs (not just the two that
were failing), all pass; a fresh npx eslint . and npm run lint:ci
across the whole repo remain clean.

* fix(#4244): correct a stale escape-hatch reference in a test comment

The comment on the "length comparison against another expression's
length" case referenced a "// allow-dirname-walk marker" that doesn't
exist -- the rule has zero comment-based escape hatches by design
(ADR-1703), and an earlier draft's marker mechanism was removed before
this branch's first commit. Spec-axis review caught the stale
reference. No behavior change; comment-only.

* chore(#4244): backfill changeset PR number (pr:0 -> pr:4246)

---------

Co-authored-by: sim <sim@local>
2026-09-03 14:14:09 -04:00
Tom Boucher
456136659d fix(#4220): terminate the Windows temp-sweep ancestor walk (and two bugs it unmasked) (#4245)
* fix(#4220): terminate the temp-sweep ancestor walk with a fixed-point check

scripts/run-tests.cjs's sweepProtectSet block walked each selected test
file's ancestor directories, stopping on `cur !== runTempRoot &&
cur.length > 1` — a POSIX-only sentinel. path.posix.dirname('/') === '/'
(length 1) correctly stops, but path.win32.dirname('C:\\') === 'C:\\'
(length 3) never satisfies the length check, so the walk spun forever on
Windows whenever a selected file lived outside runTempRoot (the common
case). This has hung every Windows CI shard since #4207.

Extract the walk into a pure, exported computeSweepProtectSet(selected,
runTempRoot, dirnameImpl) helper and replace the length sentinel with a
fixed-point check (stop when dirnameImpl(cur) === cur), which terminates
correctly on POSIX, Windows drive roots, and UNC roots alike with no
platform branch.

* fix(#4220): repoint TEMP/TMP alongside TMPDIR in run-tests-temp-root test child env

Node's os.tmpdir() on Windows never reads TMPDIR, only TEMP/TMP. The
test's runNode child-process env override only set TMPDIR, so on a
real Windows runner nested inside a run-tests invocation the child
inherited the outer process's already-repointed TEMP/TMP and its
mkdtempSync(os.tmpdir()) landed under the outer run's temp root
instead of the test's intended `outer` directory. This was masked on
gsd-test's benches and locally because Windows CI always died in the
#4220 infinite loop before reaching this test.

* fix(#4220): stop the ancestor walk from protecting the filesystem root itself

computeSweepProtectSet added `cur` to the protect set before checking
whether dirname(cur) === cur, so on the terminating iteration it
protected the filesystem root (posix `/`, and analogously a win32
drive root) instead of stopping before adding it. Caught by the
existing posix-parity regression assertion
(`!protectSet.has('/')`) on the linux-node24 gsd-test bench. Reorder
to compute the parent and check the fixed point before adding.

* fix(#4220): backfill changeset pr number to 4245

---------

Co-authored-by: sim <sim@local>
2026-09-03 13:42:14 -04:00
Tom Boucher
114dfcb739 fix(#4196): npm-audit gate blocks only NEW advisories, not pre-existing ones (#4214)
* fix(#4196): npm-audit gate blocks only NEW advisories, not pre-existing ones

The #3588 gate failed on ANY advisory in the production tree, regardless
of whether the PR/push actually introduced it. Because npm's advisory
database updates continuously and independently of repo state, a commit
could pass this gate at merge time and fail it minutes later on the
identical tree -- proven on PR #4188/dce40eeb6, which passed on all 3
OSes at 15:57-16:27 and failed the same assertion at 16:19-16:30 on the
unchanged commit, purely because GHSA-jqff-g426-hqxp was disclosed for
fast-uri in the interim.

scripts/npm-audit-baseline.cjs diffs the head tree's vulnerable-package
set against a resolved baseline (the PR's target branch, or the prior
commit on a direct push) and blocks only newly-introduced advisories.
When no baseline can be resolved, falls back to the original
zero-tolerance behavior -- fail-closed, never silently weaker.

* fix(#4196): pin the npm-audit baseline instead of using a drift-prone ref

Two orthogonal reviews found the same class of bug this repo already
fixed once for a different gate (see GSD_EMITTED_BASE's own incident
comment in test.yml): origin/<branch> is live under fetch-depth: 0 and
can advance mid-run, so resolveBaselineRef()'s fallback to
origin/${GITHUB_BASE_REF} could silently disagree with the tree
ci-rebase-check.cjs actually merged. Wire AUDIT_BASELINE_REF from the
workflow to github.event.pull_request.base.sha / github.event.before,
the same pinned values GSD_EMITTED_BASE already relies on.

Also: HEAD~1 assumed exactly one commit per push, which this repo's
allow_rebase_merge:true setting can violate (a rebase-merged PR lands
as several discrete commits in one push) -- github.event.before is
git's own record of the correct pre-push state, not an assumed offset.
HEAD~1 remains as a documented last-resort fallback for out-of-band
invocations (e.g. gsd-test) that don't set any of the above, alongside
a new local-branch fallback for gsd-test's local `next` (not
origin/next) sandbox shape.

---------

Co-authored-by: sim <sim@local>
2026-09-02 22:38:37 -04:00
Tom Boucher
2f64e6230a feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core

Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch
binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the
new pure decision-logic module (arg validation, effective concurrency,
deterministic merge order, spawn backpressure, verification/merge
routing, cleanup-entry construction — design doc rows 3-15,24,26-28,
30-36,39; property rows 51-53). quick-batch-update-items.test.cjs
covers the new updateBatchItems export on src/quick-batch.cts (rows
15,22-23, including the negative cycle-rejection case).
quick-batch-command-router.test.cjs covers the new
gsd-tools quick-batch CLI family (rows 46-47). These reference modules/
exports that do not exist yet.

* feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router

Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision
layer — CLI verbs and pure orchestration logic only; no workflow
markdown, no Agent()/git-worktree I/O.

- src/quick-batch-dispatch.cts (new): pure decision functions consumed
  by the (separate, follow-up) /gsd:quick-batch workflow markdown —
  parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder,
  computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome,
  buildCleanupManifestEntry (the last parses caller-supplied plan text
  via the existing parsePlanDocument; no filesystem access).

- src/quick-batch.cts: adds updateBatchItems, resolving the design
  doc's Open Question 1 as ONE additive export on this module instead
  of the second, independent BATCH.json writer the design doc
  originally proposed. Reuses the same withPlanningLock transaction
  shape, computeWaves, and platformWriteSync call resumeBatch/
  completeQuickItem already use; fails closed without persisting on
  an unknown item, an unknown/self dependency, or an introduced cycle.

- src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI
  family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs)
  as a first-party always-on command (like /gsd:quick), not the opt-in
  capability-registry path graphify uses. Verbs: create/update/resume/
  complete (wrap quick-batch.cts) and effective-concurrency/
  merge-eligible/spawn-plan/verification-routing/merge-routing/
  cleanup-entry/parse-args (wrap quick-batch-dispatch.cts).

Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property
rows 51-53. Rows covering workflow markdown / Agent() dispatch /
`git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the
follow-up markdown-authoring pass, per the phase brief's explicit
scope boundary.

* docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces

New-.cts-module ripple for the two Phase 4 modules (epic #3344,
ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts,
ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source,
not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json
(via `node scripts/gen-inventory-manifest.cjs --write`, after
`npm run build:lib`), and CONTEXT.md glossary entries for
"Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router
Module", plus an update to the existing "Quick-Batch Core Primitives
Module" entry documenting the new updateBatchItems export.

* test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count)

scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs
file under the quick-batch production module by longest-prefix match,
and that module is already at its 2-file cap (quick-batch.test.cjs +
quick-batch.property.test.cjs). The standalone
tests/quick-batch-update-items.test.cjs added in the prior commit
pushed it to 3 and failed `npm run lint:ci`. Fold its content into
quick-batch.test.cjs (append-only — no existing test in that file is
modified) and update the CONTEXT.md glossary reference to match.

Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after
`npm ci` (this worktree previously had no local node_modules, which
also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-*
unable to resolve typescript — resolved by npm ci, no code change
needed there). `npm run lint:ci` and
`npx tsc -p tsconfig.build.json --noEmit` are both green after this
fix.

* test(#3676): add failing tests for the quick-batch command/workflow markdown

Failing-first tests for Phase 4's markdown-authoring pass (epic #3344,
ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs
covers commands/gsd/quick-batch.md's frontmatter/objective/process,
gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR
1610 NEW_FILE_CAP) and step-fragment count, the isolation model
(rows 20-22), the executor single-writer invariant (row 18), merge
validation reusing the existing bounded primitive (row 25), the
optional research/plan-checker/verification leaves (rows 16,17,19,
30,31), planning-failure blocking execution (row 29), the submodule
guard (rows 36,44), and the new agents/gsd-planner.md quick-batch
mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers
row 48 (ordinary /gsd:quick stays byte-identical). Named
`gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's
longest-prefix bucketing doesn't fold these markdown-only tests into
the already-capped quick-batch/quick-batch-dispatch/
quick-batch-command-router production-module buckets from the CORE
pass. These reference files that do not exist yet.

* feat(#3676): author the quick-batch command, workflow, and planner mode

Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch
binding") — the orchestration layer that calls into Pass 1's CLI
verbs (src/quick-batch-command-router.cts).

- commands/gsd/quick-batch.md (new): frontmatter/objective/process,
  delegates argument validation to `quick-batch parse-args`
  (parseQuickBatchArgs) rather than re-deriving the grammar.

- gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR
  1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded
  step fragments under gsd-core/workflows/quick-batch/steps/:
  resume-mode, batch-init, research-phase (flag:--research),
  planner-wave (+ nested plan-checker-loop when --validate),
  worktree-dispatch, merge-wave, verification-wave (flag:--validate),
  completion. Covers design doc rows 3-45: capacity/isolation
  resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG-
  layer planning with full-task-catalog prompts and always-required
  depends_on/files_modified frontmatter, serialized worktree create/
  merge/cleanup via the existing worktree.cleanup-wave primitive,
  deterministic wave-order merging, verification routing
  (human_needed/gaps_found), the executor single-writer invariant,
  submodule fail-loud guard, and #1941 fork-base auto-degrade.

- agents/gsd-planner.md: additive new `load_mode_context` bullet for
  `**Mode:** quick-batch`, pointing at the new
  gsd-core/references/planner-quick-batch.md reference (documents the
  always-required depends_on/files_modified contract, reusing the
  existing frontmatter grammar — no new keys). Existing modes
  byte-identical, only a new bullet added.

- src/init.cts (+init-command-router.cts, +command-aliases.cts):
  cmdInitQuickBatch / `init.quick-batch` — model profiles,
  commit_docs, roadmap/planning existence checks, and the
  section_manifest field gating research-phase/verification-wave
  (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY
  atoms — no new atom needed).

Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the
prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43,
45 covered by construction (verb wiring, single-writer prompt
constraints, crash-window resume via unmodified Phase 3 primitives).

* docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern

npm run regen:derived output for the new command/workflow/reference
(epic #3344, ADR-1239 "Quick-batch binding"):
- skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md)
- docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md,
  planner-quick-batch.md, and the quick-batch-dispatch.cjs/
  quick-batch-command-router.cjs CLI-module rows' now-live
  `/gsd-quick-batch` cross-reference (was "(separate, follow-up)")
  + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`)
- gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`)
  — research-phase/verification-wave gsd:section entries for the new
  quick-batch workflow
- tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) —
  the new command/workflow/skill/reference files now ship to every
  runtime

scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for
gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS
word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional-
flag pattern gsd-core/workflows/quick.md already carries baselined
(e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2);
quoting would break the intended "omit this arg when the flag is
false" splitting.

* fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch

Security review pass findings, both confirmed real:

1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment
   interpolated the raw, attacker-influenced task ${description} (and
   the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw
   description into every planner's prompt in the layer) straight into
   Agent() prompt bodies with no boundary. Fixed by wrapping every such
   interpolation in a <security_context> + DATA_START/DATA_END
   boundary, matching the CONCRETE convention already implemented in
   this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md,
   gsd-core/workflows/debug.md) — commands/gsd/quick.md's own
   <security_notes> only asserts this convention in prose, so the
   debug-agent files are the real precedent followed here. Added a new
   <security_notes> block to commands/gsd/quick-batch.md (it had none)
   documenting both this fix and the one below.

2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection.
   gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both
   ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED,
   causing shell word-splitting and pathname expansion on raw task-list
   text before the parser ever saw it. Fixed at the source: added a
   `--text <string>` form to the `parse-args` verb
   (src/quick-batch-command-router.cts) that accepts the ENTIRE
   $ARGUMENTS as ONE quoted argv element and does the whitespace split
   itself, in Node — which is never glob-aware, unlike the shell.
   Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form
   is kept for direct/test callers that already have a real argv array.

The SC2086 baseline entry added for the original unquoted line is now
stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports
it) and has been removed; the two SC2046 entries for the UNRELATED,
still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style
conditional-flag splitting remain — that line only ever expands to one
of a few known-safe literal strings (never raw user text), matching
quick.md's own already-baselined convention exactly.

Tests: quick-batch-command-router.test.cjs covers the new --text form
(token splitting, glob-shaped text passing through literally
unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs
asserts the DATA_START/DATA_END boundary on every leaf prompt
(research-phase/planner-wave/plan-checker-loop/verification-wave,
including the shared task catalog) and the quoted --text call sites.

* fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35

Spec review pass findings — the test matrix claimed "yes" coverage
these assertions did not actually support:

- Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection
  only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs,
  committed alongside the security fix that touches the same file) that
  .planning/quick-batches/ is never created for any rejected value —
  createBatch is genuinely never reached.
- Row 18 (--resume <unknown-batch-id>): previously only exercised a
  hand-corrupted BATCH.json, never a genuinely nonexistent batch
  directory. Added the real nonexistent-id case (also in
  quick-batch-command-router.test.cjs).
- Row 24 (post-planning updateBatchItems racing a concurrent
  completeQuickItem for a different item, both through
  withPlanningLock): zero test existed. Added a property test
  (tests/quick-batch.property.test.cjs, appended — Phase 3's own file,
  no existing test touched) exercising both call orders and asserting
  no lost update in the final on-disk manifest — the same technique
  Phase 3's own row-15 lock-contention property test uses (sequential
  calls through the real lock; a working mutex makes any interleaving
  equivalent to some serial order, so this is the same claim a literal
  concurrent-thread test would make without OS-level threading).
- Row 34 (worktree preserved on merge_failed) and row 35 (undeclared-
  deletion detection): both were previously asserted only at the pure
  routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs
  using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs
  already establishes for executeWorktreeWaveCleanupPlan (real repo,
  real worktree, a REAL merge conflict / a REAL file deletion diffed
  against declared_deletions) — asserting the actual worktree directory
  survives on disk, not just that a pure function returns a
  preserveWorktree:true field. Named gsd-quick-batch-* so lint-test-
  file-count's bucketing doesn't fold it into any capped module bucket.

Row 48 (/gsd:quick regression) intentionally left as-is per the
reviewer's own framing: the byte-identity claim is already
mechanically proven by the changed-path diff (git diff --name-only
empty on those two paths IS byte-identity), and a genuine execution-
level regression test would require actually running the workflow —
out of scope for this repo's unit-test model (no other quick.md
regression test in this repo does that either).

* docs(#3676): add the changeset and user-facing docs the command needed

Standards review pass findings — both HARD:

- Missing changeset. None of the 6 prior #3676 commits touched
  .changeset/*. /gsd-quick-batch is a new user-facing command;
  CLAUDE.md/CONTRIBUTING.md require one. Added
  .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder —
  backfilled after the PR opens, matching CLAUDE.md's own documented
  convention and Phase 3's own precedent, #4190's
  .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen
  form `/gsd-quick-batch` throughout, never the source-artifact colon
  form (`scripts/lint-docs-command-form.cjs` confirms 0 violations;
  that check scans docs/**, not .changeset/, so it was never actually
  in scope for the fragment itself, but the wording still follows the
  doc convention for consistency, matching how Phase 3's own fragment
  named the not-yet-shipped command).
- Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis
  how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's
  existing convention for /gsd-quick /gsd-fast) covering --jobs,
  --validate, --research, --resume, --file, the capacity/isolation
  interaction, and resume/failure recovery. Cross-linked from
  docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's
  own "Related" section. Added a /gsd-quick-batch section to
  docs/COMMANDS.md (same table format as the existing /gsd-quick
  entry) and docs/features/quick-batch.md (REQ-QB-01..12, same
  frontmatter shape as docs/features/quick-mode.md) — regenerated
  docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md
  via the standard generators.

* fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught

gsd-test's real run against 155e8975b3 found 43 failures, all rooted in
this phase's own new command/workflow never being registered across
~10 independent generated/hand-maintained registries this repo keeps
in parity by convention. Root-caused each, no test weakened or
special-cased.

- help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live-
  registry.test.cjs): added a /gsd:quick-batch entry to
  gsd-core/workflows/help/modes/full.md (the real help.md content;
  gsd-core/workflows/help.md is a thin dispatcher) documenting every
  flag (--file/--jobs/--validate/--research/--resume), matching the
  existing /gsd:quick entry's format.

- gen-section-manifest.test.cjs: quick-batch.md's
  `gsd_run query init.quick-batch` invocation used inline
  `$([ ... ] && echo --flag)` substitutions, which never satisfy the
  test's exact-whitespace-token / assigned-variable detection (the
  trailing `))` glued onto `--research` in the compound substitution
  broke the "exact token" match). Rewrote to the same
  VALIDATE_PARAM/RESEARCH_PARAM two-line pattern
  gsd-core/workflows/quick.md's own Step 2 already uses.

- runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md
  fragments that call gsd_run each needed their OWN embedded copy of
  the canonical shim preamble (every workflow .md that calls gsd_run
  carries its own copy — reading one file does not persist shell state
  into another). Ran `node scripts/sync-runtime-launcher.cjs`, which
  inserted it before each file's first gsd_run call.
  plan-checker-loop.md correctly has none — it never calls gsd_run
  directly.

- Namespace routing (skill-manifest.test.cjs, install-nested-
  layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added
  `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and
  routing table (same namespace `quick` already routes through), and
  to src/clusters.cts's `utility` cluster (same cluster `quick`
  already belongs to). Verified by hand-running installRuntimeArtifacts
  + applySurface for augment/cline against a real temp install: exactly
  6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested
  under gsd-ns-workflow/skills/, never re-flattened.

- mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a
  brand-new command is a real count change, not a bug this test should
  hide).

- model-omit-when-inherit-guard.test.cjs: added the canonical
  `<!-- #2517 model-omit-on-inherit -->` marker block to
  gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/
  researcher/checker/executor/verifier — lives in a steps/ fragment,
  read combined with the host by this test's own readWorkflowCombined,
  same as quick.md's own research-phase.md carries it for its gated
  section). Also fixed a genuine pre-existing inconsistency in the
  test's own "#2711: the guarded set is derived from dispatch sites"
  check: its `nonDispatching` computation read the BARE host file while
  `derived` (the set it's checked against) reads the combined
  host+steps content — inconsistent with that same test file's own
  #2994 doc comment explaining why the combined read is necessary.
  quick-batch.md is the first workflow whose EVERY model="{...}"
  dispatch site lives in a mandatory (never gated) steps/ fragment —
  extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a
  brand-new file — which is what exposed the mismatch. Fixed by using
  the same readWorkflowCombined read in both places.

- skill-frontmatter-contract.test.cjs: shortened
  commands/gsd/quick-batch.md's frontmatter `description` from 107 to
  91 chars (<=100 budget), and added `quick-batch.md` to the hand-
  maintained KNOWN_SKILLS consolidation allowlist with a #3676
  justification comment (a genuinely new first-party command, not a
  consolidation of an existing skill).

- workflow-fragments-emission.install.test.cjs: added `quick-batch.md`
  to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is
  deliberately NOT a no-op for it — its research-phase/verification-
  wave sections are gated).

- Regenerated all downstream artifacts (npm run build:lib && npm run
  regen:derived && npm run gen:plugin-skills -- --write && npm run
  gen:features -- --write): skills/gsd-quick-batch/SKILL.md,
  skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for
  augment/cline/hermes/qwen/trae/zcode.

- emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition
  (one new `load_mode_context` bullet pointing at the new
  gsd-core/references/planner-quick-batch.md reference) grew the file
  124 bytes without an acknowledgment trailer. Acknowledged below —
  the growth is the deliberate, additive, single-bullet change from
  the earlier feat(#3676) commit, not drift.

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green
(includes lint-workflow-shellcheck, lint-test-file-count,
lint-docs-command-form). The deep install/spawn/registry tests gsd-test
actually runs (docs-parity-live-registry, gen-section-manifest,
runtime-launcher-parity, install-nested-layout,
runtime-artifact-layout-surface, skill-manifest, skill-frontmatter-
contract, mcp-server-catalog, model-omit-when-inherit-guard,
workflow-fragments-emission) are not part of lint:ci — each fix above
was independently verified by hand-invoking the exact production
function the failing test calls (installRuntimeArtifacts, applySurface,
composeWorkflow, the CLUSTERS union, the section-manifest forwarding
regex) against the real repo tree and confirming the expected shape.

Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift.

* fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget

skill-frontmatter-contract.test.cjs's "feature #3039: tiered help —
size budgets" enforces a SEPARATE line-count ceiling for
gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines,
tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's
assertTightCeiling) — independent of the skill-frontmatter description-
length budget and consolidation allowlist I touched in the prior round;
those are unrelated checks in the same test FILE, not the same check.

Root cause: the /gsd:quick-batch entry I added to full.md in the
docs-parity fix round was 17 lines, pushing the file from 834 to 851
lines — 7 over the 844 ceiling. Condensed the entry (merged the
per-flag bullet list into one dense "Flags:" line, dropped from 3
Usage examples to 1) to 844 lines exactly — at the ceiling with zero
slack, which assertTightCeiling accepts (it only fails on
actualMax > ceiling, or on slack > grace when the ceiling is too
LOOSE — zero slack triggers neither).

Verified after trimming: full.md still contains a live /gsd:quick-batch
reference (bidirectional parity) and all 5 argument-hint flags
(--jobs/--validate/--research/--resume/--file) still appear as literal
tokens (docs-parity-live-registry.test.cjs's own flag-coverage check,
re-run by hand against the trimmed content).

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green.

* docs(#3676): backfill changeset pr number to 4212

Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's
pr:0 placeholder backfilled with the real PR number now that
gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR
Number Handling convention and Phase 3's own #4190 precedent
(708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt
from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE.

* fix(#3676): resolve prompt-injection-scan false positive on test fixture

tests/quick-batch.test.cjs:232's row 11b regression proves the task-list
parser carries a prompt-injection-shaped task description through
createBatch as inert data, never interpreted. The fixture has to be a
real "ignore all previous instructions..." phrase or the test asserts
nothing, but the full-file --diff scan flagged it once unrelated edits
in the same file pulled it into the changed-file set.

Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the
sanctioned, precedented exemption already used for other legitimate
security-regression fixtures (tests/windsurf-conversion.test.cjs,
tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs)
per DEFECT.PROMPT-INJECTION-SCAN-COLLISION.

---------

Co-authored-by: sim <sim@local>
2026-09-02 22:38:31 -04:00
Tom Boucher
7d6d788b51 fix(#4020): bound the test run's temp footprint with a swept run-scoped root (#4207)
* test(#4020): the runner must bound and sweep a run-scoped temp root

* fix(#4020): bound the run's temp footprint with a swept run-scoped root

* test(#4020): isolate the env-mutating rows in child processes

* fix(#4020): gate root removal on ownership so nested runners spare the outer root

* test(#4020): pass the probe file via --files, the runner's explicit-file flag

* test(#4020): resolve the probe by basename, as --files matching requires

* test(#4020): assert root survival, not content survival, in the nested-row

* chore(#4020): changeset for the run-scoped temp root

* chore(#4020): backfill changeset pr number

* fix(#4020): the sweep spares ancestors of the runner's own selected files

* fix(#4020): TMPDIR precedence — an operator redirect beats inherited TEMP/TMP

* fix(#4020): only the root's owner sweeps — a nested runner spares live sibling fixtures

---------

Co-authored-by: sim <sim@local>
2026-09-02 20:47:34 -04:00
Tom Boucher
858bb89769 ci(#4196): exempt dependabot[bot] from issue-link, title, and unsolicited-PR gates (#4203)
Dependabot has no mechanism to link a PR it opens to a repo issue -- its
alerts live in the Security tab, not as issues -- so require-issue-link,
pr-title-validator, and auto-close-unsolicited-prs all rejected its PRs
by design (confirmed live on #4193: auto-closed for "no pre-approved
issue", then flagged again by the title gate on reopen). Exempt by
authenticated author login (github.event.pull_request.user.login /
context.payload.pull_request.user.login), which GitHub attributes and a
crafted title or branch name cannot forge -- scoped narrowly to
dependabot[bot] only, no other author gets this treatment.

Co-authored-by: sim <sim@local>
2026-09-02 15:56:43 -04:00
Tom Boucher
2131fe13f3 enhance(#3464): exec() detection widening, citation-debt cleanup — Phase 8 (#4171)
* feat(#3464): widen no-source-grep to detect regex.exec() on tracked text

Adds an execCall kind alongside the existing regexTest detection --
regex.exec(tracked) was invisible to the rule while regex.test(tracked)
was already caught, despite both reading a source-derived string through
a regex. Measured: 4 previously-invisible sites across 2 files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#3464): migrate 4 sites newly flagged by the exec() widening

docs-hooks-table-parity.test.cjs's three regex-extraction loops are
site-scoped marked (source-text-is-the-product) -- the dynamic
preToolEvent/postToolEvent dialect branching they mirror is explicitly
documented as not statically parseable, so a literal-pattern mirror is
the practical minimum-cost check.

no-bare-gsd-tools-command-position.test.cjs's readRouterVerbs() now
requires HOST_COMMAND_ROUTERS directly instead of regex-walking
gsd-tools.cjs's source text -- the same accessor three other suites
already use.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3464): pay down 6 grandfathered uncited allow-test-rule markers

Two were genuinely load-bearing (suppressing a real detected violation)
and just needed a citation added -- phase6-capstone-conformance.test.cjs,
runtime-name-policy.test.cjs, both now (#3464).

Four were dead-weight file-header markers suppressing nothing -- each
file's real effective sites are covered by separate, already-cited
markers elsewhere in the same file. Deleted outright rather than cited,
per Phase 1's own precedent (remove non-load-bearing markers instead of
grandfathering them forever) -- codex-config.test.cjs (two copies),
gsd-check-update-worker-platform-gate.test.cjs, orphaned-hooks.test.cjs,
settings-jsonc.test.cjs.

allowlist.json: 134 -> 128 entries.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3464): re-baseline effective-exemption ceiling to 84

The exec() widening's 3 newly-marked sites are now suppressed and
counted; ceiling rises 81 -> 84, the exact measured high-water mark.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3464): correct citation and restore a wrongly-deleted marker

Two review corrections, both found by the orthogonal review pass:

- docs-hooks-table-parity.test.cjs's 3 new exec() markers cited #3464
  (mechanically "the phase that widened the rule") when the file's own
  established, correct reference is #3839 (the issue this whole test
  exists to enforce, already cited in its file header) -- fixed to match.

- gsd-check-update-worker-platform-gate.test.cjs's deleted file-header
  marker was NOT dead weight: its codeOnly() helper wraps readFileSync
  and is called inline as an assert argument, a genuine source-grep
  pattern on real .cjs/.js source that the rule cannot currently see
  (helper-function indirection is a distinct blind spot from anything
  Phase 7/8 measured) -- CONTRIBUTING.md is explicit that "unverified"
  is not the same as "vestigial." Restored, site-scoped this time
  (directly above codeOnly(), not as an inert file-header comment) and
  cited (#3103, the issue the file's own docstring already references).

codex-config.test.cjs's two deletions and orphaned-hooks.test.cjs's /
settings-jsonc.test.cjs's deletions were independently re-verified and
stand: their flagged lines read generated .toml/.json OUTPUT, not
source, or have no residual pattern at all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-02 08:11:23 -04:00
Carlos Cativo
9b77320580 fix(#4076): add missing gsd-hook-version header to gsd-node-runner.sh (#4092)
* fix(#4076): add missing gsd-hook-version header to gsd-node-runner.sh

gsd-node-runner.sh was registered in MANAGED_HOOKS but shipped without a
gsd-hook-version header, so gsd-check-update-worker.js always classified it
as 'definitely stale' (a missing header is indistinguishable from a
pre-version-tracking file). Every install on an otherwise up-to-date
version showed a permanent, unclearable '⚠ stale hooks — run /gsd-update'
warning naming this one file.

Root cause: the build-hooks.js comment claimed the file is 'not a
registered hook' and 'staged verbatim — no templating', but it IS in
MANAGED_HOOKS (managed-hooks-registry.cjs:34) and install.js already
stamps {{GSD_VERSION}} into every .sh hook unconditionally, gsd-node-runner.sh
included. The comment contradicted both the registry and the installer's
actual behavior, and the header line itself was simply never added.

Fix: add the header (matching every other managed .sh hook's format) and
correct the comment so it no longer asserts the opposite of what the
registry and installer actually do.

Adds a regression test that iterates every MANAGED_HOOKS entry and asserts
it carries a header matching the worker's own detection regex, so a future
hook added to the registry without one fails CI instead of shipping
silently.

Fixes #4076

* chore(#4076): add changeset fragment for PR #4092

* fix(#4076): address review nits — drop unneeded exemption, fix blank line

Per @trek-e's review on #4092:
- tests/managed-hooks.test.cjs:96: the readFileSync call uses a loop
  variable (entry-derived hookPath), not a literal path, so
  local/no-source-grep's static literal-path detector never flags it —
  the allow-test-rule exemption comment was unnecessary. Replaced with a
  plain note explaining the source-read rationale.
- tests/managed-hooks.test.cjs:121-122: dropped a stray extra blank line
  before the bug #2136 section divider.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-02 08:04:37 -04:00
Tom Boucher
5c7243e54b fix(#3995): derive the review diff base from the phase directory (#4181)
* test(#3995): diff base keys on the phase directory, not commit subjects

All three derivation sites (Tier 3, spawn_reviewer, fallow pre-pass)
must anchor on the phase directory's first commit; the milestone-blind
repro (an archived milestone's same-numbered phase commit capturing
the base) is the failing-first row. #3191/#3503 rows reworked to the
directory-anchor contract; T6 docs-parity forbids any remaining
phase-scope message-grep site.

* fix(#3995): derive the review diff base from the phase directory

A phase number is unique within a milestone, not a repository; the
message grep had no milestone bound and tail -1 deliberately selected
the oldest same-numbered subject, dragging archived milestones phases
into the scope (7 files to 3388 plus the >50 depth downgrade). All
three lockstep sites now anchor on the first commit that added anything
under the phase own directory — the same anchor class
git-base-branch phaseStartCommit uses. ShellCheck baseline gains the
escaped fragment shifted parse signature.

Emitted-Drift-Ack-Growth: code-review.md — phase-directory anchor replaces the message-grep derivation at both sites (#3995)

* chore(#3995): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 06:03:17 -04:00
Tom Boucher
f16ff7d1b3 enhance(#3545): widen no-source-grep with fold+hooks, migrate 76 sites (#4161)
* feat(#3545): widen no-source-grep with one-hop path-fold and hooks dir

Resolve a readFileSync() path argument that is a bare Identifier one hop
back to its VariableDeclarator initializer before classification, and
recognize `hooks` as a source directory alongside bin/lib/gsd-core/src.

Measured (epic #3464 phase 7): fold+hooks together newly flag 76
unsuppressed sites across 18 files that were previously invisible to
identifier-indirected or hooks/-rooted source reads. Neither widening
alone is sufficient — hooks-only surfaces 0 new sites, confirming #3520's
prior finding that the identifier-indirection gap must close first.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#3545): migrate 76 sites newly flagged by the fold+hooks widening

Per-site classification: rewrite behaviorally (require() the real module,
assert on its actual exported behavior) wherever the read was a proxy for
code behavior; add a site-scoped `// allow-test-rule: <reason> (#3545)`
marker only where the raw source text genuinely is the product under test
(codex-config.test.cjs's adapter-header-contract checks, install.js
structural-wiring guards with no exported symbol, AST-parse fixture
inputs, etc.) — each marker cites an existing repo-sanctioned category
from CONTRIBUTING.md's allow-test-rule exception table.

Also converts two try/finally test bodies (introduced during this same
migration) to the required t.after() cleanup pattern per CONTRIBUTING.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3545): re-baseline effective-exemption ceiling to 81

The fold+hooks widening's own newly-detected sites are now suppressed by
site-scoped markers, moving them from invisible into the tightly-ratcheted
effective-exemption count. Ceiling rises from 10 to 81 (the exact measured
high-water mark, grace unchanged at 2) — a deliberate, measured re-baseline
per the widening working as intended, not an ordinary ceiling bump.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3545): use canonical allow-test-rule category tokens

4 markers added during migration cited an issue ref correctly but didn't
use one of CONTRIBUTING.md's seven recognized category tokens, unlike
every other marker in this change. Cosmetic only — same suppression
lines, same effective/live counts (81/81, 0 live).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3545): correct stale phase-artifact path in test comment

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-01 21:40:38 -04:00
Tom Boucher
4dfc46bbe7 enhance(#3348): add a context-drift pre-check gate to plan-phase (#4147)
* test(#3348): add failing-first coverage for the context-drift gate

* feat(#3348): add context-drift pre-check gate for plan-phase

Compares each phase's *-RESEARCH.md/*-PATTERNS.md/*-VALIDATION.md/*-SPEC.md
effective last-changed time (git commit time, falling back to mtime for
uncommitted edits) against *-CONTEXT.md's, so plan-phase no longer silently
reuses an upstream artifact that predates a decision added to CONTEXT.md
after that artifact was derived from it. Deterministic, no model call.

New `gsd_run verify context-drift <phase>` command, sibling to the existing
verify.codebase-drift/verify.schema-drift gates in the drift capability.
Warn-only by default (workflow.context_drift_precheck), with an opt-in
workflow.context_drift_action: block escape hatch. Wired at plan:pre in
plan-phase.md, before both the RESEARCH.md and PATTERNS.md reuse decisions.

* fix(#3348): address code-review findings — raw-text-match, stale comment, import placement, duplicated phase resolution

* fix(#3859): pin the real commit's diff.ignoreSubmodules to match the empty-diff probe

The #3859 empty-diff guard decides whether a submodule bump would land using
`--ignore-submodules=dirty`, overriding the caller's `diff.ignoreSubmodules`
config. The real `git commit -- <paths>` that follows was never given the
same override, so under a bare `diff.ignoreSubmodules=all` repo config the
two calculations disagree: driven on git 2.39.5 (Debian bookworm, the
linux-node24 test-matrix image), the guard correctly stands aside but the
scoped commit itself then silently fails (exit 1, no error text) for a
gitlink bump it had just confirmed would be recorded, surfacing as
commit_failed instead of committed:true.

Pin `-c diff.ignoreSubmodules=dirty` onto the scoped commit call too, so the
probe and the commit it protects can never diverge. Harmless when no
submodule path is involved (driven: identical outcome on an ordinary scoped
file, with and without the flag).

* fix(#3348): guard resolvePhaseDirByToken's exact-match fallback against path traversal

* fix(#3348): retarget phase-enumeration-drift exemption to the consolidated resolvePhaseDirByToken helper

cmdVerifySchemaDrift's inline readdirSync was already function-scoped-exempt
in lint-phase-enumeration-drift.cjs as a single-phase LOOKUP (not a
current-milestone enumeration). This PR's refactor pass lifted that block
into a shared helper, resolvePhaseDirByToken, also used by the new
cmdVerifyContextDrift — the guard tracks exemptions by enclosing function
name, so the readdirSync now lives in an unexempted function and started
firing. Move the exemption to resolvePhaseDirByToken (same written reason,
now covering both callers) instead of migrating to listAllPhaseDirs, which
would introduce two real behavior deltas here: it catches readdirSync
failures internally (old code let them throw) and sorts results by phase
number before matchPhaseDirs picks matches[0] (old code used raw,
OS-dependent readdirSync order).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3348): satisfy lint:ci — slash form, capability registry regen

- docs/features/context-drift-gate.md used the deprecated /gsd: colon
  form; docs are never passed through the install-time slash-form
  converters, so lint-docs-command-form requires the hyphen form.
  Regenerated docs/FEATURES.md from the corrected fragment.
- Regenerated gsd-core/bin/lib/capability-registry.cjs after editing
  capabilities/drift/capability.json (lint:generated-sync).

* fix(#3859): pin the real commit's diff.ignoreSubmodules via env, not argv -c

The prior fix pinned `-c diff.ignoreSubmodules=dirty` onto the scoped commit's
argv via `commitArgs.unshift(...)`. `-c key=val` must precede the `commit`
subcommand, so this shifted `commitArgs[0]` from `'commit'` to `'-c'` for
every scoped commit call, breaking 17 position-based assertions in the
commit-files pathspec regression suite that read `a[0] === 'commit'` to find
the commit invocation among recorded git calls.

`execGit` already accepts an `env` option merged onto `process.env` before
spawning. Git honors `GIT_CONFIG_COUNT`/`GIT_CONFIG_KEY_0`/`GIT_CONFIG_VALUE_0`
as a per-invocation config override functionally identical to `-c key=val`,
expressed via env instead of argv. Passing that env alongside the existing
commitArgs (still `['commit', ..., '--', ...stagedPaths]`, argv unchanged)
fixes the real commit's effective diff.ignoreSubmodules to match the
empty-diff guard's probe without moving anything in argv position 0. Scoped
to exactly the canScope branch, matching the probe's own preconditions and
leaving no behavior change for commits the probe never evaluated.

No test file changes needed — the 17 previously-failing assertions test
argv[0] against the array passed into execGit, which never changes.

* fix(#3348): register verify-context-drift in the check subcommand router

The drift capability's new plan:pre gate declares check.query
"verify.context-drift", which normalizes to `check verify-context-drift`,
but no such subcommand was routed — phase6-capstone-conformance's
uniform-block-field test failed with "Unknown check subcommand" for
every declared gate query.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3348): extend #1592's exact-key-list snapshot for the new context-drift config keys

tests/capability-registry.test.cjs asserted an exact, hardcoded snapshot
of the drift capability's config keys. #3348 legitimately adds two new
keys (workflow.context_drift_precheck, workflow.context_drift_action)
for its own plan:pre context-drift gate — extend the expected set
(and clarify the assertion message) without weakening the test's
exactness.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3348): reconcile E2's exemption-migration pin with the resolvePhaseDirByToken extraction

#3348 (an earlier commit on this branch, e4b80ad81) extracted
cmdVerifySchemaDrift's inline phasesDir readdirSync/matchPhaseDirs block
into the shared resolvePhaseDirByToken helper (also used by the new
cmdVerifyContextDrift), and retargeted lint-phase-enumeration-drift.cjs's
function-scoped exemption from cmdVerifySchemaDrift to
resolvePhaseDirByToken accordingly — cmdVerifySchemaDrift no longer
contains a line the guard's detectors match, so it needs no exemption.

tests/phase-locator.test.cjs's E2 test still pinned the exemption to the
old name (cmdVerifySchemaDrift), unaware of the migration. Update E2 to
match the same "migrated call site's exemption must move, not
duplicate" pattern the test already applies to cmdRoadmapAnalyze and
cmdInitMilestoneOp just below it: drop cmdVerifySchemaDrift from the
still-exempt list and add symmetric assertions that it no longer
carries the exemption while resolvePhaseDirByToken now does.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3348): fix two self-contradicting/nondeterministic tests in context-drift.test.cjs

'always exits 0 (query command contract)' included the no-phase-arg
case, which contradicts the file's own earlier
'errors with usage message on missing phase arg' test (that case
legitimately exits 1 via the Usage error) — drop it from the
always-exits-0 cases.

'degrades to mtime comparison outside a git repo' and '...in a repo
with no commits' relied on real wall-clock ordering between two
back-to-back writeFileSync calls to prove CONTEXT.md is newer than
RESEARCH.md; on a fast filesystem both can land in the same mtime
tick, producing a tie that computeContextDrift's strict `<` correctly
treats as not-stale, so stale_artifacts comes back empty. Make both
tests deterministic via explicit fs.utimesSync instead of relying on
timing (CONTRIBUTING.md: never assert elapsed wall-clock time).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3348): add context_drift_precheck:false to the plan:pre all-off fixture

The "all plan:pre when-keys false" fixture explicitly disables every
known workflow.* plan:pre toggle, but didn't yet know about the new
workflow.context_drift_precheck key (defaults to true), so the new
drift context-drift gate stayed active and broke the
empty-activeHooks assertion.

Emitted-Drift-Ack-Growth: plan-phase.md — adds the #3348 context-drift plan:pre pre-check section (new ## 4.6); this PR's own diff, not incidental drift.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3348): backfill changeset PR number (pr:0 -> 4147)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-01 21:39:37 -04:00
Dennis Alexis Valin Dittrich
b848b23861 feat(#3778): dispatch plan:pre planner contributions before quick planning (#3934)
* feat(#3778): dispatch plan:pre planner contributions in quick.md

- Add plan:pre capability gate to quick.md Step 5, mirroring plan-phase.md's
  existing render + generic contribution dispatch pattern
- Inject planner-targeted contribution fragments into the planner prompt,
  after AGENT_SKILLS_PLANNER, matching D-08 ordering
- Add tests/quick-plan-pre-capabilities.test.cjs proving the dispatch is
  generic (D-01) via real scanWiredKinds/coveredKindsInRegion functions
- Record quick.md's byte-growth rationale in this commit trailer for every
  gsd-core-verbatim runtime

Emitted-Drift-Ack-Growth: quick.md — #3778: Step 5 (Spawn planner, quick mode) gains a `plan:pre` capability gate, mirroring `plan-phase.md:420-424` and `:797`. This is shipped shell and prose read by an agent at runtime, not compiled, so the reasoning has to travel with the feature rather than being deferred to a reference doc: (1) the dispatch paragraph phrases role routing possessively ("the role each entry's `into` names") rather than as an `into ==` equality, because `coveredKindsInRegion` (scripts/gen-loop-host-contract.cjs) voids a segment's `kind == "contribution"` coverage credit when a role or capability equality shares that same segment — an equality phrasing here would silently fail the generic-dispatch proof required by D-01; (2) `activeHooks` is read directly in-context from `PLAN_PRE_HOOKS_JSON`/`HOOKS_JSON` and the unfiltered `rendered` digest is explicitly forbidden from being pasted, because `rendered` carries every kind and role — including non-planner-targeted contributions such as a `into: "checker"` twin — and pasting it would leak checker-scoped guidance into the planner's prompt (T-01-02 in the threat model); (3) the injection block sits inside `<planning_context>` AFTER `${AGENT_SKILLS_PLANNER}` and after the Project skills line, matching plan-phase's `:741` -> `:797` ordering (D-08), so agent-skills content is never shadowed by capability-contributed prose. No prose was moved into an eagerly `@`-imported reference to shrink the measured file — @gsd-core/references/loop-hook-dispatch.md already existed before this change and is deferred to for the generic contract only, exactly as plan-phase.md already does.

* test(#3778): expand quick.md plan:pre dispatch coverage to all nine locked conditions

Extend tests/quick-plan-pre-capabilities.test.cjs with D-02 (silent
omit-when-empty), D-03 (single shared planner spawn), D-06 (array-order
dispatch phrasing), D-07 (planner-only into filter), and D-08 (render call
< agent-skills placeholder < injection block < spawn ordering) assertions,
all extracted via a brace-bounded slice anchored on the literal injection
instruction rather than a naive first-brace scan (${AGENT_SKILLS_PLANNER}
and the surrounding prompt's ${VALIDATE_MODE ? ...} ternaries also contain
brace pairs).

Add a capability-registry.test.cjs describe block proving the registry-wide
D-07 exclusion is meaningful: at least one plan:pre contribution exists,
every plan:pre contribution has a non-empty into/fragment.inline, and the
registry as a whole carries at least one non-planner-into contribution.

Add a loop-host-contract.test.cjs regression pin for D-09: quick.md stays
absent from STEP_WORKFLOWS, parseLoopHostBlock still throws on quick.md's
real content, and buildContract() still yields exactly 5 entries.

Verified red-without-Task-1 by temporarily reverting quick.md to its
pre-f30de9cc content and re-running these three suites (D-08 failed as
expected), then restored via git checkout and re-confirmed green.

* docs(#3778): note quick planning also renders plan:pre in the tutorial

The tutorial's Step 6 named only /gsd-plan-phase as the trigger for the
plan:pre hook set. Since quick.md now dispatches the same hook set
(f30de9cc), the sentence understated the capability's real reach.

* feat(#3778): add changeset fragment

* chore(#3778): reference the upstream issue in the changeset fragment

The fragment was the only one of 81 in .changeset/ without a trailing
(#NNNN) reference or a bold lead-in. serializeChangelog auto-appends
only the pr: field, so the rendered CHANGELOG entry carried no link
back to issue #3778.

* test(#3778): scope the D-07 registry assertion to what it actually proves

The registry-wide non-planner check was named "D-07 exclusion is
meaningful", which overclaims: it proves only that `into` takes
non-planner values somewhere in the registry, not that anything is
excluded at plan:pre. Every plan:pre contribution is currently
into: "planner", so the filter is a forward-looking safeguard there.

Narrowing the assertion to plan:pre (as review suggested) would fail
today. Asserting plan:pre is all-planner would be brittle — it would
break the day a legitimate non-planner plan:pre contribution lands,
which is exactly when the safeguard starts doing work. So the
assertion is unchanged and only the name and comment are corrected.

* chore(#3778): point the changeset fragment at the upstream PR

The fragment carried pr: 3, the fork staging PR. changeset lint derives
the real PR number from GITHUB_EVENT_PATH, so on the upstream PR that
would read as pr-field drift. Point it at open-gsd/gsd-core#3934.

* test(#3778): require contributions in Quick revision prompts

* test(loop-host): require Quick auxiliary registration

* fix(#3778): preserve contributions in Quick plan revisions

* fix(#3778): validate Quick as a planner contribution host

* fix(#3778): tighten Quick contribution contract

* test(#3778): drop unnecessary source-contract exemption

* fix(#3778): require Quick planner target coverage

* docs(#3778): describe targeted auxiliary coverage

---------

Co-authored-by: davdittrich <davdittrich@gmail.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-01 21:16:57 -04:00
Tom Boucher
eca9c2b590 fix(#4112): portable grep in workflow markdown + submodule commit fix under diff.ignoreSubmodules (#4149)
* fix(#4112): ban GNU-only grep -P in workflow markdown to prevent macOS regressions

Add scripts/lint-portable-grep.cjs, wire it into lint:ci, and add
tests/lint-portable-grep.test.cjs. The #4112 shell-syntax fix newly exposed
a pre-existing grep -oP invocation that silently resolves to "" on stock
macOS's BSD grep (no -P support). This ratchet catches the same class of
GNU-coreutils assumption before it merges, mirroring lint-portable-timeout.cjs.

* fix(#4112): drop unneeded ls -l long-format that broke basename/dirname extraction

The previous commit on this branch replaced grep -oP 'phases/\K[^/]+'
(GNU-only, silently fails on macOS's BSD grep) with
basename "$(dirname "$phase")"), but left ls -lt (long format, -l) in
place. -l output is a full detail line (permissions, owner, size, date,
path), not a bare path, so dirname on that string throws "illegal
option -- r" (BSD) / errors under GNU coreutils too -- the extraction
never produced a usable value, on any platform. -l was never needed
here; only mtime-sort plus the bare path mattered. Dropping it to
ls -t restores one-bare-path-per-line output, which basename/dirname
actually requires. Confirmed on gsd-test's Linux bench: the #4112
regression test (tests/pause-work-context-detection.test.cjs) now
passes phase, spike, and sketch resolution.

Emitted-Drift-Ack-Growth: pause-work.md — portability fix (#4112): dropping grep -oP for a portable basename/dirname extraction is a few characters longer per line; no functional growth beyond the fix.
Emitted-Drift-Ack-Growth: sync-skills.md — portability fix (#4112): replacing grep -oP '(?<=--from )\S+' with a portable sed -E capture-group equivalent (no PCRE lookbehind available) is a longer expression; growth is the direct cost of the fix, not new functionality.

* fix(#3859): pin git commit itself against diff.ignoreSubmodules=all, not just the probe

git 2.39.x (the exact version on the CI Linux bench) resolves a
pathspec-scoped `git commit -- <path>` through the same
diff.ignoreSubmodules-gated machinery as `git diff`, so a genuinely
bumped submodule gitlink was silently dropped by the real commit even
though the #3859 empty-diff probe correctly saw the change. The commit
was misclassified as generic commit_failed because git's refusal text
never says "nothing to commit". Prefix the pathspec-scoped commit
invocation with -c diff.ignoreSubmodules=dirty, the same override the
probe already carries, so the two can no longer disagree. Confirmed on
git 2.50.1 (no-op) and git 2.39.5 via the actual gsd-tester-linux
v1.8.0-node24 bench image (turns the silent refusal into a commit).

* fix(#3859): carry the diff.ignoreSubmodules override via env, not argv

The prior commit prepended `-c diff.ignoreSubmodules=dirty` to the real
commit's argv, which shifted `commitArgs[0]` off `'commit'` and broke
several pre-existing #3859 regression tests
(tests/commit-files-pathspec.test.cjs) that assert on the raw argv
captured at the execGit seam, e.g. `gitCalls.some((a) => a[0] ===
'commit')`. Carry the same override via `GIT_CONFIG_COUNT` /
`GIT_CONFIG_KEY_0` / `GIT_CONFIG_VALUE_0` env vars instead, which git
honours identically and leaves argv untouched. Re-confirmed on the
gsd-tester-linux v1.8.0-node24 bench (git 2.39.5): the submodule commit
still succeeds.

* test(#3859): pin the commitEnv GIT_CONFIG_* override to canScope directly

The commitEnv/GIT_CONFIG_* override added in e935694fc/b3d37b929 was only
exercised indirectly through pre-existing submodule integration tests. Add
two dedicated regression tests that pin the canScope=false side of the
scoping decision: a whole-index commit (no --files) and an --amend commit
must both keep recording a bumped submodule gitlink under
diff.ignoreSubmodules=all with no override applied, since neither shape
carries a pathspec for that git internal check to consult.

* fix(#4112): match grep invocations after then/do/else/elif shell keywords

lint-portable-grep.cjs's GREP_INVOCATION_RE only anchored to line-start or
right after `| & ; ( \` {`, so a grep call positioned right after `then`,
`do`, `else`, or `elif` (e.g. `if x; then grep -oP '...'; fi`) was never
flagged, letting the GNU-only grep -P defect this lint exists to catch
reappear undetected in that shape.

* fix(#3859): apply the diff.ignoreSubmodules override to every commit, not just scoped ones

The commitEnv GIT_CONFIG_* override landed in e935694fc/b3d37b929 only when
canScope was true, on the assumption that git 2.39.5's silent refusal of a
bumped submodule gitlink under diff.ignoreSubmodules=all only affects a
pathspec-limited `git commit -- <paths>`. Reproduced directly against the
pinned CI tester image (ghcr.io/open-gsd/gsd-tester-linux:v1.8.0-node24,
git 2.39.5): a bare whole-index `git commit -m ...` with no pathspec at all
is refused identically when the only staged change is a submodule gitlink,
since git's "nothing to commit" check is a real diff against HEAD that
honours diff.ignoreSubmodules regardless of whether a pathspec narrows it.
Apply the override unconditionally instead of gating it on canScope; it
remains a confirmed no-op for --amend, which never hits this refusal at
all. Caught by the new regression test added in 579ac9f8f, which failed
against the real bench git version before this change.

* fix(#3859): apply diff.ignoreSubmodules override to commit-to-subrepo and pr-subrepo

cmdCommit already carries the GIT_CONFIG_* override that forces
diff.ignoreSubmodules=dirty on the actual git commit call so a bumped
submodule gitlink is not spuriously refused under git 2.39.5. The two
sibling multi-repo commit paths, cmdCommitToSubrepo and cmdPrSubrepo,
build the identical canScope-branched commit invocation but never
carried the override, so the same refusal there surfaces as a generic
commit_failed/error instead of a recorded gitlink. Apply the override
unconditionally to both, matching the corrected cmdCommit shape.

* fix(#3859): pin pr-subrepo's change detection against diff.ignoreSubmodules

cmdPrSubrepo discovers what to commit via `git status --porcelain`, which
(like the empty-diff probe fixed for cmdCommit) honors a local
diff.ignoreSubmodules=all config. Under that config, a genuinely bumped
submodule gitlink is invisible to the status scan, so changedFiles comes
back empty and the function reports nothing_to_commit before ever reaching
the commit call this same issue already fixed. Pin the status probe with
--ignore-submodules=dirty, mirroring the flag cmdCommit's diff probe already
uses, so a real gitlink bump is detected regardless of local config.

Found while adding a regression test for the previous #3859 follow-up fix:
the test failed not on the commit step but on this earlier detection step.

* docs(#3859): update changeset to cover all three fixed commit sites

* docs(#4112): backfill changeset PR number (pr:0 -> pr:4149)

---------

Co-authored-by: sim <sim@local>
2026-09-01 17:31:01 -04:00
Tom Boucher
80d2236ca5 fix(#4112): drop redundant $(( subshell wrapping in pause-work.md Context Detection (#4140)
* test(#4112): add failing-first regression coverage for pause-work.md Context Detection

Extracts and executes the shipped Context Detection bash block to prove the
$(( arithmetic-context misparse (SC1102/SC1106/SC2205) as a hard syntax
error, plus boundary/independence coverage for the phase/spike/sketch/
deliberation resolution the fix must preserve.

* test(#4112): fix false-green regression test — assert dash/sh failure, not bash -n

The prior version asserted bash -n syntax validity and wrapped bash -c
execution, neither of which detects this bug: bash/zsh silently fall back to
a working command-substitution reading of the ambiguous $(( construct, and
bash -n never evaluates arithmetic-context content at parse time. A gsd-test
run against the still-broken file passed 41594/41594 with this test in
place — a false green. POSIX sh/dash does not implement that fallback and
throws a real "Syntax error: Missing '))'", which is what this version
asserts against.

* fix(#4112): remove leaked tool-call tags corrupting the regression test file

A subagent's Write introduced trailing </content>/</invoke> markup at the
end of tests/pause-work-context-detection.test.cjs. gsd-test's own
tests/portability-rule-disable-ban.test.cjs (which parses every test file)
correctly caught this as a parse error. Strips the garbage lines; no
behavioral change.

* fix(#4112): drop redundant $(( subshell wrapping in pause-work.md Context Detection

phase=/spike=/sketch= opened with $((, which POSIX sh/dash parses as
arithmetic expansion and rejects (Syntax error: Missing '))') since the
enclosed text isn't valid arithmetic. bash/zsh silently retry it as command
substitution, which masked the defect there. Dropping the inner grouping
parens (matching the deliberation= line already in the same block) makes
the construct valid under bash, zsh, and dash alike, with identical
fallback-to-empty-string behavior when nothing matches.

Once the surrounding syntax parses, ShellCheck can now also analyze the
inner ls -lt calls it previously couldn't see past the parse failure,
newly surfacing the same pre-existing SC2012 ("use find instead of ls")
suggestion already baselined for the file's other ls usage. Baseline
updated to reflect the true current count; no ls-vs-find behavior change
made, as that is a separate, pre-existing, out-of-scope question.

* fix(#4112): verify shell discrimination instead of trusting the binary name

Code review finding: falling back from dash to sh could silently produce a
non-discriminating test on a host where /bin/sh is bash-compatible (e.g.
macOS) — such a shell never throws on the broken construct either, so the
test would pass whether the bug were present or not. resolvePosixShell now
verifies the candidate actually rejects a known-ambiguous $(( snippet
before using it, and skips with an explicit reason when none does. The
gsd-test Linux bench (dash as /bin/sh) is unaffected either way.

* fix(#4112): route regression test through process-seam, splitLines, cleanup

npm run lint:ci flagged 7 violations gsd-test's node:test run doesn't check:
local/no-adhoc-markdown-parsing (single fence-spanning regex),
local/no-crlf-fragile-split (bare \n split), local/no-unbounded-spawn (x4,
hand-rolled execFileSync with no timeout), and local/no-raw-rmsync-in-tests.
Rewrites the test to route every subprocess call through
tests/helpers/process-seam.cjs's runHook() (bounded by construction),
tests/helpers.cjs's cleanup() for temp-dir removal, and a line-by-line fence
scan using the text-lines.cts splitLines() seam, matching this repo's
established extraction idiom (tests/no-hardcoded-home-gsd-tools.test.cjs).
No behavioral change to what is asserted.

* docs(#4112): add changeset for the pause-work Context Detection fix

* docs(#4112): backfill changeset PR number (pr:0 -> pr:4140)

---------

Co-authored-by: sim <sim@local>
2026-09-01 13:03:33 -04:00
Tom Boucher
05092ff369 fix(#4113): add no-op to empty then-body in chunked-planning-mode.md's outline resume-check (#4125)
* fix(#4113): add no-op to empty then-body in chunked-planning-mode.md's outline resume-check

The `### 8.5.1 Outline Phase` resume-detection `if` block only contained a
comment in its `then` clause, which is a syntax error under both bash and
zsh if executed literally. The workflow ShellCheck lint added by #4109
already flags this (SC1009/1048/1072/1073/2105) but the findings were
masked by matching entries in the pre-existing-findings baseline; those
five entries are removed here so a regression re-surfaces as a new
failure instead of being silently re-absorbed.

* chore: backfill changeset PR number for #4125

* fix(#4113): fix second empty-then/continue-outside-loop defect surfaced by baseline cleanup

Removing the 5 stale baseline entries for chunked-planning-mode.md in the
prior commit unmasked a second, genuinely separate ShellCheck finding
(SC2105) in a different fenced block of the same file: a bare `continue`
outside any bash loop in the ### 8.5.2 per-plan resume-check. Same root
cause class as the original bug (pseudocode fenced as literal bash) —
fixed the same way, with a `:` no-op preserving the documentation.

---------

Co-authored-by: sim <sim@local>
2026-09-01 00:55:44 -04:00