Commit Graph

689 Commits

Author SHA1 Message Date
sim
cd58aaabf4 refactor(#4653): drain the containment duplicates and record the two rulings
Phase 3 of epic #4636, stage 3c. ADR-4650 decision 6: a wrapper may decide HOW
to degrade, never WHETHER a path is contained. Four implementations are drained
on that rule; two are retained, with the reasons recorded rather than assumed.

DRAINED — the containment decision now comes from the canonical predicate:

  scripts/check-glossary-refs.cjs   local isWithinRoot deleted outright.
  src/installer-migrations.cts      ensureInsideConfig keeps its throw and its
                                    lexical fullPath; only the decision moves.
  src/planning-inspect.cts          isPathContained keeps must-exist as its own
                                    condition; only the decision moves.

Two of those are wrappers rather than deletions, and each is a wrapper for a
reason that would have been a silent behavior change if collapsed naively:

- `isPathContained` returns FALSE for a path that does not exist, because
  fs.realpathSync throws ENOENT and its catch swallows it. The canonical
  predicate does the opposite: for a missing target it walks up to the nearest
  existing ancestor and ACCEPTS a not-yet-created path under the root. Its
  callers at planning-inspect.cts:747 and :839 guard a phaseDir immediately
  before readdirSync, so under a naive swap a missing phaseDir would stop
  reporting scope UNREADABLE and start throwing ENOENT out of readdirSync.
  Existence is therefore kept as an explicit local requirement.

- `ensureInsideConfig` returns a LEXICAL fullPath that both callers consume for
  existsSync and for journal entries. The canonical predicate realpath-resolves,
  so if configDir is itself a symlink the two differ. The decision is canonical;
  the returned value stays lexical. Its message is likewise preserved verbatim,
  which is why this uses tryWithinRoot plus an explicit throw rather than
  assertWithinRoot.

`isWithinRoot` in planning-inspect is left in place and documented: it is a pure
comparison over paths the CALLER has already resolved, which readDocument does
inline specifically to keep a third degradation shape (exists-but-unreadable vs
absent) that neither isPathContained nor the canonical predicate expresses. It
is the comparison step of one implementation, not a second implementation.

RETAINED, DELIBERATELY — gsd-core/bin/gsd-tools.cjs. My own design document said
"collapse" and that was wrong. The file carries an explicit comment forbidding
it, and the comment is correct: its three checks reject symlinks OUTRIGHT, which
is strictly stricter than the canonical predicate, not a reimplementation of it.
The canonical predicate accepts a link whose target lands inside the root — for
a restore that is still wrong, because writing through the link overwrites
whatever it points at instead of materializing a regular file. Collapsing would
have reintroduced that hole. The comment is updated to name the current exported
predicate, to record that this was reviewed under this phase and deliberately
not collapsed, and to note that isInsideDir treats target === root as NOT
contained — the one implementation in the repo that does.

THE configHome RULING — retained lexical, and a false safety claim corrected.
isPathConfined stays lexical because two of its callers must validate a
destSubpath BEFORE the mkdirSync that creates it (install-engine.cts:1608,
install-profiles.cts:880), where realpath cannot resolve and a realpath-based
predicate would reject every legitimate install.

Its docstring's justification, however, did not survive being checked. It cited
capability-source.cts:491,577,675 as the upstream symlink rejection that made
the lexical form safe. Read directly: :491 is a blank line before assertSafeId's
JSDoc and :577 is an entry-count budget check. Neither is a symlink check. The
real guards are :585-586 and :671-674. Worse than stale line numbers, the claim
that this "keeps every caller of this function's callers symlink-safe" is false:
that rejection lives in capability-source's staging path and covers only the
capability-loader route to assertDescriptorConfined. Three other callers do not
reach it, and only retired-artifact-cleanup.cts:69 carries its own defense
(its lstatSync check at :77). The docstring now states what is actually true and
cites the lines that actually exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 13:29:54 -04:00
Tom Boucher
a2331c01f1 fix(#4568): widen the phase-number regex to accept N-segment ids at 6 shell/markdown sites (#4646)
* test(#4568): pin the N-segment phase-grammar defect across all 6 shell/markdown sites

Manually traced against the current tree: the validating regex at
code-review.md rejects a 3-segment id (23.1.2), and execute-plan.md's
extraction truncates a 23.1.2-01-PLAN.md filename down to 1.2-01.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4568): widen the phase-number regex to accept N-segment ids at all 6 shell/markdown sites

Widens `?` to `*` on the dotted-segment group at all 6 sites (byte-identical
behavior for 1- and 2-segment ids, character class unchanged): code-review.md,
code-review-fix.md, gsd-code-fixer.md, gsd-code-fixer.compact.md (validating
sites, plus their comment/error-message text), execute-plan.md's plan-filename
extraction, and plan-phase.md's --research-phase flag capture.

Also disambiguates the nsegment-phase-grammar test's plan-phase.md anchor,
which was matching an unrelated earlier `--research-phase` occurrence (line
77's generic-value capture) instead of the targeted site (line 131).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): extend lint-phase-id-drift to ban the single-segment phase regex in workflows/ and agents/

Adds findSingleSegmentPhaseRegexDrift, banning the bounded
`[0-9]+(\.[0-9]+)?` shape (and its \d/doubled-backslash near-variants) on any
phase-carrying line across gsd-core/workflows/**/*.md,
gsd-core/references/**/*.md, and the newly-scanned agents/**/*.md, sanctioned
the same way as the existing shell-arith rule. Wired into scanAll; confirmed
zero violations against the real tree post-#4568 fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4568): add Fixed changeset

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: regenerate conformance-tier manifests for the new test file

The emitted-attribution gate also flags 4 files growing: code-review-fix.md
(+21 bytes), code-review.md (+21 bytes), gsd-code-fixer.compact.md (+9
bytes), gsd-code-fixer.md (+6 bytes). The growth is the fix itself: each
site's validation regex widened from a bounded single-optional-dotted-segment
shape to the unbounded form, and the accompanying comment/error-message text
grew by a few characters to mention the new 3-segment example.

Emitted-Drift-Ack-Growth: code-review-fix.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Emitted-Drift-Ack-Growth: code-review.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the error text (#4568)
Emitted-Drift-Ack-Growth: gsd-code-fixer.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4568): backfill changeset pr number to 4646

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 17:04:28 -04:00
Tom Boucher
db4d8a9bae fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic (#4644)
* fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic

$((10#${PHASE_NUMBER})) is a hard bash/zsh syntax error when PHASE_NUMBER is
decimal (01.1, from an inserted phase) or N-segment (23.1.2) — neither is
valid shell-arithmetic syntax at all, and the failed expansion aborts the
rest of the snippet in a non-interactive shell. safe_resume_gate runs
unconditionally before trusting STATE.md or dispatching any executor, so
execute-phase failed at its own gate before the first executor on any
decimal phase, regardless of workflow.tdd_mode. Regression from #4194.

Fixes all 4 sites: safe_resume_gate and the TDD gate in
workflows/execute-phase.md, the completion-signal spot-check fallback in
workflows/execute-phase/steps/completion-reconciliation.md, and the
executor gate validation example in references/tdd.md. Each now zero-strips
only the leading integer segment into a *_INT variable (via %%.* / #
parameter expansion — always valid shell syntax regardless of what follows)
and keeps the remainder as an escaped-dot string for the anchored commit-
scope regex, exactly as issue #4619 verified in both bash and zsh. A plain
integer phase (12, 01) computes byte-identically to before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4619): pin the decimal/N-segment fix and characterize the pre-fix bug

Behavioral coverage via real bash execution: the old $((10#01.1)) form
throws (characterizes the bug, matching the issue's own reproduction); the
new form resolves 01.1 -> 1\.1 and 23.1.2 -> 23\.1\.2, unchanged for plain
integers (12 -> 12, 01 -> 1); the resulting anchored ERE matches
feat(01.1-03):/test(1.1-3): and correctly rejects feat(01-03):,
feat(01.2-03):, feat(011-03):, feat(12-03): for a decimal phase — mirroring
issue #4619's own verified table exactly. Updates
safe-resume-gate-anchoring.test.cjs's 4 existing source-text assertions
(one per site) to the new fixed text.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): refine the shell-arith drift detector to distinguish safe from unsafe arithmetic

With #4619's fix in place, the guard's original "ban $((10#... outright,
match any occurrence" was too blunt: it flagged a comment merely mentioning
the pattern in prose, the now-safe $((10#$PHASE_INT)) arithmetic on an
already-%%.*-stripped integer, and the always-safe plan-id arithmetic
(plan ids are plain integers, never decimal). Refines the detector to skip
full-line comments and to only flag a captured variable/placeholder name
that contains "phase" and does NOT end in _INT/_int — the naming convention
the #4619 fix establishes at all four sites for "already reduced to a safe
integer." A plan-id variable was never phase-number arithmetic in the first
place and is excluded on the same basis.

This closes epic #4634's D6 ("lint-phase-id-drift... passes with no new
exemptions") and D7 ("a decimal and N-segment phase id survive an
end-to-end execute-phase selection without error") for real — the guard now
reports zero violations across all five .cts/.md rules.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: regenerate conformance-tier manifests for the new test file

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4619): cover the plain-padded-integer near-miss matrix too

Review found the anchored-ERE near-miss coverage only exercised the
decimal case (PHASE_NUMBER=01.1); issue #4619's own worked table also
verifies the plain padded-integer case (01 -> PHASE_N=1) against its own
near-miss set (matches 01-03, rejects 01.1-03/011-03/12-03). Adds the
missing assertion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4619): add Fixed changeset

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4619): correct JS backslash-escaping in safe-resume-gate anchoring test

The test's string-literal assertions for the PHASE_FRAC//./\\.} pattern wrote
only 2 backslash characters in JS source, which single-quoted-string parsing
collapses to 1 real backslash at runtime -- but the workflow/reference files
actually contain 2 raw backslash bytes at that position (needed so bash's
${var//pattern/replacement} produces the correct single-backslash output).
Write 4 backslash characters in the JS source at all 4 occurrences so the
runtime string matches the files' real bytes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4619): refresh the committed compact-content benchmark baseline

The new PHASE_INT/PHASE_FRAC arithmetic lines added to
gsd-core/workflows/execute-phase.md shifted its committed compaction-ratio
baseline. Regenerate via `node scripts/benchmark-compact-content.cjs --write`.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4619): note the safe_resume_gate arithmetic growth in the test header

The emitted-attribution gate flags execute-phase.md growing 91253 -> 91846
bytes (593 bytes). The growth is the fix: the safe_resume_gate and TDD RED
block now derive PHASE_INT/PHASE_FRAC before computing PHASE_N, so a
decimal/N-segment phase number (e.g. 01.1, 2.3.1) zero-strips its leading
integer segment via base-10 arithmetic instead of forcing the whole value
through $((10#...)) and hitting a hard shell syntax error on the first dot.

A blank line previously separated the Emitted-Drift-Ack-Growth trailer from
the Co-Authored-By trailer below it, which splits git's trailer-block
detection: only the last contiguous non-blank run of Key: Value lines at the
end of a commit message is recognized as trailers, so the growth ack was
silently read as ordinary body text and the differential-attribution gate
failed with the growth unacknowledged. Joining the two trailers into one
contiguous block fixes it.

Emitted-Drift-Ack-Growth: execute-phase.md — adds PHASE_INT/PHASE_FRAC derivation to the safe_resume_gate and TDD RED commit-scope grep so a decimal/N-segment phase number zero-strips its leading integer segment via base-10 arithmetic instead of failing on a non-numeric value (#4619)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4208): replace chmod-based restore-failure injection with a root-proof git shim

`tests/commit-files-deletion.test.cjs`'s two restore-failure tests simulated
an unwritable index via a `post-index-change` hook running `chmod a-w` on
the git dir. That relies on the OS enforcing the *owner's own* permission
bits against itself, which uid 0 (a routine identity inside this repo's
Docker-based gsd-test benches) does not: every DAC check short-circuits true
for root, so the write the chmod meant to block silently succeeds, the
restore comes back clean, and the disclosure/rollback behavior under test
never actually gets exercised.

This is CLAUDE.md's own named anti-pattern for I/O-failure injection
("Cross-platform test IO-failure injection" — chmod tricks fail under root
Docker/CI). It is confirmed as the actual root cause here, not a production
defect: `src/commands.cts`'s `restoreRemovedEntries`/rollback-disclosure
logic (added by #4253, merged just before this run) was hand-traced and
manually reproduced end to end on an unprivileged workstation against a
freshly built `gsd-core/bin/lib/commands.cjs`, and it already produces
exactly the `staging_failed` + "could not be restored" / "could NOT be
restored during rollback" results both tests assert. The other
`post-index-change`-based tests in this file (a `sleep` to force a timeout;
a real `update-index` to flip a restored entry's mode) are unaffected
because neither depends on a permission check — consistent with only the
two chmod-based tests failing on the real remote run.

Replaces the chmod fixture with a fake `git` placed ahead of the real one on
PATH that fails only `update-index --add --cacheinfo` — the one call the
restore makes — unconditionally, regardless of privilege level. Every other
git invocation execs straight through to the real binary, so the rest of
each scenario (`rm --cached`, the restore's own `ls-files` verification,
etc.) is exercised exactly as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4619): backfill changeset pr number to 4644

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4619): feed the bash fixture script via stdin, not argv, to fix Windows CI

Passing the script as a `-c "<script>"` argv element made it subject to
Windows' CreateProcess command-line argument encoding, which silently
dropped the escaped-dot backslashes before bash ever saw them (observed on
PR #4644's windows-latest CI shard: `1\.1` came back as `1.1`). Feeding the
same script via stdin instead removes argv entirely from the transport, so
there is nothing for Windows to re-encode. POSIX behavior is unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 15:47:14 -04:00
0xdhx
4cc2a466b5 fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec (#4253)
* fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec

`cmdCommit`'s `--files` list can stage an addition but never a deletion:
the #2014 guard skips a missing explicit entry because the filesystem
cannot tell "moved away" from "not written yet". A caller that moves a
file therefore had two forms, both wrong — a directory entry records the
move but also commits every unrelated file in that directory (a
concurrent session's in-flight todo, in the unattended execute-phase
sweep), and a file entry leaves the old path's deletion dangling with the
todo tracked at both paths.

`--files-removed <paths>` is the caller-declared delete intent. Each entry
names a file, or a directory whose tracked-but-absent files are the
removals; those paths are staged with `git rm --cached` and join the
commit pathspec. `--files` keeps its skip-if-missing contract untouched.
A file entry still present on disk fails the commit closed with the
existing staging-failure rollback; a never-tracked path is a no-op.
`--files-removed` alone is a declared scope, not the unscoped .planning/
sweep.

The dispatcher previously folded every non-flag token after `--files`
into that list, so a second list flag could not exist; each list now
runs from its flag to the next `--` token.

The execute-phase todo sweep names the moved todos on both sides from
CLOSED[@], and cleanup's archive commit moves .planning/phases/ and
.planning/quick/ under --files-removed.

Fixes #4208

Emitted-Drift-Ack-Growth: cleanup.md — the archive commit moves phases/ and quick/ under --files-removed; the growth is one paragraph stating why those two directories must not be --files entries

* chore(#4208): set changeset fragment pr to 4253

* fix(#4208): fit execute-phase.md under the ADR-857 ceiling and re-point the #2415 guard

Three CI failures, all consequences of this PR's own change.

1. gsd-core/workflows/execute-phase.md was 93,577 bytes against the
   ADR-857 Phase 6 margin gate's <= 93,400 (hard ceiling 93,600). The
   three-line rationale comment plus the four-line array-building block
   added 318 bytes to a file that had only 141 of headroom on next.

   Move the rationale to docs/CLI-TOOLS.md -- which this PR already
   extends with the --files-removed contract, and which is where the
   ADR-857 gate wants call-site detail to live rather than in the host
   workflow -- and fold the array build onto one line. 93,577 -> 93,372.

2/3. tests/close-phase-todos-stage-deletion.test.cjs pinned the #2415
   guarantee to its old MECHANISM: it regex-matched the literal
   .planning/todos/{completed,pending}/ directory pathspecs in the
   commit --files list. This PR deliberately replaced those with named
   files (a directory entry also committed an unrelated todo a
   concurrent session dropped in mid-close), so the guard failed on a
   change it should have accepted.

   Re-point it at the new mechanism without weakening it: assert the
   ADDED array reaches --files, the REMOVED array reaches
   --files-removed, STATE.md is still committed, and -- newly -- that
   the two arrays are built from $COMPLETED_DIR and $PENDING_DIR
   respectively. Verified by negative control: deleting
   --files-removed "${REMOVED[@]}" from the workflow still fails the
   test, so the #2415 regression remains caught.

Note for the merge queue: #4233 also grows execute-phase.md (+114). The
two are additive -- different regions, no textual conflict -- so with
both landed the file reaches ~93,486, over the 93,400 margin though
under the 93,600 hard ceiling. Whichever merges second will need to
reclaim ~86 bytes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0183892Y3fxxirte4WNmBKbv

* fix(#4208): reclaim execute-phase.md bytes so the PR is net-neutral under the ADR-857 margin

Rebasing onto next surfaced the byte-gate collision flagged earlier on
this PR: #4284 grew execute-phase.md by 95 bytes (93,259 -> 93,354),
so this PR's +113 landed at 93,467 against the <= 93,400 margin in
tests/claude-orchestration.test.cjs.

Compact the close_phase_todos step this PR already edits -- drop the
PHASE_NUM indirection, fold the normaliser and the match guard, print
the closed list with one printf, shorten the step's prose -- without
touching the mechanism the #2415 guard pins (ADDED/REMOVED arrays, the
plain mv). 93,467 -> 93,349: 5 bytes under the base, so the PR no
longer spends any of next's 46 bytes of headroom.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): classify absent index entries before staging a removal; restore removed entries exactly on rollback

Review of #4253 found three Majors with one root cause: the removal
side judged presence by fs.lstatSync alone, where the addition side
already reads `git ls-files -v` state. Absence from the worktree is not
removal:

- a submodule gitlink (mode 160000) whose directory was deleted by hand
  lists like a file and was `rm --cached` with no .gitmodules cleanup;
- a skip-worktree path is never materialised by a cone-mode sparse
  checkout, so a directory entry over a sparse-excluded tree dropped
  that whole tree from the index;
- an assume-unchanged path's worktree state is not something git
  itself consults;
- an intent-to-add entry (`git add -N`) renders as a plain cached entry
  on the empty blob, yet nothing tracked exists to remove and no
  rollback can restore the flag.

The index listing now carries each entry's `ls-files -v -s` tag, mode
and stage. Only a plain cached (H), stage-0, non-gitlink entry is a
removal candidate; every other state is left alone under a directory
entry (exactly like a present file) and fails closed when named
directly, with the state in the error. "Named directly" is decided on
RESOLVED paths, not strings -- realpath of the longest existing prefix
with the absent tail re-appended: an absolute path, `./x`, `--cwd`, or a
symlinked spelling of the tree (macOS `/var` ->
`/private/var`, where `process.cwd()` is the real path and the caller's
absolute path is not -- CI on this round's first push) all resolve to the
same entry, where a string compare against git's cwd-relative output
silently took the directory polarity (pre-push review, driven; the
symlink case is driven with an aliased fixture directory). The enumeration's domain is what
`ls-files -v -s` can emit for an index entry, stated at the classifier.

The third Major -- on an unborn HEAD a successful `rm --cached` was
never rolled back when a later entry failed -- is fixed differently
from the review's suggestion. Pushing the path into stagedPaths would
put it on the commit pathspec, which a root commit refuses ("pathspec
did not match", driven), and `git reset -- <path>` cannot restore an
entry with no HEAD anyway. Instead every index entry this call removes
is recorded (mode, blob) before the `rm` and put back with
`update-index --cacheinfo` on rollback. That also restores a
caller-pre-staged blob at a removed path exactly, where a reset would
have silently replaced it with HEAD's version. The rollback is
best-effort, as the addition-side reset already was, and the docs say
so.

Eight tests: gitlink under a directory entry, named directly, and named
by absolute path; skip-worktree both forms; intent-to-add both forms;
assume-unchanged named; unborn-HEAD partial failure restores the
removal; pre-staged blob survives the rollback.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): drop the empty fenced block left dangling in cleanup.md's commit step

Review nit on #4253: inserting the --files-removed rationale between the
original bash block and its closing fence left an empty ```bash``` pair
before </step>. Harmless at runtime, a formatting artifact of this PR's
own diff; removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): a boolean flag inside a commit path list no longer ends the list

Review minor on #4253: collectList stopped at the next `--` token, so a
positional wedged between a boolean flag and the next list flag
(`--files a --amend b --files-removed c`) was claimed by neither list
and silently dropped -- a regression in shape against the old
slice-to-end parse, which filtered `--` tokens and kept `b`. No current
call site interleaves that way, but the gap was real.

A list now runs to the next LIST flag (`--files` / `--files-removed`)
and skips boolean flags on the way, and a REPEATED list flag merges
its runs (`--files a --files b` -> [a, b]) as the slice-to-end parse
did -- a first cut stopped at the repeat and dropped `b`, the same
silent-drop shape one level over (pre-post comment audit). The only
change #4208 makes to parsing is that a second list flag can exist.
Tests: STATE.md wedged between --no-verify and --files-removed lands
in the commit; both runs of a repeated --files reach it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* test(#4208): drive the reappearance window with a post-index-change hook

Review nit on #4253: the defensive re-check for a file recreated between
the absence test and `git rm --cached` -- the concurrent-session race
this PR's own changeset names -- had no test. git fires
post-index-change the moment `rm --cached` writes the index, so a hook
that copies the file back exactly then exercises the window
deterministically. The call reports staging_failed / "reappeared on
disk", commits nothing, and the rollback restores the removed entry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): restore a staged removal when the call records nothing

A `git rm --cached` that succeeds mutates the index whether or not a commit
follows. Only the staging-failure rollback put those entries back, so a call
that reached `nothing_to_commit` reported no state change while the removal sat
staged -- riding along on the caller's next commit.

The review named the unborn-HEAD, removal-only shape. Keying on `headExists`
would have fixed half of it: the guard also fires with a real HEAD when the
removed path is index-only (added, never committed), because `diff HEAD` reads
clean with the path absent on both sides. Both shapes now restore, at both
`nothing_to_commit` exits. The failure exits are deliberately left alone --
they report a failure rather than no-change, and the addition side leaves its
own staged paths there too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* refactor(#4208): lift declared-removal staging out of the cmdCommit hotspot

`cmdCommit` was a critical-risk hotspot before this flag existed, and #4208 had
inlined another ~270 lines into it. `stageDeclaredRemovals(cwd, removedDeclared)`
now owns the index-state classification, path canonicalisation and entry
recording, returning the pathspec entries and the recorded removals its caller
merges.

Pure motion: no branch, message or probe changed. Only the two accumulators
became local names, and `restoreRemovedEntries` stays with the caller because
the exits that restore are the caller's. cmdCommit 888 -> 625 lines here; the
extracted helper is 277.

(Figures corrected after publication: an earlier version of this message said
854 -> 591 and claimed the result was below cmdCommit's pre-#4208 shape. Both
were wrong -- the count came from a faulty brace scanner, and `next`'s cmdCommit
is 581, so this is above it, not below.)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): property-test the two-list commit parser

RULESET.TESTS.property-based-testing asks a parser for at least one property
test asserting a domain invariant; `collectList` had only hand-picked examples,
one per shape a review round had already broken.

Hoisted it to module scope as `collectListFlagValues` and exported it in the
file's existing exported-for-tests convention -- a parser reachable only by
spawning the CLI can be tested one example at a time and no faster.

Three properties over generated argv: every positional lands in exactly the run
open at it whatever the flag order or count; no positional after the first list
flag is dropped or double-claimed; and with `--files-removed` absent the parse
equals the pre-#4208 slice-to-end parse. Controlled against two mutants -- a run
ending at any `--` token, and a repeated list flag that does not merge -- each
of which the properties catch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): pin cleanup.md's archive commit to --files-removed

execute-phase.md's rewrite is pinned by the #2415 guard in this file;
cleanup.md's equivalent was not, so reverting its routing would have been
caught by nothing -- the mechanism's unit tests never read this file and pass
either way.

Asserts the two archived directories are under --files-removed and NOT under
--files (where a directory entry sweeps in a concurrent session's in-flight
writes), and that the destinations and STATE.md stay on the additive half.
Controlled by restoring the pre-#4208 sweep, which fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): pin that a symlink to a directory is one tracked path

Review of #4253 read the `lstatSync(...).isDirectory()` test as a
symlink-following defect. Driving it says the opposite: git tracks the link as
a single blob (mode 120000) and does not traverse it, so the tracked paths
"under" it live at the real directory and were never named by the caller.
Following the link would stage those -- the directory sweep #4208 exists to
remove -- while the named entry still sat present on disk.

Pinned rather than changed, with the premise driven in the test body. Swapping
`lstatSync` for `statSync` -- the prescription as written -- fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4208): refresh the compact-content baseline for this PR's execute-phase edit

The base range added `tests/benchmark-compact-content.test.cjs` and a committed
token baseline over the compacted workflows. This PR edits
`gsd-core/workflows/execute-phase.md`, so the baseline drifts by +12 tokens on
that entry and on the aggregate.

Refreshed with `node scripts/benchmark-compact-content.cjs --write`; the diff is
those two entries and nothing else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): report a removal the call could not put back

Round review of this round found the restore itself unchecked: the helper
ignored `update-index`'s exit code, so a FAILED restore still reported
`nothing_to_commit` -- the same false "no state changed" the restore exists to
prevent, surviving one level down on the restore-failure path.

It now returns a boolean. The two no-change exits report `staging_failed`
naming the paths left staged; the staging-failure rollback still ignores it,
deliberately, because it is already reporting a failure and an unwritable index
is usually the failure being reported.

Driven with a post-index-change hook that makes the git dir unwritable the
moment `rm --cached` lands, so the restore cannot take its lock. Reverting both
guards fails the test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): disclose a removal the rollback could not restore

Round review refuted the reasoning behind leaving the rollback path's restore
unchecked. The claim was that this exit is already reporting a failure, so the
restore's result adds nothing. The counterexample is the ordinary case: the
reported failure is usually a DIFFERENT cause -- a contradictory declaration, a
reappeared path -- so a caller reading `failures` sees only that cause and
learns nothing about the removal still sitting in its index.

The rollback now appends a disclosure entry per un-restored removal, naming the
path. The reason and `file` still report the failure that caused the rollback;
the disclosure is additive.

Also moves the restore-failure test's chmod into a `finally`: `t.after` runs
AFTER the parent `afterEach`, so a throw before it left the fixture undeletable.

Both driven; reverting the disclosure fails the new test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): decide index state by observation, never by an exit code

The restore added two commits earlier keyed both its record decision and its
success verdict on git's exit code. An exit code answers "did the command
succeed", never "did the index change" -- execGit collapses a spawn timeout to
a non-zero exit, and a killed git can already have written the index. Round
review drove four failures from that one assumption, in both directions:

  - a failed `rm` still contributed an entry, so the rollback disclosed a
    removal that was never staged (stale index.lock);
  - a timed-out `rm` whose write DID land contributed none, so a real mutation
    was neither restored nor disclosed;
  - a timed-out `update-index` whose write landed reported failure, publishing
    a "could NOT be restored" disclosure that was false;
  - and the read-back that replaced it omitted `-z`, so core.quotePath rendered
    `café.md` as `"caf\303\251.md"` and an exactly-restored entry read as not
    restored -- the same quoting defect this PR already fixed for `preStaged`.

Everything now observes the index. A failed `rm` re-reads `ls-files -z` for the
path: gone means this call owns the removal and records it; still there means
nothing was staged; a probe that cannot answer becomes its own failure entry
rather than an assumption. The restore verifies the same way, comparing the
WHOLE entry (mode, blob, stage), because `--cacheinfo` restores all three and a
path-only test accepts an entry that came back as something else.

The verdict is three-valued -- `restored` / `not-restored` / `unverified` --
and the unverified wording says the restore could not be VERIFIED rather than
that it failed. The rm's own failure is pushed ahead of any probe diagnostic so
a timed-out removal keeps `timed_out: true` and its own message as the reported
cause.

Five regression cases, each negative-controlled against the shape it pins.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): treat a declared removal path as a path, not a pathspec

An index path handed back to git is parsed as a PATHSPEC, and the removal side
handed several back. Three driven harms, all of them the sweep-in this flag
exists to remove, arriving through the operand rather than through a directory
entry:

  - a tracked file literally named `.planning/*.md` made `rm --cached` GLOB: it
    removed `peer.md` and `stays.md` too, only the declared entry was recorded,
    so the rollback restored one of three and the other two rode out as staged
    deletions the result disclosed nowhere;
  - the same name reached `git commit -- <paths>`, which globbed and committed
    an undeclared `M peer.md` alongside the declared removal;
  - and the intent-to-add probe (`diff --cached` over the path) matched a
    STAGED PEER instead of itself, so an `add -N` entry was misclassified as
    ordinary content, removed, and restored by `--cacheinfo` -- which cannot
    restore the intent flag. It came back as a real staged addition.

Every operand on this path is now `:(literal)`: the `rm`, both index probes,
the intent-to-add probe, the restore read-back, the entry-level `ls-files` /
`ls-tree`, and -- for the REMOVAL-derived entries only -- the downstream
`ls-files` / dry-run / `diff HEAD` / `commit` pathspec. `--files` entries keep
whatever pathspec behaviour they have today; that is not this change's to
alter. `:(literal)` still resolves a directory to its descendants (driven), so
the directory form is unchanged.

Closes what an earlier cut of this commit declared as a residual: a filename
beginning with `:` is now removable end to end, because the commit pathspec no
longer reinterprets it.

Also fixes a MINOR from the same review: cleanup.md's contract test checked the
destinations' position relative to `--files-removed` but never that `--files`
was present at all, so deleting the flag still passed.

Un-literalising the seven sites fails three of the new tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): scope the rollback to the caller's own name space

Round review drove a rollback that destroyed the caller's own staged work. Two
causes, one of them pre-existing:

  - `git diff --cached` prints REPO-relative paths whatever the cwd, while
    `stagedPaths` holds the caller's cwd-relative names. In a project nested
    inside its repo (`<repo>/sub/.planning/...`) the two name spaces never
    intersect, so `preStaged` matched NOTHING, every path landed in `toUnstage`,
    and the reset unstaged a caller-staged deletion and modification that this
    call had never touched. `--relative` makes the two sets comparable, and is a
    no-op when the project IS the repo root. This governs the `--files` side too
    and predates this flag.
  - the rollback's `reset` was the last place a removal-derived name reached git
    as a bare pathspec; it takes `asPathspec` like every other site.

Driven on a nested fixture; dropping `--relative` fails the new test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): gate six fixtures that Windows cannot construct

CI's `test (windows-latest, 24, shard 2/3)` went red on this round. Two
primitives the new fixtures rely on do not exist on Windows, both driven on a
real Windows host rather than inferred:

  - a filename containing `*` or `:` cannot be created at all (`IOException` /
    `FileNotFoundException`), which is four of the pathspec fixtures;
  - `chmod` cannot make a directory unwritable — a write into a ReadOnly
    directory succeeds — so the two restore-failure fixtures cannot drive the
    failure they exist to drive.

Each is skipped on win32 with its measured reason, in the repo's existing
`{ skip: process.platform === 'win32' ? '<reason>' : false }` form. The
behaviours they pin are platform-independent; only the fixtures are not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): build git's index-syntax path with forward slashes

The remaining Windows red was mine, not the platform's: `git rev-parse :<path>`
takes a forward-slash path, and `path.join` yields backslashes there, so git
rejected it as an ambiguous argument. The hook in the same test already used
the slash form.

Not gated — the behaviour it pins is portable; only the argument was not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4208): refresh the compact-content baseline against the rebased base

`next` moved the `new-project` split and the aggregate under this PR's
execute-phase entry; regenerated with `scripts/benchmark-compact-content.cjs
--write` so the only leaves differing from the base's copy are the
execute-phase split and the aggregate it feeds.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh

* chore(#4208): regenerate the macOS conformance tier for this PR's fixtures

`next` gained the macOS-specific conformance tier (#4593) after this branch
was cut. Its classifier (`scripts/gen-platform-conformance-tier.cjs --target
macos`) now selects `tests/commit-files-deletion.test.cjs` on the
`chmod-mode-bit` and `symlink-keyword` signals the PR's fixtures carry (the
chmod-driven failed-restore cases and the symlink-to-directory case).
Regenerated with `--target macos --write`; the platform tier was already in
sync. The file was modified, not added, which is why the added-files check
did not surface it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-11 12:26:55 -04:00
Tom Boucher
523be34133 fix(#4282): register PATTERNS.md as a canonical .planning/ artifact (#4618)
* test(#4282): prove PATTERNS.md is unrecognized by the artifact registry

Regression test only, no fix yet: CANONICAL_EXACT in src/artifacts.cts was
never updated when workflows/graduation.md started writing .planning/
PATTERNS.md, same omission class as the already-fixed #3224 (WINDOWS.md).

Expected RED on this commit (src/artifacts.cts is unchanged).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4282): register PATTERNS.md as a canonical .planning/ artifact

CANONICAL_EXACT in src/artifacts.cts was never updated when
workflows/graduation.md started writing .planning/PATTERNS.md for the
`patterns` graduation-target category -- same omission class as the
already-fixed #3224 (WINDOWS.md). validate.health's W019 falsely flagged it
as unrecognized on every repo that has run the graduation scan.

Also backfilled 5 other pre-existing stale rows in
gsd-core/templates/README.md's artifact table (WINDOWS.md, STATE-ARCHIVE.md,
milestone.lock, state.json, skill-manifest.json) that were already in the
source registry but missing from the docs table -- found while fixing this
exact drift class, cheap to close alongside it.

RED proven on f3dd791fb8cda18196803e7144ce20e506d6490b (test-only commit,
gsd-test outcome:failed, exactly the new PATTERNS.md test failing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4282): fix stale function name in skill-manifest.json comment

Review finding: both the source comment and the new docs row said
"routeSkillManifest" -- no such symbol exists (verified via Memtrace); the
actual function is cmdSkillManifest (src/init.cts). Copied verbatim from a
pre-existing comment, not introduced by this PR, but cheap to fix alongside.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4282): add changeset fragment

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4282): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: isolate lint-vendored-deps-manifest.test.cjs's fixRow tests from the real vendor file

Genuine, pre-existing defect found and fixed per this repo's no-defer policy
(discovered while investigating a real CI failure during this PR's own
merge attempt, user-directed investigation -- not deferred to a separate
issue since it was actively blocking work and root-caused with concrete
evidence, not speculation).

Root cause: fixRow(row) (scripts/lint-vendored-deps.cjs) unconditionally
does fs.copyFileSync(upstreamCjs, vendoredCjs) as its first line. All three
tests in the #4573 describe block called fixRow(row) with the REAL js-yaml
row, so all three wrote to the real, shared gsd-core/bin/lib/vendor/
js-yaml.cjs -- a file other test files' require() calls can read at any
moment, since node --test runs files concurrently in this repo.
fs.copyFileSync's write is not atomic against a concurrent reader on every
filesystem; a concurrent require() elsewhere caught the file mid-overwrite
and read a truncated file, crashing an entirely unrelated test
(m9-statelock-write-error-orphan.test.cjs) with a SyntaxError.

Confirmed via two real CI log fetches, not assumed: the exact same shard
grouping (same 308 files) ran clean ~90 minutes earlier during PR #4615's
own final merge CI, with the identical #3660 reap-fix code already present
-- ruling out a deterministic connection to that change and confirming a
genuine, non-deterministic timing race in this pre-existing test design.

Fix: all three tests now redirect row.vendoredCjs to a private os.tmpdir()
path via a cloned row object before calling fixRow, so the real vendored
file is never touched. upstreamCjs stays pointed at the real node_modules
copy (read-only, safe to share).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: also isolate fixRow's package.json pin-rewrite from the real file

Review finding (major) on the previous race-condition fix: fixRow's pin
-rewrite path still hardcoded path.join(ROOT, 'package.json'), so the third
#4573 test still wrote the real, shared package.json -- read at module
top-level by dozens of other test files, the same concurrent-file race
class already fixed for the vendored .cjs copy.

Adds an optional pkgRoot parameter (defaults to the real ROOT) threaded
through readPinState/checkRow/fixRow -- fully backward-compatible, every
existing call site (the CLI --fix path, any other caller) is unaffected
since the default is unchanged. The pin-rewrite test now builds an isolated
temp root (its own package.json + node_modules/js-yaml/package.json) and
passes it explicitly, so the real package.json is never touched either.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: use helpers.cleanup instead of raw fs.rmSync in test cleanup

CI caught it: local/no-raw-rmsync-in-tests flagged the three t.after temp-dir
cleanup calls added for the fixRow isolation fix. helpers.cleanup() carries
the Windows-EBUSY retry budget (maxRetries/retryDelay) that raw fs.rmSync
lacks.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 22:46:18 -04:00
Tom Boucher
2cefa5a5ac enhance(#4139): Phase 8 — the toggle becomes discoverable, and the ledger closes (#4587)
* enhance(#4139): Phase 8 — the toggle becomes discoverable, and the ledger closes

ADR-4139's final phase. workflow.compact_content already defaulted to false
(Phase 1's buildNewProjectConfig hardcoded default), but nothing surfaced it:
/gsd-new-project never asked, and /gsd-settings/config had no toggle path for
an already-initialized project — config-set/config-get were the only route.

new-project.md gains a fourth question in the existing Round 2 AskUserQuestion
array (grouped with the other general-workflow-behavior toggles, not the
per-agent capability questions above it) and threads compact_content into the
config-new-project CLI JSON literal. settings.md mirrors the exact pattern
every other non-capability workflow.* key already follows: read_current bullet,
question block, update_config write, the safe-merge non-capability-keys list,
save_as_defaults, and the confirm summary table — seven edits, zero new
src/*.cts code, since Phase 1's merge logic is a generic passthrough. Its
success_criteria question-count ("24 settings") is bumped to 25 to match the
now-25-entry main AskUserQuestion batch.

settings-advanced.md deliberately does NOT get a duplicate question: no other
boolean toggle in this repo is asked in both settings.md and
settings-advanced.md, and there's no reason to start with this one.

docs/CONFIGURATION.md, docs/USER-GUIDE.md, and a new docs/features/4139-compact-
content.md fragment (regenerated into docs/FEATURES.md) document the toggle.

ADR-4139 itself: Status flips Proposed -> Accepted, the acceptance-criteria
section becomes a guard ledger — a 13-row table covering all 12 of #4139's
original checkboxes plus the shipped-content guard criterion, each with real
evidence (the merged PR that satisfied it, fetched via `gh issue view
--json closedByPullRequestsReferences` rather than asserted from phase
numbers) — and both "Open questions for the implementation phases" are
resolved rather than left dangling: discuss-phase was never converted to
spine+detail shape (verified: no detail/ subdir exists) — a genuine gap, not a
reasoned decline; the disjointness check is confirmed line-based by reading
compact-content-split.cjs's normalizeNonTrivialLines directly.

Orthogonal review (isolated Standards/Spec code-review + security-review
sub-agents) found and this fixes two real defects: the changeset fragment's
body didn't match CONTRIBUTING.md's single em-dash-sentence format (was
multi-sentence prose naming implementation file paths); and settings.md's own
success_criteria still said "24 settings" after the new question pushed the
main batch to 25. Also fixed, found by the Spec pass while confirming
commands/gsd/settings.md correctly needed no sync edit: that file and its
skills/gsd-settings/SKILL.md twin both still described "Interactive 5-question
prompt (model, research, plan_check, verifier, branching)", stale since long
before this phase (the batch has had far more than 5 questions for a while) —
replaced with a description that names the current set without hardcoding a
count that will drift again.

gsd-test (real run, sha 1da78fe2) caught a third real regression the local
sweep missed: new-project.md is a registered spine+detail split for Phase 4's
token-reduction benchmark (scripts/benchmark-compact-content.cjs), and the new
question's +167 tokens drifted the committed baseline
(tests/fixtures/compact-content-benchmark-baseline.json). The benchmark itself
is designed never to fail CI on drift, but the test asserting the COMMITTED
baseline is currently non-drifted correctly caught it. Regenerated via
`node scripts/benchmark-compact-content.cjs --write`; re-verified --check now
reports "up to date" and the test file passes 27/27.

Closes #4408.
Closes #4139.

Emitted-Drift-Ack-Growth: new-project.md — new 4th Round-2 AskUserQuestion entry (Compact Content, #4139) plus the config-new-project CLI JSON field and explanatory sentence; a new opt-in toggle needs new prose.
Emitted-Drift-Ack-Growth: settings.md — new workflow.compact_content read_current bullet, question block, update_config write, safe-merge key, save_as_defaults field, and confirm summary row (the same seven-edit pattern every other non-capability workflow.* toggle already follows), plus the 24->25 success_criteria count fix found in review.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4408): backfill changeset PR number

pr:0 -> pr:4587 now that gh pr create has returned the real number.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 00:04:51 -04:00
Tom Boucher
7fe440a838 fix(#4488): report state update as successful when the value is already correct (#4581)
* fix(#4488): report state update as successful when the value is already correct

`cmdStateUpdate` unconditionally overwrote `updateCore`'s own `updated:true`
signal with `reconcileReportedFields`'s disk-diff result. That diff reports
`[]` -- by design -- whenever `readModifyWriteStateMd`'s #948 no-op guard
fires because the transform's output was byte-identical to the input, which
happens precisely when the requested value already equals what's on disk.
The field genuinely was found and matched; there was simply nothing left to
change. Collapsing that into the same `false`/"not found" response as a
genuine miss produced an actively wrong diagnostic message and a silent
same-day no-op in gsd-ship + gsd-extract-learnings, which both write
`Last Activity` to today's date.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4488): backfill changeset pr number to 4581

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4488): untrack tdd-red-evidence.cjs, completing its ADR-457 gitignore migration

Bundled discovery from this PR's own CI run: tests/lint-compiled-artifact-
sync.test.cjs's full tsc compile (which runs whenever ANY compiled artifact
remains tracked) SIGTERM'd under shard contention. gsd-core/bin/lib/tdd-red-
evidence.cjs (introduced by #3770/PR #4279) was the sole remaining tracked
artifact -- a tenth, later, separate instance of the #2657/#2653
migration-gap defect class this test file's closed nine-item list doesn't
cover. Untracked it and added the .gitignore entry, same fix shape as the
original nine. This eliminates the slow tsc-compile path entirely (verified:
0.1s vs ~7s locally) rather than papering over a timeout. Added a generic
regression test asserting the tracked set is fully empty.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 18:32:54 -04:00
Tom Boucher
37b965c0d1 enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code (#4553)
* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code

ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills
(src/init.cts) now selects between a canonical agents/<name>.md and a
token-minimized agents/<name>.compact.md sibling based on
workflow.compact_content, resolved in code (a real function call with a real
exit code) rather than a prose config-get gate — the same precedent stream 1's
spine/detail split established for a load-bearing seam, applied here because
this seam already runs through TypeScript instead of an eager @-include.

A missing compact sibling falls back to the canonical persona and discloses
the fallback in the served payload itself (a leading HTML-comment provenance
line), so the Done-when contract — compact when on, canonical when off, never
silent or empty — holds even for an agent nobody has compacted yet.

Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md),
each an independent, complete rewrite (not an extraction — nothing is "moved"
the way spine/detail moves text) that preserves frontmatter, every @-include,
every output-format contract, and every guardrail verbatim while cutting
restatement and verbose framing. Verified mechanically: every pair registers
(a canonical sibling exists), every compact file is strictly smaller, and the
full @-include set matches canonical's — including which references are
standalone eager-load lines versus inline prose mentions, since demoting one
to inline changes what the host actually substitutes.

Traced the install path before writing any code (.gsd/phase/.../40-design.md):
stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no
stem filtering under the default full profile, so the new .compact.md files
install for free with zero installer changes — matching issue #4407's stated
scope. A tiered agent profile that doesn't stage a compact sibling degrades
through the same fallback-with-provenance path already required for an
unauthored one, so no installer change is needed there either.

Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export
(deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are
reached by a generic code construction rather than a literal path in prose,
and checkReachability's markdown-search shape has nothing to find there).
Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new
"#4407 compact payload selection" describe block spawns gsd_run agent-skills
against real compact/canonical fixture pairs and asserts on the served
payload, which can only pass if the seam genuinely wires through.

Fixed a pre-existing test whose agents/*.md glob incidentally matched the new
.compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift
guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's
roster, both real, unrelated-to-content defects the new files' mere existence
surfaced.

Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files
under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token
benchmark baseline (npm run benchmark:compact-content-variants --write).

Closes #4407.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): apply orthogonal review findings from the compact-payload seam

Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath)
to collapse the duplicated read-and-empty-check shape between the compact
and canonical branches in cmdAgentSkills, and updated the adjacent comment
enumerating flat JSON extras to name agent_payload_variant alongside
source/degraded (added by the prior commit, comment left stale).

Security review and the Spec axis found no defects requiring a code change;
their non-blocking observations (a pre-existing, unmodified path-construction
pattern; the reasoned, documented substitution of a behavioral test for the
literal reachability check) are recorded in
.gsd/phase/enhance-4407-agent-skill-seam/60-review.json.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents

Root-caused via a real gsd-test run (93 failures) rather than guessing which
tests glob agents/ naively. Two classes of defect, both genuine:

1. Identity-roster confusion (11 files/areas): many tests and one production
   script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f
   => f.endsWith('.md'))`, which incidentally matched the new .compact.md
   variant siblings too — a compact file is a rendering of an EXISTING agent
   identity, not a new one. Fixed at the shared root
   (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests
   already consolidated on) and at each independent glob that didn't use it:
   agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix
   before checking XL/LARGE membership, so a compact file inherits its
   canonical sibling's tier instead of silently falling through to DEFAULT),
   agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual
   script, not just its test), codex-config.test.cjs (confirmed directly
   against generateCodexAgentToml that a compact role's derived sandbox_mode
   is byte-identical to its canonical sibling's before excluding it — not
   assumed), and copilot-install.test.cjs (two counts that legitimately DO
   need both files — an installed-file count and a full-conversion smoke test
   — fixed to expect 70, not stay pinned to 35).

   no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix:
   two compact files reproduce descriptive prose already allowlisted at their
   canonical file's line number; added matching entries at the compact files'
   own line numbers rather than excluding them from the scan (a genuine bare
   gsd-tools command-position bug in a compact file would be as real a defect
   as in canonical).

2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree
   run): six agents' compact renditions (gsd-debugger, gsd-executor,
   gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed
   the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction —
   confirmed structural, not a compaction-quality gap: each is dominated by
   content this phase's own rules require verbatim (the ~2.6 KB gsd_run
   bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in
   every agent that calls gsd_run, output-format contracts, guardrails).
   ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing
   spot in cmdAgentSkills's single-file synchronous read. Removed these 6
   compact files rather than ship an over-cap file or invent a multi-part
   read mechanism out of scope for this phase; recorded by name with the
   reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and
   50-test-matrix.md, per #4407's own "or explicitly recorded as not worth
   covering" allowance. Their canonical personas are served correctly today
   via the fallback-with-disclosed-provenance path this phase's own Done-when
   #2 already requires — 29 of 35 agents now have a compact variant.

Also fixes an unrelated, genuinely pre-existing defect this gsd-test run
surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own
DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in
this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md
| wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed:
two meaning-preserving trims in the <success_criteria> block (a repeated
parenthetical replaced with a same-exception reference; one redundant
qualifier dropped) bring it to 40,940 bytes.

Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant
benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6
now-orphaned roster rows removed alongside them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): make .compact.md-aware roster checks resilient to partial coverage

Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent
has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception),
breaking once 6 stems legitimately have none.

- tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was
  picking up the "### Compact Payload Variants" subsection's rows as
  phantom/uncounted entries in the primary/advanced/inventory-only
  classification this test validates — a compact row documents an existing
  agent's alternate rendition and never gets its own AGENTS.md heading, so it
  was never meant to participate in that classification. Excluded at the
  parser, not per-assertion.
- tests/copilot-install.test.cjs: the derived expected-file-list generator
  assumed every listAgentFiles() stem has a .compact.md source sibling;
  checks disk per stem now instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4407): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 12:38:59 -04:00
dependabot[bot]
c504e715c6 chore(deps-dev): bump js-yaml from 4.3.1 to 4.3.2 in the npm_and_yarn group across 1 directory (#4565)
* chore(deps-dev): bump js-yaml

Bumps the npm_and_yarn group with 1 update in the / directory: [js-yaml](https://github.com/nodeca/js-yaml).


Updates `js-yaml` from 4.3.1 to 4.3.2
- [Changelog](https://github.com/nodeca/js-yaml/blob/4.3.2/CHANGELOG.md)
- [Commits](https://github.com/nodeca/js-yaml/compare/4.3.1...4.3.2)

---
updated-dependencies:
- dependency-name: js-yaml
  dependency-version: 4.3.2
  dependency-type: direct:development
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>

* chore: refresh vendored js-yaml to 4.3.2 (#4565)

lint-vendored-deps caught the drift: this PR's lockfile-only bump left
gsd-core/bin/lib/vendor/js-yaml.cjs and the package.json pin behind the
new js-yaml 4.3.2 resolved by package-lock.json (merge-key CPU-limit
backport, GHSA for excessive merge-key processing). Refreshes the
vendored copy from node_modules and bumps the manifest pin to match.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs: add Security changeset for js-yaml 4.3.2 vendor bump (#4565)

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:46:35 -04:00
Tom Boucher
ba26aa065d docs(#4467): document fallow's structural-pre-pass has no upper-bound scope (#4574)
* docs(#4467): document fallow's structural-pre-pass has no upper-bound scope

structural-pre-pass.md's FALLOW_SCOPE_ARGS=(--changed-since "$FALLOW_BASE")
derives a correct, phase-anchored LOWER bound (lockstep with Tier 3's own
scope step, #3995), but fallow's --changed-since is one-sided by design --
verified against fallow 2.70.0's own --help: the only other scoping flags
are --changed-workspaces (workspace selection, not a file range) and
--diff-file (its own help text scopes it to line-range refinement within
the hot-path-touched verdict, not general file selection; gsd-core never
uses it). Reviewing an earlier phase after a later one has landed pulls
the later phase's files into the earlier phase's structural audit.

Not fixable inside this file: fixing fallow itself is a third-party
concern, and working around it (e.g. auditing from a temporary worktree
checked out at the phase tip) is disproportionate machinery for what is
supplementary structural-analysis context, not a blocking gate -- both
routes the issue's own analysis already ruled out. Documented the
asymmetry at the point the scope is derived instead, so a future reader
does not assume this step's tip agrees with Tier 3's just because the
base does.

No regression test: documentation-only, no runtime behavior change.

A prior revision of this commit carried an Emitted-Drift-Ack-Growth
trailer for this growth -- gsd-test's own emitted-attribution check
rejected it as stale ("written or reworded in THIS diff, but nothing
here needed them"), meaning this file (nested under
gsd-core/workflows/code-review/steps/, unlike a top-level
gsd-core/workflows/*.md file) is not tracked by that specific growth
conservation law. Removed the now-confirmed-unnecessary trailer rather
than guess again -- the test's own verdict is authoritative here, not
a re-derivation of its tracked-path rules.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4467): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 10:45:20 -04:00
Tom Boucher
59da9f016b fix(#4466): bound quick.md's post-execute review scope tip at the task's own last commit (#4571)
* fix(#4466): bound quick.md's post-execute review scope tip at the task's own last commit

The review-scoping step computed CHANGED_FILES as `git diff --name-only
"${DIFF_BASE}..HEAD"`. DIFF_BASE is correctly bound to the quick task's
start (via the oldest QUICK_COMMITS entry's parent), but the tip was bare
HEAD -- unbounded. Anything landing on the shared tree between the task's
own commits and this review step running (a worktree merge-back, another
session sharing the tree) got folded into the quick task's own review
scope.

QUICK_COMMITS (newest-first) already holds the correct tip as its first
line -- read QUICK_TIP from the value already computed, diff against
that instead of HEAD. No new derivation, no new git call.

Added tests/quick-review-scope-tip-bound.test.cjs: extracts the scoping
fence verbatim from quick.md and runs it against a real git fixture
matching the issue's own scenario (quick task's own commit, then a later
unrelated commit on the shared tree). Manually verified watch-it-fail
(bare HEAD includes the unrelated file) / watch-it-pass (bounded tip
excludes it) via direct bash execution before wiring the test file,
since this repo blocks local node --test.

Independent code review caught one drive-by finding: an allow-test-rule
marker copied from a sibling test's pattern was unnecessary here (and
there) -- local/no-source-grep's looksLikeSourcePath only matches
readFileSync targets ending in .cjs/.cts/.js/.mjs/.mts/.ts, never .md,
so the rule can never fire regardless of the marker. Confirmed by
reading eslint-rules/no-source-grep.cjs directly; removed.

Emitted-Drift-Ack-Growth: quick.md — the fix adds a QUICK_TIP line and its explanatory comment; not a regeneration artifact, a hand-authored bug fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4466): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 10:12:06 -04:00
sim
d3e3a8f535 fix(#4460): restore compact-file reachability + fix test stdout capture
Two more gsd-test-surfaced findings:

1. The previous execute-plan.md trim removed the literal
   `summary.compact.md` filename mention, breaking
   tests/compact-content-variant-guard.test.cjs's reachability check
   (ADR-4139 Phase 6): every registered .compact.md variant must be
   named by at least one workflow "spine" file, and execute-plan.md was
   apparently the only spine naming this one. Restored the bare
   filename (kept the shortened surrounding wording) -- read
   tests/helpers/compact-content-variant.cjs's checkReachability/
   isUnprefixedMatch directly to confirm the fix rather than guessing.
   40926 bytes, still 34 under the size cap.

2. The redesigned test (previous commit) still failed: both tiers'
   diagnostic `echo`/`printf "Warning: ..."` lines were mixing into the
   captured stdout the assertions parse as the file list, so
   "--files=src/alpha.js" appeared to produce 2 lines instead of 1.
   Wrapped both tier fences in a `{ ...; } > /dev/null` brace group
   (not a subshell -- REVIEW_FILES still persists to the enclosing
   shell) so only the final printf reaches stdout. Manually re-verified
   both cases against a real git fixture before re-running the suite.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 08:07:33 -04:00
sim
db7a349a8c fix(#4460): trim execute-plan.md under its size budget (unrelated regression)
gsd-test surfaced a SEPARATE, unrelated failure while re-verifying this
branch: tests/workflow-size-budget.test.cjs found execute-plan.md at
40981 bytes, 21 over the 40960 DEFAULT hard cap. Root-caused (not
assumed): already-merged PR #4540 (enhance(#4139), unrelated to
#4460/#4459/#4461) added two near-identical explanatory parentheticals
about .compact.md template variants across two nearby steps
(user_setup, create_summary), pushing the file over. Confirmed directly
against origin/next independent of any merge with this branch --
`next` itself already carries this.

This branch's fork point predated PR #4540's merge, so gsd-test's
merge-testing against the current next only now surfaced it (merged
origin/next into this branch in a separate commit first -- 0 conflicts,
after discovering and fixing that this worktree's git clone was
SHALLOW, via `git fetch --unshallow`, which is what made a plain `git
merge origin/next` fail with "refusing to merge unrelated histories").

Fixed by trimming the SECOND (of two near-identical) parentheticals in
the create_summary step to a short back-reference to the first -- same
information, no duplication, no cap raised (the test explicitly warns
against raising it). 40931 bytes, 29 under the cap.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 08:07:33 -04:00
sim
5946926b94 fix(#4460): rework test to not depend on Tier 2's broken bash (#4461)
A fresh code-review pass found the test's original approach (concatenate
and execute Tier 1 + Tier 2 + Tier 3 verbatim, matching the issue's own
reproduction) cannot run: Tier 2's own fence -- untouched by this diff --
is not currently parseable bash. Two unescaped `"` inside its embedded
`node -e "..."` regex literal (`raw.replace(/^['"]|['"]$/g, '')`)
terminate the outer double-quoted string early, which breaks bash's
PARSE of the whole concatenated script even though Tier 2's body never
executes under --files. Independently confirmed via manual extraction
and execution before accepting the finding.

This is a real, separately-filed, already-queued sibling issue (#4461,
filed by #4460's own reporter specifically to avoid folding it in here)
-- not fixed in this PR. Instead reworked the test to run only Tier 1 +
Tier 3 verbatim, seeding the Tier-2-equivalent REVIEW_FILES state
directly for the "without --files" case (documented in the module
docblock, explaining why Tier 2 isn't sourced and pointing at #4461).

Also fixed a nit from the same review pass: a code comment overstated
Tier 2's guard as "immediately above" when it's ~150 lines away.

Manually re-verified both test cases against a real git fixture with a
GNU-realpath-compatible `realpath` (matching gsd-test's Linux bench --
this Mac's BSD realpath lacks the `-m` flag Tier 1 uses, a SEPARATE
pre-existing portability gap surfaced during this check, masked on Linux
CI, not touched by this fix) before re-running the full suite.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 08:07:33 -04:00
sim
b1c78f0d2e fix(#4460): gate Tier 3's #2666 cross-check on FILES_OVERRIDE
code-review.md states (line 144) "Skip SUMMARY/git scoping entirely when
--files is provided." Tier 2 honors this via `if [ -z "$FILES_OVERRIDE"
]`, but Tier 3's #2666 SUMMARY/diff cross-check had no FILES_OVERRIDE
reference at all -- reached via `elif [ -n "$DIFF_BASE" ]` whenever
REVIEW_FILES was already non-empty (true under --files, since Tier 1
fills it), so it silently appended the whole phase's changed files onto
an explicit user-supplied file list. --files is documented as the
highest-precedence scoping tier (D-08) and is the flag Tier 3's own
fail-closed path recommends when no reliable diff base is found; a user
narrowing a review to two files silently got the whole phase instead,
and the reviewer agent spent its budget on files nobody asked about.

Gated the elif on the same condition Tier 2 already uses:

  elif [ -z "$FILES_OVERRIDE" ] && [ -n "$DIFF_BASE" ]; then

The issue's own narrowest suggested form, reasoned through against two
alternatives (wrapping the whole Tier-3 fence, or changing the stated
invariant instead) -- both explicitly rejected there for good reasons
concurred with after reading the surrounding code.

Added tests/code-review-tier3-files-override-scoping.test.cjs, mirroring
the issue's own verified reproduction methodology: extracts the Tier
1/2/3 fences VERBATIM from code-review.md (never reimplemented) and runs
them against a real constructed git fixture matching the issue's own
scenario exactly (5 files, a SUMMARY listing only 1). Confirms --files
stays scoped to exactly the requested file, and separately confirms the
#2666 cross-check still widens a genuinely partial SUMMARY scope when
--files is absent (proving this is a gate, not a blanket disable).

Emitted-Drift-Ack-Growth: code-review.md — #4460 gates the Tier-3 #2666 cross-check on FILES_OVERRIDE, matching Tier 2's own guard, net +453 bytes
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 08:07:33 -04:00
Norman Yee
42c02a00c0 enhance(#3418): report real codebase drift instead of the whole repository (#4124)
* fix(#3418): write the codebase-drift baseline from code instead of agent prose

writeMappedCommit shipped correct and callerless, so no full map-codebase run ever wrote last_mapped_commit. The gate then read null and diffed HEAD against the empty tree, reporting every tracked file as newly added on every run.

Adds the stamp-codebase-map leaf verb and calls it from the map-codebase workflow and the execute-phase auto-remap path, replacing the prose instruction that asked the mapper agent to stamp its own output. An agent that concludes its work is already done skips a prose step silently, which is the failure the stamp exists to detect.

The gate now reports an absent or unresolvable baseline as skipped, with reason no-mapped-commit or unresolvable-mapped-commit, rather than as whole-repo drift. Files under .planning/ are excluded from the diff so the map's own commit does not read as seven new directories on the next run.

Emitted-Drift-Ack-Growth: map-codebase.md — adds the stamp_codebase_map step and its rationale, new workflow content this change requires

* test(#3418): cover the stamp writer and the absent-baseline gate

* docs(#3418): document how the drift baseline is written and skipped

* docs(#3418): note that a manual stamp reflows the map's whitespace

writeMappedCommit writes through platformWriteSync, which normalizes markdown whitespace on .md targets. Run in its workflow position the stamp lands on documents the mapper just wrote, so the normalization is folded into the same commit, but a hand-run stamp over an already-committed map reflows that map as a side effect. Reported on the issue thread.

* chore(#3418): add changeset fragment

Typed Changed to match the enhancement route the linked issue's label sets. The docs-required lint is satisfied by the ARCHITECTURE.md update already in this branch.

* fix(#3418): anchor the planning-artifact filter to the repo root

git diff --name-status prints repo-root-relative paths whatever the cwd, so computing the exclusion prefix against cwd yielded ".planning/" while git printed "sub/.planning/" and the filter silently matched nothing from a subdirectory.

* fix(#3418): derive the planning prefix from git, not from path arithmetic

Anchoring the exclusion prefix with path.relative() against `rev-parse --show-toplevel` broke on Windows, where os.tmpdir() hands back the 8.3 short form and git resolves the long one, so relative() produced a "../.." chain that matched nothing. `rev-parse --show-prefix` gives the cwd's root-relative prefix from the same producer as the diff paths, so the two sides cannot disagree.

* fix(#3418): take the planning lock around the codebase-map stamp

Stamping seven documents is seven frontmatter read-modify-writes, and two stampers can run at once: the full map-codebase run and the execute-phase auto-remap. Wrap the write loop in withPlanningLock, the same lock the other .planning/ writers take, so a concurrent pair cannot lose an update.

Also corrects the path-arithmetic comment, which read as if the Windows short-path hazard applied to the .planning half of the prefix. It applies to the rejected --show-toplevel alternative; both sides of the surviving relative() call are the same cwd string.

* fix(#3418): read HEAD and the map file list under the planning lock

The stamp resolved HEAD and listed the present codebase-map documents before it acquired the planning lock, so a stamper that then waited on the lock could write its now-stale sha over a newer one, or recreate a document deleted while it waited as a frontmatter-only stub. Both reads now happen inside the lock, matching the read-and-write-in-one-lock pattern config.cts and phase.cts already use. An empty --files value is refused as well instead of silently widening the stamp to all seven documents.

* fix(#3418): narrow the map stamp to the documents an update run refreshed

An "Update - only update specific documents" run reached the new stamp step with no --files narrowing, so the six documents the user did not select were stamped at HEAD and read as freshly mapped. The selection now threads through to --files, the same way the auto-remap path already does.

A bare --files (an unquoted empty shell variable drops the token) parsed to null, indistinguishable from an absent flag, so it skipped the empty-filter refusal and stamped all seven. Presence is now read off argv.

* fix(#3418): require the drift baseline to resolve to a commit, not any object

`git cat-file -t` exits 0 for a tree or blob sha and for a ref name, and `git diff <tree> HEAD` is valid, so an exit-code-only probe accepted a baseline that is not a commit and reported the resulting diff as real drift. Check the reported type instead of the exit code alone, which routes every non-commit stamp to the same `unresolvable-mapped-commit` skip.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-09 05:08:20 +00:00
Behruz Nassre Esfahani
3ad75a6d59 enhance(#4285): resolve context-monitor fire-points from .planning/config.json (#4366)
* enhance(#4285): resolve context-monitor fire-points from .planning/config.json

The monitor's WARNING (35%) and CRITICAL (25%) fire-points were module
constants, so the only way to tune them was editing gsd-context-monitor.js —
a file in the MANAGED hooks registry, whose body the next install re-stages,
silently discarding the edit. The alternative was turning the safety net off.

Both are now readable from the config block the hook already opens:
hooks.context_warning_threshold and hooks.context_critical_threshold. Absent
keys resolve to today's 35/25, so every existing project is byte-identical.

Resolution is total and never throws — this hook must not block the tool call
it rides in on. A value is usable only if Number.isFinite (type-strict, so the
string "30" and true are rejected) and inside the 0-100 domain of the
remaining_percentage it is compared against; anything else falls back to the
default. The PAIR falls back together: critical >= warning has no coherent
reading, and honouring one side silently picks which of the operator's two
numbers to discard. That also covers a single override contradicting the other
key's default.

config-set validates the domain per key so accept and honour agree, but
deliberately does not enforce the pair — it writes one key per call, so a
two-step retune is transiently inconsistent on disk and refusing it there
would block a legitimate configuration.

Registration follows the statusline.show_git precedent: schema manifest plus
src/config.cts validation, not config-defaults.manifest.json and not
buildNewProjectConfig — emitting 35/25 into every new project would pin the
defaults at creation time for a setting nobody has tuned.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2

* enhance(#4285): address Codex review — per-key fallback docs, discriminating tests

Codex full-PR review (gpt-6-astra, read-only) returned five findings. Each was
verified against source before acting; all five are real.

1. docs/CONFIGURATION.md described the wrong fallback. An out-of-domain value
   falls back PER KEY; both defaults apply only when the RESOLVED pair violates
   critical < warning. warning 150 with critical 30 resolves to 35/30, not
   35/25 — at remaining 28 that difference changes the severity emitted. The
   table now states the two rules in the order they compose, and
   docs/context-monitor.md gains the same worked example.

2. The inconsistent-pair test could not prove the CRITICAL side reverts: its
   pair was 20/25, and 25 is already the default, so an implementation that
   reset only `warning` passed it. A 45/50 pair — both halves away from their
   defaults — now pins each side with its own reading, and an equal 45/45 pair
   pins that the rule is strict (`<`, not `<=`).

3. The rejection table's rows could not tell rejection from acceptance: an
   accepted -5 pairs with the default critical 25, trips the pair check, and
   produces the same silence. Two rows now separate those: a below-domain
   critical must escalate remaining 20 to CRITICAL (proving -5 was rejected,
   not honoured), and an unusable critical beside a usable warning 45 must
   still fire WARNING at remaining 40 (proving per-key fallback rather than
   reset-both). The over-claiming comments are narrowed to what each row
   actually shows.

4. Scope, reproduced rather than assumed: config-set writes through
   planningDir(), so under GSD_WORKSTREAM it lands in
   .planning/workstreams/<name>/config.json while this hook reads only
   <cwd>/.planning/config.json. That is the pre-existing root-only scope
   hooks.context_warnings has always had, but this PR advertises the setter
   route, so both docs now say the keys are root-project settings.

5. Four other English docs still stated 35/25 as fixed: the REQ-CTX-02/03
   requirements fragment, ARCHITECTURE.md's hook table and threshold table,
   and INVENTORY.md's hook row. All now name them as defaults and point at the
   config keys; docs/FEATURES.md is regenerated from its fragment via
   scripts/gen-features.cjs --write, not hand-edited.

Four new mutations, each reverted after: resetting only the warning half on an
inconsistent pair (1 red), resetting both on any unusable key (1), dropping the
>= 0 bound (1), and accepting critical == warning (1). perf-317 is 116/0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2

* enhance(#4285): tighten claims after Codex round 2 — scoped paths, one more discriminator

Confirmation round found no runtime defect and confirmed the five round-1 fixes
landed. Four precision items, all real, all fixed here.

1. The scoped-write note named the wrong path for GSD_PROJECT. planningDir()
   composes three distinct shapes, confirmed by running config-set under each:
   .planning/<project>/config.json, .planning/workstreams/<ws>/config.json, and
   .planning/<project>/workstreams/<ws>/config.json. docs/context-monitor.md
   now tabulates all four cases instead of collapsing them into one.

2. The 45/50 silence row asserted empty stdout without pinning the exit code.
   runMonitorRaw turns a spawn failure, a non-zero exit or a timeout into empty
   stdout as well, so the row could have passed on a dead child. It asserts
   exitCode === 0 first now, like the equal-pair row already did.

3. The sibling row's message claimed it proved critical fell back to 25. It
   does not: coercing '30' to 30 yields WARNING at remaining 40 too, so the row
   pins the WARNING side surviving and nothing more. Message narrowed, and a
   new row reads the same config at remaining 28, where the two candidate
   resolutions diverge — rejected gives (45, 25) and WARNING, coerced gives
   (45, 30) and CRITICAL. Mutation-verified: swapping Number.isFinite for the
   coercing global reds it.

4. "Accept and honour must agree" was too absolute in the src/config.cts and
   tests/config.test.cjs comments. The agreement holds on the DOMAIN and per
   key: an accepted value can still lose to the hook's pair check at read time,
   and a scoped write never reaches the hook at all. Likewise a two-step retune
   only CAN be transiently inconsistent — 35/25 to 20/10 is valid throughout if
   critical moves first — so the docs now say what a setter-side pair check
   would actually cost: rejecting that intermediate write and forcing an order.

The same over-absolute phrasing is in b7d179c89's message, which is left as
written rather than rewriting history; this commit and the PR body carry the
precise claim.

perf-317 117/0, config 192/0, config-field-docs 47/0, features-index-gate 84/0,
lint:ci clean cold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2

* chore(#4285): add changeset

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2

* enhance(#4285): address review — planning-config rows, resolveThresholds properties

Two Minor findings from the maintainer review, no behaviour change.

Minor 1: gsd-core/references/planning-config.md's "Hook Fields" table gains
rows for hooks.context_warning_threshold and hooks.context_critical_threshold,
in that table's 5-column form, carrying the same per-key-fallback,
pair-reversion and root-config-scope claims docs/CONFIGURATION.md already
makes. hooks.workflow_guard's absence from that table is pre-existing and
out of scope here.

Minor 2: resolveThresholds() gets fast-check property coverage, which ADR 456
requires of a threshold/limit contract. Reaching it needed a require-time
seam: the resolver was previously observable only by spawning the hook, and a
subprocess per case cannot drive 200 runs — the same conclusion CONTEXT-INDEX
records for the ROADMAP Requirements parser. The stdin adapter therefore moves
into main() behind `require.main === module`, mirroring
gsd-cursor-subagent-start.js and gsd-statusline.js, and module.exports exposes
the resolver plus both default constants so a test asserts fallback against
the source of truth rather than a second copy of 35/25. Spawned behaviour is
unchanged: the 10s stdin timeout still arms per invocation (stdinTimeout is
now a module-scope let assigned in main(), still cleared by the end handler),
and the try/catch crash(ON_CRASH) path is untouched.

Seven properties: totality, ordering, exactness, togetherness, non-vacuity,
per-key fallback, non-object argument. Exactness is stated PER KEY — a mixed
result (one key honoured, one fallen back) is legal and is the documented
contract; the property falsified a per-pair phrasing of it in 4 runs.

Verified: cold lint:ci 0; perf-317 file 125/0; seven mutations killed and
restored, one of which (upper bound widened to 120) is invisible to the 17
hand-written cases and caught only by a property.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte

* enhance(#4285): close the Codex-found gap in the property coverage

Codex whole-PR review of round 3 returned no Blocker and no Major. Two items,
both in the tests added this round, both verified against source before acting.

Minor — the per-key fallback property was asymmetric: it required a usable
warning to survive an unusable critical, but never the reverse. A resolver
that reverted BOTH keys the moment warning was unusable passed all seven
properties. Reproduced exactly: that mutant answers 35/25 for
{warning: 150, critical: 30} where the resolver answers 35/30, and the file
stayed green at 125/0. The mirrored property closes it — with the mutant
re-applied it is now the single failing row, and it is the only row that
fails, so it is load-bearing rather than incidental.

Nit — the ordering property's comment credited it with catching a
half-honoured pair, which it does not: 45/50 "repaired" by resetting only
critical yields 45/25, perfectly ordered. That case belongs to togetherness.
The same comment claimed the behavioural rows sample an inconsistent pair at
exactly one point; stale — they cover 20/25, 45/50 and the 45/45 equality
boundary. Both claims corrected in place.

Verified: cold lint:ci 0; perf-317 file 126/0; the mutant above killed by the
new property alone and the hook restored byte-identical afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte

* enhance(#4285): name the installed-monitor prerequisite; close the negative-critical gap

Second Codex whole-PR pass, run because the base moved: the author's three
"Update branch" merges pulled ~26 upstream commits in, so the previously
reviewed diff sat on a base that no longer exists. No Blocker, no Major, two
Minor — both verified against source before acting.

Minor 1, and only reachable because of what the merge brought in: #2586
(03738824d) landed in that window and stops staging
hooks/gsd-context-monitor.js for Codex, since the metrics bridge it reads is
written only by hooks/gsd-statusline.js, which Codex never installs
(bin/install.js: "gsd-context-monitor.js is deliberately NOT copied for
Codex"). These two keys are read by that hook and nothing else, so on such a
runtime config-set stores and validates them and nothing consumes them — a
claim the docs this PR adds did not make. docs/context-monitor.md now carries
the explanation and both key tables carry a clause pointing at it; the FEATURES
and INVENTORY entries already link through to those two files, so they are not
edited again. The changeset says it too, because it is user-facing.

Accepting the keys on every runtime is kept deliberately: config is shared
across runtimes, so validation stays runtime-independent and the runtime
caveat lives in documentation rather than in the setter.

Minor 2: the per-key fallback property's junk generator had no negative arm,
though its mirror did — and that asymmetry hid a gap. A resolver reverting
BOTH keys whenever critical is negative answers 35/25 for {45, -5} where the
resolver answers 45/25, and it passed all 126 tests. With the negative arm
added it is the single failing row.

Verified: cold lint:ci 0; perf-317 126/0; both mutants above killed and the
hook restored byte-identical; 538/0 across the config, changeset, doc-parity
and emitted-attribution gates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte

* enhance(#4285): refuse the two dead threshold endpoints; resolve absent keys

Maintainer review round 2 raised two Minors and a nit.

Minor 1 — `hooks.context_warning_threshold: 0` was accepted and stored but can
never take effect: `critical < warning` must hold and both sides are clamped to
0-100, so nothing can sit below a warning of 0. Verifying it surfaced the MIRROR
case the review did not name: `critical: 100` is equally dead, since nothing can
sit above it. Both confirmed against the real resolver for partners {absent, 0,
50, 100}, with 0.001 and 99.999 honoured as controls.

`config-set` now refuses both, because storing a value the reader always
discards is the accept-then-discard shape this codebase refuses elsewhere. The
hook is unchanged and still total — it degrades to defaults rather than
throwing, so a project that already carries one of these on disk still loads.
The old "accepts the domain bounds 0 and 100" row asserted the misleading half
and is replaced by tables that make the asymmetry the point (0 is legal for
critical and illegal for warning; 100 is the reverse), plus a control row so
"refuse both endpoints outright" would not pass in its place.

Minor 2 — the keys are absent from config-defaults.manifest.json /
buildNewProjectConfig where the sibling `hooks.context_warnings` lives. Kept
that way: buildNewProjectConfig writes a hooks object into every NEW project's
config.json, which would freeze today's fire-points as an explicit per-project
override everywhere — the opposite of this PR's premise. But the underlying
complaint was real, so the actual symptom is fixed: `config-get` on an absent
key returned "Key not found" while the hook silently used 35/25. It now resolves
through SCHEMA_DEFAULTS. Restated rather than derived because CONFIG_DEFAULTS is
re-exported flattened and has no `hooks` member at runtime; the one resulting
copy of 35/25 outside the hook is pinned against the hook's exported constants
by a drift test (red-checked: moving the literal to 40 reds it).

Nit — PR-body counts unverifiable from the diff. Noted, no code change.

Codex round 3 then found a broken doc link (`context-monitor.md` resolved
inside gsd-core/references/, where it does not exist; the emitted tree's own
convention is `../../docs/...`) and a stale comment still describing the
manifest-derived approach I had backed out. Both fixed. It also corrected my
rationale on a point of fact: manifest entries alone would NOT have reached new
project configs, since buildNewProjectConfig builds its own literal — the
freezing argument applies to that function, not to the manifest. The comment now
says so rather than running the two together.

Verified: cold lint:ci 0; full suite 36,082 / 0 fail before these two fixes,
config + perf-317 321/0 after; drift pin red-checked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-09 04:31:14 +00:00
Lorenz Leslie Espinosa
b33df03726 enhance(#4089): add minimum-solution reasoning check (#4118)
* enhance(planning): add minimum-solution reasoning check

* chore: add changeset for planning guidance

* chore: bind changeset to PR 4118

* docs: document planning sufficiency check

* docs: distinguish planning sufficiency guidance

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-09 00:58:23 +00:00
sim
8bcf633e21 fix(#4554): trim execute-plan.md 21 bytes under its DEFAULT size-tier cap
gsd-core/workflows/execute-plan.md sat 21 bytes over its own DEFAULT-tier
hard cap (40,960 bytes, tests/workflow-size-budget.test.cjs, ADR-1610) at
next@a27cb6b2fa — introduced by #4540's call-site wiring for the summary.md/
user-setup.md .compact.md variants, which nobody caught crossing this exact
margin before merge. This trips next's own Tests run on every shard/OS
combination, which in turn blocks the repo's Base branch health PR gate
(#4422/#4428) for every open and future PR regardless of that PR's own diff.

Two meaning-preserving trims in the <success_criteria> block: a repeated
parenthetical ("— unless parallel mode (orchestrator handles)", appearing
twice) replaced with a "— same exception" back-reference on its second
occurrence, and one redundant qualifier ("prominently") dropped — its
behavioral content (surface the USER-SETUP.md warning at the TOP of output)
is already fully specified earlier in the same file. 40,981 -> 40,940 bytes,
20 bytes of headroom under the cap. No procedural content lost.

Fixes #4554.

Emitted-Drift-Ack-Hash: gsd-core/workflows/execute-plan.md — deliberate content trim to clear the DEFAULT size-tier cap (#4554); not a regeneration artifact, a hand-authored byte reduction.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 17:52:35 -04:00
Tom Boucher
a27cb6b2fa enhance(#4139): Phase 6 — the lazily-read remainder and the artifact templates (#4540)
* enhance(#4406): the lazily-read remainder and the artifact templates

ADR-4139 Decision 3, Phase 6 of the #4139 Compact Content epic. Covers stream 1b
(gsd-core/workflows/<name>/{modes,steps,templates}/*.md) and stream 4
(gsd-core/templates/**) with a variant-swap mechanism, confirmed with the user:
two independent, complete files per covered path (canonical + .compact.md
sibling), with the gate picking which one gets Read at the call site. This is
a different shape from Phase 5's spine+detail partition, and is safe here
specifically because these files are already reached only by a runtime Read —
a missed Read already means zero overlay content today, with or without
workflow.compact_content, so selecting between two independently-complete
files introduces no new failure mode (documented in
gsd-core/references/compact-content-gate.md's new "Streams 1b and 4" section).

Disposition, after inspecting every candidate rather than trusting a byte-size
threshold (same rigor Phase 5 applied to review.md):

- Stream 1b: 1 of 78 files compacted (help/modes/full.md, a user-facing
  reference doc emitted verbatim, not orchestrator instruction). The other 9
  size-threshold candidates are dominated by fail-closed guards, exact CLI
  invocations, or output-format contracts (AskUserQuestion blocks) — recorded
  not-worth-compacting, same reasoning as Phase 5's review.md.
- Stream 4: a ground-truth reachability audit replaced the initial size-only
  candidate list. Two files (summary.md, user-setup.md) got compact variants;
  a third (spec.md) was drafted, then dropped after discovering its only two
  call sites are eager @-includes, not a runtime Read — stream-1 material
  hiding under gsd-core/templates/, not stream-4's actual mechanism. summary.md
  itself has 3 eager call sites and only 1 genuine runtime-Read call site
  (execute-plan.md); only that one was wired, so the compact variant's savings
  apply to the sequential single-plan execution path only.
- Discovered while auditing reachability: 12 gsd-core/templates/** files with
  zero references anywhere in workflow/agent/command prose, compiled source,
  or tests — dead scaffolding predating this phase. Deleted in this same PR
  per this repo's no-defer policy, after re-verifying against a computed
  path.join(...) pattern (not just a plain-string search) that nearly caused
  two genuinely load-bearing templates (user-profile.md, dev-preferences.md)
  to be misclassified as dead.

New checker (tests/helpers/compact-content-variant.cjs): registration,
reachability, protected-content-preserved, size-smaller — replacing Phase
3/5's disjointness/completeness checks, which assume a partition rather than
two deliberately-overlapping documents. The reachability check's own
"unprefixed match" guard had a real bug (rejected the repo's own
`~/.claude/gsd-core/...` convention), caught by running it against the
already-wired help/modes/full.compact.md pair rather than only synthetic
fixtures — fixed to anchor on the nearest `gsd-core` path segment instead.

Template consumer parity (tests/compact-content-template-variant-parity.test.cjs):
proves each compact variant's `## File Template` fenced block — the actual
output-format contract a generated SUMMARY.md/USER-SETUP.md is parsed
against — is byte-identical to the canonical file, then runs the one real
deterministic consumer (gsd-core/bin/lib/coverage.cjs's classifyContent,
backing `gsd-tools uat classify-coverage`) against content built from that
shared contract.

Added a sibling benchmark script (scripts/benchmark-compact-content-variants.cjs)
rather than extending the existing spine/detail one — different data shape,
and the existing script's own contract deliberately isolates it from a
test-only helper's shape changing.

Emitted-drift acknowledgement: not needed. Every changed/added path in this
diff is hand-authored and present in the diff itself, so diffEmitted's
attribution loop resolves `via` to the path's own source before reaching the
ack-lookup branch (same reasoning Phase 5 verified for its own diff).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* enhance(#4406): address code-review findings on the variant-swap gate

- docs/CONFIGURATION.md and gsd-core/references/planning-config.md's
  workflow.compact_content entries described only the spine+detail mechanism
  (Phase 5) and were missing this phase's variant-swap mechanism and its
  benchmark:compact-content-variants script entirely — required since this
  PR's changeset is type Added (CLAUDE.md's "Missing Docs for Changesets"
  rule). Both now describe both mechanisms and which call sites are wired.
- Added the missing RED^-1/no-op fixture for checkProtectedContentPreserved:
  a canonical file with zero <!-- gsd:protected --> blocks must be a
  no-op, not a violation — the only branch of that function the existing
  fixtures didn't exercise.
- Collapsed findCompactFiles/findMarkdownFiles in
  tests/helpers/compact-content-variant.cjs into one findFilesWithSuffix
  helper — the two were identical recursive walks differing only in the
  extension predicate (minor Duplicated-Code finding).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4406): restore copilot-instructions.md, a false-positive dead-template classification

gsd-test caught this, not static analysis: 10 real failures in
tests/copilot-install.test.cjs, tests/installer-migration-install.integration.test.cjs,
and tests/repo-layout.test.cjs — all downstream of bin/install.js's Copilot install
path, which does
fs.readFileSync(path.join(targetDir, 'gsd-core', 'templates', 'copilot-instructions.md'))
after copying gsd-core/templates/** into the target project, then merges it into both
.github/copilot-instructions.md and (local installs) AGENTS.md. The reachability audit
that flagged this file as dead checked src/*.cts and gsd-core/bin/*.cjs but never the
repo-root bin/install.js — a separately maintained installer bundle outside the
src/-to-gsd-core/bin/lib/ compiled-output convention. The fs.existsSync guard around
that read degrades to a silent skip rather than a crash when the template is missing,
which is why this surfaced only once the real E2E install test ran, not from any
static check.

Re-verified the remaining 11 deleted filenames against bin/install.js specifically
(plain substring and quoted-filename search) before trusting that list — all 11 have
zero hits there, confirmed dead by the same standard this one file failed.

Regenerated the installer emitted-tree goldens (tests/fixtures/install-tree/*.json) to
reflect the restored file, and corrected the "Removed" changeset (jolly-lynx-sprint.md)
and the phase design doc from 12 to 11 deleted files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: execute-plan.md — call-site wiring for the summary.md and user-setup.md .compact.md variants
Emitted-Drift-Ack-Growth: help.md — call-site wiring for full.compact.md, same variant-resolution rule

* docs(#4406): backfill changeset PR numbers

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4406): resolve removed-but-needed lint findings on the dead-template deletion

CI's own full-test matrix (not gsd-test's matrix, which does not run this
check) caught 4 more false-positive dead-template classifications via
tests/removed-but-needed-lint.test.cjs / scripts/lint-removed-but-needed.cjs
— a literal, word-boundary basename check across .github/workflows/,
gsd-core/, and docs/ (excluding docs/adr/** and docs/research/**) for every
file a PR deletes. It has no semantic awareness, so a deleted template's
basename colliding with something else entirely still fires:

- claude-md.md: gsd-core/templates/README.md had a stale table row claiming
  /gsd-profile reads this template to generate CLAUDE.md. Verified false (no
  code reads it anywhere, same search that already covered bin/install.js) —
  fixed the row to *(inline)*, matching every other command-generated
  artifact in that table. File stays deleted.
- codebase/testing.md: collided with docs/guides/testing.md, an illustrative
  example row in docs-update.md's sample output table (an unrelated real
  generated-docs path). Swapped the example topic to "contributing" — the
  row is illustrative, any topic works. File stays deleted.
- codebase/architecture.md, codebase/stack.md: collided with docs/reference/
  planning-artifacts.md's directory listing of a user's own generated
  .planning/codebase/architecture.md and stack.md output — the same
  semantic mismatch already investigated and dismissed as unrelated earlier
  in this phase's audit, now caught by a gate instead of judgment. That
  listing repeats across 5 locale copies of the doc.
- continue-here.md: collided with the real .continue-here.md pause-work
  artifact, referenced across 15+ locale and workflow files.

For the last two, the lint's own error message offers "restore the file or
update every consumer in the same commit." Rewording 15+ files across
languages I cannot verify translation quality for, to shave 2 already-tiny
templates that were merely presumed dead, is disproportionate to this PR's
actual scope — restored codebase/architecture.md, codebase/stack.md, and
continue-here.md instead, and corrected docs/ARCHITECTURE.md's Templates
section accordingly.

Final confirmed-dead set: claude-md.md, codebase/concerns.md,
codebase/conventions.md, codebase/integrations.md, codebase/structure.md,
codebase/testing.md, debug-subagent-prompt.md, discovery.md — 8 files, down
from the original 12. Verified locally: GSD_REMOVED_BUT_NEEDED_BASE=next
node scripts/lint-removed-but-needed.cjs now passes clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: docs-update.md — swapped an illustrative example-table topic (testing -> contributing) to avoid a removed-but-needed basename collision with the deleted codebase/testing.md template; net +10 bytes

* fix(#4406): split codex-config.test.cjs to fix a genuine Windows CI timeout

Root cause of the `full test (windows-latest, 24, shard 2/3)` failure the
user asked to be actually fixed, not just re-run past: PR #4497 (landed
2026-09-07, one day before this PR's CI run) isolated
tests/codex-config.test.cjs into its own dedicated chunk because its
measured weight (17.87, ~45% of the post-cut Windows budget) made it unsafe
to share a chunk with any other file. That isolation was necessary but not
sufficient — even alone, with zero companion-file contention, the file's
real Windows execution time sits right at the 600s per-chunk ceiling. Two
independent CI runs on two unrelated PRs (this one and #4154) were both
killed within ~1.4s of the identical 600000ms mark — not random contention,
a deterministic near-miss the isolation fix couldn't address because it
never reduced the file's own cost, only removed the risk of a companion
file's cost stacking on top of it (which the PR #4497 comment explicitly
anticipated: "if a future profiling pass genuinely speeds up
codex-config.test.cjs itself, this isolation can be revisited").

The file itself explains why it's this heavy: 11,262 lines / 433 tests / 79
describe blocks, accumulated over dozens of bug-fix PRs (#2695, #2760,
#3245, #3285, #3346, #3426, #3427, #3562, #3566, #3582, #3808, and more),
several of which are explicitly documented as "folded" in from separate
files that were never actually split back out ("Verified non-duplicate
against both the pre-existing target and the other three folded sources").

Split into 4 files by top-level AST statement boundaries (never a naive
column-0 regex — an early attempt at that overcounted 79 apparent
"describe(" matches when only 21 are genuinely top-level; the rest are
nested inside a handful of large folded-in blocks, which a regex can't tell
apart from real top-level statements). Verified lossless twice: the split
script asserts byte-for-byte reconstruction of every source character, and
independently, total test()/describe() call counts match exactly between
the original file and the sum across all 4 new files (433/79 both sides).
Each new file carries the complete original shared header (imports/helpers)
for safety; per-file unused-import warnings from that duplication are
resolved via ESLint-precise alias renames (`{ foo: _foo }`, the standard
form for an intentionally-unused destructured binding — never a bare `{
_foo }`, which would destructure a different, nonexistent property).

No change needed to scripts/run-tests.cjs's ISOLATED_HEAVY_FILES or its
pinned test in tests/run-tests-harness.test.cjs: the file that keeps the
original name (tests/codex-config.test.cjs) is now only ~28% of the
original's size and safely isolated in its own chunk as before; the other
three new files re-enter normal weight-balanced packing, none individually
close to disproportionate. Confirmed no other file hardcodes the hardcoded
filename anywhere that would silently stop these tests from running (the
CI test-selection scripts determine scope algorithmically, not by literal
filename).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 14:31:04 -04:00
Tom Boucher
66dbb104a0 fix(#4459): anchor update_codebase_map's diff base on the phase directory (#4549)
* fix(#4459): anchor update_codebase_map's diff base on the phase directory

execute-plan.md's update_codebase_map step derived its diff base from:

  git log --oneline --grep="feat({phase}-{plan}):" ... --reverse | head -1

A phase number is unique within a MILESTONE, not a repository (#3995).
`--reverse | head -1` deliberately selects the OLDEST matching commit
subject, so on a milestone that reuses a phase number, the diff base
lands in the PREVIOUS milestone's same-numbered phase -- silently
widening the file list that then drives which .planning/codebase/*.md
files get amended, with no warning and nothing downstream that would
notice.

This is the same defect class already fixed at two other sites in this
repo (code-review.md, structural-pre-pass.md) via a phase-DIRECTORY
anchor instead of a commit-subject grep: PHASE_START = the first commit
that ADDED anything under the phase directory, diffing from its parent
(or the commit itself on a root commit). Mirrored that exact pattern
here rather than inventing a new one.

Added tests/execute-plan-update-codebase-map-diff-base.test.cjs:
static regression guards (old grep gone, new #3995-shaped anchor
present) plus a real-execution test reproducing the issue's own
scenario -- two milestones reusing a phase number with a real
constructed git fixture, extracting and running the step's actual bash
fence, asserting the resulting diff is scoped to the current
milestone's files only.

Emitted-Drift-Ack-Growth: execute-plan.md — #4459 replaces the unbounded commit-subject grep with the phase-directory anchor already used by code-review.md/structural-pre-pass.md, net +687 bytes
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4459): cite this issue in the new file's allow-test-rule marker

gsd-test caught tests/lint-allow-test-rule-refs.test.cjs failing: the new
test file's `// allow-test-rule: source-text-is-the-product` comment
(copied from the two sibling precedent files) was missing the required
issue-ref suffix -- ADR-456 requires a NEW exemption to cite an issue via
`#NNN` on the same comment line. Added `(see #4459)`. Verified via
`node scripts/lint-allow-test-rule-refs.cjs` directly (clean) before
re-running the full suite.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4459): backfill changeset PR number

pr: 0 -> pr: 4549

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 14:25:56 -04:00
Tom Boucher
511c900052 fix(#4458): reuse detectSubRepos for new-project.md's sub-repo detection (#4548)
* fix(#4458): reuse detectSubRepos for new-project.md's sub-repo detection

new-project.md's Step 5.1 (Sub-Repo Detection) ran its own bash predicate:

  find . -maxdepth 1 -type d -not -name ".*" -not -name "node_modules" \
    -exec test -d "{}/.git" \; -print

`test -d` requires .git to be a DIRECTORY. A linked git worktree's .git is
a FILE (a `gitdir: <path>` pointer), so this predicate silently excluded
valid linked-worktree children while still finding ordinary clones.

src/core-utils.cts's detectSubRepos(cwd) already handles this correctly
(fs.existsSync, type-agnostic) but had zero callers anywhere in the
codebase -- orphaned logic the workflow never actually used, despite
duplicating a narrower version of the same check inline.

Wired detectSubRepos into cmdInitNewProject's JSON output as a new
sub_repos_detected field (matching the file's existing pattern of
similar directory-scan-derived fields like has_existing_code/
is_brownfield/has_codebase_map) and replaced the workflow's raw find
fence with a gsd_run query init.new-project call reading that field --
removing the duplicate, narrower detection logic entirely rather than
patching it in place, per the issue's own "reuse a central policy"
framing.

Added the missing .git-as-FILE test case to the existing
tests/core-utils.test.cjs detectSubRepos coverage (proving the helper
was already correct -- the defect was entirely in the unwired workflow
predicate) plus CLI-level end-to-end coverage in
tests/init-manager.test.cjs using a REAL `git worktree add` fixture,
matching the issue's own reproduction steps, alongside an ordinary
child-clone case and a non-repository-directory negative case.

Refreshed tests/fixtures/compact-content-benchmark-baseline.json (gsd-test
caught the drift from new-project.md's byte-count change; the benchmark
script itself always exits 0 -- report, not gate -- but the wrapper test
enforces the committed baseline stays in sync).

Emitted-Drift-Ack-Growth: new-project.md — #4458 replaces the raw find predicate in Step 5.1 with a gsd_run query call reading the new sub_repos_detected field, net +140 bytes
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4458): backfill changeset PR number

pr: 0 -> pr: 4548

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4458): bound the new git worktree add spawn with a named timeout

CI's lint-tests caught two ESLint findings my local gsd-test run
couldn't see (gsd-test's matrix doesn't run npm run lint:ci -- same gap
already observed on #4456's PR):

- local/no-unbounded-spawn: the new execFileSync('git', ['worktree',
  'add', ...]) call had no timeout, an indefinite-hang risk.
- local/no-adhoc-timeout-literal: my first fix (a bare `timeout: 15_000`
  literal) was itself flagged -- two independent hardcoded copies of a
  guessed timeout can silently drift or collide (this repo hit exactly
  that on 2026-09-06, PR #4428).

Fixed by importing GIT_FIXTURE_TIMEOUT_MS from tests/helpers/timeouts.cjs
-- `git worktree add` checks out files into a new working tree, the same
"construction" weight class as init/config/add/commit that constant
already covers, not plain plumbing (GIT_TIMEOUT_MS's class).

Verified via `npm run lint` directly (clean) before re-running gsd-test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4458): reduce redundant git subprocess overhead in new tests

CI's full-test Windows shard 2/3 failed: chunk 1/9 (306 files) exceeded
its internal 600s budget and was force-killed, with an unrelated file
(codex-config.test.cjs) in flight at the moment of the kill -- meaning
the chunk's AGGREGATE runtime, not any single hang, blew the budget.

This PR's own three new tests each independently called
createTempGitProject() (git init + a commit), and one of them also runs
git worktree add -- real subprocess spawns, each Defender-scanned on
Windows CI (tests/helpers/timeouts.cjs's own documented rationale for
why Windows spawn classes get generous budgets). That's a genuine,
quantifiable overhead addition to the exact chunk that timed out, not
something to wave off as unrelated flake without checking.

Two of the three tests never actually needed a real git repo --
detectSubRepos only inspects a CHILD directory's own .git, never the
root's git state, and the existing SUBCOMMANDS loop earlier in this same
file already proves `init new-project` succeeds against a plain,
non-git createTempProject() fixture. Switched those two tests to the
lighter fixture, leaving only the one test that genuinely needs a real
repo (git worktree add requires one) on createTempGitProject().

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 12:39:34 -04:00
Tom Boucher
147c89a9b8 fix(#4456): forward --ws to every downstream new-milestone.md call (#4545)
* fix(#4456): forward --ws to every downstream new-milestone.md call

new-milestone.md's Step 1 parses --ws <name> into GSD_WS, but each
workflow step's bash fence is a separate shell invocation — GSD_WS set in
Step 1 never survived to Steps 5, 6, or 7. Four call sites never
forwarded it: init.new-milestone (both calls), state.milestone-switch,
and both phases.clear branches. Under GSD_WORKSTREAM env or a stored
session pointer differing from the explicitly requested --ws, every
downstream operation silently operated on the wrong workstream (or root)
instead of the one the caller asked for.

Confirmed --ws is a universally-parsed CLI flag (gsd-core/bin/gsd-tools.cjs:
4867, resolveActiveWorkstream) — stripped from argv and written into
process.env.GSD_WORKSTREAM for the rest of that process, so appending it to
ANY gsd_run query call works uniformly. Fixed by persisting GSD_WS to
.planning/.gsd-ws-arg right after Step 1 parses it (mirroring the
established .gsd-outgoing-milestone round-trip idiom this same file
already uses for the identical cross-fence problem), reading it back in
each later step, and appending it unquoted (matching the ${GSD_WS}
splicing convention documented in workstream-flag.md). Cleaned up after
its last use in Step 7.

Bundled, in-scope fixes found while implementing the above (per this
repo's no-defer policy):

- Step 6's phase-archive `git add .planning/milestones/ .planning/phases/`
  hardcoded literal ROOT paths — both directories are workstream-scoped
  (matching cmdMilestoneComplete's established #1911 precedent), so under
  a workstream this staged nothing real. Added phases_dir/archive_dir
  fields to cmdInitNewMilestone and resolved through them instead.
- Step 6's milestone-start commit hardcoded .planning/STATE.md — also
  workstream-scoped, so it would commit the wrong (or a stale) file under
  a workstream. Resolved through init.new-milestone's existing state_path
  field instead; PROJECT.md correctly stays a literal-shaped-but-resolved
  root path (shared, per the #4455 follow-up already merged).
- cmdInitNewMilestone's config_path field: config.json is ALSO a shared
  file (marked `# Shared` in workstream-flag.md's directory diagram, same
  as PROJECT.md) but was resolved via the workstream-aware planningDir —
  fixed alongside cmdInitNewProject's identical instance of the same bug
  (found via grep, matching the precedent from the #4455 follow-up of
  fixing every occurrence of an identically-evidenced bug uniformly).

Verified: direct CLI invocation confirms phases_dir/archive_dir/state_path
resolve into the workstream while project_path/config_path stay root under
GSD_WORKSTREAM=alpha. Manual bash-fence execution of every modified fence
(Step 1 parse+persist, Step 5 forwarding, both phases.clear branches, the
git add fence, Step 7's forward+cleanup, the commit fence) confirms correct
behavior in both flat and --ws modes, including two flags composing
together (--archive-version + --ws; --reset-phase-numbers + --ws).

Rewrote the pre-existing "step 6: commit stages PROJECT.md" test, which
asserted the literal (buggy) --files string verbatim — it now asserts the
resolved paths via a JSON-returning stub, and gained isolated per-test tmp
dirs (the prior version ran with no explicit cwd, at real risk of writing
a stray .gsd-ws-arg into this repo's own .planning/ once Step 1's fence
started performing a real write).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4456): Steps 9/10 also commit workstream-scoped files via literal root paths

A fresh isolated code-review pass on the first version of this fix found
the identical bug in two more places, missed in the initial sweep:

- Step 9's requirements commit (`gsd_run query commit ... --files
  .planning/REQUIREMENTS.md`) and Step 10's roadmap commit (`--files
  .planning/ROADMAP.md .planning/STATE.md .planning/REQUIREMENTS.md`)
  both hardcoded literal ROOT paths for files that are workstream-scoped.
- Worse: `.planning/.gsd-ws-arg` was being deleted at the end of Step 7,
  but Steps 9 and 10 run AFTER Step 7 and still needed to re-read it —
  the round-trip mechanism this fix builds was already gone before its
  two remaining consumers ran.

Fixed by moving the `.gsd-ws-arg` cleanup to Step 10 (its true last
consumer, after the roadmap commit) and adding the same
fetch-then-_gsd_field-extract pattern already used in Step 6 to Steps 9
and 10, resolving `requirements_path`/`roadmap_path`/`state_path` through
`init.new-milestone $GSD_WS_ARG` instead of literal paths.

Also fixed (MEDIUM, same review pass): Step 1's `.gsd-ws-arg` write had
no `2>/dev/null || true`, unlike every other round-trip write in this
same file — brought into line with the established idiom.

Verified: reproduced the pre-fix bug directly (Step 9/10 fences echoing
the literal root paths regardless of --ws), confirmed both fences now
resolve the workstream-scoped paths correctly, and confirmed the
round-trip file survives Step 7 and is only removed after Step 10.
Updated the Step 7 test that previously asserted premature cleanup
(inverted to assert the file survives); added new coverage for Steps 9
and 10 in both flat and --ws modes.

Two remaining LOW/pre-existing findings from the same review pass,
deliberately left as-is: `phase_archive_path` (src/init.cts, untouched by
this diff) resolves via the same root-only `getLatestCompletedMilestone`
this fix's earlier commit already declined to touch, for the same
genuine-product-intent-ambiguity reason (workstream-scoped vs
project-pooled "latest completed milestone" is not resolvable from the
code alone). `.planning/research/` staying root-scoped in the #222
self-heal prose is consistent with the existing (unchanged) `research_dir`
field, not a new inconsistency.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4456): revert wrong config_path change; fix isolated-cwd test env

gsd-test caught two real regressions from this fix's earlier commits:

1. config_path is NOT shared like PROJECT.md. The prior commit's
   grep-and-replace ("fix six more functions with the identical bug")
   also touched cmdInitExecutePhase's config_path (a fourth call site
   beyond the two I'd manually checked) — but tests/init.test.cjs's
   pre-existing, ADR-0006-governed "init handlers honor GSD_WORKSTREAM"
   coverage explicitly asserts config_path IS workstream-scoped for
   execute-phase/plan-phase/phase-op/milestone-op. workstream-flag.md's
   "# Shared" marking for config.json is stale (the same class of
   staleness already found for milestones/ during the #4455 follow-up);
   ADR-0006 plus its real, passing tests is the authoritative source.
   Reverted config_path to the plain workstream-aware planningDir(cwd)
   in all four functions it was wrongly changed in.

2. Isolating cwd to a tmpDir (needed once Step 1's fence started
   performing a real .gsd-ws-arg write) broke the runtime-launcher
   preamble's own gsd-tools.cjs discovery — no git repo at an isolated
   tmpDir, no global gsd_run on the CI bench's PATH. Fixed by passing
   RUNTIME_DIR explicitly in every isolated-cwd test's env, matching
   the preamble's own documented override precedence.

Verified: direct CLI invocation confirms execute-phase's config_path is
workstream-scoped again under GSD_WORKSTREAM=wsx; the RUNTIME_DIR fix
confirmed against a stripped PATH (no global gsd_run), matching the
bench condition that surfaced the original failure.

Emitted-Drift-Ack-Growth: new-milestone.md — #4456 forwards --ws to every downstream gsd_run call across 7 fences (Steps 1/5/6x3/7/9/10), adding a persisted round-trip file plus resolved-path fetches that replace several hardcoded literal paths
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4456): backfill changeset PR number and correct final scope

pr: 0 -> pr: 4545, and removed the changeset's claim that config.json
is a shared file -- that was the change this same PR later reverted
after gsd-test caught it contradicting ADR-0006's established,
workstream-scoped config_path contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4456): baseline the 10 new SC2086 findings from --ws forwarding

The lint-tests CI job failed with a hard exit 1. Diagnosis (not assumed):
the log's two `fatal: ambiguous argument 'origin/next...HEAD'` git errors
(lines 244/248) are a red herring — both belong to
lint-removed-but-needed.cjs, which prints its own "could not resolve
origin/next, skipping" message and exits gracefully, exactly like the
already-handled two-dot-form error from lint-fix-has-regression-tests
earlier in the same log. Neither contributes to the actual failure.

The real cause is lint-workflow-shellcheck: this fix's new fences append
$GSD_WS_ARG unquoted to gsd_run calls (deliberately, so it splits into 0
or 2 argv tokens — the same idiom gsd-core/workflows/verify-work.md
already uses for ${GSD_WS} and already has baselined). ShellCheck
correctly flags each as SC2086, and lint-workflow-shellcheck.cjs's
baseline is a deliberate ratchet (#4109) requiring new findings to be
explicitly accepted, not auto-passed. new-milestone.md previously had
zero baselined SC2086 findings, so all 10 new (correct, intentional)
occurrences were reported as new and failed the gate.

Added 10 {file, code, message} entries to
scripts/lint-workflow-shellcheck-baseline.json for
gsd-core/workflows/new-milestone.md's SC2086 findings, matching the
established, already-accepted precedent for the identical pattern in
verify-work.md. No source or workflow file changed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 09:23:20 -04:00
Tom Boucher
c6df4e1e46 fix(#4455): autonomous.md and complete-milestone.md resolve STATE/ROADMAP/MILESTONES/PROJECT/REQUIREMENTS through the workstream-scoped init fields (#4542)
* fix(#4455): thread workstream-scoped paths through autonomous and complete-milestone workflows

autonomous.md and complete-milestone.md read/wrote hardcoded literal
`.planning/STATE.md` / `.planning/ROADMAP.md` / `.planning/milestones/...`
paths in their shell fences, bypassing workstream scoping entirely. With
GSD_WORKSTREAM=alpha set, planningDir(cwd) correctly resolves into
workstreams/alpha/, but a literal `cat .planning/STATE.md` still read the
ROOT file (or silently returned empty if root state was absent) --
reproduced deterministically in the issue's own repro.

Root cause: each workflow step's bash fence is a separate shell
invocation, and cmdInitManager/cmdInitCompleteMilestone's JSON payloads
never carried resolved state_path/roadmap_path/archive_dir fields for the
workflows to extract -- unlike cmdInitPlanPhase, which already does this
correctly and is the pattern this fix mirrors.

- src/init.cts: cmdInitManager and cmdInitCompleteMilestone now emit
  state_path/roadmap_path (workstream-scoped via planningDir(cwd),
  existence-checked, toPosixPath'd, null when absent -- identical to
  cmdInitPlanPhase's existing contract) and archive_dir (the milestone
  archive directory, composed the same way milestone.cts's already-correct
  archive helper does per #1911).
- autonomous.md: discover_phases and iterate now extract state_path via
  the already-fetched INIT_MANAGER payload instead of hardcoding
  `.planning/STATE.md`; iterate's second, previously-separate hardcoded
  read is folded into the same fence (no double-fetch); lifecycle step 5b
  checks the resolved archive_dir instead of a hardcoded milestones path.
- complete-milestone.md's reorganize_roadmap_and_delete_originals step
  (which previously called no init command at all) now fetches
  init.complete-milestone and uses the resolved roadmap_path/state_path/
  archive_dir for the backlog read, the write-guard sentinel's armed
  content, the Write-tool target for the reorganized ROADMAP.md (the
  sentinel fence now echoes the resolved path so the executing agent can
  see it), and the safety-commit --files list. `.planning/MILESTONES.md`
  and `.planning/PROJECT.md` stay literal root paths -- documented shared
  files, per the issue's explicit "not a blanket replacement" scope.

Regression tests extract and execute the real bash fences (with a stubbed
gsd_run) rather than string-matching the markdown, covering flat mode
(unaffected), an active workstream (the issue's own repro shape, now
correctly resolving), the no-double-fetch requirement, and a dedicated
guard locking MILESTONES.md/PROJECT.md as shared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4455): add changeset for workstream-scoped autonomous/complete-milestone fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4455): close write-guard gap on workstream-scoped curated paths

Isolated security review of the #4455 fix (workstream-scoped STATE/
ROADMAP/milestone-archive path resolution in autonomous.md and
complete-milestone.md) flagged that hooks/gsd-write-guard.js's
CURATED_PATTERNS only matched root-level .planning/ paths, never
.planning/[<project>/]workstreams/<ws>/... — meaning the catastrophic-
shrink guard silently never engaged for a workstream-scoped write.
This is directly relevant here: the #4455 change makes a workstream-
scoped ROADMAP.md Write reachable via complete-milestone.md's own
explicit sentinel-hatch instructions, which assume guard protection
that did not actually exist for that path shape. Extended
CURATED_PATTERNS with the three workstream-scoped equivalents;
consumeSentinelFor's own path-derivation logic needed no change since
it derives from the actual write target. Verified empirically (a
293->16 line workstream ROADMAP.md shrink now correctly returns
exit 2 / decision:"block") and with 5 new regression tests.

Also addressed a code-review nit on the core #4455 fix:
cmdInitCompleteMilestone called planningDir(cwd) three separate
times instead of caching it once.

Accepted as-is (not fixed): complete-milestone.md's
reorganize_roadmap_and_delete_originals step re-fetches
`gsd_run query init.complete-milestone` three times across its
fences rather than merging the first two (no state-changing Write
between them, unlike autonomous.md's iterate step which does merge).
This is an efficiency nit, not a correctness bug — merging risks
disrupting the step's prose flow and its existing binding test for a
non-functional gain.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4455): add changeset for the write-guard workstream-scope fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4455): fix gsd-test-surfaced regressions from workstream-path fix

Running gsd-test against the full #4455 diff (including the write-guard
security fix and the cmdInitCompleteMilestone caching nit) surfaced four
real, non-flaky failures, all direct consequences of editing
gsd-core/workflows/autonomous.md and complete-milestone.md:

1. tests/autonomous-converge.test.cjs pinned the OLD hardcoded
   `STATE_CONTENT=$(cat .planning/STATE.md ...)` read in both
   discover_phases and iterate. That is exactly the literal-path
   behavior #4455 fixes, so the test needed updating to assert the new
   init.manager-resolved `STATE_PATH` read instead (with an explicit
   doesNotMatch guard against regressing to the old literal).

2. tests/workstream-scoped-paths.test.cjs's own "no-double-fetch" test
   counted gsd_run invocations via a shell variable incremented inside
   the stub function — but `INIT_MANAGER=$(gsd_run ...)` runs gsd_run
   inside the command-substitution SUBSHELL, so that increment never
   survives back to the parent shell and the counter always read 0.
   Switched to a file-based call log (one byte appended per call),
   which survives the subshell boundary.

3. tests/compact-content-partition-guard.test.cjs's disjointness check
   flagged the reorganize_roadmap_and_delete_originals step's new
   `INIT_CM=$(gsd_run query init.complete-milestone)` fetch (added 3x,
   per the accepted-as-is disposition in the prior commit) as
   byte-identical to a pre-existing, unrelated fetch already present in
   complete-milestone/detail/elaboration.md's handle_branches section
   (§2). Same idiom, same conventional variable name, coincidentally
   colliding across the spine/detail split boundary. Renamed the new
   step's local variable to INIT_REORG — a distinct, purpose-specific
   name is arguably better practice anyway for two logically unrelated
   fetches, and it removes the literal collision honestly rather than
   restructuring the split.

4. tests/benchmark-compact-content.test.cjs reported real byte-count
   drift in the committed baseline (autonomous.md and
   complete-milestone.md both grew from the #4455 content). Refreshed
   via `node scripts/benchmark-compact-content.cjs --write`.

Verified: node scripts/benchmark-compact-content.cjs --check now
reports the baseline up to date; a standalone invocation of
checkDisjointness() against the real repo state now reports zero
violations across all 6 registered splits; manual bash-fence execution
of both the autonomous.md iterate fence (call count = 1) and the
complete-milestone.md backlog fence (with INIT_REORG) confirms correct
behavior.

Emitted-Drift-Ack-Growth: autonomous.md — #4455 workstream-scoped STATE.md path resolution replaces hardcoded literal reads
Emitted-Drift-Ack-Growth: complete-milestone.md — #4455 workstream-scoped STATE/ROADMAP/archive path resolution replaces hardcoded literal reads
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4455): MILESTONES.md/PROJECT.md/REQUIREMENTS.md are workstream-scoped too, and so is project-only mode

Fresh isolated code-review and security-review passes against the full
diff (run after the previous gsd-test-surfaced fixups landed) each
found one real, confirmed defect:

Code review: the safety-commit `--files` list and the REQUIREMENTS.md
`git rm` step both hardcoded `.planning/MILESTONES.md`,
`.planning/PROJECT.md`, and `.planning/REQUIREMENTS.md` as literal
root paths — but src/milestone.cts's cmdMilestoneComplete writes
MILESTONES.md via `planningPaths(cwd).planning` (the workstream base)
and PROJECT.md/REQUIREMENTS.md resolve the same way through
`planningPaths().project`/`.requirements` (src/planning-workspace.cts).
Only `todos` is the documented root-scoped exception (#4256); an
earlier version of this fix wrongly generalized that exception to
MILESTONES.md/PROJECT.md too, and the now-corrected test previously
enshrined that wrong behavior as intended. Under an active workstream,
the safety commit would have silently missed the actual files
`milestone complete` just wrote, and the git-rm step would have
targeted the wrong (root) REQUIREMENTS.md entirely. Fixed by exposing
`milestones_path`/`project_path`/`requirements_path` from
init.complete-milestone (src/init.cts) and resolving all three through
them, the same pattern already used for state_path/roadmap_path/
archive_dir. The four remaining literal MILESTONES.md/PROJECT.md
mentions elsewhere in complete-milestone.md (lines ~12-13, ~441, ~607,
~662) are display-only prose in status/summary message templates, not
actual file operations — left as-is; they are a cosmetic path-display
inaccuracy under an active workstream, not a data-integrity bug like
the two fixed here.

Security review: confirmed the write-guard fix from the prior commit
is correct and complete for workstream scoping, and independently
surfaced the same project-only gap the code-review pass above also
caught structurally: `CURATED_PATTERNS` had no pattern for
`.planning/<project>/...` (GSD_PROJECT set, GSD_WORKSTREAM unset) —
planningDir(cwd) supports that shape independently of workstream
nesting, so it is reachable, not hypothetical. Fixed by adding three
more patterns, verified empirically (a project-scoped 292->16 line
ROADMAP.md shrink now correctly returns exit 2 / decision:"block")
and with 6 new regression tests.

Verified: manual bash-fence execution of the corrected commit-files
and requirements-rm fences (both flat mode and GSD_WORKSTREAM=alpha)
resolves to the right paths in both cases; a standalone invocation of
checkDisjointness() against the real repo state still reports zero
violations; the benchmark baseline was refreshed again for the further
size change (already covered by the existing Emitted-Drift-Ack-Growth
trailer on complete-milestone.md two commits back — that trailer is
read over the whole merge-base..HEAD range, not per-commit, so it
still applies here).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4455): backfill changeset PR numbers and correct final scope

pr: 0 -> pr: 4542 for both fragments, and updated both bodies to
reflect the final fix scope (MILESTONES/PROJECT/REQUIREMENTS are
workstream-scoped too, not shared-root exceptions; the write-guard fix
also covers project-only scoping, not just workstream nesting).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4455): lifecycle-5b archive-path assertions use the fence's own separator, not path.join

PR CI's windows-latest shard 3/3 failed: "expected ls to find the root
archive file, got: ...\milestones-root/v1.0-ROADMAP.md". The
autonomous.md lifecycle step 5b fence composes the checked path with a
literal bash `/` (`"${ARCHIVE_DIR}/v${milestone_version}-ROADMAP.md"`),
which on Windows yields a MIXED-separator path — Windows backslashes
from archiveDir plus one trailing `/`. My test's assertion used
path.join(archiveDir, 'v1.0-ROADMAP.md') instead, which on a Windows
Node process produces an all-backslash path that never matches the
fence's mixed-separator output. Both assertions in that describe block
now mirror the fence's own literal `/` concatenation
(`${archiveDir}/v1.0-ROADMAP.md`) instead of path.join — matching the
style the other two describe blocks in this same file (safety-commit
--files list) already used correctly for the identical archive-dir
pattern, so this brings the one outlier into line rather than
introducing a new idiom.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4455): write-guard sentinel comparison now realpath-resolves the token, not just the target

PR CI's macos-latest full-test shard 2/3 failed a #4455 test: "the
sentinel hatch ... unblocks a workstream ROADMAP.md write" got status
2 (still blocked) instead of 0.

Root cause, unrelated to the Windows fix in the previous commit:
hooks/gsd-write-guard.js's main flow realpath-resolves the Write
TARGET before the curated-pattern match (round 9 Minor 1's
symlink-before-match fix, `filePath = fs.realpathSync(filePath)`), but
consumeSentinelFor resolved the sentinel TOKEN's absolute path via
plain path.resolve() with no realpath step. On macOS, os.tmpdir()
resolves through a /var -> /private/var symlink, so a test's cwd
(lexically under /var/folders/...) and its realpath'd target
(/private/var/folders/...) diverge — an armed, correct sentinel then
never matches the realpath'd target string, and the guard stays
incorrectly blocked. This is not macOS-specific in principle: ANY cwd
sitting under a symlink (a symlinked project checkout, a symlinked
worktree) hits the same asymmetry — gsd-test's Linux bench runs never
caught it because /tmp there is not a symlink.

Fixed by applying the same fs.realpathSync (with the same
keep-lexical-on-failure fallback the caller already uses) to the
token's resolved path before comparing. The named file is already
known to exist at this point (the caller only reaches consumeSentinelFor
after successfully reading the target), so realpath is expected to
succeed in the legitimate case; a garbage/mismatched token still fails
safe (verified — falls back to the lexical path, still mismatches,
stays blocked).

Verified: reproduced the exact bug locally (macOS) via os.tmpdir()
before the fix, confirmed it resolves after; the negative case
(sentinel armed for a DIFFERENT file) still correctly blocks; the
pre-existing relative-token sentinel tests (predating #4455) still
pass; a garbage/non-existent token still fails safe. Added a
deterministic, cross-platform regression test using an explicit
symlink (skipped on Windows, matching the existing round-9 symlink
test's own skip condition) so this class of bug is caught by
gsd-test's Linux bench too, not only by a real macOS CI run.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 04:56:11 -04:00
Tom Boucher
78013b3b74 fix(#4447): classify mixed structural+transient planning commits as a 5th arm (#4537)
* fix(#4447): classify mixed structural+transient planning commits as a 5th arm

pr-branch.md's analyze_commits step computed only NON_PLANNING and
STRUCTURAL per commit -- never a total planning-file count -- so its four
classification arms assumed every planning-only commit was either wholly
structural or wholly non-structural. A commit touching both a structural
.planning/ path and a transient/other one matched no arm, and the
ambiguous prose let an LLM executing the workflow silently drop it,
breaking STATE.md's per-commit revision chain in default mode.

Adds an explicit PLANNING_COUNT variable and rewrites the four arms into
five, each with an exact computable condition. The new "mixed planning
commit" arm (structural + transient/other, no code) gets the same
treatment mixed code+planning commits already get: INCLUDE, relying on
create_pr_branch's existing universal per-commit filter to strip the
transient/other paths -- no new filtering logic needed.

tests/helpers/pr-branch-filter.cjs's classifyCommit already returned
'include' for this shape (no upper bound on its structural check); the
defect was entirely in the workflow's own prose spec, which is what an
executing agent actually reads. New tests pin both: classifyCommit's
already-correct behavior (tests 49-50), and a failing-first assertion
that analyze_commits computes an explicit planning-total signal (test
51, fails against the pre-fix text).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4447): address code-review findings on the mixed-planning arm

- Correct the mixed-planning arm's prose: create_pr_branch's universal
  filter only strips the TRANSIENT_DIRS subset, not the "other" bucket
  (config.json, intel/, etc.) -- that subset is preserved, not filtered,
  same as default mode already does for it on any commit.
- Fix the "Mixed planning commits" display line to use the same
  mode-conditional bracket form as "Structural planning commits" --
  it was hardcoding "included" even though the arm is EXCLUDE in strict
  mode, which would have misled a strict-mode user.
- Tighten test 51's regex from unanchored /PLANNING_COUNT=/ to
  /^PLANNING_COUNT=\$\(/m so it requires the real shell-assignment
  shape, not just the substring appearing anywhere in prose.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: pr-branch.md — growth is this PR's own #4447 fix (5th classification arm with explicit computable conditions), not incidental drift

* fix(#4447): bound test 51's regex quantifier (local/no-unbounded-quantifier)

lint:ci flagged the unbounded [\s\S]*? over readFileSync content as a
catastrophic-backtracking risk (CWE-1333 class). Bounded to {0,20000},
comfortably larger than the analyze_commits step's actual size.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4447): add changeset for the pr-branch mixed-planning classification fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: eliminate SIGPIPE race in gsd-validate-commit.sh subject/config extraction

Discovered while verifying an unrelated PR (#4447): tests/hooks-opt-in.test.cjs's
"a git-GENERATED subject is never measured against the supplied message
(round 7)" test intermittently got r.status===141 instead of the expected 2
for the --fixup=HEAD case, on a run where the identical code had passed
cleanly moments earlier -- confirming a timing race, not a deterministic
bug in the test's own assertions.

Root cause: gsd-validate-commit.sh runs under `set -euo pipefail` and
extracted the commit subject via `SUBJECT=$(echo "$MSG" | head -1)` (two
call sites) and the opt-in ENABLED flag via `$(printf '%s\n' "$CONFIG_OUT"
| head -1)`. `head -1` closes its read end as soon as it has one line; a
real commit message or multi-command-type CONFIG_OUT is multi-line, so the
writer can receive SIGPIPE (exit 128+13=141) if its write lands after that
close. Under pipefail this is NOT suppressed -- it aborts the whole hook
instead of the intended exit-2 rejection.

Fix: replace both patterns with pure bash parameter expansion
(`${VAR%%$'\n'*}`) -- zero subprocesses, zero pipe/race surface, and
behaviorally identical to `head -1` for single-line, multi-line, and
trailing-newline input (verified directly). The third similar pipe
(`tail -n +2` feeding a `while read` loop that drains to EOF) is a
different, race-free shape and was left alone.

Regression test is a static, by-construction assertion (per this repo's
policy against forcing scheduling races to reproduce deterministically):
the vulnerable pipe patterns must be absent from the shipped script, and
the parameter-expansion forms must be present.

This overrides one-concern-per-PR per CLAUDE.md's Defects & Warnings
policy -- a genuine defect discovered mid-work is fixed inline, not
deferred to a separate issue.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs: add changeset for the gsd-validate-commit.sh SIGPIPE race fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: harden SIGPIPE-race regression test against reformatted reintroduction

Code review found the test's original exact-string regexes would miss a
cosmetically-reworded reintroduction of the same dangerous head-1 pipe
(extra whitespace, an appended 2>/dev/null). Broadened to content-tolerant
but still $(...)-wrapped regexes (bounded quantifiers per
local/no-unbounded-quantifier) -- verified against both the current file
(no false positive, including the fix's own explanatory comments that
quote the bare unwrapped pattern in prose) and a synthetic reformatted
reintroduction (correctly caught).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4447): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 00:19:03 -04:00
Tom Boucher
e03921c7d8 enhance(#4405): split the rest of the eager-window workflows worth splitting (#4536) 2026-09-08 00:17:22 -04:00
Dennis Alexis Valin Dittrich
18c899def5 enhance(#4209): optional external source reviewer lanes for /gsd:code-review (#4323)
* test(01-01): define reviewer-support trait contract

Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02):
validator rejects non-boolean values with an exact field path, accepts
missing/true/false, and the real code-review capability.json steps
must declare supportsReviewerLanes: true. Add loop-resolver projection
coverage proving the trait reaches activeHooks verbatim for a
provider-neutral synthetic step (not code-review-specific), and that
omitted/false values stay inert (no key on the active hook).

All 8 new assertions fail today: the validator has no such field, and
loop-resolver has nothing to project. RED before GREEN.

* feat(01-01): declare reviewer-capable steps

Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional
boolean opt-in trait, step-scoped (not capability-wide). Only a
literal true validates and projects; false/omitted stay inert (no
key on the projected active hook), and every non-boolean type fails
capability-validator.cjs with an exact field-path error.

Opt both existing code-review steps (execute:post, execute:wave:post)
into the trait in capabilities/code-review/capability.json. Project
the validated field through src/loop-resolver.cts into activeHooks
so a provider-neutral generic interpreter can read it without any
code-review-specific knowledge. Document the field in
docs/reference/capability-manifest.md and regenerate
gsd-core/bin/lib/capability-registry.cjs via the generator (never
hand-edited).

Makes all 8 RED assertions from the prior commit pass.

* test(01-02): define shared reviewer dispatch

- Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes:
  inert when the supportsReviewerLanes trait is off or nothing is selected,
  exactly-once plan/invoke per selected lane, duplicate-alias dedup, the
  bounded metadata-only source-review prompt (repo root, paths+baseSha,
  depth, four fixed prohibitions), and capability-neutral reuse via a
  second synthetic step context.
- RED: module under test (src/reviewer-step-dispatch.cts) does not exist
  yet, so require() fails and every assertion is unreached.

* feat(01-02): dispatch reviewers for opted-in steps

- Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps),
  ONE interpreter for a step's supportsReviewerLanes trait. Reuses
  resolveReviewerSelection for selection and resolveLanePlan for planning
  (both already-existing, pure building blocks); invocation is the one
  required, caller-injected seam (deps.invoke) since runLane needs
  OS-aware spawn plumbing this module does not own.
- trait !== true, or a selection resolving to zero lanes, dispatches
  nothing (zero plan/invoke calls). Each selected lane is planned and
  invoked exactly once, in the selector's deduped/sorted order.
- buildSourceReviewPrompt assembles a metadata-only bounded prompt
  (repo root, canonical paths + base SHA, depth, four fixed
  prohibitions) — never file contents — written once per dispatch and
  shared across every invoked lane.
- GREEN: tests/reviewer-step-dispatch.test.cjs now passes.

* test(01-02): define reviewer dispatch failures

- Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed
  matrix: an explicitly requested lane the selector could not resolve
  still lets the OTHER resolved lane run, but the aggregate result must
  never read as a clean success (and 'every explicit lane unavailable'
  must be distinguishable from the plain no-flags-passed inert case);
  request-level validation (path traversal, absolute paths outside
  repoRoot, empty/non-string paths, missing depth/base SHA) halts the
  whole dispatch before any lane is planned or invoked; a per-lane
  prompt-budget overflow hard-fails only that lane before invoke while
  its sibling still runs.
- RED: src/reviewer-step-dispatch.cts does not yet implement any of
  these guards, so 9 of the new assertions fail against the current
  (Task 1) implementation.

* fix(01-02): fail closed in reviewer dispatch

- src/reviewer-step-dispatch.cts: add the fail-closed guards the prior
  commit deliberately left out. An explicitly requested lane the
  selector could not resolve no longer lets the aggregate read as a
  clean success — lanes that DID resolve still run and keep their
  results (never narrow the requested set), but selection.errors now
  flips the aggregate ok to false, and 'every explicit lane
  unavailable' is now distinguishable (SELECTION_FAILED) from the
  plain no-flags-passed inert case (NO_LANES_SELECTED).
- Add request-level validation (validatePaths, depth/baseSha presence)
  that halts the WHOLE dispatch before any lane is planned or invoked:
  path traversal, absolute paths outside repoRoot, empty/non-string
  paths, and missing provenance are all rejected up front.
- Add per-lane prompt-budget enforcement (resolveBudget, mirroring
  gsd-tools.cjs's budgetFor convention including budget 0 = unbounded):
  a lane whose resolved budget the prompt exceeds hard-fails before
  invoke runs for it, without cancelling a sibling lane already
  planned.
- Document the supportsReviewerLanes trait and its dispatch-step
  interpreter in gsd-core/references/loop-hook-dispatch.md.
- GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass;
  no regressions in the review-lane/reviewer-selection/prompt-budget
  suites (356 passing).

* test(01-03): define optional source reviewer flow

RED: assert code-review.md dispatches roster-derived reviewer-lane flags
through a single review-lane dispatch-step call (DISP-01..05), that the
no-flag path stays byte-for-behavior unchanged (COMP-01), and that
external evidence reaching the internal reviewer prompt is marked
unverified (CONS-02). Also covers the CLI contract directly: no-op with
no explicit selection, and fail-closed on an explicit unknown lane
(SAFE-07) via real gsd-tools.cjs subprocess calls.

* feat(01-03): route optional source reviewers

GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches
canonical reviewer-lane flags against the merged first-party + installed
roster (never a hand-maintained list) and, only when at least one is
present, calls the shared reviewer-step interpreter exactly once with the
already-resolved repo root, file scope, depth, and base SHA. Its evidence
paths are appended to the internal reviewer prompt via
${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane
flag leaves the internal-only dispatch byte-for-behavior unchanged
(COMP-01).

Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane
dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI
route `dispatchReviewerLanes` wires through, but never implemented the
gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add
it to the existing review-lane router, reusing the same effort-aware plan
building and runner deps `plan`/`invoke` already use (factored into
buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard
the CLI's own `detected` set on whether an explicit flag was passed:
resolveReviewerSelection's no-explicit-selection fallback is "select every
detected reviewer" (the correct default for /gsd:review), and passing it
an unconditionally non-empty detected set would silently invoke the whole
roster on every no-flag code review, violating COMP-01.

* test(01-03): define external finding consolidation

RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as
untrusted input — independently re-verifies every claim against the actual
current source, resists a prompt-injection attempt embedded in evidence
text, and folds a verified claim into the existing Narrative Findings
section with no second REVIEW.md schema (CONS-01..03). Also assert
code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* feat(01-03): consolidate external review evidence

GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence>
as untrusted data, independently re-verifies every cited claim against the
actual current source before it can appear in REVIEW.md, and explicitly
resists prompt injection embedded in evidence text (never a command, no
matter what it claims to be). A verified claim folds into the existing
Narrative Findings section with (external: {slug}) provenance — one
REVIEW.md schema only, no separate external-findings section.
code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* fix(01-02): gitignore the reviewer-step-dispatch build artifact

01-02 added src/reviewer-step-dispatch.cts but never added its
npm run build:lib output to .gitignore, unlike every sibling
gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked
noise in git status.

* docs(01-04): publish user and command contract for reviewer-lane source review

- Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md
  and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on
  failure, findings independently consolidated into the single REVIEW.md
- Add the same contract to the docs/features/code-review-pipeline.md
  fragment and regenerate docs/FEATURES.md from it
- Preserve /gsd-review as the plan-review command; cross-reference it
  rather than duplicating the reviewer roster
- Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md
  drift owned by source already shipped in Plans 01-01/01-03 but never
  regenerated (npm run regen:derived had not been run in this worktree)

* docs(01-04): align architecture and agent ownership docs for reviewer-lane trait

- ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes)
  through the shared dispatchReviewerLanes interpreter to the existing
  review-lane plan/invoke machinery, ending at gsd-code-reviewer as the
  sole REVIEW.md consolidator
- AGENTS.md: document gsd-code-reviewer's full-context verification scope
  and its treatment of external reviewer evidence as unverified input
- No new diagram, abstraction, or config key; docs/CONFIGURATION.md is
  unchanged since the feature adds no setting or default

* fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact

Same gap as the earlier .gitignore fix: 01-02 added
src/reviewer-step-dispatch.cts but never added its generated
gsd-core/bin/lib/reviewer-step-dispatch.cjs output to
eslint.config.mjs's ignore list like every sibling generated file,
so tsc's emitted __importDefault CommonJS-interop var tripped
no-var.

* fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md

01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists
cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster
row in docs/INVENTORY.md — required by design, since a role sentence
cannot be generated — was never added.

* fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes

refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook
strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added
supportsReviewerLanes: true to that step and this fixture was not updated.

* chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md

Both files grew as a direct, intended consequence of wiring optional
reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes
step and the untrusted-evidence consolidation contract) — not
incidental drift.

Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209)
Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209)

* test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes

From internal code review: dispatched must be false when zero lanes
actually reached plan(), and a throwing plan()/invoke() for one lane
must not discard results already collected for a sibling lane —
matching the fail-closed pattern gsd-tools.cjs already uses for the
same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794).

Refs: gsd-core-dks.16, gsd-core-dks.17

* fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review

- WR-01: dispatched now tracks whether any lane actually reached
  plan(), not results.length — an unresolvable selected slug no
  longer reports dispatched:true.
- WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a
  throw for one lane can never discard results already collected
  for a sibling lane, matching the same guard gsd-tools.cjs already
  has around the identical resolveLanePlan call.
- IN-01: documents the intentional budget===0-is-unbounded
  convention (#2797) the caller already relies on.
- IN-02: review-lane dispatch-step no longer blocks indefinitely on
  an un-piped interactive TTY; fails closed to empty paths instead.

Refs: gsd-core-dks.16, gsd-core-dks.17

* docs(01-05): add changeset fragment for PR #17

* fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract

agents/gsd-code-reviewer.md's untrusted-evidence section and its
pinning regression test both quote injection phrases as the exact
attack they defend against/detect — same
DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing
allowlist entries, not an actual injection vector.

* test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip

From CodeRabbit review: WR-02's earlier fix only wrapped plan() —
writePromptFile()/deps.invoke() still ran unguarded, so a throw
there still aborted every later selected lane. Also covers the
dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit
lanes silently not running when no prior review and no phase-start
commit exist).

* fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved

Previously an explicit reviewer-lane request with no prior review and
no resolvable phase-start commit reached dispatch-step with an empty
--base-sha, which fails closed via missing_provenance — correct, but
silent about why explicitly requested lanes didn't run. Now skip
dispatch entirely in that case with a stderr warning naming the
actual cause.

* fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan()

WR-02's original fix only guarded plan() — a throw from
writePromptFile() or deps.invoke() still aborted the whole dispatch,
discarding results already collected for lanes processed earlier in
the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it.

* fix(01-05): WR-02b mock must throw only on the first writePromptFile() call

The committed mock threw unconditionally, so codex's retry also threw and
failed for the same reason as claude's — the test could not distinguish
'sibling still runs' from 'sibling also breaks'. Gate the throw to the
first call, matching WR-02/WR-02c's single-failure intent.

* fix(#4209): close review findings from adversarial + critical-code-reviewer pass

Two independent reviews (agy adversarial review, Opus critical-code-reviewer +
ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch
wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and
fixed here:

- dispatch-step's reducer silently swallowed whole-dispatch rejections
  (invalid paths, missing provenance, etc); it now checks parsed.ok/reason.
- spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the
  LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review;
  now shares the single compute_file_scope derivation.
- the external reviewer prompt had no actual review request or citation
  requirement, only prohibitions; added both.
- removed the supportsReviewerLanes trait plumbing (capability registry,
  validator, loop-resolver, docs, tests) — it was never consulted by the
  real dispatch path, which gates on explicit CLI flags instead.
- flag-resolution require() was a fragile cwd-relative literal that failed
  silently on non-vendored installs; now resolves via GSD_TOOLS's own
  directory and warns instead of swallowing failure.
- reducer didn't unwrap the @file: overflow protocol for large payloads.
- deduplicated resolveBudget/budgetFor into one resolveLaneBudget.
- lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a
  second dispatch can't overwrite prior evidence.
- validatePaths rejects control characters, closing a markdown-injection
  vector into the external prompt via crafted filenames.
- reworded the one line that tripped prompt-injection-scan.sh instead of
  allowlisting the whole production prompt file.
- fixed a stale docstring range and a dispatched-field ordering bug.
- added 3 integration tests executing the actual reducer against synthetic
  dispatch-step JSON, replacing markdown-substring-only assertions.

771/771 tests pass across every touched suite; tsc --noEmit clean.

* fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait

The maintainer's approval on issue #4209 explicitly redirected implementation
shape: reviewer-lane dispatch must be a reusable capability/step-dispatch
trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call
itself. My previous commit (e2558326) deleted that trait entirely after
finding it declared-but-never-consulted, which was backwards — the fix was to
wire it, not remove it.

Restores the trait (capability.json, generated registry, validator,
loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes
now resolves its own active hook via `gsd_run loop render-hooks` for the
configured workflow.code_review_point and only proceeds to CLI-flag matching
when supportsReviewerLanes reads true. Explicit flags no longer bypass the
trait; a matching flag with the trait false resolves zero slugs (proven by a
new integration test executing the real fence with both trait states).

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the
dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer
redirect requires the capability layer, not the workflow, own the opt-in
decision).

* fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point

Both an agy adversarial review and an Opus critical-code-reviewer pass
independently found the same gap in my previous commit (9b2c3773d): the trait
check I wired into code-review.md only protected code-review's OWN
invocation — gsd-tools.cjs's dispatch-step handler still hardcoded
`trait: true` unconditionally, so a second capability declaring
supportsReviewerLanes would get zero enforcement from the shared CLI unless
it correctly re-implemented the ~15-line render-hooks scrape itself. That is
exactly the "each workflow.md hand-wiring the call" the maintainer's redirect
said to eliminate.

Moves the trait check into dispatch-step itself: given --cap-id/--point, it
self-invokes `loop render-hooks <point>` (relocating the one subprocess
code-review.md used to spawn for this, not adding a new one) and derives the
real trait from that capId's active hook, rather than trusting a
caller-passed boolean. code-review.md now only passes
--cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or
gates on the trait itself — the ~20-line scrape it previously carried is
gone. Any other capability opts into the identical enforcement by declaring
the trait and passing the same two flags.

Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input
variable (they proved a bash branch honors a variable, not that the variable
reflects the real capability manifest) with three integration tests that
invoke the real dispatch-step CLI against the real first-party capability
registry: the real code-review trait resolves true, an unknown --cap-id
resolves false (trait_not_enabled, fail-closed), and omitting
--cap-id/--point entirely resolves false (no context means no opt-in).

Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check
(agy-F1 was incomplete), and delete the promptWritten per-lane coupling
flag — the prompt write is idempotent, so writing it once per lane instead
of gating on "did any lane write it yet" removes a latent bug where a
deps.plan override that ever varies promptPath per lane would silently skip
writing for a later lane.

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count
drops (the trait scrape moved into dispatch-step), but the file still grew
this session across multiple commits; acknowledging per the growth-tracking
convention.

* fix(#4209): remove per-run token waste from the shipped prompts

Runtime prompt content, not session tokens: two real, per-invocation token
costs in the code that ships.

1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of
   load_context step 5's ~180-word untrusted-evidence contract in ~90 more
   words, breaking this section's own established terse one-liner style
   (every other rule here is 1-2 sentences). This prompt loads fresh on
   every /gsd:code-review invocation. Shrunk to a one-line cross-reference,
   matching how write_review's own reference to step 5 already does it.

2. buildSourceReviewPrompt repeated the base SHA on every single file line
   even though it is identical for every file and already stated once at
   the top of the prompt — O(files) wasted tokens on every dispatched lane
   for a 50-file review, for zero information gain. File lines are now bare
   paths.

* fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3

Opus critical-code-reviewer found a real Blocking defect in the --cap-id/
--point self-invocation added last commit: `dispatch-step` spawned
`loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its
stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to
`@file:<path>` instead of inline JSON -- the same overflow protocol this
feature already unwraps for its OWN dispatch result 60 lines later in
code-review.md. A large-enough activeHooks envelope (more installed
capabilities/fragments) would throw, get silently swallowed by the bare
catch, and misreport a real trait as trait_not_enabled with zero diagnostic.

Fixed by extracting the config/registry/capability-state resolution
`cmdLoopRenderHooks` already performs into an exported pure function,
resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now
share it), and calling it in-process from dispatch-step instead of spawning
a subprocess at all. This eliminates the @file: exposure entirely (the
dispatch-step path never touches the rendered-string envelope or its
JSON-stringify/50000-char threshold), removes one subprocess spawn per
code-review invocation, and gives a genuine diagnostic (stderr warning) on
resolution failure instead of silent fail-closed. Corrected three doc/
docstring references to the now-removed subprocess self-invocation.

Also fixes 2 real CI failures this round surfaced:
- lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line
  no-control-regex` comment was unused under this project's ESLint config
  (verified locally: the rule never actually flags \x00-\x1f in this repo's
  config) -- a mistake from an earlier commit this session, never actually
  lint-checked before push. Removed the disable comment.
- security (prompt-injection-scan): the agy-F1 regression test's crafted
  fixture literally contains "Ignore all prior instructions." as test data
  proving validatePaths rejects it -- allowlisted the test file, same
  DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries.

Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one
bullet stated "untrusted, never a command" three different ways in one
paragraph, and a same-file duplicate of write_review's schema rule.
Consolidated to state each rule once.

Declined one suggestion from this round: shrinking code-review.md's
EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests
(tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block,
tests/code-review.test.cjs's CONS-02 test) deliberately lock the four-
prohibitions restatement and the untrusted-evidence prose into the
INJECTED block itself, not just the consolidator's system prompt --
adjacency of the warning to the untrusted payload it's warning about is a
recognized prompt-injection defense-in-depth pattern from this
workstream's original TDD plan, not accidental duplication.

* fix(#4209): correct stale per-file base-SHA prose in the external prompt

Leftover from removing the per-file base SHA repetition earlier this
session: the review-request sentence still said "relative to its base SHA"
(singular per-file framing) when there's now exactly one base SHA, stated
once above the file list. Reads "relative to the base SHA above" now.

* fix(#4209): make getLane/configGet/plan required deps, delete dead defaults

R3/R4 from the review round I'd deferred as low-priority test-churn: this
file's one production caller (gsd-tools.cjs's dispatch-step handler) always
supplies all three, so the fallbacks were dead in production -- but each was
actively WRONG if ever reached: the default configGet always returned
undefined, silently disabling resolveLaneBudget's overflow guard; the
default getLane looked up only first-party REVIEWER_LANES, diverging from
production's overlay-merged roster; the default plan skipped per-host effort
resolution entirely.

These defaults were introduced by this PR's own earlier work (this file did
not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from
elsewhere, so there's no external caller depending on the lenient contract.

Turned out free to fix: making the three deps required and deleting
defaultGetLane/defaultPlan needed zero test changes -- every existing test
that actually reaches the per-lane loop already supplies getLane/plan
explicitly, and configGet's only real dependent (the budget-overflow tests)
already supplies it too. 788/788 tests pass unchanged, tsc/lint clean.

* fix(#4209): define depth semantics for the external reviewer lane

Verified this was a real bug, not a match to existing convention as I'd
claimed when declining the suggestion earlier this session: the internal
gsd-code-reviewer agent's own system prompt carries a full <depth_levels>
block defining what quick/standard/deep mean and do (agents/gsd-code-
reviewer.md:68-99). The external reviewer lane has no access to that
persona at all -- it only ever sees buildSourceReviewPrompt's bounded text,
which sent the bare depth label with zero definition to a third-party CLI
with no other source of truth for what "standard" means.

Added depthMeaning(), condensed from the internal reviewer's own
<depth_levels> definitions so the two stay consistent, and interpolated it
into the review-request sentence. 150/150 tests pass, tsc/lint clean.

* fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation

CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the
roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a
SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide
whether to dispatch at all. This file's own documented rule (its
depth-resolution guard, stated explicitly a few hundred lines earlier) is
that a guard and the extraction it protects must run as one shell
control-flow decision, because markdown-fenced blocks do not share shell
state -- this step violated its own file's rule for the entire feature's
gating condition.

Merged the roster-resolution fence and the dispatch-decision fence into one
continuous bash block, removing the intervening prose that split them.
Fixed the stderr-based failure detection in the same edit (RQ-01: checking
whether stderr is non-empty misfires on any benign Node warning; now checks
the actual exit status of the roster-resolution command).

Verified by extracting the merged fence and executing it standalone, driving
both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a
real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty,
SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests
pass, tsc/lint clean.

* fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write

Batch of Required/Suggestion fixes from the Opus critical-code-reviewer +
writing-for-agents pass:

- CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch
  blocks, commented-out code) and deep (error propagation, state mutation
  consistency, circular dependencies) relative to the real <depth_levels>
  block, and had zero test coverage. Restored full accuracy and added tests
  that read the real agents/gsd-code-reviewer.md file directly, so drift
  between the two can't recur silently. Unrecognised depth now normalizes to
  standard's definition, matching that agent's own documented rule, instead
  of rendering an undefined bare label.

- RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt
  `paths` does, but weren't checked for control characters like paths were
  (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and
  applied it to all four fields at the same provenance-check boundary.
  runDir previously had zero validation at all.

- S1: deleted the dead `identity` parameter on `invoke` -- the one production
  caller already ignores it, no test read it by name.

- S2: hoisted the shared prompt write above the per-lane loop -- promptPath
  is derived from runDir alone (constant across lanes by construction), so
  writing it once is both correct and cheaper than the per-lane write R1
  introduced earlier this session. Discovered and fixed a real regression
  from the naive version of this hoist: an unguarded throw would have
  escaped dispatchReviewerLanes as an uncaught exception instead of a clean
  per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason,
  matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with
  a dedicated regression test.

- S3: moved `planned = true` past the budget-overflow gate, so `dispatched`
  only reports true once a lane has cleared BOTH plan and budget checks.

- S5: relayed gsd-code-reviewer.md's own "performance issues are out of
  scope unless also correctness issues" policy into the external-lane
  prompt, which previously had no such guidance and could return findings
  the internal reviewer's own contract excludes.

- RQ-05 (partial): shrunk this file's own header docstring's restatement of
  the trait-reuse architecture to a pointer at
  gsd-core/references/loop-hook-dispatch.md, the canonical home.

234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion

RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the
SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke`
already share. code-review.md's ~18-line inline `node -e` reimplementing
`loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in
gsd-tools.cjs) is now a single call to this subcommand -- the exact
violation code-review-flags.cjs's own header warns against ("this is the
canonical flag-parsing surface -- do not replicate inline bash parsing").

RQ-03: an empty --cap-id XOR --point now warns distinctly from the
legitimate no-context opt-out (both absent) -- a caller that named a
capability without its point was silently indistinguishable from a correct
opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only
ever fires when the config-get COMMAND ITSELF fails (config-get already
resolves the manifest's own schema default in the normal case), but that
failure was previously silent.

RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait
resolved inside dispatch-step" explanation was restated in full in 5
places across this session's own review cycles. Consolidated to ONE
canonical statement in gsd-core/references/loop-hook-dispatch.md; the other
4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md,
code-review.md's step-opening comment) now point at it instead.

W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two
inert cases when capability-validator.cjs already rejects non-boolean at
load -- restated as the two cases that actually reach this code. Removed a
"do not hand-roll trait resolution" prohibition whose target no longer
exists once the positive description precedes it.

W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing
block means proceed as normal") -- an absent optional block already means
proceed as normal without being told.

W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the
made-up compound "byte-for-behavior [un]changed" with the token this
session's own docs already coined for this concept (inert) and the word
that means what byte-for-behavior was reaching for (unchanged).

W-10: dispatch_reviewer_lanes had no completion criterion -- added one
sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set,
either populated or empty). This exact sentence would have caught the
cross-fence bug fixed two commits ago at authoring time.

Declined from this round, with reasoning: W-02/W-03 (trim the
untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) --
two tests deliberately lock this as intentional adjacency-based
prompt-injection defense-in-depth, not accidental duplication (see this
branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site
`trap ... EXIT`) -- would fire at the end of the CREATING fence, before
spawn_reviewer's agent ever reads the evidence files, given this file's own
documented fenced-block execution model; the existing named cross-reference
between creation and cleanup already satisfies the co-location concern
without introducing that regression.

853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex

Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed
for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get
fallback lived in an earlier, separate fence from the fence that consumes it
via --point, split only by prose (not a guard, per this step's own documented
rule). Merged into the single continuous fence and added a structural test
asserting exactly one bash fence in the step.

The new end-to-end regression test for this used --codex, which drives the
fence's real `review-lane dispatch-step` call and, with the codex binary
present on PATH, spawns the real external CLI — which then blocks on
interactive auth with no stdin (BL-01). Stubbed gsd_run for
`review-lane dispatch-step` only (captures argv instead of executing),
keeping the real config-get/explicit-from-argv calls the test is actually
about.

* fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment

Round-5 review (Opus) warning-tier findings:

- WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but
  a control-character injection attempt" — a caller distinguishing a config
  problem from a security event couldn't tell them apart. Split into
  MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid).
- WR-05: validatePaths' containment check was lexical only (path.resolve),
  so a symlink whose own path sits inside repoRoot could still point outside
  it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can
  legitimately name a file already deleted in a stale worktree), realpathing
  repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't
  false-positive-reject its own real children.
- WR-08: a comment in the per-lane loop still said a throwing writePromptFile()
  was caught there — stale since the prompt write was hoisted above the loop
  in an earlier round.

WR-03 (validate depth against the quick/standard/deep enum) was considered
and declined: this dispatcher is deliberately capability-neutral (see the
existing "synthetic step context" test, which passes a non-code-review depth
label on purpose to prove no code-review-specific special-casing exists).
WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and
WR-07 (reason omitted on the aggregate return) were verified against source
and are not bugs — see review notes.

* docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap

Round-5 review (Opus, BL-03) flagged that an early exit between
dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A
trap-based cleanup was considered and rejected: if a step genuinely runs as
a separate process, a trap set at creation time would fire at the end of
that SAME fence, deleting the directory before spawn_reviewer/commit_review
ever read it — worse than the leak it would fix.

review.md's own gather_context/cleanup pair for the identical resource class
(a run-scoped reviewer temp dir) already makes and documents this exact
trade-off: cleanup runs only on a documented success path, and a leftover
$TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording
that precedent here so this isn't re-raised as a live gap in a future review.

* fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path

reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes
paths: ['docs/spec.md'] as a synthetic, never-read path proving the
dispatcher has no code-review-specific special-casing. lint-docs-guard-
registration correctly flagged this as an unregistered docs/ path reference —
add the docs-guard-exempt marker and its pinned baseline entry, the same
pattern every other synthetic docs/ literal in this test suite already uses.

* fix(#4209): backfill changeset pr: field with the real upstream PR number

changeset-lint's fail_pr_field_drift caught the fragment still pointing at
the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this
branch is now also open against.

* docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam

trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every
decision in ADR-2782 (D1-D9) and every prior dated amendment governs the
`role: "reviewer"` capability body and its one consumer, /gsd:review. This
PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary
feature capability's `steps[]` entry, projected through loop-resolver.cts
and resolved in-process via resolveActiveHooksForPoint - is a different
capability axis (steps/gates/contributions) that the ADR's own scope note
explicitly places out of reach. Per docs/contributor-standards.md's
"Amending an accepted ADR", an in-place dated section is the established,
lighter-weight path for an addition that stays within the ADR's existing
decisions - used twice already in this same file - so this appends a third
dated entry documenting the new seam, its consumer, and why it reuses the
existing D1-D9-governed plan/invoke machinery rather than adding a second
one. No decision is reversed; no new Amends/Amended-by pair is needed since
the steps/gates/contributions axis already carries reciprocal links to
ADR-857 and ADR-894.

* fix(#4209): close two test-quality gaps trek-e's review found

Minor 1: validatePaths (a path-shape parser guarding the prompt-
injection/path-traversal trust boundary) had only example-based coverage,
violating ADR-456's rule that parsers/budget limits carry at least one
fast-check property test. Adds three: safe-segment paths are never
rejected, a single leading "../" always escapes the one-segment repoRoot,
and a control character anywhere is always rejected - one property per
rejection reason validatePaths owns.

Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only
ever exercised far below budget or at budget:0 (unbounded), never at the
exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds
three exact-boundary tests using the real estimateTokens/
buildSourceReviewPrompt the module calls internally, so the resolved
token count is exact rather than approximated: budget == estimate (must
pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must
pass).

Also extracts okPlan()'s fixture timeoutMs into a named constant -
local/no-adhoc-timeout-literal (#4446) landed on next after this branch
was authored and flagged the pre-existing literal on rebase; it is fixture
data for a synthetic plan object dispatchReviewerLanes never waits on, a
distinct class from tests/helpers/timeouts.cjs's real subprocess norms.

* fix(#4209): update docs-guard-registration baseline for the new ADR citation

reviewer-step-dispatch.test.cjs's new fast-check property tests cite
docs/adr/456-test-rigor-architecture.md in a justifying comment (never a
real read). lint-docs-guard-registration fingerprints every docs/ path
string an exempted test file mentions and fails on drift so a human
re-confirms the exemption still holds - re-confirmed, and the baseline is
updated to match.

* fix(#4209): point changeset pr: field at the fork PR for CI validation

changeset-lint's fail_pr_field_drift check compares the fragment's pr:
field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH),
not a fixed target. Rehearsing this branch on fork PR
davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's
pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but
fails here. Backfill to 4323 happens again, as the last commit, immediately
before the approved push to open-gsd#4323 - never leaving pr: 17 on the
branch that ships upstream.

* fix(#4209): reject promptChannel:none lanes from source-review dispatch

CodeRabbit found a real scope mismatch: coderabbit's lane declares
promptChannel: 'none' and reviews the working tree on its own terms,
fed nothing (review.md:367). Silently dispatching it through
dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope
buildSourceReviewPrompt promises and let the lane review whatever it
independently sees fit, violating this interpreter's own scoped,
metadata-only contract. Reject before plan()/invoke(), same as an
unresolved slug.

* fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file

CodeRabbit found the whole-file match on workflowContent would still
pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts
of this 1000+-line workflow, proving nothing about the actual evidence
block's contract. Line-filtered via splitLines (not a bare-\n regex
spanning readFileSync content) so this stays CRLF-portable and passes
local/no-unbounded-quantifier and local/no-crlf-fragile-split.

* fix(#4209): guard DISPATCH_JSON substitution and capture its stderr

CodeRabbit found the dispatch-step command substitution unguarded: a
non-zero exit could leave DISPATCH_JSON empty (or halt the step under
errexit with no warning), and the downstream reducer would only ever
report the generic unparseable_dispatch_output reason, discarding the
command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/
EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface
it in a warning on failure, and fall back to a parseable dispatch_
command_failed JSON stub so the reducer's existing reason-reporting
path still fires.

* docs(#4209): fix byte-for-behavior wording and missing colon, regenerate

CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the
established repo term for output-identical unchanged behavior) and a
missing colon after the bold "Optional external reviewer lanes (#4209)"
lead-in in docs/features/code-review-pipeline.md. Fixed in the two
hand-authored sources (commands/gsd/code-review.md, docs/features/
code-review-pipeline.md) and regenerated the two derived projections
(skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/
FEATURES.md via gen-features.cjs) so they stay in sync.

* fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI)

The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds
double-quoted JSON keys inside a single-quoted shell literal. That
extra quote density, inside an already quote-heavy ~8KB driver string,
passed bash -n and the full local suite on Linux but broke Windows
Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end
to end (#4209 round 5)` failed on two Windows CI shards with `bash -c:
unexpected EOF while looking for matching '''` — a Windows argv-to-
command-line re-quoting edge case, reproducible on rerun, not a flake.
Root-caused via gh api job logs plus a byte-identical local
reconstruction of the test's own driver script.

Fix: drop the fabricated stub. The downstream node -e reducer already
falls back to reason `unparseable_dispatch_output` on any JSON.parse
failure, so an empty/partial DISPATCH_JSON on command failure is still
handled correctly, with zero new quoting risk.

* revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI)

Two materially different mechanisms for the same CodeRabbit Nitpick
("Trivial | Quick win") both broke Windows Git-Bash reproducibly:
a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ...
matching '''") and, after removing that, a plain `head -1 "$VAR"`
inside a nested command substitution ("unexpected EOF ... matching
'"'"). Both passed bash -n and the full local suite on Linux every
time; both failed the SAME test deterministically on Windows CI. Two
attempts at the same class of fix (nested-quote construction near
this exact step) is the retry limit - reverting to the original,
already-shipped, Windows-verified unguarded form rather than
continuing to guess at a third quoting mechanism for a Trivial-
severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for
anyone attempting this again: the fix belongs outside this specific
markdown-fence-driver test harness (e.g., a real .sh helper script)
if it's worth doing at all.

* fix(#4209): backfill changeset pr: field to the real upstream PR before push

Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy
changeset-lint's PR-number check while rehearsing there; this is the
last commit before the approved push to the real upstream PR
(open-gsd/gsd-core#4323), so the field points at that PR number again.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 22:52:33 -04:00
Tom Boucher
423f38e655 fix(#4444): honor config-set --dry-run instead of silently ignoring it (#4504)
* test(#4444): failing-first regression coverage for config-set --dry-run

config-set --dry-run is currently parsed nowhere -- routeConfigSet
(gsd-core/bin/gsd-tools.cjs) never checks args for it, and cmdConfigSet
has no dry-run parameter, so the flag is silently swallowed and the
command always writes for real. Reproduces the issue's own repro
(sequential --dry-run calls where the second's previousValue proves
the first persisted), plus coverage for validation-still-runs,
secret-masking, and the sibling unset (config-set <key> null) branch,
which has the identical defect. This commit adds the regression
coverage only; the fix lands in the next commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4444): honor config-set --dry-run instead of silently ignoring it

routeConfigSet (gsd-core/bin/gsd-tools.cjs) never read args for
--dry-run, and cmdConfigSet had no dry-run parameter at all -- so the
flag was silently accepted (as any unrecognized trailing argument is)
and the command always wrote for real. A second "dry run" then showed
previousValue reflecting the first one, proving it had persisted.

Threads a dryRun option through cmdConfigSet, gating BOTH mutating
branches: the null/unset path (unsetConfigValue) and the real-set path
(setConfigValue) -- the unset branch had the identical defect,
undiscovered until auditing every mutation site while designing this
fix. Each gains a previewConfigValue/previewUnsetConfigValue
counterpart that reuses the real function's exact traversal/creation
logic (_setNestedValue/_unsetNestedValue) on a throwaway in-memory
config copy that is never written -- so the preview can never diverge
from what the real write would compute. All validation (unknown key,
enum/number/boolean checks, secret masking) runs identically whether
or not --dry-run is passed; only the final write is skipped, replaced
with a `{ dry_run: true, would_update / would_unset: true, ... }`
preview payload matching the precedent established by `milestone
complete --dry-run` (#2118) and `todo complete --dry-run` (#4096/#4325).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* refactor(#4444): extract loadConfigJson to stop a 5th copy-paste of the same load/parse block

Code review flagged that setConfigValue, unsetConfigValue,
setConfigValues, and the two new preview functions each repeated the
identical "load .planning/config.json, JSON.parse, catch ->
CONFIG_PARSE_FAILED" block -- exactly CLAUDE.md's own
"Generative Fix Divergence" known-defect pattern. Extracted a single
loadConfigJson(cwd) helper; behavior is unchanged (verified: build,
tsc, and the dry-run/real-write smoke test all pass byte-identical to
before).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4444): changeset for the config-set --dry-run fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4444): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4444): raise per-chunk CI test timeout to 800s for Windows headroom

install-minimal-hooks.test.cjs (weight=24.45, the heaviest file in the
suite) sits alone in its own chunk yet still occasionally brushed the
600000ms per-chunk ceiling on Windows -- observed on PR #4504's first
CI run for this change (passed clean on rerun, consistent with the
"legitimately too slow for the budget" cause the chunk-timeout
diagnostic already names, not a leaked handle).

Raised RUN_TESTS_CHUNK_TIMEOUT_MS's default from 600000ms to 800000ms:
~33% more margin, still comfortably below the 900000ms regen:derived
fixture timeout that fragment-single-edit-propagation.install.test.cjs
deliberately keeps ABOVE the chunk ceiling, and far under the 45-minute
job cap -- Windows shards currently finish in ~19-20 minutes total, so
there is ample headroom. Updated every dependent mirror/assertion in
lockstep (tests/helpers/emitted-runtime.cjs's duplicated
CHUNK_TIMEOUT_CEILING_MS constant, its lock test in
tests/emitted-attribution.test.cjs, the Windows-skip prose in
fragment-single-edit-propagation.install.test.cjs, and
docs/TESTING-SUITES.md's reference table) so nothing describes a stale
value.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Revert "fix(#4444): raise per-chunk CI test timeout to 800s for Windows headroom"

This reverts commit 394aadaf6f5af6fd700bf0f444c9fbd686285a4f.

* test(#4444): consolidate redundant installer spawns in install-minimal-hooks.test.cjs

This file's real, unrelated pre-existing cost (dated 2026-09-06, PR #4428)
is what tipped a Windows CI shard over the per-chunk timeout backstop on
PR #4504 (issue #4444's own diff never touches this file or the
installer). Rather than raise the timeout, cut the file's actual spawn
count: several describe blocks independently re-installed the IDENTICAL
runtime/scope/flag configuration just to assert different things about
the same install output. Merged each such group onto a single shared
install, with every original assertion preserved:

- --help x3 -> x1
- the three per-runtime/scope --minimal E2E loops (global, local, and
  on-disk-matches-manifest) merged into one loop over
  SKILL_RUNTIMES x [global, local]: 44 spawns -> 22
- the --minimal manifest-mode/backcompat triple-install -> one shared,
  memoized install via sharedMinimalManifestInstall()
- .sh hooks existence checks (5 tests) -> 1, executable-bit check (its
  own Windows-conditional skip) left separate
- Codex #4087 hook-helper tests (3) -> 1
- Windsurf #4087 hook-helper tests (2) -> 1
- pi shared-hooks-bundle tests (3 per scope) -> 1 per scope

Net: ~65 real installer spawns in this file down to ~29, no assertion
dropped or weakened.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 21:08:06 -04:00
Tom Boucher
8dcdcb253e fix(#4443): register hooks.commit_types (and sibling hooks.community) in config schema (#4501)
* test(#4443): failing-first regression coverage for hooks.commit_types config key

isValidConfigKey('hooks.commit_types') currently returns false and
config-set hooks.commit_types rejects with "Unknown config key",
because the key was never added to config-schema.manifest.json's
validKeys when it shipped (#3811/#4340, 1.13.0) despite being
documented (docs/COMMANDS.md) and consumed by
hooks/gsd-validate-commit.sh. This commit adds the regression coverage
only; the manifest fix lands in the next commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4443): avoid false-positive docs-guard registration trip

The assert message for the new hooks.commit_types test mentioned
"docs/COMMANDS.md" literally, which happened to land between two
unrelated pre-existing backticks and tripped
lint-docs-guard-registration.cjs's template-literal co-occurrence
detector (a known, documented false-positive shape for that lint).
Rephrased to drop the literal docs/ path from the message; the test's
intent (documenting why the key must be valid) is unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4443): register hooks.commit_types (and sibling hooks.community) in config schema

config-schema.manifest.json's validKeys never got hooks.commit_types
added when the feature shipped (#3811/#4340, 1.13.0) despite it being
documented (docs/COMMANDS.md) and consumed by
hooks/gsd-validate-commit.sh -- so config-set hooks.commit_types
rejected with "Unknown config key", and the only way to configure a
documented feature was hand-editing .planning/config.json.

While auditing every hooks.* key actually read by shipped code against
validKeys (CLAUDE.md's no-deferrals rule: a defect found anywhere in
the tree while working an issue is fixed in the current change, not
filed separately), hooks.community -- gsd-validate-commit.sh's own
opt-in gate -- turned out to have the exact same gap. Both are added
here; an audit of every hooks.* read site confirmed these are the only
two missing entries.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4443): e2e coverage for hooks.community + changeset

Closes the coverage-rigor gap the Standards review flagged: hooks.community
had only a unit-level isValidConfigKey assertion, not the same real
config-set CLI round-trip hooks.commit_types already got. Also adds the
changeset fragment the same review flagged as a missing hard requirement.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4443): use PROBE_TIMEOUT_MS instead of a bare 15000 literal

local/no-adhoc-timeout-literal (lint:ci) correctly flagged both new
spawnSync calls' bare timeout: 15000 -- this call class (a short CLI
probe against a temp fixture) is exactly what tests/helpers/timeouts.cjs's
PROBE_TIMEOUT_MS documents.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4443): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 17:29:22 -04:00
Tom Boucher
a0f8f956c4 enhance(#4139): Phase 3 — partition rules + the five checks (#4497)
* enhance(#4139): Phase 3 — partition rules + the five checks

ADR-4139 Decision 5, epic #4139 Phase 3. Issue #4403's own "Proposed behavior"
section lists four checks; the ADR's Decision 5 and its own phase table ("partition
rules + the five checks") list five — the same four plus "boundary moves are
declared, ongoing". Same issue-vs-ADR drift Phase 2 hit on the detail.md vs
detail/*.md layout: the ADR is the locked, reviewed document, so it wins. This PR
implements all five.

docs/PARTITION-RULES.md (new) is the partition-rules document: the partition rule
itself, the protected-content list and <!-- gsd:protected --> sentinel syntax
(relocated unchanged from gsd-core/references/compact-content-protected-content.md,
now deleted — it was never referenced by any runtime workflow Read, only by the
predecessor test as documentation, so nothing at runtime regresses, and removing it
from gsd-core/references/ also drops it from all 19 installed-project shipped-content
trees for a file nothing ever read), and the five checks explained for a human
reader. Referenced from a new CONTRIBUTING.md subsection under "Editing shipped
content".

tests/helpers/compact-content-split.cjs (new) is the shared mechanics: split
discovery (any gsd-core/workflows/<name>/detail/*.md paired with <name>.md — no
registry file, a pair is registered by existing on disk), line normalization
(carries forward Phase 2's bare-label-line isTrivial fix and the canonical
gsd_run-launcher-preamble exclusion), sentinel extraction, and a
Boundary-Move-Declared commit-trailer reader that is a direct structural port of
tests/helpers/emitted-runtime.cjs's Emitted-Drift-Ack-Hash/-Growth trailer reader
(ADR-3942) — same merge-base range, same fail-closed throw on an uncomputable range,
same dedupe/conflict rules.

tests/compact-content-partition-guard.test.cjs (new) is the actual guard, superseding
tests/plan-phase-compact-split.test.cjs (deleted — its per-pair checks are now the
general guard's job for plan-phase specifically). Checks 2 (disjointness) and 3
(registration + size cap) run unconditionally against every registered split. Checks
1 (completeness, fires once per split on the PR that introduces a new detail/ path),
4 (protected content — no trailer can ever excuse this one, unlike check 5) and 5
(boundary moves declared) are PR-diff-scoped against the resolved base ref and skip
cleanly when there's nothing to compare (a fresh clone, no PR in flight) — a
deliberate asymmetry from check 5's trailer reader, which must throw rather than
silently pass when ITS range is uncomputable, since that function is answering "did
this PR declare its moves" rather than "is there even a diff to look at". Each of the
five checks carries a RED (deliberately broken fixture) / GREEN (fixed) test pair,
built against synthetic temp files or real throwaway git repos, per this repo's rule
that a guard nobody has seen go red is not yet a guard. Building the real fixtures
caught and fixed one real bug before it shipped: check 4's line-presence test was
using the trivial-line-filtered normalizer, so a byte-identical spine falsely
reported its own protected code-fence line as "deleted" — fixed with a
non-filtering membership check.

Extending docs/INVENTORY.md's "Workflow Sub-Files" table for `detail` surfaced a
pre-existing, unrelated gap in the SAME area: gsd-core/workflows/<name>/templates/*.md
is a fourth workflow sub-file kind that already existed on disk and was already known
to lint-response-language-coverage.cjs's FRAGMENT_DIRS, but was invisible to
gen-inventory-manifest.cjs and undocumented in that table. Fixed alongside it, same
pattern, same PR, rather than deferred.

Also, mechanically required by the new fourth sub-file kind:
- scripts/lint-response-language-coverage.cjs: `detail` added to FRAGMENT_DIRS
  alongside modes/steps/templates — a detail/<part>.md inherits its parent's
  response_language coverage through the same per-file proof, not a parallel one.
- tests/workflow-size-budget.test.cjs: explicit regression test locking that
  detail/ files are governed solely by the hard, non-waivable NEW_FILE_CAP
  (tests/helpers/emitted-diff.cjs) and never by the XL/LARGE/DEFAULT spine tiers —
  true by construction (measureWorkflows/listWorkflowStems don't recurse), made
  explicit per the issue's own Done-when item rather than left true-by-omission.
- scripts/gen-inventory-manifest.cjs: `workflow_detail` and `workflow_templates`
  NESTED_FAMILIES entries; docs/INVENTORY-MANIFEST.json regenerated
  (plan-phase/detail/elaboration.md, discuss-phase/templates/*.md now tracked);
  docs/INVENTORY.md's table updated to four kinds.

Verified: `npm run lint:ci` clean with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4403): review findings + a real gsd-test failure in the new guard

Two orthogonal review passes (Standards + Spec, isolated sub-agents) plus a
separate security review ran against the prior commit. Fixed everything each
surfaced:

- Security (Low, path-traversal existence oracle): checkRegistration's
  dangling-reference check extracted detail-path-shaped substrings from spine
  PROSE via a regex that permits `.`/`/` freely, then joined them onto repoRoot
  and probed fs.existsSync with no containment check — a spine file containing
  `../../../etc/detail/passwd.md`-shaped text could make the guard test file
  existence outside the repo. Added a path.relative-based containment check
  before the fs.existsSync call; anything that resolves outside repoRoot is now
  reported as a dangling reference directly, never probed on disk.
- Standards (Boundary Coverage): the size-cap fixtures covered NEW_FILE_CAP and
  NEW_FILE_CAP-1 but not NEW_FILE_CAP+1 — added the third boundary-point case
  CLAUDE.md's TEST RULES require (limit-1/limit/limit+1).
- Standards (Property-Based Testing): extractProtectedBlocks (a sentinel
  parser) and the new parseBoundaryMoveTrailerValues (a declare/dedupe/conflict
  parser, bijective-shaped) had no fast-check property test. Added three: a
  render/parse bijectivity property for the trailer parser (mirroring the exact
  ADR-3942 sibling test's alphabet/idiom), a dedupe-is-idempotent property for
  the same parser, and a well-formed-sentinel-round-trips property for
  extractProtectedBlocks.

Then dispatched gsd-test on the resulting commit. It found a real bug the
reviews couldn't have caught (none of them can run inside gsd-test's sandbox):
checks 4/5's real-repo assertion failed against plan-phase's own split,
reporting DISK_PLANS/#3218-comment lines as "undeclared boundary moves" —
content Phase 2 (#4402) legitimately moved into detail/elaboration.md months
before this PR's Boundary-Move-Declared mechanism existed to require a
trailer for it. Root cause: `resolveBase()`'s own doc comment already documents
that no `origin/*` remote-tracking ref exists inside the gsd-test sandbox
container, and its fallback candidate (a bare `next` branch) can resolve to a
point in history that predates an already-merged, already-reviewed split —
making that split look "newly introduced" from the sandbox's vantage point.
Check 1 (completeness) already scopes itself correctly to only genuinely-new
detail paths (git diff status 'A'); checks 4 and 5 did not share that scoping,
so a stale base made them re-litigate a settled split retroactively. Fixed by
having checks 4/5 skip any split name check 1 already counted as newly-split —
their own premise ("did an EXISTING split shed/undeclare something") does not
apply to a split that is, from the resolved base's vantage point, brand new;
that is check 1's domain alone. Verified locally (25/25 tests pass via a
direct `node -e` require, since `node --test` is blocked in this repo) and via
re-reasoning through the exact real-repo scenario the gsd-test failure showed.

Also regenerated all 19 tests/fixtures/install-tree/*.json goldens — the
prior commit's deletion of gsd-core/references/compact-content-protected-content.md
was never reflected there, which is what golden-install-tree.test.cjs's other
19 failures in the same gsd-test run were.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4403): backfill changeset pr number to 4497

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4403): isolate codex-config.test.cjs into its own chunk, root-causing the Windows CI failure

PR #4497's "full test (windows-latest, 24, shard 2/3)" job failed: run-tests
killed chunk 3/8 at the 600s per-chunk backstop, with codex-config.test.cjs
(weight 17.87, by far the chunk's dominant cost) packed alongside 39 other
files. Traced, not assumed:

- scripts/run-tests.cjs's own timeout-headroom comment for the OUTER
  per-shard timeout documents that "adding one test file reshuffled 115 of
  268 unit files between shards" — shard/chunk composition is architecturally
  known to be unstable to single-file additions, which is exactly what this
  PR's own new tests/compact-content-partition-guard.test.cjs is.
- A second comment, dated 2026-09-06 (one day before this PR, PR #4428's own
  CI), already documents the SAME chunk hitting the SAME 600s backstop with
  the SAME file (codex-config.test.cjs, "a genuinely MEASURED weight of
  17.87 — not a stale-table miss") dominating it — the fix then was cutting
  the Windows per-chunk budget from 60 to 40. That cut clearly was not
  enough: two documented incidents in two days, at two different budget
  settings, both centered on one file that alone consumes ~45% of even the
  reduced Windows budget.
- tests/test-timings.json's own header confirms its source data
  (test-events-linux-node22/24.jsonl) is Linux-only, and run-tests.cjs's own
  chunk-timeout diagnostic already prints "real Windows cost runs ~2.2x the
  recorded figure" — the packer's weight-balancing is working off data that
  is both stale (table last regenerated 2026-08-07) and known to
  underestimate the platform where the failure occurs.

Given codex-config.test.cjs is disproportionately heavy AND every companion
sharing its chunk is decided by a packing algorithm already documented as
reshuffling unpredictably on any new file, tuning the shared budget a third
time only moves the marginal line to wherever the next new file happens to
land — it does not remove the gamble. Isolating codex-config.test.cjs into
its own dedicated single-file chunk, unconditionally and on every platform,
removes it at the source: the file never enters the pool packChunks balances,
so no other file's packing changes, and no future single-file addition
(mine or anyone else's) can silently reintroduce this exact failure by
landing in its chunk.

Extracted as a small pure function, partitionIsolatedFiles (mirroring this
file's existing pattern of pulling packing/analysis logic out of main() for
in-process unit coverage — see computeSweepProtectSet, analyzeChunkEvents),
with 6 new tests in tests/run-tests-harness.test.cjs covering basename
matching across path separators, near-miss non-matches, the empty-list case,
and the isolated-set contents.

Root cause is now closed rather than papered over with a retry: this failure
is a property of one specific heavy file's chunk placement, not something
that recurs randomly. If codex-config.test.cjs itself is ever genuinely sped
up, this isolation can be revisited — this is a packing-side mitigation for
a known file's cost, not a claim the cost is irreducible.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 16:44:53 -04:00
Tom Boucher
c3a18b5ba0 docs(#4440): stop telling agents to grep .env files the secret guard denies (#4500)
* docs(#4440): stop telling agents to grep .env files the secret guard denies

verification-patterns.md's <environment_config> and user-setup.md's
three per-service Verification examples documented reading .env/.env.local
directly via grep. Every covered runtime's secret-read guard denies
that (Claude Code deny-rules since #768/v1.4.0; the always-on
gsd-secret-read-guard hook since #4236/#4221 in 1.13.0) -- verified by
piping each documented command through the shipped hook.

verification-patterns.md now checks the environment (printenv) instead
of the file, with a case statement replacing a broken grep -v
alternation (grep's BRE `|` is literal, so the old placeholder filter
matched nothing -- PLACEHOLDER/TODO_fill values passed the "substantive"
check as real). Verified under sh (dash) against real/placeholder/empty/
unset values. Existence check ([ -f ".env" ] || [ -f ".env.local" ])
is untouched -- it was never denied.

user-setup.md's three grep <SERVICE> .env.local lines are removed
outright rather than swapped for printenv: those examples describe a
Next.js shape where the framework loads .env.local at runtime without
exporting it to the shell, so a printenv substitute would wrongly
report "not set" on a correctly configured project. Each block's
existing service-level check (build/webhook/connection/email test)
already verifies the setup.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4440): changeset for the secret-guard verification-examples fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4440): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 14:10:08 -04:00
Tom Boucher
8b7a0b696b enhance(#4139): Phase 2 — one shared gate, one pilot split, one accuracy spot-check (#4471)
* enhance(#4402): split plan-phase into a spine + detail, add the shared compact-content gate

ADR-4139 Decisions 3-5, Phase 2 of the #4139 Compact Content epic. Pilot
split for plan-phase.md, the largest of the 58 eagerly-@-included workflow
files (98,290 bytes): the spine keeps every happy-path step, every
protected-content block (planner/checker prompt templates, quality gates,
the failing-direction few-shot example, the two ScheduleWakeup guardrail
paragraphs — each marked with a <!-- gsd:protected --> sentinel), and
condensed one-paragraph summaries of five rare/opt-in fallback paths
(planner and checker filesystem-hang recovery, phase-split recommendation,
source-audit gaps, the thinking-partner conditional, and plan bounce). The
full text of those five moves verbatim to gsd-core/workflows/plan-phase/detail.md
(9.9KB, well under the 32,768-byte NEW_FILE_CAP), read by the spine only
when workflow.compact_content is false (the default) — the exact same
resolution rule now stated once in the new shared
gsd-core/references/compact-content-gate.md, which every future split
references instead of restating.

Verified mechanically (tests/plan-phase-compact-split.test.cjs, scoped to
this one split — Phase 3/#4403 owns the generalized guard): the union of
spine + detail contains every non-trivial line the parent commit carried
(0 missing), no non-trivial line is duplicated between them (0 duplicated),
and every declared protected block is well-formed and non-empty. The spine
shrinks from 98,290 to 93,206 bytes (-5.2% of the eager-window cost this
epic exists to reduce); detail.md's 9,853 bytes are only ever paid by a
project that has NOT opted in.

Verified live, end to end, twice, against this actual repo (not a
synthetic fixture) — real gsd-planner and gsd-plan-checker subagent
spawns, real PLAN.md output:
- workflow.compact_content=false: planned a real disposable phase
  (a docs/how-to page for enabling the key itself); planner returned
  PLANNING COMPLETE, checker returned VERIFICATION PASSED, all fact-checks
  against real repo state confirmed.
- workflow.compact_content=true (detail.md never read): planned a second
  real disposable phase; planner returned PLANNING COMPLETE with
  frontmatter.validate and verify.plan-structure both clean, again fully
  grounded against real repo state. The five condensed fallback sections
  were independently re-read spine-only and confirmed sufficient to act on
  correctly without detail.md's elaboration.

Also drafts gsd-core/references/compact-content-protected-content.md — the
protected-content category list and <!-- gsd:protected --> sentinel syntax
ADR-4139 Decision 5 calls for, written to move to Phase 3 (#4403) unchanged
once it lands there.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4402): move detail.md into the ADR-4139-mandated detail/ subdirectory

Two independent review sub-agents (Standards and Spec axes of /code-review)
caught the same structural defect: ADR-4139 Decision 6 mandates
gsd-core/workflows/<name>/detail/*.md ("one or more parts... individually
skippable"), and this PR had shipped a flat plan-phase/detail.md instead,
copying issue #4402's own (inconsistent) restatement rather than the
locked ADR text. Fixed by git-mv to plan-phase/detail/elaboration.md and
updating every cross-reference (the spine's step 0.5 gate pointer, the
shared compact-content-gate.md's own resolution-rule wording, and the
completeness test's path constants).

Also, from the same review pass:
- docs/CONFIGURATION.md and gsd-core/references/planning-config.md's
  workflow.compact_content rows said "nothing branches on it yet" — no
  longer true now that plan-phase.md's spine does. Updated both to name
  plan-phase as the pilot and note the rest of the corpus is still pending.
- Regenerated all 19 tests/fixtures/install-tree/*.json golden fixtures
  (npm run gen:install-tree) — the three new shipped files were missing
  from the installer emitted-tree goldens.
- Found via a cache-busted `eslint . --max-warnings 0` (this repo's
  eslint --cache has produced false-greens before): the split test's
  `git show` call had a bare `timeout: 10000` literal, tripping
  local/no-adhoc-timeout-literal. Extracted to the existing GIT_TIMEOUT_MS
  constant from tests/helpers/timeouts.cjs instead of a second guessed
  copy of the same class of timeout.

Verified NOT needed, by tracing the actual mechanism rather than asserting
(tests/helpers/emitted-provenance.cjs's gsd-core-verbatim rule attributes
every gsd-core/{workflows,references}/** path to itself as an identity
source): an Emitted-Drift-Ack-Hash/-Growth trailer. Every changed/added
path in this diff is hand-authored and present in the diff itself, so
diffEmitted's attribution loop resolves `via` to the path's own source
before ever reaching the ack-lookup branch — there is no unattributed
delta to acknowledge. The spine also shrank (98,290 to 93,206 bytes), so
the growth ratchet has nothing to ack either.

Re-verified after these changes: the completeness/disjointness self-check
(0 missing, 0 duplicated) still holds against the relocated detail file,
and a full `npm run lint:ci` passes clean with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4402): restore literal content the pre-existing drift guards pin on

The first gsd-test run against this split (19 failures) surfaced real
regressions: several pre-existing structural guards pin the EXACT text
of the sections this split condensed, and paraphrasing broke them.

- tests/plan-phase-drift-guard.test.cjs expects the literal
  `DISK_PLANS=$(gsd_run query find-phase ...)` bash assignment inside
  plan-phase.md itself, not a prose description of the same check.
  Restored the exact line into both §9a and §11a's spine summaries.
- tests/thinking-partner.test.cjs expects plan-phase.md to literally
  offer "No, I'll decide" as the skip option. Restored that exact
  phrase into the condensed thinking-partner paragraph.
- Both restores would have duplicated the same text into
  plan-phase/detail/elaboration.md (which still carries the full
  elaboration). Removed the now-redundant restatements from the
  detail file instead of leaving them duplicated — the spine already
  computes DISK_PLANS before the detail elaboration is ever read, so
  the detail file references it rather than recomputing it.
- Re-running scripts/sync-runtime-launcher.cjs after that edit found
  the canonical gsd_run preamble had also become an unintentional
  spine/detail duplicate (both files call gsd_run and each is
  required, by runtime-launcher-parity's own contract, to carry its
  own copy). That's sanctioned duplication under a DIFFERENT
  contract, not lost/copy-pasted content, so
  tests/plan-phase-compact-split.test.cjs now excludes it from the
  disjointness check the same way it already excludes trivial
  fences/headings.
- Applied the adversarial-review finding on tests/plan-phase-compact-split.test.cjs's
  own isTrivial(): a blanket `line.length <= 15` cutoff silently
  swallowed real content (e.g. the 14-char `<quality_gate>`
  sentinel). Replaced it with a specific bare-label-line pattern
  (`Options:`, `Display banner:` etc.) — verified 0 missing / 0
  duplicated against the actual split, an improvement over both the
  original cutoff and a naive full removal (which produces
  false-positive "duplicates" on generic recurring labels).
- gsd-core/references/planning-config.md's own workflow.compact_content
  row used `/gsd-plan-phase` (hyphen). That file is Claude-facing
  source text (gsd-core/references/), which tests/slash-command-namespace.test.cjs
  requires in colon form; docs/CONFIGURATION.md's use of the hyphen
  form is correct as-is since docs/ is human-facing and outside that
  test's scanned directories. Fixed to `/gsd:plan-phase`.
- tests/plan-phase-compact-split.test.cjs's own `git show` of the
  parent commit failed inside the gsd-test sandbox ("detected dubious
  ownership") because the checkout is mounted under a UID the
  invoking user doesn't own. Scoped `-c safe.directory=<repo-root>`
  to that one git invocation rather than touching global git config.
- docs/INVENTORY.md still had one outstanding "detail.md part" wording
  fix from the earlier adversarial-review pass, staged now.

Re-verified locally against the exact assertions in all four affected
test files (all pass) before dispatching a fresh gsd-test run — no
change here should have broken any of the other 18 gates; `npm run
lint` is clean with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4402): restore the full marker enumeration to §9a's spine trigger line

The isolated Spec-axis review flagged that §9a's "Triggered when" line was
condensed to "Agent() returns but the return contains no recognized
marker" — dropping the literal `## PLANNING COMPLETE` / `## PHASE SPLIT
RECOMMENDED` / `## ⚠ Source Audit` / `## CHECKPOINT REACHED` /
`## PLANNING INCONCLUSIVE` enumeration, which is exactly the "machine-
parsed structural headings" category compact-content-protected-content.md
lists as protected. The load-bearing use of that same list (the
gsd_stall_watch call and the Handle Planner Return bullets a few lines
above) was never touched — only this one descriptive restatement was
genericized — but leaving any instance of a protected category
unsentineled is the silent erosion ADR-4139 Decision 4(c) warns
sufficiency isn't machine-checkable enough to catch on its own. Restored
the full enumeration into the spine.

That reintroduced an exact duplicate into plan-phase/detail/elaboration.md,
which still stated the same trigger sentence verbatim. Reworded the
detail file's version to reference the spine's trigger condition instead
of restating it, since the spine is now the single place that sentence
lives in full — mirroring the DISK_PLANS/"already computed above" pattern
from the previous commit.

Re-verified locally: completeness/disjointness (0 missing, 0 duplicated)
and all previously-fixed literal-content assertions still hold.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4402): backfill changeset pr number to 4471

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 12:18:38 -04:00
Tom Boucher
476394689a fix(#4254): pin sequential executor to the orchestrator's validated root (#4476)
* test(#4254): sequential executor root pin — failing-first regression + matrix

The new suite executes the shipped supplied-root-pin guard against real git
fixtures (drifted primary-checkout cwd halts before the write and the FATAL
names both roots; matching cwd permits it; unexpanded/empty pins halt;
normalization forms; submodule and sibling boundaries; metacharacter quoting;
drive-letter form gate) and locks the dispatch contract across execute-phase.md,
its sequential-root-pin step fragment, and worktree-path-safety.md. The #2772
per-plan serialization assertion retargets to the fragment that now carries
those rules (ADR-857 Phase 6 ceiling), plus the host-step wiring.

* fix(#4254): pin sequential executor to the orchestrator's validated root

Sequential-mode dispatch told the executor to self-derive PROJECT_ROOT from its
own cwd; every existing guard is worktree-mode-only or self-referential, so an
executor spawned with a drifted cwd committed onto the wrong checkout silently.

- worktree-path-safety.md step 0p: mode-agnostic supplied-root pin guard,
  composed by the orchestrator at build time with the literal $ORCHESTRATOR_WT
  (git-vs-git comparison on both sides — representation-safe on Windows, the
  #4296 lesson), fail-closed on empty/unexpanded pins, registered-submodule
  allowance, warn-and-proceed only when the dispatch carries no pin block.
- execute-phase.md sequential branch: build-time embed of the bound
  <project_root_pin> via the new execute-phase/steps/sequential-root-pin.md
  fragment (ADR-857 Phase 6 frozen ceiling — the host step cannot grow; the
  wave serialization rules move with the fragment, verbatim in substance) plus
  the per-write/commit pin instruction in <sequential_execution>. Worktree-mode
  dispatch untouched (its self-derived toplevel IS correct there).
- INVENTORY rows (5 locales) + INVENTORY-MANIFEST + install-tree goldens
  regenerated for the new fragment; changeset added.

* chore(#4254): backfill changeset PR number

* fix(#4254): accept backslash-separated Windows drive pins

CI on windows-latest showed every permit-path test failing with
"Actual root: <none>": pins composed from Node's path.join arrive in the
backslash drive form (C:\Users\RUNNER~1\...), which the guard's absolute-form
gate rejected before the cwd-side root was ever computed — a legitimate
matching pin could never pass. The gate now accepts either separator
([A-Za-z]:[\\/]); git -C resolves both forms (and 8.3 short names) to the
same canonical toplevel, so the git-vs-git comparison is unaffected. Form-gate
tests cover the emitted (C:/…) and produced (C:\…) spellings plus short names.

* fix(#4254): portable drive-form gate for MSYS bash

The bracket class [\\/] that accepted backslash drive pins parses
inconsistently on MSYS bash (the Windows CI leg still rejected C:\ pins —
every permit-path test red with "Actual root: <none>"). Replace it with
standard pattern escaping outside brackets: [A-Za-z]:/*|[A-Za-z]:\\* —
the escape form is version- and build-portable. Verified across all forms:
both drive spellings accepted; bare "C:", relative, empty, and unexpanded
rejected.

* fix(#4254): runtime-generated backslash comparator + self-describing FATAL

The Windows CI legs failed every #4254 permit-path row with
'Actual root: <none>' across two prior pattern spellings ([\\/] and \\*).
Stage misattribution: <none> appears whenever the FATAL fires BEFORE the
cwd-side capture assigns ACTUAL_ROOT — the absolute-form gate was what fired.

Mechanism: the test harness spawns bash -c <script> through the Windows
command-line boundary; that round-trip applies one extra shell-quoting pass
with double-quote semantics — a backslash written twice in the script text
arrives halved, while a lone backslash survives (the pin displays intact;
row 9's pure-bash gate independently showed the halved pattern rejecting
C:\ while C:/ still passed its surviving arm). On windows-latest every pin
carries backslashes (os.tmpdir() is the 8.3 short form C:\Users\RUNNER~1\...),
so the gate ate every pin before the actual root was ever computed.

Fix, robust by construction:
- the drive-form gate generates its backslash comparator at RUNTIME
  (BS=$(printf '\134'); match [A-Za-z]:"$BS"*) — the shipped guard now
  contains no doubled backslash anywhere, enforced by a regression
  assertion on the extracted guard text;
- the FATAL self-describes: Guard stage (pin-unbound / form-gate /
  actual-capture / pinned-capture / root-mismatch) plus a Diagnostic line
  carrying git's own stderr for capture failures and both compared values
  for mismatches — future platform failures name their stage in the log;
- row 9's hand-rolled duplicate case gate (transit-fragile copy, #4296
  Minor 1 duplication smell) is replaced by driving the SHIPPED guard and
  asserting the stage; rows 2/4 pin the new stage machinery.

Validated on darwin across drift/match/relative/unbound/empty/bare-drive/
forward-and-backslash drive forms, each also re-run under a simulated
Windows transit (every doubled backslash halved) with identical outcomes.

* fix(#4254): close the empty-comparator fail-open seam in the drive-form gate

Self-review of the runtime-generated backslash comparator: if printf's
octal escape ever returned empty, the drive arm [A-Za-z]:"$BS"* would
widen to drive-RELATIVE pins (C:foo) — the construction's one theoretical
fail-open path. Fail closed with a self-describing diagnostic instead of
trusting the shell's printf.

---------

Co-authored-by: sim <sim@local>
2026-09-07 10:54:30 -04:00
Brenden Smerbeck
e54d3aa159 enhance(#4401): register workflow.compact_content as a validated config key (#4441)
* feat(#4401): register workflow.compact_content as a validated config key

- Add compact_content: false to the nested workflow object in
  gsd-core/bin/shared/config-defaults.manifest.json
- Add 'workflow.compact_content': false to SCHEMA_DEFAULTS in src/config.cts
  so an absent key resolves to false via config-get --raw
- validKeys entry in config-schema.manifest.json already present

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(#4401): behavioral and boundary tests for workflow.compact_content

- 19 behavioral tests covering config-set/config-get round trip, invalid-shape
  rejection (banana, 42, empty string), the corrected null-unset semantics
  (#2046), absent-key resolution against config-defaults.manifest.json,
  config-new-project wiring, and doc-row shape assertions
- Drops the install-tree fixture-parity block (and its docstring item) that
  asserted gsd-core/references/compact-content-gate.md and
  gsd-core/workflows/compact/map-codebase.md fixture entries — those paths
  belong to #4402 and do not exist on this filtered branch

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(#4401): document workflow.compact_content in both config references

- One 4-cell row in docs/CONFIGURATION.md (workflow.* run)
- One 5-cell row under Workflow Fields in gsd-core/references/planning-config.md
- Both cross-reference ADR-4139

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(#4401): add changeset

- Added-type fragment, pr: 4401 (issue number; backfill to the real PR number
  is a required follow-up once the PR is opened, per D-08 and CHANGESET-PR-
  FIELD-DRIFT)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(#4401): backfill changeset pr field to #4441

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(#4401): derive workflow.compact_content default from CONFIG_DEFAULTS

SCHEMA_DEFAULTS['workflow.compact_content'] hardcoded the literal false
instead of deriving it from CONFIG_DEFAULTS the way 3 of its 8 sibling
entries do (smart_zone_tokens, pr_strict, inline_plan_threshold), leaving
a single-source-of-truth drift risk: a future manifest-only edit to the
default could silently diverge from this literal, only caught later by
the D-03 test if it ever happened to manifest.

Adds compact_content to CONFIG_DEFAULTS in src/config-loader.cts and
derives SCHEMA_DEFAULTS from it in src/config.cts, matching the majority
sibling pattern. Found during maintainer review (review-open-prs) of
this PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4401): map compact_content in config-field-docs NAMESPACE_MAP

The previous commit added compact_content to CONFIG_DEFAULTS in
src/config-loader.cts but missed the matching entry in
tests/config-field-docs.test.cjs's NAMESPACE_MAP, which maps flat
CONFIG_DEFAULTS keys to their namespaced doc form before checking
gsd-core/references/planning-config.md for a match. Without it, the
test looked for a bare `compact_content` doc reference instead of the
actual `workflow.compact_content` row, and failed:
"CONFIG_DEFAULTS keys missing from planning-config.md: compact_content".

Found by actually running gsd-test against the branch rather than
trusting the plausible-looking fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4401): register compact-content-4139 test in the docs-guard lane

tests/compact-content-4139.test.cjs's D-06 tests read docs/CONFIGURATION.md
directly (fs.readFileSync) to assert the workflow.compact_content doc row's
shape, which makes it a doc-reading test file under the #3753 docs-guard
lane. It was never added to scripts/docs-guard-registry.cjs's
DOCS_GUARD_TESTS map and carries no docs-guard-exempt marker, so
tests/ci-docs-guard-registry.test.cjs's registration lint correctly failed:
"compact-content-4139.test.cjs reads a docs/ path but is not registered in
the docs-guard lane and carries no docs-guard-exempt marker".

Registers it with ['docs/CONFIGURATION.md'] (the only real docs/-prefixed
path it reads; gsd-core/references/planning-config.md is outside this
registry's docs/ scope, matching the sibling config-field-docs.test.cjs
entry's existing convention).

Found by actually running gsd-test against the branch — this gap predates
the maintainer's config-loader.cts fix and was already present in the
original PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: sim <sim@local>
2026-09-06 19:52:59 -04:00
Tom Boucher
2cf119f57e fix(#4217): reconcile artifacts before classifying an abnormally-ended executor (#4442)
* fix(#4217): reconcile artifacts before classifying abnormal ends

* test(#4217): pin the completion-reconciliation contract

* chore(#4217): regen derived inventory and install-tree fixtures

* test(#4217): follow the #4003 anchoring pins into the reconciliation fragment

Emitted-Drift-Ack-Growth: execute-phase.md — the runtime-neutral completion-reconciliation pointer, the two Codex wait-rule bindings, and the step-7 reconcile-first gate net +33 bytes over the extracted fallback block (#4217)

* chore(#4217): add changeset fragment

* chore(#4217): backfill PR number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-09-06 18:57:38 -04:00
Michel Moreira
19b66c3ec8 fix(#4218): stop the orchestrator steering an executor that is still working (#4391)
* fix(#4218): stop the orchestrator steering an executor that is still working

An executor with recent RED/GREEN/REFACTOR commits and passing verification had
not yet written its SUMMARY because it was finishing closeout. The parent saw no
local OS test/build process, inferred an "idle tail", and sent "Finalize
immediately" into a working child; in CLI runs the same inference interrupted an
executor before GREEN, leaving a RED commit and an uncommitted edit.

The stall block said only "if no completion signal, no SUMMARY.md, and no
expected-branch commits appear for N minutes" — it never said what to do when
commits DO exist and only the SUMMARY is outstanding, never defined the
threshold as a period without progress rather than a total runtime, and never
ruled out a process listing as an idleness signal. Four rules close that:

- the threshold measures a period WITHOUT MEANINGFUL PROGRESS, from the last
  sign of progress, not from dispatch — a long verification tail is not a stall;
- commits + missing SUMMARY + recent activity resolves to KEEP WAITING, with
  steering, interrupting and re-dispatching each named and forbidden;
- urgency/finalization messages ("Finalize immediately" and family) are
  forbidden outright — they arrive mid-verification and truncate a correct run.
  The existing user-facing pause is the only sanctioned stop, and `kill and
  retry` is a clean restart, not a nudge;
- the absence of a local OS test/build process is NOT idleness: a native
  subagent runs in the runtime's own session, and an executor between two tool
  calls shows no process at all. Progress is judged only by the signals this
  workflow names.

Five prose-contract assertions in tests/execute-phase-wave.test.cjs, all red on
next.

* fix(#4218): extract the progress policy to a step fragment

CI's #1168 gate caught it: execute-phase.md sits 77 bytes under a frozen 93600
ceiling and the four rules added ~2.3 KB. "Extract, not bump" is the repo's
stated remedy, and this workflow already carries policy detail that way.

execute-phase/steps/executor-progress-policy.md owns the policy. The
worktree-recovery arm moved with it — `kill and switch to inline execution`
qualifies the stop this policy governs, so it belongs beside the rule about when
stopping is sanctioned at all, not stranded in the host. The #3212 recovery
OPTIONS stay in the host, where tests/config.test.cjs pins them.

The host keeps what must be read before the orchestrator acts: the verdict, the
threshold definition, and a pointer that fires before any message is sent to the
child. execute-phase.md is now 93475 bytes — 48 SMALLER than next.

* chore: add changeset for #4218

* chore(#4218): regenerate the inventory manifest for the new step fragment

docs/INVENTORY-MANIFEST.json is the authoritative per-file list behind
INVENTORY.md's `<workflow>/steps/*.md` row, so a new fragment has to appear
there or gen-inventory-manifest --check reds the lint-tests lane.

* chore(#4218): restore the issue ref on the allow-test-rule marker

ADR-456 requires a #NNN on a new exemption; the block rewrite that moved the
policy into the fragment dropped it.

* chore(#4218): regenerate the install-tree fixtures for the new step fragment

The fragment ships with the workflow, so every runtime's golden install tree
gains one path — gen:install-tree is the generator that owns those fixtures.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-06 17:51:34 -04:00
Tom Boucher
1c0acb2359 feat(#4422): block merging into next/main while the base branch's Tests run is red (#4428)
* feat(#4422): block merging into next/main while the base branch's Tests run is red

Adds a next-health job to test.yml that checks the base branch's own last
push-triggered Tests run via the GitHub API and fails the existing "Required
tests" required check when it's red, with a maintainer-applied "fix-next"
label as the explicit escape hatch for the fix-forward PR itself. No
branch-protection config change needed — it rides the already-required
check. The job is deliberately not gated behind preflight, same reasoning
as the changes job: a compute-free API read has nothing to save by waiting.

Documents the fix-next label in CONTRIBUTING.md and adds a property test
locking the CLEAN/RED/INDETERMINATE classification's iff-relationship.

This closes the second half of the 2026-09-06 RCA: three unrelated PRs
merged on top of an already-broken next before anyone noticed it was red.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: close two zero-margin CI timing gaps found while verifying #4422

Discovered while watching this branch's own CI, root-caused via /diagnose
rather than dismissed as Windows flakiness:

1. tests/gsd-check-update-worker-atomic-cache.test.cjs's outer timeout
   (15000ms) exactly matched the inner npm-view timeout the worker wraps
   (NPM_VIEW_TIMEOUT_MS, gsd-core/bin/check-latest-version.cjs). A slow
   registry response raced two SIGKILLs at the same instant, killing the
   worker before it could catch its own timeout and degrade gracefully.
   Windows's shell-wrapped npm subprocess made the race lose more often
   there, but the zero margin was platform-agnostic. Fixed by giving the
   test real headroom (+10s) beyond the named constant it wraps, plus an
   invariant test so the two values can't silently collide again.

2. scripts/run-tests.cjs's per-chunk weight budget (MAX_FILES_PER_CHUNK)
   let a Windows full-matrix chunk that was well under budget by the
   Linux/macOS-calibrated weight table (~32/60 units) still exceed the
   600s wall-clock backstop — codex-config.test.cjs's genuinely-measured
   weight (17.87) doesn't transfer 1:1 to Windows's slower install/
   subprocess overhead. Windows now gets its own lower cap (40 vs 60).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 17:29:54 -04:00
Tom Boucher
b7917882bb fix(#4398): render the pending-todo bullet link repo-relative (#4416)
* test(#4384): failing-first regression rows for the macOS long-base todo-cap failure

The 240-char pending-todo bullet cap must be deterministic w.r.t. where the
repo is checked out. Deterministic long-base-path fixtures (a single 110-char
segment, no real macOS dependency) reproduce next's own macos shard 3/3
failure (run 34038716700) on every OS: with an absolute link the bullet
exceeds the cap and the documented needs-first truncation drops the
'Needs <solution>' clause. Rows cover the determinism property (byte-identical
bullets under short and long bases), the CLI surface, relative-path stability,
legacy no-projectRoot behavior, drop-order preservation, and adversarial
edges (outside-root, path===root, non-string path).

* fix(#4384): render the pending-todo bullet link repo-relative

renderPendingTodosMarkdown gains an optional projectRoot; when given and the
todo's path is absolute, the bullet's [todo file](…) target becomes
toPosixPath(path.relative(projectRoot, path)) — the idiom already used for
project_exists. cmdInitTodos passes cwd.

The JSON todos[].path field stays absolute (#2376). Only the rendered display
link changes: embedding the machine-variable absolute base let macOS's
/private/var/folders/… temp paths consume the 240-char budget and drop the
'Needs' clause on long-path machines only — next's own macos-latest shard 3/3
went red on exactly this (run 34038716700), Linux's short /tmp passed. The
240-char whole-bullet cap and the needs→title→area drop order are unchanged;
this matches PR #4384's own canonical example, docs, and unit tests, which all
show repo-relative links. Docs updated at all three surfaces that describe the
bullet (COMMANDS.md, templates/state.md, reference/state-md.md — the last was
still pre-#4384 'count and reference' prose).

Fixes the macOS regression introduced by #4384; next is red on its own CI.

* test(#4384): fix substring false positive in the outside-root regression row

The ../-form relative link legitimately contains the absolute path as a
substring, so !line.includes(absolutePath) fired on correct output (caught by
the first remote verify run, linux-node24 44018/44019). Assert the property
itself instead: extract the link target and require it to be non-absolute and
not equal to the absolute path.

* chore(#4398): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-06 13:23:26 -04:00
Tom Boucher
fd4aac5670 fix(#4192): honor explicit model pins on the claude runtime (#4396)
* fix(#4192): honor explicit model pins on the claude runtime

Two documented model-configuration contracts did not hold on the claude
runtime (confirmed-bug scope from the issue triage):

Finding 1 — model_profile_overrides.claude.<tier> was inert. Step 3 of
resolveModelInternal gated runtime-aware tier resolution on
configRuntime !== 'claude', so the key's only reader was never consulted,
while workflows/settings-advanced.md writes it for claude-runtime users.
A new step 4.5 resolves ONLY the user's override entry (never the builtin
claude tier map, so unpinned installs keep resolving aliases). An
override value that maps to a current tier alias collapses to that alias
(byte-equivalent, the #2041 protection); anything else — a pinned older
generation, a bare alias repoint, a non-Anthropic id — resolves verbatim.
It sits after the resolve_model_ids:'omit' gate so an explicit project
omit still wins (#2297) and before the alias return so
resolve_model_ids:true cannot re-materialize the pin to the latest id.

Finding 2 — fully-qualified claude-* ids in model_overrides were
warn-dropped to tier resolution (mapClaudeOverrideForRuntime unmappable
branch, #2041), while the docs promise any fully-qualified model id is
valid. The unmappable branch now passes the pin through verbatim with a
warn-once breadcrumb (text describes the pass-through). Dropping it
silently unpinned the operator's explicit choice — the exact 'profile
can misrepresent what actually runs' defect of #4192. Mappable ids and
non-claude values behave exactly as before; resolveModelForTier shares
the mapping; the tier honesty signal is unchanged (raw ids still report
'unknown'); the model_policy path is untouched.

Docs updated to the agreed contract (CONFIGURATION.md false 'Claude
example' corrected; how-to + shipped reference document the pin
semantics, the fable alias, and the tier-override composition).

* test(#4192): pin explicit model pin resolution on the claude runtime

28 failing-first rows across the resolver seam and the resolve-model CLI:
pinned-generation fidelity (tier override + per-agent verbatim pins,
object form, explicit runtime), unpinned controls byte-stable (no
override, other runtime/tier, inherit, project omit, precedence),
adversarial rows (prototype-chain keys, malformed values, warn-once
dedupe, 64-char stderr cap), and behavioral AC1/AC2 rows through
runGsdTools. The stale #2041 fall-through assertions now pin the
pass-through contract; mappable-id collapse assertions unchanged.

* chore(#4192): add changeset fragment

* chore(#4192): backfill PR number in changeset fragment

---------

Co-authored-by: ZCode <zcode@localhost>
2026-09-06 10:17:50 -04:00
Tom Boucher
b7406b293f enhance(#2618): render pending todos as one bounded bullet per todo (#4384) 2026-09-06 08:06:39 -04:00
Tom Boucher
03738824de enhance(#2586): stop installing Codex context-monitor hooks without metrics (#4367) 2026-09-06 05:46:49 -04:00
Tom Boucher
0aa4202f6a fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening (#4376)
* test(#4135): regression rows for pristine regen coverage collapse

RED skeleton: src/pristine-baseline.cts exports findPristineInGit as a
null-returning stub (wired into verifyFile after the #4145 orphan tier,
behavior-neutral) so the git-history rows fail behaviorally, not at require
time. Failing-first rows: baseline_covered aggregate on a 1-of-13
multi-version fixture, coverageHeadline typed renderer, the opt-in
--min-baseline-coverage gate (exit 3, >= threshold semantics, vacuous-pass
and malformed-value boundaries), git-history baseline recovery (dropped-line
catch + surviving-line verify + older-commit hop), findPristineInGit unit,
Step 5a workflow headline contract, and the installer-side
describeBaselineCoverage honest N-of-M summary with the collapse disk-state
pinned. Negative-space rows pin today: non-git ok_no_baseline posture,
no-match-no-adoption, #3657 drift never rescued, canonical precedence, and
no git tier without --pristine-dir.

* fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening

The #3407 promotion rule regenerates gsd-pristine/ baselines from the
INCOMING release source and keeps only candidates byte-identical with the
OUTGOING recorded hash — correct in isolation, but on a multi-version jump
the surviving set is precisely the files upstream did NOT change. The
verifier then reports ok_no_baseline (advisory, exit 0) for everything
else, and no surface distinguishes a 12-of-13-unverified green run from a
fully-verified one: the human summary printed Checked/Failures only, the
JSON had no coverage aggregate, and the installer's update output gave
per-bucket counts without N-of-M framing.

All three issue directions, none exclusive:

- Report coverage prominently: --json gains an additive baseline_covered
  aggregate; the human summary leads with 'Baseline coverage: N of M
  file(s)...' on every run plus an advisory section naming each skipped
  file and reason; the installer prints an honest covered-of-modified line
  via the exported describeBaselineCoverage helper (typed return, exact
  contract); workflow Step 5a computes and prints the headline before any
  pass/fail framing.
- Fail louder on low coverage: opt-in --min-baseline-coverage <0..1>
  exits with new documented code 3 when coverage falls below the
  threshold (>= semantics; empty run vacuously passes; content failure
  exit 1 outranks it; malformed values are usage errors, exit 2).
  Default posture unchanged — no_baseline stays advisory per #934.
- Widen the promotion rule (its only trustworthy form): when no baseline
  resolves under gsd-pristine/ and a hash is recorded, the verifier now
  recovers the baseline from the config dir's own git history — the
  workflow's documented Option A — anchored by the same authority every
  tier trusts, exact pristine_hashes sha-256 equality. Read-only
  (git log/git show, windowsHide per #685), bounded (100 commits/file,
  10s/subprocess), null-on-any-failure so ok_no_baseline remains the
  universal fallback. Tier order: canonical join -> #4145 orphan scan ->
  git history -> OK_NO_BASELINE; #3657 drift and canonical precedence
  untouched.

Hash validation in saveLocalPatches is NOT relaxed — the collapse is
legitimate conservatism; hiding it was the bug. Measured on the issue's
shape (13 files, 12 changed upstream, 1.10->1.12): non-git installs report
baseline_covered 1/13 with the headline and can gate at exit 3; a
git-managed config dir with the outgoing bytes in history verifies 13/13.

Review fixes folded in: workflow headline derives the unverified count
from checked - baseline_covered (not the drift+no_baseline sum), and the
new site-scoped allow-test-rule annotation carries its ADR-456 see-ref on
the marker line.

Emitted-Drift-Ack-Growth: reapply-patches.md — #4135 — +20 lines / ~1.5 KB, prose and bash only: two additive parse lines (BASELINE_COVERED, CHECKED_COUNT), a Step 5a coverage-headline block printed BEFORE any pass/fail statement (documents the opt-in --min-baseline-coverage exit-3 gate), and one Option B sentence noting the verifier's read-only git-history fallback. No step ordering, gate, tool-invocation, or dispatch shape changed; 5a's fail/drift/advisory handling is unchanged, the headline only precedes it.

* chore(#4135): backfill PR number into changeset fragment

---------

Co-authored-by: agent-4135 <agent-4135@gsd.local>
2026-09-06 05:26:05 -04:00
Tom Boucher
c95b734145 fix(#4136): compute the Incorporated status; stop re-grafting superseded patches (#4373)
* test(#4136): failing-first rows for the unreachable incorporated status

Folded block bug-4136-reapply-incorporated-status locks the --classify
contract (incorporated / needs_merge / unknown, never-incorporated
guards, cycle end-to-end) plus the workflow-contract rows; REASON gains
OK_UNVALIDATED_BASELINE in both shape-locks; the #2994 invocation count
moves 1 -> 2 (classify + gate). All red until the verifier grows
--classify and the workflow consumes it.

* fix(#4136): compute the incorporated status; stop re-grafting superseded patches

Add --classify pre-merge mode to the deterministic verifier: with a
hash-validated pristine baseline, a file whose every significant
user-added line is already present verbatim in the freshly installed
version is classified incorporated (new frozen CLASSIFICATION enum,
structured --json report, always exit 0 — the binding gate stays the
post-merge run). Drifted (#3657), absent (#934), unvalidated (new
OK_UNVALIDATED_BASELINE), and no-baseline runs classify unknown — a
false incorporated silently retires a live customization, so only a
confirmed baseline may ever confirm adoption.

The baseline-resolution block moves out of verifyFile into a shared
resolvePristineBaseline so the gate and the classifier cannot drift on
what counts as a usable baseline; gate behavior is byte-identical.

reapply-patches.md step 4 gains the pre-flight classifier invocation
and the not-re-grafted contract: incorporated files are left exactly as
shipped (their hash then re-converges with the manifest, ending the
backup cycle), statuses feed steps 3/7, and the merge rules gain the
already-present-verbatim arm.

Emitted-Drift-Ack-Growth: reapply-patches.md — step 4 pre-flight classifier section, the documented Incorporated status needed its deterministic contract

* fix(#4136): address review findings on the workflow contract

Drop the unused INCORPORATED_COUNT shell variable (standards pass) and
close the all-files-incorporated gap in the hunk-table guidance: emit a
header row plus a note line so the step 5b absent-table halt is not
tripped when nothing was merged (spec pass).

Emitted-Drift-Ack-Growth: reapply-patches.md — step 4 pre-flight classifier section, the documented Incorporated status needed its deterministic contract

* chore(#4136): backfill changeset fragment with PR 4373

* test(#4136): lock that a 4145-recovered baseline can confirm incorporation

The hash-first orphan recovery merged with next (PR #4364) lands in the
shared resolvePristineBaseline as a validated resolution; this row pins
the composition so a future change cannot quietly downgrade recovered
baselines to unknown and silently disable incorporated detection for
prefix-less installs.

---------

Co-authored-by: sim <sim@local>
2026-09-06 02:55:41 -04:00
Tom Boucher
6adf3098ac fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans (#4364)
* test(#4145): regression rows for hash-matching prefix-less pristine baselines

RED skeleton: src/pristine-baseline.cts exports findPristineByHash as a
null-returning stub so the new rows fail behaviorally, not at require time.
Failing-first rows: verifier resolution (no_baseline must drop to 0 when an
exact-hash orphan exists), findPristineByHash unit row, and the two
saveLocalPatches relocation rows. Negative-space rows pin today's behavior:
missing baselines still report ok_no_baseline, mismatching orphans are never
adopted or deleted, canonical precedence and the #3657 drift posture are
untouched.

* fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans

Both pristine readers joined the manifest-keyed path strictly, so a snapshot
stored without the gsd-core/ prefix (an earlier release's writer) was reported
as ok_no_baseline by the verifier and pushed into regeneration by
saveLocalPatches — where incoming-release candidates can never satisfy the
recorded outgoing hash, leaving the correct baseline permanently unconsumed.

- src/pristine-baseline.cts (new, ADR-457): shared findPristineByHash —
  deterministic sorted scan of gsd-pristine/, exact sha-256 equality with the
  recorded pristine_hashes entry (the same authority the #3657 drift guard
  trusts), symlink-skipping, canonical path excluded via skipRel.
- verify-reapply-patches.cjs verifyFile(): on canonical miss with a recorded
  hash, adopt byte-identical content found anywhere under gsd-pristine/ before
  reporting OK_NO_BASELINE. Drift posture (#3657), canonical precedence, and
  the frozen REASON/report shapes are untouched; the verifier stays read-only.
- install.js saveLocalPatches(): preserve-check rescue — relocate a
  hash-matching orphan to the canonical path (copy, hash-verify, then remove
  the orphan) so the state self-heals on the next update instead of repeating
  forever. Honest accounting: new non-overlapping rescued counter.
- Workflow doc: one-sentence note on hash-based snapshot resolution.
- Derived ripples: INVENTORY-MANIFEST.json regen, eslint ignore + .gitignore
  entries for the compiled artifact, seedFixture mkdir fix in the new rows.

Emitted-Drift-Ack-Growth: reapply-patches.md — one-sentence note on hash-based pristine snapshot resolution (#4145)

* fix(#4145): review follow-up — orphan scan never consumes a canonical path

Adversarial review finding: with two modified files sharing byte-identical
outgoing content, recoverOrphanedPristine could adopt the OTHER file's
canonical pristine as its rescue source — relocating it (copy + delete at
its home path) and ping-ponging the single baseline between the two files
across updates. findPristineByHash's skip parameter now accepts a Set, and
saveLocalPatches passes the normalized manifest keys so every canonical
path is excluded; only genuine non-canonical orphans are eligible for
removal (no strict-join reader ever consults those). Adds the
canonical-theft regression row, a Set-skip unit assertion, and tightens the
workflow doc sentence the same pass flagged as overstated.

* fix(#4145): INVENTORY roster row + symlink-fixture correction

Two leftovers from the ab17b7a1e5 bench run, both root-caused:
- docs/INVENTORY.md roster row for cli_modules/pristine-baseline.cjs
  (#3762 gate: every manifest entry carries a row).
- The findPristineByHash symlink unit fixture placed its symlink target
  INSIDE the scanned root, so the walk legitimately matched the real target
  file. The implementation skips the symlink itself; the fixture now keeps
  the target outside the scanned tree so the assertion tests what it claims.

* changeset(#4145): fixed fragment for pristine baseline hash resolution

---------

Co-authored-by: gsd-agent <agent@gsd.local>
2026-09-06 02:04:18 -04:00
github-actions[bot]
b3906c66f6 chore: sync next package version to 1.13.0 2026-09-06 02:10:29 +00:00
Tom Boucher
f9f72cb54c enhance(#3777): opt-in concurrent per-plan planners in chunked mode (#4346)
* test(#3777): add failing-first coverage for concurrent per-plan planner dispatch

Extracts and executes the real bash blocks this PR is about to add to
plan-phase.md and chunked-planning-mode.md (CHUNKED_PARALLEL resolution and
the BATCH_PLAN_IDS dedup guard), plus config-set/config-get coverage for the
new planning.chunked_parallel key. Expected RED against the current shipped
workflow text — the extraction anchors do not exist yet.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(#3777): dispatch chunked mode's per-plan planners concurrently within a Wave

Adds opt-in planning.chunked_parallel (default false, byte-identical to the
existing serial loop). When true and the runtime's negotiated dispatch
capacity (dispatch-capacity, #3673) is greater than 1, chunked planning's
per-plan Tasks that share one outline Wave are issued together instead of
one at a time; a later Wave still waits for the current one to be verified
on disk and committed. A host with no declared maxConcurrency (most
non-Claude runtimes today) stays serial regardless of the setting.

Resolution and the Plan-ID dedup guard live in chunked-planning-mode.md
itself (gated on the section's own CHUNKED_MODE skip-check) rather than in
plan-phase.md, so a non-chunked run pays no extra gsd_run calls.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#3777): repoint extraction at chunked-planning-mode.md after the move

CHUNKED_PARALLEL resolution moved out of plan-phase.md into
chunked-planning-mode.md itself (see the preceding commit); update the
test's extraction path and header comment to match.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3777): relocate the canonical runtime-launcher preamble before its first use

The CHUNKED_PARALLEL resolution block's two gsd_run calls landed earlier in
the file than the sole existing preamble (in the commit step), which
tests/runtime-launcher-parity.test.cjs's (B) check requires to precede every
gsd_run call in the file. Move the preamble (not duplicate it) to the top of
the resolution block; the commit step's fenced block now just calls
gsd_run directly.

Caught by the GREEN checkpoint gsd-test run before push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3777): strip the canonical preamble from the extracted resolution block

The CHUNKED_PARALLEL resolution fence now carries the relocated
runtime-launcher preamble as its first line (previous commit). Extracting
the whole fence and running it after the test's own gsd_run stub let the
embedded preamble's own resolver logic `unset -f gsd_run` and exit 1 before
reaching the resolution logic, since no real gsd-tools.cjs exists in the
temp script dir — every test calling runChunkedParallelResolution() failed.

Strip the preamble (sourced from gsd-core/workflows/_runtime-launcher.snippet.sh,
the same file scripts/sync-runtime-launcher.cjs treats as canonical) before
splicing in the stub, so this suite tests only the resolution logic it is
actually about.

Caught by the post-rebase gsd-test run before push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3777): add the How-To page the phase gate requires

Enablement is 2 commands (config-set, then --chunked), which this repo's
own doc-quadrant gate flags as how-to-owed: a reference table cannot carry
a sequence. Covers enablement, the dispatch-capacity gate's honest
"most runtimes today: no effect" case, and the two accepted trade-offs.

An earlier reasoning pass (recorded in .gsd/phase/.../70-docs.json before
this commit) had incorrectly claimed #3034 shipped with no equivalent
how-to page, as precedent for skipping one here. That claim was false —
docs/how-to/enable-parallel-reviewer-lanes.md exists and is indexed. The
phase gate caught the omission before merge; corrected here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3777): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:57:17 -04:00
Tom Boucher
1db726ebbf feat(#3806): canonize the Review Dispositions Ledger contract (#4345)
* test(#3806): add parity tests for the Review Dispositions Ledger contract

Failing-first: asserts references/planner-reviews.md, workflows/plan-phase.md,
and agents/gsd-plan-checker.md agree on a single canonical "Review Dispositions
Ledger" heading, its round-scoping, L##@{sha} anchor format, and append-only
supersession rule. These fail until the canon and its two references are added.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(#3806): canonize the Review Dispositions Ledger contract

Promote the existing planner-reviews.md Step 4 return-payload tables
(Review Feedback Addressed/Deferred) into a canonical `## Review
Dispositions Ledger` PLAN.md section, stated once in planner-reviews.md
and referenced (not restated) from plan-phase.md's
<review_incorporation_contract> and gsd-plan-checker.md's Review
Incorporation dimension. Adds round-scoping (`### Round {N} —
{REVIEWS_sha}`), a `L##@{sha}` line-anchor format so a REVIEWS.md
reference survives the file being rewritten each round, and an
append-only supersession rule. Scoped to part 1 only per the
maintainer's approved-feature verdict — the deterministic lint/check
verb (part 2) is explicitly deferred to a follow-up.

Also: ADR-3806 recording the decision, a docs/features/ fragment
(FEATURES.md is generated), and a changeset fragment.

Closes #3806

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): fenced-example count bug and lint findings from review

- tests/plan-review-convergence.test.cjs: the "heading exactly once"
  test counted the canonical heading text globally, so it also matched
  the illustrative fenced-code example in planner-reviews.md that shows
  the same heading as sample content, always failing 2 !== 1. Rewritten
  as a bounded line scanner that skips fenced blocks (found by an
  isolated adversarial review pass). Also bounded an unbounded regex
  quantifier over readFileSync content flagged by
  local/no-unbounded-quantifier.
- docs/features/review-dispositions-ledger.md: match house fragment
  style (bold-lead paragraphs, not #### headings) per the Standards-axis
  review; regenerated docs/FEATURES.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): fit reference-cite fix within size hard caps; ack growth

Trims the plan-phase.md / gsd-plan-checker.md reference-cite text to a
single short clause pointing at gsd-core/references/planner-reviews.md
(also fixes the bare `references/planner-reviews.md` cite the #3576
shipped-reference-cites gate rejects), bringing both files back under
their SIZE hard caps and the plan-phase.md phase6 shrink-only baseline.
Both files still grow slightly versus origin/next, acknowledged below
per ADR-2719's emitted-drift-ack contract.

Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): correct malformed Emitted-Drift-Ack-Growth trailer block

The previous commit's two Emitted-Drift-Ack-Growth trailers were
separated from the Co-Authored-By trailer by a blank line, so git's
own trailer parser (which tests/helpers/emitted-runtime.cjs reads via
`%(trailers:key=...)`) only recognized the last contiguous block
(Co-Authored-By) and treated the Ack-Growth lines as ordinary body
text — invisible to the emitted-attribution gate, not malformed data.
Restating them here as one contiguous trailer block, git log over the
PR range aggregates trailers from every commit, so this is additive.
Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): isolate the ack-trailer paragraph as its own trailer block

Git's trailer parser requires the trailer paragraph to be the message's
final paragraph, preceded by a blank line, and to contain nothing but
trailer-shaped lines. The prior commit's blank line before the trailer
lines was missing, which folded the leading Emitted-Drift-Ack-Growth
lines into an ordinary prose paragraph.

Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3806): backfill PR #4345 into changeset and ADR

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:50:27 -04:00
Tom Boucher
c20675cc4d fix(#3819): widen executor's pre-commit guard beyond worktree mode (#4343)
* fix(#3819): widen executor's pre-commit guard beyond worktree mode

The pre-commit protected-branch assertion in the executor agent (#2924)
only fired inside a Claude Code worktree and matched a hardcoded
five-name branch list. It never ran in an ordinary checkout and never
covered this repo's own default branch ("next"), so gsd-executor could
commit planning-repo documents directly onto a shared checkout's
default branch with no PR ever created.

Widen the guard to run in every isolation mode, and resolve the
protected branch via the repository's actual default branch (with the
existing five-name list retained as a fallback when the resolver
itself cannot be invoked) plus any configured git.protected_branches.
Add a git.allow_default_branch_commits escape hatch for projects that
intentionally execute on their default branch. Also point the
separate <final_commit> commit helper back at the same guard, so it
cannot be sidestepped by that path.

Emitted-Drift-Ack-Growth: gsd-executor.md — widened pre-commit protected-branch guard (#3819); tightened comments to stay under the size cap.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3819): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:35:42 -04:00
Tom Boucher
2f920c5af3 fix(#4097): evidence preservation no longer sweeps the run's own input copies into .review-diagnostics/ (#4329)
* test(#4097): failing-first regression — run input copies swept into .review-diagnostics

The present_results preserve+cleanup glob `_DIAG_MD=( "$RUN_DIR"/gsd-review-*.md )`
matches both lane outputs and the run's own assembled input copies (prompt,
instructions, roadmap, per-plan copies, project/context/research/requirements,
per-lane trimmed prompts) because both share the gsd-review- prefix.

Seed the full input-copy set into the existing runWriteReviewsFlow fixture and
assert (a) the diagnostics dir holds exactly the lane report + .err sidecar and
no input basenames, and (b) an inputs-only run creates no diagnostics dir at all
and still cleans up. Both RED against the shipped block.

* fix(#4097): preserve lane output only — exclude the run's input copies from evidence

The preserve+cleanup glob treated every gsd-review-*.md in RUN_DIR as lane
evidence, but the workflow itself writes the run's assembled INPUTS there under
the same prefix (prompt, instructions, roadmap, per-plan copies, project/
context/research/requirements, per-lane trimmed prompts). Filter _DIAG_MD by
basename against that closed input set instead.

Direct glob iteration + case filter — identical under bash and zsh (#4099/
#4109), nullglob-safe (#2962), and the exclusion list is closed and owned in
this step: a future input basename cannot silently rejoin the evidence set.
Lane reports, diagnostic stubs and non-empty .err sidecars are unchanged.

Emitted-Drift-Ack-Growth: review.md — #4097 narrows the present_results evidence-preservation glob: the closed input-basename exclusion list (case filter) plus its rationale note are deliberate additions so a future input basename cannot silently rejoin the evidence set.

* changeset(#4097): fixed — review diagnostics no longer sweep run input copies

* changeset(#4097): backfill PR number 4329

---------

Co-authored-by: sim <sim@local>
2026-09-05 16:45:43 -04:00
Dennis Alexis Valin Dittrich
1017898cb9 fix(#3771): make remediation examples non-binding and surface revision conflicts (#3916)
* fix(#3771): separate the binding property from the advisory remediation

Checker findings fused "what property failed" with "how to fix it" into a
single `fix_hint` and never marked which half binds. The checker rendered
every hint under a "must fix" heading, the orchestrators injected the issues
verbatim and ordered targeted updates, and the shared revision references
mapped each hint to a prescriptive strategy — so a contract-following planner
applied a hint literally even when a smaller mechanism satisfied the same
property, or when the hint contradicted a locked decision. There was no
channel to report that conflict, and every attempt burned a revision
iteration.

Checker side: every issue now carries a binding `required_property` (the
invariant that failed) plus its evidence and severity, and `fix_hint` is
labelled non-binding wherever it appears — including the human-facing blocker
rendering, so "must fix" unambiguously names the property and never the
example.

Planner side: revision re-checks locked decisions, capability guidance and
existing plan constraints before editing; satisfying a blocker through a
smaller valid alternative counts as addressing it; and a hint that conflicts
with any of those returns `REVISION_CONFLICT` carrying the conflict and the
alternatives considered. Orchestrators route that to user choice or the
configured plan-review convergence loop without consuming retry budget.

Also applied to the UI-spec revision loop and the gap-plan hint, and the
generic pattern's stray `suggested_fix` field name is reconciled to the
plan-checker's `fix_hint`.

Nothing legitimately binding is weakened: blockers still block, severity
still gates, iteration caps and stall escalation still fire, and required
task fields and decision coverage still hold.

Refs #3771

* test(#3771): pin the binding/advisory split across the revision chain

Locks the separation at every link that carries it: the checker's issue
schema and blocker rendering, the planner's constraint re-check and
REVISION_CONFLICT return, the generic pattern's reconciled field names, and
each orchestrator's conflict routing without retry-budget consumption. Also
pins what must not have been weakened — blockers, severity gating, iteration
caps and stall escalation.

Red against the pre-fix prose: 32 of 34 assertions fail (the 2 that pass are
the preservation checks, correctly).

Refs #3771

* chore(#3771): add changeset fragment for the remediation-binding fix

* chore(#3771): acknowledge the remediation-binding growth

Five runtime-loaded files grow: the two checkers carry the binding/advisory
split where the model reads it (a `required_property` on every dimension
example, since a schema the examples contradict teaches the examples), and
the three orchestrators carry the REVISION_CONFLICT route, which has to live
with the `iteration_count`/`revision_count` state it declines to spend.

Deletes tests/emitted-drift-acks/3172-stated-failing-direction.json: it is
fully spent on next and still owned plan-phase.md, so it walls off a key it
can no longer clear (#3078). Its removal is the documented remedy for the
duplicate-key collision, not drive-by cleanup.

* fix(#3771): close the review gaps in the conflict contract

Adversarial review (Codex, Antigravity) found four real defects in the first
pass, each confirmed against the source before acting:

- The UI checker's structured return still ordered `Fix: {exact fix required}`
  and "list each BLOCK dimension with exact fix required". The dimension
  examples had been marked non-binding but the rendering the researcher
  actually reads had not — the same omission this issue is about.
- `ui-phase` and the canonical `revision-loop` flow incremented their counter
  BEFORE dispatching the reviser, so "do NOT increment on REVISION_CONFLICT"
  was unreachable prose: the iteration was already spent. The increment now
  sits on the return path in both.
- The conflict gate offered "accept as-is", which is an early exit from a
  still-failing blocker — a weakening the brief explicitly forbids. The three
  options are now adopt an alternative / override the constraint / amend the
  constraint; every one resolves the conflict. Accepting an unaddressed blocker
  remains available only at the unchanged iteration-cap escalation.
- The convergence route was declarative: nothing in
  plan-review-convergence.md could receive a conflict. plan-phase now records
  it in REVIEWS.md — the channel that loop already consumes — convergence
  refuses to declare convergence over an open entry, and routing back into a
  run convergence itself started is explicitly excluded as a cycle. `quick` has
  no REVIEWS.md and no phase, so its convergence branch was dead prose and is
  deleted in favour of asking the user.

Also reconciles the last two drifted field names (`finding`, `affected_field`)
to the plan-checker schema, and repairs a silent no-op: the few-shot
`required_property` insertion never applied because those lines are
blockquoted, and the test's own block filter was anchored on indentation only,
so a vacuous loop passed over zero blocks. Both are fixed and the filter now
asserts it found blocks.

Refs #3771

* chore(#3771): extend the growth acknowledgment for the review round

plan-review-convergence.md joins the list: the conflict route needed a
receiving end, and it lands on the seam that loop already reads (REVIEWS.md)
rather than a new mechanism. The plan-phase, ui-phase and gsd-ui-checker
entries gain the second-pass reasoning — an executable convergence branch, the
increment moved onto the return path, and the structured return that still
ordered an exact fix.

* fix(#3771): make the conflict route bounded, ordered, and owned

Round-2 adversarial review found five more defects, each confirmed in the
source before acting:

- The convergence gate sat AFTER `gsd_run state planned-phase` and the success
  banner, so a run could write and announce convergence over an unresolved
  conflict. OPEN_CONFLICTS is now read from REVIEWS.md and is part of the
  converged CONDITION, evaluated before any write.
- plan-phase's cycle-exclusion ("unless this run was invoked by convergence")
  was not a question the orchestrator can answer at runtime. plan-phase now
  never invokes convergence at all — it records the conflict when a phase
  REVIEWS.md exists and resolves it with the user in-place, which removes the
  cycle instead of describing it.
- Closure had no owner. plan-phase writes the row, so plan-phase strikes it
  resolved; convergence only reads. An open row is a live blocker, never a
  stale artifact.
- Declining to increment the counter removed the only bound on the conflict
  path: an agent returning the same conflict forever would loop unattended. A
  conflict naming the same `required_property` twice in a row is now a stall
  and escalates through the existing gate.
- verify-work's gap-plan revision loop hands `<revision_context>` to
  gsd-planner and so inherits the whole contract, but stated none of it and
  could not handle the conflict return. It is now covered like the others, and
  is in the test's orchestrator table.

Refs #3771

* chore(#3771): acknowledge the round-2 growth

verify-work.md joins the list — the flow the second review found missed — and
the plan-phase, ui-phase and plan-review-convergence entries gain the
round-2 reasoning: the gate moved ahead of the state write, the convergence
hand-off replaced with a runtime-checkable record-and-resolve, and the
recurrence bound that replaces the counter the conflict path stopped spending.

* fix(#3771): make the convergence gate countable and stop the conflict fall-through

Third adversarial round (Antigravity) found three defects:

- The OPEN_CONFLICTS pipeline had no `grep -v '~~'` despite its own comment
  claiming one, and `grep -c '^| '` also counts a markdown table's header and
  separator rows — every resolved conflict would have read as open and
  convergence would have deadlocked instead of converging. plan-phase now
  records each conflict as a `- [ ]` checklist line and flips it to `- [x]`, so
  the gate is an exact fixed-string match with no table parsing.
- "then continue below" fell through to the checker spawn, so a SECOND
  REVISION_CONFLICT would have been handed to the checker as though it were a
  revised plan. plan-phase, quick and verify-work now re-evaluate the return
  from the top of the conflict handler; ui-phase already looped back.
- revision-loop.md still described plan-phase routing a conflict to the
  convergence loop instead of asking — the behaviour round 2 removed. Recording
  is now stated as being in addition to asking, never instead of it.

Refs #3771

* chore(#3771): bring the changeset in line with what shipped

Two review rounds widened the change after the fragment was written:
verify-work's gap-plan revision and the convergence loop are covered, two
more drifted field names are reconciled, and the conflict path carries an
explicit recurrence bound.

* fix(#3771): declare and emit the REVISION_CONFLICT marker

check:contract-drift on CI caught what local lint never reached: four
workflows dispatch on `## REVISION_CONFLICT`, but no agent declared or emitted
it — an orphan consumer, matching a marker nothing produces. The shared
reference (planner-revision.md Step 7b) described the return; the agent
definitions did not carry it.

gsd-planner and gsd-ui-researcher now emit the marker in-fence alongside their
other return markers, and both registry rows in agent-contracts.md declare it.
gsd-planner's Consumed by gains the two workflows that dispatch on it and were
missing from the row.

The gate is right: a return contract belongs where the agent is defined, not
only in a reference the agent happens to load.

Refs #3771

* chore(#3771): acknowledge the return-marker growth

gsd-planner.md and gsd-ui-researcher.md each gain the REVISION_CONFLICT
marker that check:contract-drift requires them to emit.

* fix(#3771): hoist the shared conflict protocol out of the workflows

Two CI failures, both correct gates:

- tests/few-shot-calibration.test.cjs pins the plan-checker calibration file
  at exactly 4 examples (2 positive, 2 negative). The example added in the
  first pass broke that balance — and described PLANNER behaviour in the
  CHECKER's calibration set, which is the wrong surface for it. Removed; the
  smaller-alternative rule is already normative in gsd-plan-checker.md and
  planner-revision.md, and pinned by the regression suite.
- tests/phase6-capstone-conformance.test.cjs (ADR-857 phase 6, #1168) requires
  plan-phase.md to stay BELOW its pre-phase-6 baseline of 94519 bytes. The
  inline conflict block pushed it to 94988.

The fix for the second is the one that should have been made first: the
record/resolve/close protocol and the recurrence bound were identical in four
workflows, and revision-loop.md — which plan-phase already @-imports — is what
a shared contract is for. The protocol now lives there once; plan-phase states
only its bindings (which counter, which artifact, which next step) and points
at it. plan-phase.md: 94988 -> 92739, under the ratchet with headroom, and the
four-way duplication is gone.

quick, ui-phase and verify-work do not import the reference, so they keep their
inline statements. The suite asserts each rule against what the runtime
actually loads for that orchestrator, not against the file in isolation.

Refs #3771

* docs(#3771): state the shared-protocol relationship accurately

Three of the four revision-bearing workflows do not @-import revision-loop.md,
so 'follows it verbatim' overstated the coupling. Only plan-phase defers; the
others restate the rules inline and this section is the authority they must
agree with.

* refactor(#3771): name the authority instead of restating it four times

Self-review finding: the same ~700-byte reviser paragraph was inlined in four
prompts while gsd-planner already loads planner-revision.md whenever
<revision_context> is present (agents/gsd-planner.md:555, :588) — a fifth copy
of the same contract. Four places to edit in lockstep is precisely the drift
class this PR exists to fix; the generic pattern calling the field
suggested_fix while the checker emitted fix_hint is what that looks like after
a year.

Each prompt now carries only the load-bearing clauses and names the authority
it summarises. Deliberately NOT reduced to a bare pointer: these are LLM
prompts, and a contract stated only in a file the reader is supposed to fetch
is the failure mode of this very bug. The saving is modest (~50 bytes each) —
the point is the named source of truth, not the bytes.

Suite: dropped seven assertions that pinned heading text and bold-lead
phrasing a reword would break without changing what the runtime is told; their
neighbours already pin the same contract by content. 519 -> 499 lines, 61
tests. Red gate against origin/next: 56 of 61 fail.

Refs #3771

* fix(#3771): sanitize agent-authored conflict text and bound total conflicts

Cross-AI review (agy/Gemini) raised two security findings. I confirmed the
first by running the gate's own awk against a crafted file and reproduced it
exactly; the second is a real hole in my bound.

1. FAIL-OPEN in the convergence gate. The conflict text is agent-authored and
   is appended to REVIEWS.md, which the gate scans with an awk that stops at
   the next '## ' heading. One line of agent text beginning '## ' ends that
   scan early, so conflicts below it are never counted and convergence declares
   success over a live blocker. Measured: 3 open conflicts, awk returned 2.

   Fixed at the write boundary, which is the trust boundary: every field has
   newlines and tabs collapsed to spaces and a leading '#', '-', '|' or fence
   stripped, so one conflict is exactly one line. Both producing agents now
   declare their fields single-line plain text, and the reader states the
   invariant it depends on so a later edit cannot silently break it. Verified:
   3 open + 1 resolved now counts 3; missing file and absent section count 0.

2. The recurrence bound was 'same required_property twice in a row', which an
   agent alternating property names never trips, leaving the un-incremented
   conflict path unbounded. Now bounded twice: the repeat rule catches the
   common case, and the THIRD conflict return of a loop escalates whatever
   property it names. A conflict still never consumes a revision iteration;
   this cap is separate from and additional to the revision cap.

Rejected from the same review: deleting 'a planner that reaches
required_property by a smaller or different mechanism has addressed the issue
in full' from the CHECKER prompt as misplaced. It is load-bearing exactly
there. A checker that does not know a different mechanism counts will re-flag
the issue on re-check, which is the revision loop that never terminates. The
argument offered for deleting it, that the checker evaluates the new state
independently, describes the failure mode.

Refs #3771

* fix(#3771): fail closed on an unverifiable convergence gate

Second cross-AI pass (agy, this time with the full files rather than the diff)
found two more, both real:

1. The gate read REVIEWS_FILE with `2>/dev/null || echo 0`, so an unreadable or
   empty path counted as ZERO open conflicts and converged. That path is
   resolved a few lines earlier by a pre-existing unquoted
   `ls ${phase_dir}/${padded_phase}-REVIEWS.md` (line 346, not touched by this
   PR), which yields an empty string rather than an error when the path
   contains a space. Unverifiable is not clean: the gate now tests -z and -r
   first and BLOCKS. Verified both branches.

   The unquoted ls itself is left alone deliberately — it predates this change
   and belongs to the reviews lookup, not the conflict gate. Fixing it at my
   own boundary removes its effect on this gate without widening scope.

2. REVIEWS.md is writable by the review agent, which could flip a `- [ ]` to
   `- [x]` or delete the section and forge the state of a blocking gate. The
   section now declares a single writer: /gsd:plan-phase appends and closes,
   every other agent leaves it byte-for-byte alone, readers read.

Also trimmed a clause that explained the increment ordering by reference to
what the file said before this PR. Commit history is not instruction, and
these files are prompts.

Rejected: the claim that quick's conflict gate deadlocks autonomous pipelines
by asking the user. Its existing max-iteration escalation in the same file
already asks the user the same way; this adds no new interaction class.
Noted but out of scope: the per-dimension YAML example blocks and the shim
boilerplate duplicated across agent prompts both predate this change.

Refs #3771

* fix(#3771): count conflicts by line shape, not by section

CodeRabbit review on the rehearsal PR. Five findings, all valid, all applied.

The best one is a deletion. The convergence gate scanned between
'## Plan-Revision Conflicts' and the next '## ' heading, and that scan stops at
the FIRST heading it meets — so one stray '## ' line hid every conflict beneath
it and returned 0, converging over a live blocker. Reproduced: section-scan 0,
shape-scan 1. Sanitizing at the write boundary does not cover a hand-edited,
legacy, or foreign-written REVIEWS.md, so the reader needed its own guarantee.

It now matches the conflict line SHAPE anywhere in the file:

  grep -c '^- \[ \] .*required_property:'

No section bookkeeping, nothing a heading can truncate, and it composes with the
writer's sanitization (which strips a leading '-' from agent text, so agent prose
cannot forge the shape). Verified: injected heading -> 1, all resolved -> 0.

The other four:

- Both checkers told the author never to emit a contradictory fix_hint, then
  offered an escape hatch that put the forbidden route in the hint anyway. They
  now name NO route in that case and state only that the property conflicts with
  the constraint. A hint carrying a forbidden route is applied by anyone who
  trusts hints.
- The REVISION_CONFLICT marker description in gsd-planner.md was narrower than
  planner-revision.md: it covered a contradictory hint but not an unreachable
  required_property. A planner reading only the agent file would have burned
  retry budget on the case the reference routes to a conflict.
- The few-shot calibration examples used uppercase BLOCKER/INFO while the schema
  defines blocker/warning/info. Pre-existing, but it is the same schema-vs-example
  disagreement this PR exists to end, and the file was already being edited.
- verify-work's re-entry instruction existed but sat after the Bounded clause, so
  the paragraph read "re-spawn ... stop re-spawning ... after re-spawning". The
  re-entry now immediately follows the re-spawn, and states that only a
  non-conflict return may reach the checker or increment iteration_count.

Refs #3771

* fix(#3771): resolve the contradictory scope_sanity severity examples

Sixth CodeRabbit finding, posted outside the diff range and missed on my first
read — I had claimed all findings were addressed after reading only the five
inline comments. This one was in the review body.

agents/gsd-plan-checker.md carried TWO scope_sanity examples with identical
metrics (5 tasks, 12 files) and OPPOSITE severities: warning in Dimension 5,
blocker in <examples>. Line 872 states "2-3 tasks/plan good, 4 warning, 5+
blocker" and the severity table lists warning as "Scope 4 tasks (borderline)",
so the warning example contradicted both.

ADR-2629 Decision 5's "over budget is a WARNING, never a blocker" does not
excuse it: that rule governs the smart-zone TOKEN estimate (the estimate-check
verb, lines 299-306), which is a different axis from task count. Verified in
source before touching it.

The contradiction is pre-existing but this PR made it binding and visible:
severity is now declared part of the binding payload, and both examples were
given the same required_property, so they now disagree on the severity of an
identical finding about an identical property.

Deviating from the proposed correction, which was warning -> blocker: that
would duplicate the <examples> entry outright (same tasks, files, severity).
The Dimension 5 example is instead made a genuine 4-task borderline warning, so
the file keeps one worked example per severity and the thresholds, the severity
table and both examples finally agree.

Refs #3771

* fix(#3771): stop laundering a grep error into zero open conflicts

Seventh CodeRabbit finding — from a SECOND review round my own CR-4 push
triggered, which I had not looked for. This one is a regression I introduced
while fixing the previous fail-open.

CR-4 replaced the truncatable section scan with:

  OPEN_CONFLICTS=$(grep -c '^- \[ \] .*required_property:' "$REVIEWS_FILE" || true)

`|| true` masks every grep failure. grep exits 1 for "no matches" (a legitimate
zero) but 2 for a read error, and `|| true` turns both into an empty capture
that `${OPEN_CONFLICTS:-0}` renders as 0. If REVIEWS.md is removed or becomes
unreadable between the -r check and the scan, the gate reports no conflicts and
convergence proceeds. Proven: unreadable file -> captured empty -> 0.

The status is now inspected, and only exit 1 counts as zero; anything else
blocks.

My first attempt at this fix was itself wrong and my own harness caught it: I
wrote `if ! grep ...; then grep_status=$?`, but `!` inverts the status, so `$?`
in that branch is 0 and every failure reads as success — the clean-file case
printed "BLOCKED (grep exit 0)". The status must be read in the ELSE branch of a
non-negated `if`, which is what CodeRabbit proposed. Both traps are now pinned
by tests.

Verified end to end: all resolved -> 0, no conflicts at all -> 0, injected
heading -> 1, unreadable file -> BLOCKED with grep exit 2.

Refs #3771

* test(#3771): execute the conflict gate instead of reading it

CodeRabbit round three: 0 actionable, 1 nitpick — "these assertions inspect
Markdown source only; they do not prove that grep status 1 produces zero
conflicts or that a scan error exits before convergence." Rated Trivial. It is
the most valuable finding of the three rounds.

This gate has been wrong three times: a section scan a heading could truncate, a
`|| true` that laundered grep's error status into zero, and an `if !` whose `$?`
reported the negation rather than the command. Every one of those passed the
text assertions that existed at the time. I proved each fix by hand in a shell,
and none of that proof lived in the suite.

The gate is one self-contained fenced block, so the test now extracts it from
the workflow — located by content, not line number — writes it to a script and
RUNS it against fixtures: two open plus one resolved counts 2; no matches counts
0 and does not fail; a conflict below an injected `## ` heading still counts; an
unreadable path and an empty path both BLOCK with a non-zero status and no zero
count on stdout.

Non-vacuity proven by mutation rather than asserted. Reverting the gate to each
of its three historical broken forms reds the suite:

  section-scan awk  -> 7 failures (5 in the gate cases)
  || true           -> 4 failures (3 in the gate cases)
  if ! (negated $?) -> 4 failures (3 in the gate cases)
  restored          -> 69 pass, 0 fail

The prose assertions stay: they are the right instrument for a prompt. This
covers the one part of the change that is real shell an orchestrator executes.

Refs #3771

* test(#3771): route the gate harness through the shared test helpers

ESLint's project rules caught three violations in the new harness: an unbounded
execFileSync (DEFECT.UNBOUNDED-SUBPROCESS — an unbounded spawn is an indefinite
hang, and on macOS CI that is how a stuck run stops reporting instead of failing)
and two raw fs.rmSync calls, which skip the Windows-EBUSY retry budget that
helpers.cleanup carries.

Now uses createTempDir/cleanup from tests/helpers.cjs and passes an explicit
30s timeout. Suppressing the rules was available and would have been the wrong
call: both exist because of real CI failure modes on platforms I am not testing
on.

* chore(#3771): backfill the changeset PR number

The pr: field is drift-checked against the PR event payload, so it cannot be
written before the PR exists. Set to 3916.

* fix(#3771): close revision conflict persistence gaps

Use the authoritative review artifact, keep conflict and normal retry paths disjoint, and enforce one writer-reader grammar so malformed state fails closed.

Emitted-Drift-Ack-Growth: diagnose-issues.md — #3771 marks the gap-plan remediation hint non-binding while keeping root_cause authoritative
Emitted-Drift-Ack-Growth: gsd-plan-checker.md — #3771 separates binding required_property evidence from advisory fix_hint examples across the checker contract
Emitted-Drift-Ack-Growth: gsd-planner.md — #3771 declares the REVISION_CONFLICT return used when remediation contradicts governing constraints
Emitted-Drift-Ack-Growth: gsd-ui-checker.md — #3771 applies the same binding-property and advisory-hint split to UI review findings
Emitted-Drift-Ack-Growth: gsd-ui-researcher.md — #3771 defines the UI revision producer's structured REVISION_CONFLICT return
Emitted-Drift-Ack-Growth: plan-phase.md — #3771 routes and persists bounded revision conflicts before spending the normal retry budget
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #3771 adds the fail-closed owned-block parser and prevents convergence over open conflicts
Emitted-Drift-Ack-Growth: review.md — #3771 emits and preserves the canonical writer-owned conflict block across review regeneration
Emitted-Drift-Ack-Growth: ui-phase.md — #3771 routes UI revision conflicts to resolution before consuming revision_count
Emitted-Drift-Ack-Growth: verify-work.md — #3771 gives gap-plan revision the same bounded conflict route before iteration_count

* test(#3916): guard rebases against schema drift

Load the current-base progressive-disclosure examples so every integrated issue remains bound by required_property after branch reconciliation.

* test(#3916): skip the extracted-gate suite's bash spawns on win32

Third review round's sole survivor: runConflictGate()/withReviews() spawn
bash against a Node-native temp path built by createTempDir(), which is
backslash-separated on the Windows CI lane and not a path Git Bash is
guaranteed to accept (DEFECT.WINDOWS-TEST-PORTABILITY, matching the
observed CI failure at revision-remediation-binding.test.cjs:844,
ENOENT on a path Windows read as a directory separator). No eslint rule
catches it since the call has neither a chmod nor a `bash -c` form.

Guards the four call sites with the repo's existing skipOnWin32
convention (describe/test `{ skip: IS_WINDOWS }`) rather than
normalizing the harness path to forward slashes, which would defeat the
one test whose purpose is proving the production gate does NOT rewrite
a literal backslash in a POSIX filename.

* fix(#3916): backfill changeset pr field to the fork validation PR number

* fix(#3771): forbid silently accepting an open plan-revision conflict at max-cycles escalation

The max-cycles escalation prompt only surfaced HIGH_COUNT and ACTIONABLE_COUNT; an open
plan-revision conflict (OPEN_CONFLICTS > 0) was never disclosed there, and "Proceed anyway"
could exit successfully over it — exactly the failure mode this PR exists to close (a success
banner over an unresolved conflict nobody resolved). Blockers still block: withhold "Proceed
anyway" and route to Manual review whenever a conflict is open.

* fix(#3771): do not hard-block REVISION_CONFLICT persistence when no REVIEWS.md exists yet

A phase's first-ever revision cycle can return REVISION_CONFLICT before any REVIEWS.md has been
written — REVIEWS_PATH is then legitimately empty, not a corrupt or deleted file. The persistence
gate's own accompanying prose already says the record channel applies 'when REVIEWS_FILE is
non-empty', but the bash condition never checked that, so it hard-blocked every conflict on a
brand-new phase regardless of whether persistence was even expected to run. Require a non-empty
REVIEWS_FILE before treating a missing file as an error.

* chore(#3771): raise the plan-phase.md ADR-857 host-loop ceiling to 96700

The frozen pre-phase-6 ceiling (94519) collided on rebase: this PR's own
REVISION_CONFLICT persistence/routing gate is core planner control flow, not an
un-extracted optional feature, and landed alongside an unrelated, already-merged
same-file growth (the #4.6 context-drift pre-check) already on next. Same
rationale #1298 already established for execute-phase.md's ceiling.

* chore(#3916): backfill changeset pr field to the upstream PR number

* fix(#3771): make the writer-side REVISION_CONFLICT sanitize step real shell

The Conflict Return record channel sanitized agent-authored fields via a
prose instruction ("Sanitize each agent-authored field before appending")
for the orchestrator LLM to apply by hand, while the reader-side gate in
plan-review-convergence.md parses the same slot with real, executed awk.
Flagged Minor across two review rounds (round 4, round 6) since no code
performed the sanitize anywhere.

plan-phase.md's Conflict Return step now runs a real bash gate: sanitize
each field (collapse newline/tab to space, strip a leading #/-/|/fence),
build the one-line record, skip the append if an identical line already
exists (idempotent), insert before the writer-owned end delimiter, and
fail closed if that delimiter is missing rather than silently dropping
the conflict.

tests/revision-remediation-binding.test.cjs extracts and RUNS the new
fence (matching how the reader gate is already tested), composing it
with the existing reader gate: hostile-field fast-check fuzzing, a
repeated-conflict idempotency check, and a missing-delimiter fail-closed
check that the file is left byte-for-byte unchanged on failure.

Emitted-Drift-Ack-Growth: plan-phase.md — #3916 turns the writer-side
REVISION_CONFLICT sanitize+insert step into real, executed shell instead
of a prose instruction, matching the reader gate's existing rigor

* fix(#3916): backfill changeset pr field to the fork validation PR number

Fork CI's changeset-lint reads the real PR number from its own event
payload; the fragment still carried the upstream number from the last
sync, so the DEFECT.CHANGESET-PR-FIELD-DRIFT check failed on this fork
PR. Re-backfill to the upstream number before the final push.

* fix(#3771): close the awk -v forgery and same-session close gaps agy found

Adversarial review (gemini-3.8-flash-high via the internal agy review
lane) on the full PR found two BLOCKERs against the just-added
writer-side conflict gate:

1. `awk -v line="${LINE}"` decodes escape sequences in its argument, so
   a literal two-character `\n` in agent-authored text became a real
   newline inside awk, splitting the appended record across two
   physical lines. `tr` only strips actual control bytes, so it never
   saw this — it defeated the exact forgery the gate exists to
   prevent, both the reader's zero-count and the writer's own
   idempotency check. Fixed by passing LINE/END through awk's
   ENVIRON, which is not escape-decoded.

2. A conflict resolved and re-spawned within the same plan-phase
   session was never flipped from `- [ ]` to `- [x]` — the record
   channel bullet said "plan-phase closes it," but no step did. Only
   a *separate* `--reviews` re-entry (line ~622, still prose-only)
   closes conflicts; the in-session resolve path left them open
   forever, permanently blocking convergence. Fixed by carrying the
   just-written line in `PENDING_CONFLICT` and closing it in the
   `Otherwise` branch before the checker re-spawns.

Also fixed a MAJOR: docs/COMMANDS.md described the `--max-cycles`
escalation gate as uniformly offering "proceed or review manually,"
but the code (this PR's own change) withholds "Proceed anyway"
specifically when a plan-revision conflict is open — only manual
review is offered in that case. Docs now say so.

Not applied: the reviewer's `\r` truncated to plain tr from a MINOR
that also asked for temp-file permission preservation across `mktemp`.
Applying `chmod --reference` is not portable to macOS/BSD `chmod`, so
this is left as a documented low-severity tradeoff — the temp file now
sits alongside REVIEWS.md (same filesystem, atomic `mv`), which was
the same finding's more substantive half. Also not applied: a
suggested `gsd_run review record-conflict` CLI subcommand to
deduplicate the two `awk` blocks — a new command plus wiring is out of
scope for a review-remediation fix.

tests/revision-remediation-binding.test.cjs adds regression coverage
for both BLOCKERs: a literal-backslash-n hostile field composed with
the reader gate, and a close-gate extraction that verifies the flip to
`[x]`, the reader's count dropping to 0, and a fail-closed path when
the pending line is missing.

Emitted-Drift-Ack-Growth: plan-phase.md — #3916 fixes an awk -v escape-
decoding forgery and adds the missing same-session conflict-close step
an adversarial review found in the writer-side gate

* chore(#3916): backfill changeset pr field to the upstream PR number

Fork-validation CI needed pr: 1 to pass its own changeset-lint; restore
pr: 3916 before this push reaches open-gsd/gsd-core.

* fix(#3771): trim plan-phase.md prose back under the XL byte cap

Merging origin/next's unrelated growth pushed plan-phase.md 473 bytes
past the workflow-size-budget XL cap and the ADR-857 phase-6 baseline,
both tripped by CI after review approval. Removed an unpinned inert
bash comment and tightened connective prose in three REVISION_CONFLICT
bullets; no executable shell or test-pinned substring changed.

* chore(rehearsal): pin changeset pr field to fork rehearsal PR #25

Scratch-only commit for the rehearsal branch's own CI. Will not be
carried onto the branch backing upstream #3916 — that keeps pr: 3916.

* fix(#3771): address CodeRabbit findings on the REVISION_CONFLICT protocol

Fork rehearsal PR #25's first CodeRabbit pass surfaced 7 findings against
the already-approved #3916 diff; each verified against current code
before fixing (none hallucinated):

- plan-phase.md: writer-side awk gates now strip a trailing \r before
  comparing lines, matching the reader gate (plan-review-convergence.md)
  -- a CRLF REVIEWS.md previously made both writer gates fail closed.
- plan-phase.md: the close-fence's REVIEWS_FILE/PENDING_CONFLICT/
  CONFLICT_RESOLUTION were read without ever being (re)defined in that
  fence -- shell state does not survive across separate fenced blocks
  (same convention already documented in review.md). Added the explicit
  recompute/set instruction.
- revision-loop.md: previous_conflict_property was never reset after a
  normal (non-conflict) revision, so a later, unrelated conflict on the
  same property could be misread as a repeat and escalate prematurely.
- gsd-plan-checker.md / few-shot-examples/plan-checker.md: two example
  required_property strings were unconditionally binding in a way their
  own dimension's rules aren't (no-analog RESEARCH.md fallback; tasks
  that create no functions), now scoped to match.
- quick/steps/plan-checker-loop.md: added the same disjoint
  "Otherwise (not REVISION_CONFLICT)" branch plan-phase.md already had,
  closing an ambiguity between the conflict and non-conflict return paths.
- revision-remediation-binding.test.cjs: the REVIEWS_PATH init-order
  assertion used indexOf() without checking for -1, so it would pass
  vacuously if either anchor were renamed away.

Also restores an "Export the row's CONFLICT_*" instruction I had cut in
the prior byte-budget trim -- checked non-pinned by tests, but it was the
only text telling the agent to set those vars before the awk block reads
them via ENVIRON.

Net growth from these fixes required reclaiming bytes elsewhere in
plan-phase.md (verified against every pinned substring in
revision-remediation-binding.test.cjs) to stay under the XL tier's
hard 98304-byte cap; final size 98245 bytes.

* fix(#3771): resync the #4079 shrink-only mirror to the current PRE_PHASE6 line

tests/plan-phase-background-wait-wakeup.test.cjs (landed on next via an
unrelated #4079 PR, merged in by this branch's next-sync) mirrored
plan-phase.md's phase6 shrink-only ceiling as a hardcoded local constant
(94519) rather than reading tests/phase6-capstone-conformance.test.cjs's
PRE_PHASE6 value. That value has since been legitimately raised twice
during this PR's own review (94519 -> 96700 -> 98300) to accommodate the
REVISION_CONFLICT persistence/routing gate. The two branches' independent
histories left the mirror stale post-merge -- not a textual git conflict,
but the same class of thing. Resynced to 98300.

* fix(#3771): address round-2 CodeRabbit findings on the conflict gates

CodeRabbit's re-review of the previous remediation commit found two real
issues in what it had already flagged:

- Both writer-side awk CRLF fixes used \`sub(/\r$/, "")\` directly on \`\$0\`,
  which mutates it in place -- \`{ print }\` then emitted the CR-stripped
  copy for every passed-through line, silently rewriting an unrelated
  CRLF REVIEWS.md to LF on any insert or close. Now compares against a
  separate \`cur\` copy and prints the original, untouched \`\$0\`.
- The close-fence's "recompute REVIEWS_FILE/PENDING_CONFLICT" prose
  implied in-fence derivation, but the fence has no such code and the
  test harness (\`runCloseGate\`) deliberately supplies all three as
  pre-set env vars -- matching how the open fence's "Export the row's
  CONFLICT_*" instruction already works. Reworded to "export ... in the
  same invocation", matching that established, test-verified pattern
  instead of promising logic that isn't there.

Added a regression test proving the CRLF fix no longer touches
passthrough lines (red against the mutate-in-place version, green now).

* fix(#3771): use a CRLF-safe check in the new passthrough regression test

local/no-crlf-fragile-split forbids splitting readFileSync content on a
literal \n (Windows git-autocrlf checkouts yield \r\n). My CRLF
passthrough-preservation test from the previous commit did exactly that
to inspect the first line. Replaced with a direct startsWith() check
against the known CRLF-terminated header, which needs no split.

* test(#3771): assert the record itself is inserted in the CRLF passthrough test

CodeRabbit nitpick (round 3): the passthrough-preservation test checked
gate status and the pre-existing line's CRLF ending, but never asserted
the new REVISION_CONFLICT record was actually written.

* fix(#3771): address agy/gemini-3.8-flash-high adversarial review findings

Full-PR adversarial review (internal /gsd-review antigravity lane,
gemini-3.8-flash-high) surfaced 9 findings; each verified against current
code before fixing (none hallucinated):

HIGH:
- quick-batch/steps/plan-checker-loop.md never received the
  required_property/fix_hint binding language or REVISION_CONFLICT
  handling this PR added everywhere else -- a genuinely unmigrated
  producing context. Migrated to match quick/steps/plan-checker-loop.md,
  and added it to the ORCHESTRATORS consistency battery in
  revision-remediation-binding.test.cjs so future drift is caught
  automatically.
- The close-fence's PENDING_CONFLICT was an agent-supplied env var that
  had to exactly reconstruct a five-field sanitized line across a
  multi-minute subagent dispatch -- fragile, and a scalar var also meant
  a second simultaneous conflict silently dropped the first on overwrite.
  Redesigned to match the open conflict by CONFLICT_DIMENSION/
  CONFLICT_PLAN identity instead: the agent re-supplies two short,
  already-tracked identifiers rather than reconstructing the full
  sanitized text, and each conflict resolves independently regardless of
  how many are open. Updated the test harness's runCloseGate contract to
  match, and added a two-open-conflicts regression test.

MEDIUM:
- plan-phase.md's `--reviews` replanning path told the reader to "flip
  the matching line to [x]" in prose only, with no executable path to
  it -- pointed it at the same close gate used in step 12.
- plan-review-convergence.md's reader-gate awk tolerated a blank line
  before the opening delimiter but not before the heading that follows
  it; a formatter or LLM writer inserting one would hard-abort
  convergence on an otherwise well-formed REVIEWS.md. Added the same
  tolerance already granted above it, with a regression test.

LOW:
- Clarified that the escalation destination for a stalled conflict is
  the same iteration/revision-count cap gate already defined in each of
  quick, quick-batch, ui-phase, and verify-work, rather than an
  undefined "stall" concept.
- Clarified "twice in a row" means no successful revision intervened,
  matching revision-loop.md's now-explicit previous_conflict_property
  reset.
- Fixed gsd-ui-researcher.md's stale rationale text, copied verbatim
  from planner-revision.md: ui-phase presents the conflict table
  directly to the user, it does not persist to a shared file scanned by
  heading.

Net growth again required reclaiming bytes in plan-phase.md (verified
against every pinned substring in revision-remediation-binding.test.cjs)
to stay under the XL tier's hard 98304-byte cap; removed a now-dead
PENDING_CONFLICT assignment in the process. Final size 98258 bytes.

* fix(#3771): scope row 48's quick/steps guard away from plan-checker-loop.md

tests/gsd-quick-batch-quick-regression.test.cjs's row 48 (#3676) flagged
this branch's quick-batch/steps/plan-checker-loop.md migration (the agy
HIGH finding) as a violation, because it also edits
quick/steps/plan-checker-loop.md for the same underlying #3771 protocol
fix.

Verified against git history before scoping: 2f64e6230 (#3676's own
landing commit) CREATED quick-batch/steps/plan-checker-loop.md as a new,
independent 119-line file, never a call-site into quick/'s copy. Row
48's "shared primitives, never edits the ordinary quick command" premise
was never about this specific file -- it was always meant to carry its
own per-flow copy of whatever revision-loop contract applies, same as
ui-phase.md/verify-work.md throughout this PR. This is the same
false-positive class the row's own comments already document scoping
away twice (#3730, #2529 round 40); excluded plan-checker-loop.md from
its touched-quick-steps check with the same evidence trail.

* chore(#3771): point changeset pr field at upstream PR 3916

---------

Co-authored-by: davdittrich <davdittrich@gmail.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 15:16:38 -04:00