Commit Graph

1129 Commits

Author SHA1 Message Date
Rezolv
7bcfbe4542 docs(#4629): add ADR-4629 — STATE.md write intent beyond frontmatter (#4645) 2026-09-12 11:43:16 -04:00
Tom Boucher
241646a43a fix(#4651): classify .env names by final extension, and close the trailing-dot alias bypass — Phase 1 of #4636 (#4659)
* test(#4651): failing-first coverage for final-extension classification

Phase 1 of epic #4636, absorbing #4580. Tests only; no fix. These MUST fail.

The guard classifies a name by comparing everything after `.env.` as one
token against a set whose members are FINAL EXTENSIONS. So `.env.local.example`
yields suffix `local.example`, which is not a member, and a committed
secret-free template is refused. That is a category error, not strictness.

Two arms are covered because the same classification is hand-rolled twice in
one file: `isSecretBasename` for Read/Bash, and `globAltSelectsSecret`
(`lit.startsWith('.env.')`) for Grep globs. Fixing one alone would ship a
guard that allows `cat .env.local.example` while refusing
`Grep --glob '.env.local.example'` — the same file, the same hook, opposite
answers. A cross-arm parity loop over one shared list asserts the two cannot
drift.

Rows that exist because they are the ones nobody enumerates:

- `.env.example.local` must stay BLOCKED. Final extension is `local`; this is
  dotenv's documented local-override convention and a real secret. Any fix
  shaped as "contains example" admits it.
- `.env.local.` must stay BLOCKED — empty final extension is not a member.
- `.env.` must stay ALLOWED. Note #4580's proposed patch adds
  `if (suffix === '') return true;`, which flips it to blocked; that breaks the
  existing `allows` assertion in this suite and broadens the protected set,
  which epic #4636's non-goals forbid. Not applied.
- `.env.local.exam*` (partial glob literal) must stay BLOCKED — it can select
  `.env.local`, and a partial literal cannot be classified.
- `*.example` and `*` must stay ALLOWED — regression protection on the arm
  that already works.

Local behavioral repro of the current guard, confirming the tests fail for the
right reason rather than by construction:

  .env.local.example  rc=2 (blocked)   <- the defect
  .env.example        rc=0 (allowed)
  .env.local          rc=2 (blocked)
  .env.example.local  rc=2 (blocked)
  .env.               rc=0 (allowed)
  glob .env.local.example  rc=2        <- the second arm

Regressions are folded into the owning module's suite rather than a new
tests/fix-NNNN-*.test.cjs file, per scripts/lint-regression-test-names.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): classify by final extension so .env.<name>.example is readable

Phase 1 of epic #4636, absorbing #4580. Implements ADR-4650 decision 5.

The guard compared everything after `.env.` as ONE token against a set whose
members are FINAL EXTENSIONS. `.env.local.example` yielded `local.example`,
which is not a member, so a committed, secret-free template was refused — the
guard blocked the one file that exists so nobody has to open the real `.env`.

That is a category error, not strictness. The fix is not "add local.example to
the set"; it is to compare the right token. hooks/lib/filename-classification.js
now owns that distinction and is the only place it is expressed.

Both arms are fixed, because the same classification was hand-rolled twice in
this one file:

  - isSecretBasename (Read/Bash) now tests finalExtension(suffix).
  - globAltSelectsSecret (Grep --glob) split its first branch. With no
    wildcard the alternative IS a whole filename, so it is classified exactly
    via isSecretBasename. With a wildcard present the literal is only a
    PARTIAL prefix (`.env.local.exam*` can still select `.env.local`) and
    cannot be classified, so the original conservative rule stays.

Fixing only the first would have shipped a self-contradicting guard: `cat
.env.local.example` allowed while `Grep --glob '.env.local.example'` refused —
same file, same hook, opposite answers. A cross-arm parity loop over one shared
list now asserts the two cannot drift.

Two deliberate departures from #4580's suggested patch, both verified:

  - Its `if (suffix === '') return true;` is NOT applied. That flips `.env.`
    from allowed to blocked, breaking an existing assertion in this suite and
    broadening the protected set, which epic #4636's non-goals forbid.
  - `fullSuffix` was drafted alongside finalExtension and removed before
    commit: zero production consumers, and none planned (Phases 2-4 are
    containment, duplicate draining and the path-join ratchet, none of which
    classify filenames). A zero-caller export is dead code. The distinction is
    pinned instead by a test asserting finalExtension('local.example') is
    'example' and explicitly NOT 'local.example'.

The protected set is unchanged. `.env.example.local` stays BLOCKED — its final
extension is `local`, dotenv's local-override convention and a real secret;
any fix shaped as "contains example" admits it.

Scoped out by measurement, not assumption: src/validate.cts:395 and
src/phase.cts:1674 also hand-roll lastIndexOf('.'), but both parse phase
identifiers (`3.2` -> parent `3`), owned by the phase-id.cts seam. Folding
them in would repeat this same category error in the opposite direction.

Checkpoint 1 (prove RED) on the tests-only commit 91d3d6e1: outcome=failed,
26 failures / 45330, all 26 in the two new test files, zero pre-existing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): document the widened template exemption and cover the Bash arm

Two findings from the isolated adversarial review, both fixed in place.

1. The header's "Stated cost" passage named only the four literal template
   names, but since this change the exemption keys on the FINAL EXTENSION, so
   the trusted set is `.env.<anything>.{example,sample,template,dist}` — an
   unbounded family. The reviewer demonstrated it: `.env.prod-real-secrets.example`
   is allowed. That is the deliberate and necessary cost of fixing #4580, but
   it was materially larger than what the header disclosed, and a silent
   expansion of a security guard's trusted set is not acceptable. The passage
   now states the family, the concrete bypass, and that it applies across
   Read, Grep and Bash alike.

2. The cross-arm parity loop asserted Read and the exact-literal Grep glob but
   not Bash, whose `namesSecret` -> `isSecretBasename` path is genuinely
   distinct. The Bash arm was covered only by two one-off tests outside the
   shared table, so the table could not have caught a drift there. The loop now
   drives all three arms from the same TEMPLATES/SECRETS arrays.

No classification logic changed. The Read-arm behavioral table is byte-identical
before and after: rc=0 for .env.local.example / .env.example / .env. ; rc=2 for
.env.local / .env.example.local / .env / .secrets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): treat trailing dots and spaces as aliases of the protected file

Closes a Windows path-alias bypass surfaced by the isolated adversarial review
of this phase. Maintainer-approved as in scope.

Win32 strips trailing dots and spaces from every path component, so `.env.`,
`.env..`, `.env `, `.env. `, `.env .`, `.secrets.` and `.secrets ` all resolve
to the real `.env` / `.secrets` on Windows. The guard allowed every one of them
— a bypass of a file it already protects, reachable from Read, Grep and Bash
alike. `isSecretBasename` now normalizes the basename before classifying.

The whole class is fixed, not the reported name. `.env.` alone would have left
`.secrets.` and the trailing-space forms open, which is the same
one-cause-explains-every-failure trap this epic exists to close.

Two consequences, both measured rather than assumed:

  - `.env.example.` flips blocked -> ALLOWED. It aliases the already-trusted
    `.env.example` template, so this is correct; it was previously blocked only
    because the trailing dot broke final-extension parsing.
  - A Bash token that is exactly `.env` plus trailing whitespace flips
    allowed -> BLOCKED. Verified this is CONSISTENCY, not a new false-positive
    class: the bare `.env` token was ALREADY blocked as an operand in the same
    position before this change, so the alias now simply behaves like the thing
    it aliases.

The header's "No whitespace trimming" guarantee is preserved and now stated
precisely: leading and interior whitespace is still never trimmed, so prose
like a commit message mentioning `.env` in a sentence stays prose and stays
allowed. Only TRAILING dots and spaces are stripped. Two tests pin that.

This lands at the same behavior #4580's proposed `if (suffix === '') return
true;` would have produced for `.env.`, which this phase earlier rejected. The
rejection was correct on its stated grounds — that line broadens the protected
set, which epic #4636's non-goals forbid. The Windows framing is different:
normalizing an alias of an already-protected file is not a broadening, and the
fix is reached by normalization rather than by special-casing an empty suffix,
so it generalizes to `.secrets.` and the space forms.

Cannot be reproduced on this host — the remote matrix is Linux-only and Windows
coverage arrives from CI — so this ships on the Win32 path-normalization
contract plus the CI lane, and that limitation is stated rather than implied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): one owner for path segmentation, closing a Read/Grep divergence

Four findings from the two-axis review, all fixed in place.

The real one: the guard had TWO path-segmentation rules. `lastSegment` (used
by Read and Bash via `namesSecret`) splits on both `/` and `\`, while
`classifyGrepGlob` hand-rolled its own on `/` only. Measured:

  Read  of `config\.env`      rc=2  BLOCKED
  Grep  --glob 'config\.env'  rc=0  ALLOWED

Same logical file, opposite answers — precisely the divergence this epic
exists to remove, sitting inside the file this phase was already fixing.
`lastSegment` now lives in hooks/lib/filename-classification.js and both arms
call it. All five path-bearing cases (both separators) now agree.

Note on how this was nearly missed: the first measurement of it reported
"both allow", which looked like the reviewer was wrong. That reading was a
measurement artifact — `config\.env` inside a printf'd JSON payload is an
invalid escape, so the hook fails open at rc=0 and the test was observing
JSON breakage rather than the predicate. Re-measured with correct escaping,
the divergence is real. The tests added here use properly escaped literals
and were verified by running, not by reasoning about the escaping.

Also fixed:

  - Both fast-check properties were satisfied by a degenerate
    always-return-'' implementation: "never contains a dot / is a suffix" and
    "never ends with dot-or-space / is a prefix" are both trivially true of
    the empty string. They now additionally pin content preservation — the
    removed tail must match /^[. ]*$/, and a name with nothing to strip must
    come back unchanged.
  - The cross-arm parity loop used only bare basenames, so it could not have
    caught the divergence above. It now covers path-bearing names with both
    separators.
  - That loop's description overclaimed: Read and Bash BOTH route through
    `namesSecret`, so they are not independent paths; only the Grep glob arm
    is genuinely separate. The description now says so rather than implying
    three-way independence.
  - `normalizeWindowsBasename` runs on every platform, not only Windows. Its
    doc now states that explicitly: the guard must answer identically
    everywhere, and a name is judged by what Win32 would resolve it to.

No classification logic changed; the 12-name regression sweep is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4651): regenerate install-tree goldens, correct the guard's user-facing docs

Three things, all consequences of the fix rather than new behavior.

1. Install-tree goldens. `hooks/lib/filename-classification.js` is a SHIPPED
   file — package.json `files` includes `hooks` — so every per-runtime install
   tree gains a path. Checkpoint 2 failed on exactly this: 11 failures, all in
   tests/golden-install-tree.test.cjs, against 45356 passing. Regenerated via
   scripts/gen-install-tree-fixtures.cjs; 11 goldens changed, matching the 11
   failures one-for-one.

   This ripple was identified at design time and then not acted on. Fleet's
   impact preview named golden-install-tree.test.cjs before any code was
   written, and 40-design.md records it under "Ripples identified". Writing a
   risk down is not the same as discharging it, and a full matrix run was spent
   discovering something already known.

2. docs/USER-GUIDE.md made a precise and now-false claim about the guard's
   protected set: it named `.env.example` / `.sample` / `.template` / `.dist`
   as the four exempt names. The exemption keys on the FINAL EXTENSION, so the
   exempt set is the unbounded family `.env.<anything>.{example,sample,template,dist}`.
   The page now states that family, the widened residual, that order matters
   and only the last segment counts (`.env.example.local` is a secret), and
   that trailing dots and spaces are stripped because Windows resolves them to
   the protected file. A wrong user-facing model of what a security guard
   protects is worth correcting even though Fixed/Security changesets are
   exempt from the required-docs rule.

   docs/ARCHITECTURE.md and docs/INVENTORY.md say "templates such as
   `.env.example` exempt" — non-exhaustive, still true, deliberately left
   alone. Same for the ja-JP / zh-CN / ko-KR / pt-BR rows, which carry the same
   hedged phrasing; hand-translating a security description unreviewed is not
   something to do silently.

3. Two changeset fragments, not one. A refusal corrected is `Fixed`; a bypass
   closed is `Security`. Folding the second into the first would under-report
   it in the release notes. Both carry `pr: 0` for backfill once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4651): backfill changeset PR number to 4659

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 08:42:43 -04:00
Tom Boucher
8fa2c3dbcf docs(#4650): record the path-containment and filename-classification design lock (#4655)
Phase 0 of epic #4636. ADR-4650 fixes the decisions Phases 1-4 inherit, so that
four phases do not each invent them independently.

The load-bearing decision is that the engine and the exported shape are
separable. A resolver-based, symlink-safe containment predicate already exists
as validatePath, and building the epic's literal assertWithinRoot() from scratch
would create a sixth implementation of the very thing this epic consolidates --
while risking silent loss of behavior validatePath acquired as bug fixes (a
closed dangling-symlink existence oracle, ancestor canonicalization for
non-canonical roots, a separator-aware boundary test).

But the epic's other clause is correct and lands on the current export:
validatePath returns a boolean a caller can forget to check, and populates
`resolved` with the escaping path precisely on the traversal branch. While that
form stays exported, the Phase-4 ratchet could only assert that a helper was
called -- validatePath(x, root).resolved would pass the rule.

So: preserve the engine, narrow the export. assertWithinRoot becomes the only
export and yields a branded ContainedPath.

Also recorded, each found by measurement rather than from the epic text:

- The rejection message text is a real contract. tests/quick-batch.test.cjs
  asserts a user-facing `reason` field matches /escapes allowed directory/, so
  the string reaches CLI consumers and Phase 3 must preserve it verbatim.
- Two further unconfined boundaries the epic does not enumerate: resolvePath
  and gap-analysis.plan-post, both in check-command-router.cts.
- --phase-dir also interpolates into ${PHASE_DIR} for command-exit-zero, so
  confining at the boundary covers both predicate kinds; the evaluator stays
  fs-free.
- opts.allowAbsolute is a per-call-site liberality knob, which is an acceptance
  policy living exactly where this ADR says it must not.
- planning-inspect's isWithinRoot is deliberately pure-string with no I/O; its
  contract differs, so Phase 3 decides rather than assumes.

The acceptance policy is stated once: conservative about the resource, exact
about the classification. #4580's guard was not too strict, it was wrong -- a
category error comparing a whole suffix against a set of final extensions.

ADR opens as Proposed; ratified at Phase 4 closeout per docs/adr/README.md.
No changeset: the diff touches docs/adr/ only, which is outside
USER_FACING_PREFIXES in scripts/changeset/lint.cjs.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 23:45:04 -04:00
Tom Boucher
4d65c248e5 fix(#4641): make test-conformance the sole Windows selector and narrow the tier to 28.5% (#4643)
* test(#4641): failing-first tests for the tier ceiling and a single Windows selector

Tests only, committed ahead of the implementation so the RED run is real.

- tests/platform-conformance-tier.test.cjs: tier-size ceiling asserted as a
  ratio against a live denominator (Windows 33%, macOS 25%); per-helper negative
  cases proving seam calls and path-call-plus-slash-literal are not platform
  signals; positive pins that genuine platform content, seam-bypassing spawns,
  chmod and symlink still classify in; macOS signal set and generated list
  unchanged.
- tests/ci-full-lane-sharding.test.cjs: the test job has zero windows-latest
  rows and test-conformance still has 3 windows + 1 macOS.
- tests/ci-test-scope.test.cjs: windows_tests is absent rather than empty, a
  non-tier test file no longer forces full_matrix, a RULE-pulled windows-hint
  test does, and resolveSelection rejects the retired windows scope.

Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): delete the second Windows selector and narrow the conformance tier

Epic #4589's goal — the OS-agnostic bulk on Linux, a small explicitly-scoped
conformance tier on real Windows/macOS — was not met. Measured on PR #4640
(run 34618834118): 7 non-Linux jobs, a 546/930 (58.7%) "tier", and 5 of 7
changed test files running on a real Windows runner twice.

Two selectors, only one in the epic's scope. The test job's three scope:windows
shards predate the epic (#494, sharded #3057) and gate on product_changed, not
full_matrix, so they fire on every product PR whatever Phase 3's classifier
decides. They are deleted; test-conformance becomes the sole Windows selector,
as it already was for macOS. Non-Linux jobs 7 -> 4.

Gating the lane instead was rejected as provably redundant: for a test file
reachesConformanceTierOrSeam is literally CONFORMANCE_TIER_FILES.includes(file),
and that same predicate sets full_matrix, which turns test-conformance on. Every
file a gated lane would run is already covered in the same run. The lane's one
non-redundant residue -- RULE-pulled tests matched by the isWindowsHint filename
heuristic -- is ported into reachesConformanceTierOrSeam so it sets full_matrix
instead of feeding a parallel lane.

Two detectors matched the repo's own test idiom rather than any platform signal
and carried 226 of the tier's sole-signal membership against 41 for the other
eight: process-seam-subprocess (335 files, 118 unique) matches the
tests/helpers.cjs entry points nearly every CLI test uses, and going through the
seam is the opposite of a platform signal since shell-command-projection takes
platform as an injected parameter; hardcoded-path-vs-path-call (328, 108) needs
only a path call anywhere plus a slash literal anywhere, and that class is
already enforced by ADR-1703's Linux-runnable ESLint rules. Both are removed.
Tier 546 -> 254 (27.3%). src/ reachability is unchanged at 28 files, measured.

Adds the size gate Phase 2 never had, as a ratio against a live denominator so
it cannot stop binding as the suite grows.

292 files leave real-OS Windows execution. The drop-out set was audited: 14 have
a platform-suggestive filename and all 14 are static source-text analyses or
seam-mediated CLI tests. raw-child-process was investigated as a suspected false
negative and left unchanged -- relaxing it adds 13 files, all false positives.

macOS is untouched: MACOS_CATEGORIES is a separate array and the regenerated
macos-conformance-tier.generated.cjs is byte-identical at 196 files.

Fixes #4641
Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): register the new ADR path in the docs-guard exempt baseline

tests/ci-test-scope.test.cjs references docs/adr/4641-windows-selector-consolidation.md
in a comment justifying the retired windows scope; lint-docs-guard-registration
tracks that reference set, so the baseline needs the new path. Verified the
exemption still holds: the path is prose, not a filesystem read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): make the escalation tier-backed and drop every hardcoded count

Three follow-ups from measuring the first pass rather than trusting it.

The windows-hint escalation now requires tier membership as well as the
filename hint. Setting full_matrix runs test-conformance, which runs only the
tier; escalating on a test that is NOT in the tier costs four jobs and still
never runs that test on Windows. Measured over the 16 RULES entries the
narrowed predicate fires on exactly the same rules today, so this is
correct-by-construction rather than a behavior change. The broader variant --
escalate on any tier member a rule pulls in, ignoring the hint -- was measured
at 14/16 rules and rejected as over-broad.

Removes the hardcoded counts. A hardcoded macOS tier length of 196 broke as
soon as the rebase pulled in one new test file from #4253, which is the whole
argument against them: the ceilings are ratios against a live denominator, the
committed lists are pinned by comparison against a fresh classification of the
live tree, and the three named probe files now assert on their SIGNAL rather
than on membership in a literal list -- asserting by filename is the exact
error this PR fixes in the classifier.

Regenerates both lists against the rebased tree. Same-tree figures are now
547 -> 255 of 931 eligible (58.8% -> 27.4%), 292 entries removed and none
added; macOS is unchanged at 197 with a zero-line diff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): restore real-shell-spawn coverage and repair assertions the narrowing broke

An isolated adversarial review found a real false negative. Removing the
blanket process-seam-subprocess detector also removed the only coverage for
tests that spawn a REAL shell: tests/helpers/process-seam.cjs's runHook
spawns options.interpreter via real spawnSync, so
runHook('-c', [script], { interpreter: 'bash' }) runs a real bash binary
executing a shell script extracted from workflow markdown. The seam argument
holds for src/shell-command-projection.cts, which takes platform as an
injected parameter; it does NOT hold for the test helpers, which spawn real
binaries. Conflating the two is what made the blanket detector look purely
noisy -- it was 99% noise wrapping a real signal.

Adds a narrow shell-interpreter-spawn category keyed on a real interpreter
option. Measured 2026-09-11: 33 files match, 9 were outside the tier and are
added back, taking it 255 -> 264 of 931 (27.4% -> 28.4%), still under the 33%
ceiling. All 9 confirmed by reading the matching source line, zero comment or
fixture matches. runGit-alone and non-node-spawnSeam alternatives were measured
and rejected -- each adds 9 files but misses the counterexample entirely.

Fixes a real bug the suite caught: jobs.test is ubuntu-only now that its
scope:windows rows are gone, so it must wire GSD_STRICT_LIVE_CONFIG_GUARD
strictly rather than carrying the Windows report-only carve-out. The carve-out
now lives solely on test-conformance, whose matrix does include windows.

Repairs seven pre-existing assertions the category removal invalidated,
preserving each case's purpose rather than deleting coverage, and converts the
last hardcoded tier bounds to live-derived ratios -- including the macOS
sanity range that was still a magic [100, 350].

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): keep the confinement test on a real OS via a documented allowlist

A security review found tests/external-descriptor-confinement.test.cjs had
dropped out of the Windows tier. It must stay in, and no content signal can
express why: it exercises isPathConfined (src/external-descriptor-trust.cts),
which uses the AMBIENT path module -- path.resolve(root, target) and path.sep
-- with no injection. Its win32 semantics (drive letters, UNC, separator) are
only reachable by actually running on Windows, and it is a security-relevant
write-confinement gate. A content classifier cannot see 'this module reads the
ambient path module', so no regex belongs here.

Adds ALWAYS_REAL_OS, a Map of path -> recorded reason, unioned into the Windows
tier only. A Map rather than a list so an entry without a reason is impossible
by construction, and tests assert every entry names a file that exists on disk
so a stale entry fails loudly instead of rotting. This is the centrally-
enumerated single source of truth epic #4589 Phase 2 asked for and ADR-1703's
portability-vocab.cjs already models -- deliberately not a heuristic.

Windows tier 264 -> 265 of 931 (28.5%), still under the 33% ceiling. macOS is
untouched and byte-identical: the win32 concern does not apply to a POSIX
runner, and a test asserts the allowlist does not leak into that tier.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): inject the path impl into isPathConfined and correct the ADR count

Two review findings, both fixed rather than dispositioned.

A security review found tests/external-descriptor-confinement.test.cjs had left
real-OS execution. The allowlist pinned it back, but that only restored
INCIDENTAL coverage: isPathConfined used the ambient path module, and its test
carried POSIX-only literals, so a win32 confinement escape was unverified on
every platform including Windows. isPathConfined now takes an optional third
parameter carrying the path implementation, defaulting to the ambient module.
Blast radius is CRITICAL -- 53 affected symbols across 19 files -- so the change
is purely additive and every existing two-argument caller is byte-identical.

Tests now inject path.win32 and path.posix, covering a different drive letter,
a cross-drive absolute, backslash and forward-slash traversal, UNC, and the
startsWith prefix-boundary bug (.gsdEVIL against root .gsd) on both separators.
Proved load-bearing: dropping the + p.sep from the prefix check fails exactly
the two boundary cases and nothing else. Callers' suites 149/149.

The spec review caught an off-by-one: the ADR narrated a 264-file tier while the
committed list holds 265. The ADR now records the full chain 547 -> 255 -> 264
-> 265 (28.5%).

Also corrects a stale comment in scripts/docs-guard-registry.cjs that narrated
classify() as zeroing windows_tests, a key this change removes -- kept as
historical narration but labelled as such.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#131): make the unwritable-HOME test actually test something

Found by sweeping for the root-bypass class after fixing commit-files-deletion.
This one is the silent variant, and it was broken twice over.

First, the condition: the test made a fake HOME unwritable with chmod 0o500.
The gsd-test Docker bench runs as root, root bypasses mode bits, so HOME stayed
writable and the hostile condition never existed. Replaced with a HOME whose
PARENT is a regular file, so every write under it fails ENOTDIR at the VFS
layer for every uid -- no permission check is involved at all.

Second, and more fundamental: the probe was npm --version, which on npm 11.19.0
performs zero filesystem I/O against HOME. Proven rather than assumed --
neutralizing runNpm()'s isolation turned the sibling test red while this one
stayed green, so its assertion could never detect the regression it guards, on
any uid, with or without the condition fix. npm config get cache was tried next
and proved vacuous the same way (it only string-resolves the path). The probe is
now npm cache verify, which really does mkdir _cacache under HOME.

Re-proved load-bearing after the change: with isolation neutralized the test now
fails with ENOTDIR on <blocker>/home/.npm/_cacache. tests/helpers.cjs was
restored and verified diff-clean; suite 13/13.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the net drop-out figure in ADR-4641

The Consequences section still said 292 files leave real-OS Windows execution.
That was the count before the narrow shell-interpreter-spawn replacement
restored 9 and ALWAYS_REAL_OS pinned 1. Net is 282. Also names both real-binary
categories rather than only raw-child-process, and clarifies that the 14-file
filename audit was against the 292 initially dropped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the rejected concentration ceiling and its measurement

Applying Goodhart's own question to the new ceiling -- how would you make this
metric look good without improving what it represents -- surfaces a real
weakness: a ratio can be satisfied by inflating the denominator, so adding
OS-agnostic tests loosens it without narrowing the tier.

The obvious companion gate was a sole-signal concentration ceiling, since the
original defect was one detector carrying half the tier. Measured and rejected:
peak concentration post-fix is raw-child-process at 53/265 = 20.0%, against the
historic offenders at 21.6% and 19.8%. Any threshold above 20% misses the
original defect; any threshold below it fails on a legitimate category. The
discriminator is whether a signal is platform-meaningful, which no threshold
encodes. Weakness disclosed rather than covered by a gate that does not bind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4641): add the changeset fragment for the confinement-check change

changeset-lint failed on PR #4643: the PR touches user-facing paths and carried
no fragment. The earlier no-changeset call matched #4604's CI-only precedent and
was correct then; it was not revisited once the PR grew a src/ change, which is
my miss.

The fragment describes the real user-visible improvement: the external-descriptor
write-confinement check's Windows semantics are now verified deterministically
rather than only when the suite happened to run on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the tier count in TESTING-SUITES.md

Said the tier narrowed from 546 to 254. The final committed list is 265 of 931
eligible (58.8% -> 28.5%) after the shell-interpreter-spawn replacement restored
9 files and ALWAYS_REAL_OS pinned 1. Same error class the spec review caught in
the ADR, in a live reference page rather than a dated record, so it states the
current truth rather than carrying an amendment note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the measured aggregate from real CI job lists

Epic #4589's closeout asserted its reduction from a static count; #4641's
acceptance criterion asks for a figure read off a real run. Recorded here:
test.yml job count 21 -> 15 and non-Linux 7 -> 4, comparing PR #4640's run
against this PR's own. Against the true pre-epic baseline of 9, that is 9 -> 4.

Also states the caveat that a PR's total CHECK count is not a clean before/after
comparison, since many gates are path-scoped and this change touches a broader
path set -- the like-for-like figure is the test.yml job count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): compare job totals the same way on both sides

The measured-aggregate table put #4640's COMPLETED run total (21) against this
run's count at matrix-expansion time (15). Those are not the same measurement:
the completed total includes the post-test Coverage gate and baseline-publisher
jobs. Counted identically, it is 21 -> 17. The load-bearing figure, non-Linux
jobs 7 -> 4, was correct and is unchanged.

Called out in the table rather than silently corrected -- comparing two
differently-derived numbers is exactly the error class this ADR is about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record measured conformance wall-clock and date the stale counterfactual

Adds the per-job durations from both runs. The honest read is that this is a
correctness win more than a speed one: file count fell 52% but wall-clock only
9-29%, because what was removed were the cheap static tests and what remains is
concentrated in expensive spawn-heavy work. Stated explicitly so nobody expects
a future narrowing to buy time proportional to file count.

The load-bearing figure is windows shard 3/3: 40m24s against a 45-minute cap on
the 547-file tier -- 90% of the cliff #869 and #3057 were both filed about --
pulled back to 31m27s. macOS moved the wrong way (17m48s -> 21m02s) while its
tier was UNCHANGED at 197 files, which fixes that as runner variance and is
noted as a caution against reading a single duration as signal.

Also dates the symlink-keyword counterfactual, which cited a 254-file tier from
before the replacement category and allowlist took it to its final 265.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): re-measure against the rebased tree and disclose the allowlist's zero

next gained #4644 mid-flight, so every absolute count shifted. Re-measured on
the tree this actually ships against (932 eligible): 548 -> 257 by detector
removal, 257 -> 266 once shell-interpreter-spawn restores 9. Net 282 removed,
9 restored. macOS 198, unchanged by this PR.

The percentages did not move across three rebases (58.8% -> 28.5%), which is
the whole argument for expressing the ceilings as ratios rather than counts --
noted in the ADR since it is now evidence rather than assertion.

Also discloses that ALWAYS_REAL_OS now contributes ZERO files: this PR's own
win32 test cases introduced the literal win32 into the pinned file, so it
classifies in on content via win32-darwin-literal. The entry stays and the
reason is written down, because the file's real-OS need is a property of the
code under test (isPathConfined reads the ambient path module), not of the
test's text -- the text that currently saves it is incidental and could be
refactored away silently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 17:00:11 -04:00
Tom Boucher
5e0a7b1b56 fix(#4433,#4569,#4126): consolidate the phase-identity seam at name-validity, allocation, and branch-slug (#4640)
* fix(#4433): apply the name-validity guard symmetrically to every milestone-name capture

extractMilestoneHeadingName already refused a punctuation-only captured name
(#4134), but its two sibling capture sites in getMilestoneInfo — the
STATE.md-anchored 🚧-bullet match and the no-STATE.md in-progress 🚧-bullet
fallback — skipped straight to a bare truthiness check, so a malformed bullet
whose only content past the version was punctuation passed through as a real
milestone name.

Extracts the existing inline /[\p{L}\p{N}]/u check into a single shared
hasNameableContent predicate and applies it at all three capture sites, so
the guard is one owner rather than a copy that happened to land at only one
of them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4433): pin the name-validity guard at all three milestone-name capture sites

Failing-first coverage for the hasNameableContent extraction: a
punctuation-only 🚧-bullet name must not surface as a real milestone name,
either on the STATE.md-anchored path or the no-STATE.md in-progress
fallback, while a real name (including a digits-only one) still resolves
COMPLETE exactly as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4569): consolidate decimal-phase-number allocation into one function

cmdPhaseInsert allocated its next decimal sub-phase number by scanning only
on-disk phases/ directories and ### Phase N.M: headings, never the roadmap
summary checklist — so a decimal that existed only as a checklist bullet
(no heading yet, no on-disk directory yet) was invisible, and phase insert
could silently reallocate an already-used number. It also always nested one
level deeper under afterPhase, with no way to request a sibling.

cmdPhaseNextDecimal had its own separate, near-identical two-source scan
(missing the checklist source too) — the exact "duplicate implementations
kept in sync instead of deleted" pattern this issue exists to close.

Extracts scanExistingDecimalPhaseNumbers (directories + headings + checklist
bullets, in one place) and migrates both cmdPhaseInsert and
cmdPhaseNextDecimal onto it — deleting cmdPhaseNextDecimal's own copy rather
than patching it in parallel. Adds an allocation: 'nested' | 'sibling'
argument to cmdPhaseInsert (default 'nested', matching every existing
caller's behavior); a top-level phase with no existing decimal segment falls
back to nested since there is no sibling level to join. No CLI flag wires
'sibling' yet — that is a separate, disclosed follow-up.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4569): pin decimal-allocation coverage across phase insert and next-decimal

Failing-first coverage for scanExistingDecimalPhaseNumbers: a checklist-only
decimal must not be reallocated by phase insert; a decimal present in
heading, checklist, and on-disk directory simultaneously must count once;
an unrelated phase family's checklist bullet must not cross-pollute; and
phase next-decimal (migrated onto the same shared helper) must see a
checklist-only decimal too, closing the same gap in a second command.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): extend the phase-id drift guard for name-validity and shell arithmetic

The epic's ratchet requirement: lint-phase-id-drift.cjs must cover the two
new predicates this PR introduces, and must also scan shell inside
gsd-core/workflows/**/*.md and gsd-core/references/**/*.md for
integer-coercing phase-number arithmetic ($((10#...)) and friends), which
neither the canonical TypeScript module nor a source-only lint can reach.

Adds findNameValidityDrift (bans re-deriving /[\p{L}\p{N}]/u outside
hasNameableContent's owner file) and findShellPhaseArithDrift +
scanMarkdownShellArith (bans $((10#...)) in workflow/reference markdown,
sanctioned via <!-- phase-id-owner: --> on the preceding line). scanRepo
keeps its existing, narrower contract (src/**/*.cts only) so the
already-passing "the live repo is clean" test is untouched; a new scanAll
merges both for the CLI's full report.

Running the guard directly against this tree correctly reports the 7
pre-existing #4619 shell sites (workflows/execute-phase.md x4,
workflows/execute-phase/steps/completion-reconciliation.md x2,
references/tdd.md x1) as violations — demonstrating the ratchet works, not
fixing them. #4619 is a live regression tracked and fixed separately; this
PR does not touch those markdown files. A characterization test pins the
current count of 7 so a future change to that number is investigated rather
than silently absorbed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4569): wire --sibling through phase insert's CLI so the argument is reachable

cmdPhaseInsert's allocation parameter had no CLI path to 'sibling' — shipped,
untested, unreachable code (code-review finding: a guaranteed surviving
mutant). Adds --sibling to phase insert's argument parsing, threads it
through, and documents the flag in docs/CLI-TOOLS.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4569): exercise --sibling end-to-end through the real CLI

Confirms --sibling joins afterPhase's parent decimal level rather than
nesting, and falls back to nested when afterPhase has no existing decimal
segment (no sibling level to join).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4634): demonstrate the two new drift detectors end-to-end via a planted violation

The epic asks for the guard to be "demonstrated by watching it go red" on a
reintroduced copy. The two new detectors (name-validity, shell-arith) had
only unit-level fixture tests; mirrors the existing bracket-rule's
planted-violation-in-a-temp-tree test for both, proving they're actually
wired into scanRepo/scanMarkdownShellArith end-to-end, not just correct in
isolation.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): consolidate the drift guard's own owner-sanction-check logic

Standards review flagged the "walk to nearest preceding non-blank line,
check for a phase-id-owner comment" logic as duplicated across all four
detector functions in a PR whose whole point is eliminating exactly that
pattern. Extracts isSanctionedByPrecedingComment, shared by all four;
behavior-preserving (verified: identical output before/after, same 7 known
#4619 violations, zero token/bracket/name-validity).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): add Fixed changeset for the name-validity guard and allocation consolidation

pr:0 placeholder — backfilled once the real PR number exists.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4126): consolidate branch-name slug substitution into one shared renderer

cmdCommit (commands.cts) and cmdInitExecutePhase (init.cts) each
independently implemented branch-name template substitution, and both
substituted the literal string 'phase' when phase_slug was empty or
undeliverable — producing a non-identifying branch name (gsd/phase-08-phase)
that contradicted the honestly-reported phase_slug: null in the same
payload. Same structural defect as the other three gaps in this epic: two
consumers reimplementing one concept independently instead of sharing an
owner.

Adds renderPhaseBranchName (src/phase-id.cts) as the sole owner: a real slug
substitutes normally; an empty/undeliverable one drops the {slug} token plus
one adjacent separator (collapsing/trimming the result) rather than
substituting a placeholder word, for the shipped default template and any
user-configured shape alike. Both call sites now delegate to it; the old
inline duplicates are deleted, not kept in sync. {project} substitution
stays a separate step in init.cts, unchanged, since it is a config-level
field with its own fallback contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4126): pin renderPhaseBranchName and both migrated call sites

Property-based coverage for the shared renderer's degrade-path invariant
(output, when non-null, never contains {slug} and never starts/ends with a
separator), plus example coverage for real-slug substitution, empty/null/
non-string slug, token position at either edge, a doubled-separator
template, and the only-{slug} -> null case. One regression test each in
commands.test.cjs and init.test.cjs confirms a phase with no derivable slug
no longer produces a branch name ending in the literal '-phase'.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: route scanExistingDecimalPhaseNumbers through the canonical enumeration owner

Caught by an actual gsd-test run, not a hypothesis: the new decimal-scan
helper (fix(#4569)) enumerated phases/ directories via a raw
fs.readdirSync, which the pre-existing phase-enumeration drift guard
(#3185/#3882) correctly flags as an unsanctioned re-derivation outside its
canonical owner (listAllPhaseDirs / isSentinelPhaseId). Ironic given this
epic's own thesis, and exactly why the guard exists: consolidating one seam
can reintroduce drift in an adjacent one if the new code doesn't route
through what's already there. Migrates the enumeration to listAllPhaseDirs;
identical decimal-detection output for every existing case.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): extend the drift guard for branch-slug fallback; fix a real regex bug

Adds the fourth detector the epic's ratchet section names ("both
branch-name sites"): bans a `.replace('{slug}', ... || 'phase')` call
outright, sanctioned via renderPhaseBranchName or a dedicated comment.
Wired into scanRepo (no per-file exemption — this is a banned anti-pattern
everywhere, not a grammar with one legitimate owner). Now that #4126's fix
(prior commit) has landed, scanRepo reports zero violations across all four
.cts-scanning rules, restoring the simple "the live repo is clean" assertion
instead of a pinned-known-count characterization.

Also fixes a real bug an actual gsd-test run caught: findNameValidityDrift's
regex didn't tolerate the doubled-backslash template-string form its own
test claimed to cover (0 !== 1) — widened to \{1,2} matching
TOKEN_DRIFT_RE's existing tolerance for the same two forms.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4126): document the {slug} degrade behavior; update changeset for the full seam

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: detectPhaseNumberFromFiles wrongly rejected bare, slug-less phase directories

Caught by an actual gsd-test run on the #4126 regression test, not a
hypothesis: a bare phase directory with no slug remainder (e.g.
.planning/phases/01/) has extractPhaseToken correctly return "01" — which is
simply identical to the directory name in that case, not its no-match
fallback. A stale `token !== phaseDir` check treated that equality as "no
numeric token found" and rejected it regardless, leaving phaseNum null and
silently skipping cmdCommit's phase-branching block entirely (the commit
proceeded on whatever branch was already checked out instead of the
phase branch).

phaseTokenShape.test(normalized) already excludes every genuine non-phase
case on its own: extractPhaseToken's real no-match fallback only fires for a
dirName that doesn't start with a digit or short letter+digit prefix, and
normalizePhaseName's leading-\d+ requirement rejects those regardless. The
equality check was redundant for real rejections and actively wrong for
bare-numeric directories.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: backfill changeset PR number to 4640

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 12:42:04 -04:00
0xdhx
4cc2a466b5 fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec (#4253)
* fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec

`cmdCommit`'s `--files` list can stage an addition but never a deletion:
the #2014 guard skips a missing explicit entry because the filesystem
cannot tell "moved away" from "not written yet". A caller that moves a
file therefore had two forms, both wrong — a directory entry records the
move but also commits every unrelated file in that directory (a
concurrent session's in-flight todo, in the unattended execute-phase
sweep), and a file entry leaves the old path's deletion dangling with the
todo tracked at both paths.

`--files-removed <paths>` is the caller-declared delete intent. Each entry
names a file, or a directory whose tracked-but-absent files are the
removals; those paths are staged with `git rm --cached` and join the
commit pathspec. `--files` keeps its skip-if-missing contract untouched.
A file entry still present on disk fails the commit closed with the
existing staging-failure rollback; a never-tracked path is a no-op.
`--files-removed` alone is a declared scope, not the unscoped .planning/
sweep.

The dispatcher previously folded every non-flag token after `--files`
into that list, so a second list flag could not exist; each list now
runs from its flag to the next `--` token.

The execute-phase todo sweep names the moved todos on both sides from
CLOSED[@], and cleanup's archive commit moves .planning/phases/ and
.planning/quick/ under --files-removed.

Fixes #4208

Emitted-Drift-Ack-Growth: cleanup.md — the archive commit moves phases/ and quick/ under --files-removed; the growth is one paragraph stating why those two directories must not be --files entries

* chore(#4208): set changeset fragment pr to 4253

* fix(#4208): fit execute-phase.md under the ADR-857 ceiling and re-point the #2415 guard

Three CI failures, all consequences of this PR's own change.

1. gsd-core/workflows/execute-phase.md was 93,577 bytes against the
   ADR-857 Phase 6 margin gate's <= 93,400 (hard ceiling 93,600). The
   three-line rationale comment plus the four-line array-building block
   added 318 bytes to a file that had only 141 of headroom on next.

   Move the rationale to docs/CLI-TOOLS.md -- which this PR already
   extends with the --files-removed contract, and which is where the
   ADR-857 gate wants call-site detail to live rather than in the host
   workflow -- and fold the array build onto one line. 93,577 -> 93,372.

2/3. tests/close-phase-todos-stage-deletion.test.cjs pinned the #2415
   guarantee to its old MECHANISM: it regex-matched the literal
   .planning/todos/{completed,pending}/ directory pathspecs in the
   commit --files list. This PR deliberately replaced those with named
   files (a directory entry also committed an unrelated todo a
   concurrent session dropped in mid-close), so the guard failed on a
   change it should have accepted.

   Re-point it at the new mechanism without weakening it: assert the
   ADDED array reaches --files, the REMOVED array reaches
   --files-removed, STATE.md is still committed, and -- newly -- that
   the two arrays are built from $COMPLETED_DIR and $PENDING_DIR
   respectively. Verified by negative control: deleting
   --files-removed "${REMOVED[@]}" from the workflow still fails the
   test, so the #2415 regression remains caught.

Note for the merge queue: #4233 also grows execute-phase.md (+114). The
two are additive -- different regions, no textual conflict -- so with
both landed the file reaches ~93,486, over the 93,400 margin though
under the 93,600 hard ceiling. Whichever merges second will need to
reclaim ~86 bytes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0183892Y3fxxirte4WNmBKbv

* fix(#4208): reclaim execute-phase.md bytes so the PR is net-neutral under the ADR-857 margin

Rebasing onto next surfaced the byte-gate collision flagged earlier on
this PR: #4284 grew execute-phase.md by 95 bytes (93,259 -> 93,354),
so this PR's +113 landed at 93,467 against the <= 93,400 margin in
tests/claude-orchestration.test.cjs.

Compact the close_phase_todos step this PR already edits -- drop the
PHASE_NUM indirection, fold the normaliser and the match guard, print
the closed list with one printf, shorten the step's prose -- without
touching the mechanism the #2415 guard pins (ADDED/REMOVED arrays, the
plain mv). 93,467 -> 93,349: 5 bytes under the base, so the PR no
longer spends any of next's 46 bytes of headroom.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): classify absent index entries before staging a removal; restore removed entries exactly on rollback

Review of #4253 found three Majors with one root cause: the removal
side judged presence by fs.lstatSync alone, where the addition side
already reads `git ls-files -v` state. Absence from the worktree is not
removal:

- a submodule gitlink (mode 160000) whose directory was deleted by hand
  lists like a file and was `rm --cached` with no .gitmodules cleanup;
- a skip-worktree path is never materialised by a cone-mode sparse
  checkout, so a directory entry over a sparse-excluded tree dropped
  that whole tree from the index;
- an assume-unchanged path's worktree state is not something git
  itself consults;
- an intent-to-add entry (`git add -N`) renders as a plain cached entry
  on the empty blob, yet nothing tracked exists to remove and no
  rollback can restore the flag.

The index listing now carries each entry's `ls-files -v -s` tag, mode
and stage. Only a plain cached (H), stage-0, non-gitlink entry is a
removal candidate; every other state is left alone under a directory
entry (exactly like a present file) and fails closed when named
directly, with the state in the error. "Named directly" is decided on
RESOLVED paths, not strings -- realpath of the longest existing prefix
with the absent tail re-appended: an absolute path, `./x`, `--cwd`, or a
symlinked spelling of the tree (macOS `/var` ->
`/private/var`, where `process.cwd()` is the real path and the caller's
absolute path is not -- CI on this round's first push) all resolve to the
same entry, where a string compare against git's cwd-relative output
silently took the directory polarity (pre-push review, driven; the
symlink case is driven with an aliased fixture directory). The enumeration's domain is what
`ls-files -v -s` can emit for an index entry, stated at the classifier.

The third Major -- on an unborn HEAD a successful `rm --cached` was
never rolled back when a later entry failed -- is fixed differently
from the review's suggestion. Pushing the path into stagedPaths would
put it on the commit pathspec, which a root commit refuses ("pathspec
did not match", driven), and `git reset -- <path>` cannot restore an
entry with no HEAD anyway. Instead every index entry this call removes
is recorded (mode, blob) before the `rm` and put back with
`update-index --cacheinfo` on rollback. That also restores a
caller-pre-staged blob at a removed path exactly, where a reset would
have silently replaced it with HEAD's version. The rollback is
best-effort, as the addition-side reset already was, and the docs say
so.

Eight tests: gitlink under a directory entry, named directly, and named
by absolute path; skip-worktree both forms; intent-to-add both forms;
assume-unchanged named; unborn-HEAD partial failure restores the
removal; pre-staged blob survives the rollback.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): drop the empty fenced block left dangling in cleanup.md's commit step

Review nit on #4253: inserting the --files-removed rationale between the
original bash block and its closing fence left an empty ```bash``` pair
before </step>. Harmless at runtime, a formatting artifact of this PR's
own diff; removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): a boolean flag inside a commit path list no longer ends the list

Review minor on #4253: collectList stopped at the next `--` token, so a
positional wedged between a boolean flag and the next list flag
(`--files a --amend b --files-removed c`) was claimed by neither list
and silently dropped -- a regression in shape against the old
slice-to-end parse, which filtered `--` tokens and kept `b`. No current
call site interleaves that way, but the gap was real.

A list now runs to the next LIST flag (`--files` / `--files-removed`)
and skips boolean flags on the way, and a REPEATED list flag merges
its runs (`--files a --files b` -> [a, b]) as the slice-to-end parse
did -- a first cut stopped at the repeat and dropped `b`, the same
silent-drop shape one level over (pre-post comment audit). The only
change #4208 makes to parsing is that a second list flag can exist.
Tests: STATE.md wedged between --no-verify and --files-removed lands
in the commit; both runs of a repeated --files reach it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* test(#4208): drive the reappearance window with a post-index-change hook

Review nit on #4253: the defensive re-check for a file recreated between
the absence test and `git rm --cached` -- the concurrent-session race
this PR's own changeset names -- had no test. git fires
post-index-change the moment `rm --cached` writes the index, so a hook
that copies the file back exactly then exercises the window
deterministically. The call reports staging_failed / "reappeared on
disk", commits nothing, and the rollback restores the removed entry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): restore a staged removal when the call records nothing

A `git rm --cached` that succeeds mutates the index whether or not a commit
follows. Only the staging-failure rollback put those entries back, so a call
that reached `nothing_to_commit` reported no state change while the removal sat
staged -- riding along on the caller's next commit.

The review named the unborn-HEAD, removal-only shape. Keying on `headExists`
would have fixed half of it: the guard also fires with a real HEAD when the
removed path is index-only (added, never committed), because `diff HEAD` reads
clean with the path absent on both sides. Both shapes now restore, at both
`nothing_to_commit` exits. The failure exits are deliberately left alone --
they report a failure rather than no-change, and the addition side leaves its
own staged paths there too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* refactor(#4208): lift declared-removal staging out of the cmdCommit hotspot

`cmdCommit` was a critical-risk hotspot before this flag existed, and #4208 had
inlined another ~270 lines into it. `stageDeclaredRemovals(cwd, removedDeclared)`
now owns the index-state classification, path canonicalisation and entry
recording, returning the pathspec entries and the recorded removals its caller
merges.

Pure motion: no branch, message or probe changed. Only the two accumulators
became local names, and `restoreRemovedEntries` stays with the caller because
the exits that restore are the caller's. cmdCommit 888 -> 625 lines here; the
extracted helper is 277.

(Figures corrected after publication: an earlier version of this message said
854 -> 591 and claimed the result was below cmdCommit's pre-#4208 shape. Both
were wrong -- the count came from a faulty brace scanner, and `next`'s cmdCommit
is 581, so this is above it, not below.)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): property-test the two-list commit parser

RULESET.TESTS.property-based-testing asks a parser for at least one property
test asserting a domain invariant; `collectList` had only hand-picked examples,
one per shape a review round had already broken.

Hoisted it to module scope as `collectListFlagValues` and exported it in the
file's existing exported-for-tests convention -- a parser reachable only by
spawning the CLI can be tested one example at a time and no faster.

Three properties over generated argv: every positional lands in exactly the run
open at it whatever the flag order or count; no positional after the first list
flag is dropped or double-claimed; and with `--files-removed` absent the parse
equals the pre-#4208 slice-to-end parse. Controlled against two mutants -- a run
ending at any `--` token, and a repeated list flag that does not merge -- each
of which the properties catch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): pin cleanup.md's archive commit to --files-removed

execute-phase.md's rewrite is pinned by the #2415 guard in this file;
cleanup.md's equivalent was not, so reverting its routing would have been
caught by nothing -- the mechanism's unit tests never read this file and pass
either way.

Asserts the two archived directories are under --files-removed and NOT under
--files (where a directory entry sweeps in a concurrent session's in-flight
writes), and that the destinations and STATE.md stay on the additive half.
Controlled by restoring the pre-#4208 sweep, which fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): pin that a symlink to a directory is one tracked path

Review of #4253 read the `lstatSync(...).isDirectory()` test as a
symlink-following defect. Driving it says the opposite: git tracks the link as
a single blob (mode 120000) and does not traverse it, so the tracked paths
"under" it live at the real directory and were never named by the caller.
Following the link would stage those -- the directory sweep #4208 exists to
remove -- while the named entry still sat present on disk.

Pinned rather than changed, with the premise driven in the test body. Swapping
`lstatSync` for `statSync` -- the prescription as written -- fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4208): refresh the compact-content baseline for this PR's execute-phase edit

The base range added `tests/benchmark-compact-content.test.cjs` and a committed
token baseline over the compacted workflows. This PR edits
`gsd-core/workflows/execute-phase.md`, so the baseline drifts by +12 tokens on
that entry and on the aggregate.

Refreshed with `node scripts/benchmark-compact-content.cjs --write`; the diff is
those two entries and nothing else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): report a removal the call could not put back

Round review of this round found the restore itself unchecked: the helper
ignored `update-index`'s exit code, so a FAILED restore still reported
`nothing_to_commit` -- the same false "no state changed" the restore exists to
prevent, surviving one level down on the restore-failure path.

It now returns a boolean. The two no-change exits report `staging_failed`
naming the paths left staged; the staging-failure rollback still ignores it,
deliberately, because it is already reporting a failure and an unwritable index
is usually the failure being reported.

Driven with a post-index-change hook that makes the git dir unwritable the
moment `rm --cached` lands, so the restore cannot take its lock. Reverting both
guards fails the test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): disclose a removal the rollback could not restore

Round review refuted the reasoning behind leaving the rollback path's restore
unchecked. The claim was that this exit is already reporting a failure, so the
restore's result adds nothing. The counterexample is the ordinary case: the
reported failure is usually a DIFFERENT cause -- a contradictory declaration, a
reappeared path -- so a caller reading `failures` sees only that cause and
learns nothing about the removal still sitting in its index.

The rollback now appends a disclosure entry per un-restored removal, naming the
path. The reason and `file` still report the failure that caused the rollback;
the disclosure is additive.

Also moves the restore-failure test's chmod into a `finally`: `t.after` runs
AFTER the parent `afterEach`, so a throw before it left the fixture undeletable.

Both driven; reverting the disclosure fails the new test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): decide index state by observation, never by an exit code

The restore added two commits earlier keyed both its record decision and its
success verdict on git's exit code. An exit code answers "did the command
succeed", never "did the index change" -- execGit collapses a spawn timeout to
a non-zero exit, and a killed git can already have written the index. Round
review drove four failures from that one assumption, in both directions:

  - a failed `rm` still contributed an entry, so the rollback disclosed a
    removal that was never staged (stale index.lock);
  - a timed-out `rm` whose write DID land contributed none, so a real mutation
    was neither restored nor disclosed;
  - a timed-out `update-index` whose write landed reported failure, publishing
    a "could NOT be restored" disclosure that was false;
  - and the read-back that replaced it omitted `-z`, so core.quotePath rendered
    `café.md` as `"caf\303\251.md"` and an exactly-restored entry read as not
    restored -- the same quoting defect this PR already fixed for `preStaged`.

Everything now observes the index. A failed `rm` re-reads `ls-files -z` for the
path: gone means this call owns the removal and records it; still there means
nothing was staged; a probe that cannot answer becomes its own failure entry
rather than an assumption. The restore verifies the same way, comparing the
WHOLE entry (mode, blob, stage), because `--cacheinfo` restores all three and a
path-only test accepts an entry that came back as something else.

The verdict is three-valued -- `restored` / `not-restored` / `unverified` --
and the unverified wording says the restore could not be VERIFIED rather than
that it failed. The rm's own failure is pushed ahead of any probe diagnostic so
a timed-out removal keeps `timed_out: true` and its own message as the reported
cause.

Five regression cases, each negative-controlled against the shape it pins.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): treat a declared removal path as a path, not a pathspec

An index path handed back to git is parsed as a PATHSPEC, and the removal side
handed several back. Three driven harms, all of them the sweep-in this flag
exists to remove, arriving through the operand rather than through a directory
entry:

  - a tracked file literally named `.planning/*.md` made `rm --cached` GLOB: it
    removed `peer.md` and `stays.md` too, only the declared entry was recorded,
    so the rollback restored one of three and the other two rode out as staged
    deletions the result disclosed nowhere;
  - the same name reached `git commit -- <paths>`, which globbed and committed
    an undeclared `M peer.md` alongside the declared removal;
  - and the intent-to-add probe (`diff --cached` over the path) matched a
    STAGED PEER instead of itself, so an `add -N` entry was misclassified as
    ordinary content, removed, and restored by `--cacheinfo` -- which cannot
    restore the intent flag. It came back as a real staged addition.

Every operand on this path is now `:(literal)`: the `rm`, both index probes,
the intent-to-add probe, the restore read-back, the entry-level `ls-files` /
`ls-tree`, and -- for the REMOVAL-derived entries only -- the downstream
`ls-files` / dry-run / `diff HEAD` / `commit` pathspec. `--files` entries keep
whatever pathspec behaviour they have today; that is not this change's to
alter. `:(literal)` still resolves a directory to its descendants (driven), so
the directory form is unchanged.

Closes what an earlier cut of this commit declared as a residual: a filename
beginning with `:` is now removable end to end, because the commit pathspec no
longer reinterprets it.

Also fixes a MINOR from the same review: cleanup.md's contract test checked the
destinations' position relative to `--files-removed` but never that `--files`
was present at all, so deleting the flag still passed.

Un-literalising the seven sites fails three of the new tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): scope the rollback to the caller's own name space

Round review drove a rollback that destroyed the caller's own staged work. Two
causes, one of them pre-existing:

  - `git diff --cached` prints REPO-relative paths whatever the cwd, while
    `stagedPaths` holds the caller's cwd-relative names. In a project nested
    inside its repo (`<repo>/sub/.planning/...`) the two name spaces never
    intersect, so `preStaged` matched NOTHING, every path landed in `toUnstage`,
    and the reset unstaged a caller-staged deletion and modification that this
    call had never touched. `--relative` makes the two sets comparable, and is a
    no-op when the project IS the repo root. This governs the `--files` side too
    and predates this flag.
  - the rollback's `reset` was the last place a removal-derived name reached git
    as a bare pathspec; it takes `asPathspec` like every other site.

Driven on a nested fixture; dropping `--relative` fails the new test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): gate six fixtures that Windows cannot construct

CI's `test (windows-latest, 24, shard 2/3)` went red on this round. Two
primitives the new fixtures rely on do not exist on Windows, both driven on a
real Windows host rather than inferred:

  - a filename containing `*` or `:` cannot be created at all (`IOException` /
    `FileNotFoundException`), which is four of the pathspec fixtures;
  - `chmod` cannot make a directory unwritable — a write into a ReadOnly
    directory succeeds — so the two restore-failure fixtures cannot drive the
    failure they exist to drive.

Each is skipped on win32 with its measured reason, in the repo's existing
`{ skip: process.platform === 'win32' ? '<reason>' : false }` form. The
behaviours they pin are platform-independent; only the fixtures are not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): build git's index-syntax path with forward slashes

The remaining Windows red was mine, not the platform's: `git rev-parse :<path>`
takes a forward-slash path, and `path.join` yields backslashes there, so git
rejected it as an ambiguous argument. The hook in the same test already used
the slash form.

Not gated — the behaviour it pins is portable; only the argument was not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4208): refresh the compact-content baseline against the rebased base

`next` moved the `new-project` split and the aggregate under this PR's
execute-phase entry; regenerated with `scripts/benchmark-compact-content.cjs
--write` so the only leaves differing from the base's copy are the
execute-phase split and the aggregate it feeds.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh

* chore(#4208): regenerate the macOS conformance tier for this PR's fixtures

`next` gained the macOS-specific conformance tier (#4593) after this branch
was cut. Its classifier (`scripts/gen-platform-conformance-tier.cjs --target
macos`) now selects `tests/commit-files-deletion.test.cjs` on the
`chmod-mode-bit` and `symlink-keyword` signals the PR's fixtures carry (the
chmod-driven failed-restore cases and the symlink-to-directory case).
Regenerated with `--target macos --write`; the platform tier was already in
sync. The file was modified, not added, which is why the added-files check
did not surface it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-11 12:26:55 -04:00
Tom Boucher
1e47560e34 feat(#4593): add a macOS-specific conformance tier, final phase of epic #4589 (#4607)
test-conformance's macos-latest leg (Phase 2, #4591) has been running the
same 546-file, Windows-oriented conformance-tier list as windows-latest --
built from signals like windows-shell-token/windows-env-var that have
nothing to do with macOS. Issue #4593 asked for macOS coverage sized to
its own evidence-backed surface (zsh dispatch, case-sensitivity, darwin-
specific behavior) instead.

Issue #4593 was filed before Phase 5 (#4603) existed and referenced
updating test-full's macOS legs -- that job is gone. Corrected the issue's
body before any code was touched: the "shrink from full replay" half of
the original ask was already done by Phase 5; what remained was narrowing
the still-Windows-oriented tier macOS was inheriting.

Two design assumptions were measured and rejected before accepting a
design (documented in docs/adr/4593-macos-conformance-tier-architecture.md):
- Reusing the general tier's signals minus its 3 Windows-specific
  categories barely narrows anything (546 -> 424, 78% retained) -- most
  files match multiple signals and only need one to survive exclusion.
- A standalone CRLF/autocrlf signal, despite the issue naming
  "CRLF-checkout behavior": even narrowed to /\bCRLF\b|autocrlf/i it hit
  143/930 files. Root cause: CRLF is primarily a Windows checkout concern
  in this codebase (ADR-1703 files it under DEFECT.WINDOWS-TEST-
  PORTABILITY), so the signal was really re-selecting Windows-relevant
  files already covered by the general tier, not narrowing macOS
  specifically.

Built 5 new, genuinely macOS-specific signals instead: darwin-literal
(darwin alone, not the general tier's win32-OR-darwin), zsh-dispatch,
case-sensitivity, plus chmod-mode-bit and symlink-keyword reused verbatim
from the general tier (genuinely Unix-relevant, not Windows-motivated).
Measured against the real tree: 196 of 930 eligible unit-suite files
(21%), versus the general tier's 546 (59%) -- a real, evidence-backed
narrowing.

scripts/gen-platform-conformance-tier.cjs gains classifyMacosContent/
classifyMacosTree/renderMacosGeneratedFile and a --target windows
(default, unchanged)/--target macos CLI flag, so the same generator
produces two independent, gated outputs rather than needing a second
script. New committed output: scripts/lib/macos-conformance-tier.
generated.cjs. .github/workflows/test.yml's test-conformance job: only
the macos-latest leg's file-list source changes; windows-latest is
byte-for-byte untouched. New shipped-file ripples handled proactively
(19 install-tree fixtures regenerated, bin/install.js registered).

An isolated code-review pass found one real defect: the ADR's per-
category count table had drifted by 1 (zsh-dispatch, case-sensitivity)
because the new test file's own fixture strings joined the tree it
classifies after the table was authored -- fixed, with the union total
(196, what CI actually gates on) confirmed unaffected. An isolated
security-review pass found no qualifying findings.

The ADR also records an explicit requirement for any future widening
proposal: check whether the motivating regression is already covered by
Phase 1's no-rendered-text-length-assert lint rule (#4590) before
re-proposing full macOS/Linux parity, since that is exactly what #4421's
root cause was (a rendered-text-length assertion, not a real behavioral
divergence).

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 14:37:15 -04:00
Tom Boucher
181c4c8659 chore(#4603): retire the test-full CI job (#4604)
* chore(#4603): retire the test-full CI job

Phase 2 (#4591) added test-conformance but left test-full (the pre-existing
full-suite Windows/macOS replay) running unchanged, gated on the same
full_matrix flag, downgraded only from a hard gate to a non-blocking
::warning:: -- framed as "a non-gating safety net for one release cycle."
No phase or issue ever retired it. Result: every full_matrix=true PR ran
10 OS-specific jobs (test-full's 6 + test-conformance's 4, purely
additive) instead of the original 6 -- the epic's own goal (reduce
runner-minutes) was measurably regressing, not improving, for the
majority of PRs.

This phase was missing from the original 4-phase epic decomposition; the
epic (#4589) has been amended to add it as Phase 5 (see its comment
thread), and this issue was filed as the tracked sub-issue.

Deletes the test-full job from .github/workflows/test.yml entirely, along
with every reference to it: required-tests' needs/FULL_TEST_RESULT
warning branch, ci-timeout-report.cjs's JOB_RULES entry,
ci-test-job-timeout-budget.test.cjs's LANE_COSTS/staticLanes/testFullRule
entries, ci-test-scope.test.cjs's test-full-specific tests (preserving
three unrelated tests that were nested in the same describe block, moved
under a renamed describe rather than deleted), and docs mentions.
test-conformance is now the sole gating signal for real-OS coverage.

Two separate defects found and fixed while auditing every test-full
reference:
- tests/ci-pr-mergeability.test.cjs's GATED['test.yml'] safety-critical
  array (jobs that must needs: the mergeability preflight) had test-full
  but was missing test-conformance entirely -- Phase 2 never added it.
  Verified the real workflow wiring was already correct (test-conformance
  does have needs: [changes, preflight]); this was a test-coverage gap,
  not a live defect. Fixed by swapping the array entry.
- docs/TESTING-SUITES.md's "## CI matrix" section was substantially stale
  independent of this phase (predating even #2952's coverage-gate split).
  Rewritten against the real, current job topology, verified directly
  against test.yml rather than trusted from memory.

An isolated code-review pass found and fixed two minor inaccuracies in the
rewritten docs table (two jobs' "Gated on" column didn't match their real
if: condition exactly). An isolated security-review pass found no
qualifying findings -- every compute-provisioning job already carries
needs: preflight directly, unaffected by this deletion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(ci): isolate 7 more heavy test files from chunk-weight packing

`next`'s own push-triggered Tests run failed: `conformance test
(windows-latest, 24, shard 2/3)` chunk 3/6 was killed after 600019ms.
Root cause: state.test.cjs (weight 21.35, measured) was packed alongside
companions by run-tests.cjs's LPT chunk packer, the same failure mode
that previously hit codex-config.test.cjs (weight 17.87) twice and got a
dedicated fix (ISOLATED_HEAVY_FILES, #4497) -- but state.test.cjs was
never added to that set.

This is a direct, unintended consequence of epic #4589 Phase 2: the new
platform-conformance-tier job packs only ~546 files per shard (vs. the
~950-file full suite the packer used to balance against), so the same
absolute-weight outlier now represents a larger share of a smaller, more
homogeneous pool -- the LPT packer has fewer light files to pad around
it with. This was a real, foreseeable side effect of shrinking the
packing pool that nobody checked for when Phase 2 shipped.

A first attempt at this fix hand-picked 4 candidates by eyeballing a
truncated weight list and missed 3 heavier ones -- caught by an isolated
code-review pass (blocker: emitted-attribution.test.cjs at 66.2% of the
Windows chunk budget, install-minimal-hooks.test.cjs at 61.1%,
install.test.cjs at 47.1%, all above codex-config.test.cjs's own
44.7% -- the ratio that already proved dangerous twice). Corrected by
systematically computing weight/budget for every unit-suite file and
isolating everything at or above that same ratio: 7 files total, plus
the pre-existing codex-config.test.cjs (8 total).

Added a durable regression test (tests/run-tests-harness.test.cjs) that
re-derives this exact computation from the live tests/test-timings.json
on every run, so a future heavy file crossing this threshold fails the
test instead of silently reintroducing this failure -- not just a
one-time manual sweep.

Verified end-to-end: simulated the real 3-way windows shard split of the
actual conformance-tier file list with the real packing functions. Max
packable-chunk weight across all 3 shards is now 27.04 / 24.10 / 23.91
(shard 2 is the exact shard that failed on next), comfortably under the
40 budget -- versus 40+ and a 600s kill before this fix.

A second isolated code-review + security-review pass on the corrected
diff found nothing further.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 12:46:04 -04:00
Tom Boucher
bcd99696d3 chore(#4591): add platform-conformance-tier classifier + gate CI on it (#4598) 2026-09-10 08:36:13 -04:00
Tom Boucher
2cefa5a5ac enhance(#4139): Phase 8 — the toggle becomes discoverable, and the ledger closes (#4587)
* enhance(#4139): Phase 8 — the toggle becomes discoverable, and the ledger closes

ADR-4139's final phase. workflow.compact_content already defaulted to false
(Phase 1's buildNewProjectConfig hardcoded default), but nothing surfaced it:
/gsd-new-project never asked, and /gsd-settings/config had no toggle path for
an already-initialized project — config-set/config-get were the only route.

new-project.md gains a fourth question in the existing Round 2 AskUserQuestion
array (grouped with the other general-workflow-behavior toggles, not the
per-agent capability questions above it) and threads compact_content into the
config-new-project CLI JSON literal. settings.md mirrors the exact pattern
every other non-capability workflow.* key already follows: read_current bullet,
question block, update_config write, the safe-merge non-capability-keys list,
save_as_defaults, and the confirm summary table — seven edits, zero new
src/*.cts code, since Phase 1's merge logic is a generic passthrough. Its
success_criteria question-count ("24 settings") is bumped to 25 to match the
now-25-entry main AskUserQuestion batch.

settings-advanced.md deliberately does NOT get a duplicate question: no other
boolean toggle in this repo is asked in both settings.md and
settings-advanced.md, and there's no reason to start with this one.

docs/CONFIGURATION.md, docs/USER-GUIDE.md, and a new docs/features/4139-compact-
content.md fragment (regenerated into docs/FEATURES.md) document the toggle.

ADR-4139 itself: Status flips Proposed -> Accepted, the acceptance-criteria
section becomes a guard ledger — a 13-row table covering all 12 of #4139's
original checkboxes plus the shipped-content guard criterion, each with real
evidence (the merged PR that satisfied it, fetched via `gh issue view
--json closedByPullRequestsReferences` rather than asserted from phase
numbers) — and both "Open questions for the implementation phases" are
resolved rather than left dangling: discuss-phase was never converted to
spine+detail shape (verified: no detail/ subdir exists) — a genuine gap, not a
reasoned decline; the disjointness check is confirmed line-based by reading
compact-content-split.cjs's normalizeNonTrivialLines directly.

Orthogonal review (isolated Standards/Spec code-review + security-review
sub-agents) found and this fixes two real defects: the changeset fragment's
body didn't match CONTRIBUTING.md's single em-dash-sentence format (was
multi-sentence prose naming implementation file paths); and settings.md's own
success_criteria still said "24 settings" after the new question pushed the
main batch to 25. Also fixed, found by the Spec pass while confirming
commands/gsd/settings.md correctly needed no sync edit: that file and its
skills/gsd-settings/SKILL.md twin both still described "Interactive 5-question
prompt (model, research, plan_check, verifier, branching)", stale since long
before this phase (the batch has had far more than 5 questions for a while) —
replaced with a description that names the current set without hardcoding a
count that will drift again.

gsd-test (real run, sha 1da78fe2) caught a third real regression the local
sweep missed: new-project.md is a registered spine+detail split for Phase 4's
token-reduction benchmark (scripts/benchmark-compact-content.cjs), and the new
question's +167 tokens drifted the committed baseline
(tests/fixtures/compact-content-benchmark-baseline.json). The benchmark itself
is designed never to fail CI on drift, but the test asserting the COMMITTED
baseline is currently non-drifted correctly caught it. Regenerated via
`node scripts/benchmark-compact-content.cjs --write`; re-verified --check now
reports "up to date" and the test file passes 27/27.

Closes #4408.
Closes #4139.

Emitted-Drift-Ack-Growth: new-project.md — new 4th Round-2 AskUserQuestion entry (Compact Content, #4139) plus the config-new-project CLI JSON field and explanatory sentence; a new opt-in toggle needs new prose.
Emitted-Drift-Ack-Growth: settings.md — new workflow.compact_content read_current bullet, question block, update_config write, safe-merge key, save_as_defaults field, and confirm summary row (the same seven-edit pattern every other non-capability workflow.* toggle already follows), plus the 24->25 success_criteria count fix found in review.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4408): backfill changeset PR number

pr:0 -> pr:4587 now that gh pr create has returned the real number.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 00:04:51 -04:00
Tom Boucher
9770258558 chore(#4590): add no-rendered-text-length-assert ESLint rule (#4595)
* test(#4590): add no-rendered-text-length-assert ESLint rule

Enforces ADR-456's typed-surface mandate for one specific bug shape: a test
assertion whose pass/fail depends on the length/substring content of a
template literal that interpolates an OS-derived path (os.tmpdir(),
os.homedir(), path.join/resolve/..., or a PATH_RETURNING_FNS resolver).
Because macOS's default tmpdir prefix is longer than Linux's, such an
assertion can pass on one runner and fail on another -- the defect class
behind #4421's incident (git show 4e75b836e9), already fixed there by
pinning to a typed field per ADR-456 Sec(c) before this rule existed to
catch a recurrence.

Two repo-wide sweeps against the real tests/ tree narrowed the rule to a
sound scope: an initial design that traced call arguments (to approximate
the historical incident's cross-file render-function shape) produced false
positives on ordinary fs.readFileSync(path.join(...)) + assert.match
patterns; a second design that matched any bare direct path-returning call
produced 45 false positives on path suffix/prefix/non-emptiness checks. The
shipped rule matches only a path-returning expression interpolated into a
template literal, directly or via one identifier hop -- disclosed in the
rule's own "Known boundaries" as not covering the literal cross-file
incident shape, which would require tracing into a callee's body.

Phase 1 of epic #4589 (CI test-matrix Linux-primary migration) -- Phase 2's
safety argument depends on this class of OS-dependent test assertion being
enforced going forward, not merely fixed once.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4590): address code-review findings on no-rendered-text-length-assert

Reletter the "Known boundaries" doc-comment list (a)-(e), fixing a gap left
by an earlier edit pass and every stale cross-reference to it. Collapse
isDirectPathTaint/isTaintedInterpolation's duplicated TemplateLiteral-walk
into one recursive relationship (isTaintedInterpolation now delegates a
nested-template-literal case back to isDirectPathTaint instead of
re-implementing the .some() traversal) -- behavior unchanged, confirmed by
re-running the repo-wide sweep (still zero false positives).

Found by the Standards-axis /code-review pass on this PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 21:55:51 -04:00
Tom Boucher
385ed619f1 fix(#4487): stamp broken-windows ledger entries with the resolved milestone (#4583)
* enhance(#4487): stamp windows-ledger entries with the resolved milestone

Broken-windows ledger entries (`.planning/WINDOWS.md`) carry `phase` as
a bare number. Phase numbers are unique only within one active phases/
directory -- `milestone complete` archives phases and frees their
numbers for reuse, so two milestones routinely produce entries sharing
the same phase value with nothing distinguishing them. Since
`/gsd-ship` blocks while any entry is open, an already-archived
milestone's open entries could silently block shipping the CURRENT
milestone, with no supported way to attribute which entry belonged to
which milestone short of manually cross-referencing MILESTONES.md
timestamps against decision IDs that happened to appear in description
prose.

Added an optional `milestone: string | null` field to WindowEntry,
stamped by `windows append` (cmdWindowsAppend, which already does file
I/O) from the workstream's resolved milestone version. Reused the
existing `readCurrentMilestoneVersion` (workstream-inventory.cts --
STATE.md `milestone:` frontmatter first, ROADMAP.md in-progress marker
as fallback) rather than writing a parallel implementation: exported it
via that module's existing `export = {...}` CJS-interop convention
(matching the `import ... = require(...)` pattern already used in
workstream.cts/init.cts). appendWindow itself stays pure -- it accepts
milestone as an optional input field and passes it through; only the
CLI-facing cmdWindowsAppend resolves it from disk.

Backward compatible by construction: validateEntryShape does NOT add
`milestone` to its required fields, so an existing ledger entry with no
milestone key at all parses without error and reads back as null --
exactly "recorded before this change," no migration needed. The
rendered markdown table is deliberately left unchanged (the issue's own
words: "the JSON is the source of truth"); adding a table column would
be a separate, larger change than adding an optional JSON field.

Two smaller gaps the issue itself flags as separable ("happy to split
them out") are explicitly NOT addressed here: no verb to amend an
entry's description, and the table/JSON drift-repair advice that can
destroy table-only edits on a parse failure.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4487): preserve absent-vs-null milestone through parse/render roundtrip

validateEntryShape stamped an explicit `milestone: null` onto every
entry lacking the key, so a pre-#4487 ledger entry gained permanent
JSON churn ("milestone": null) the first time ANY entry in the ledger
was touched -- breaking the pure parse/render roundtrip-identity
property test (render(parse(render(ledger))) must equal render(ledger))
and, in real usage, contaminating unrelated entries' diffs on every
append/waive/fixed of an old ledger.

Fixed by distinguishing "key genuinely absent" (undefined -- JSON.
stringify drops it, matching pre-#4487 behavior exactly) from "recorded
but unresolvable" (explicit null, the real signal appendWindow stamps
on brand-new entries). WindowEntry.milestone is now optional
(`milestone?: string | null`) so returning undefined type-checks.

Updated tests/broken-windows.test.cjs's roundtrip property generator to
exercise all three states (absent/null/string) -- its prior silence on
this field is exactly what let the regression through. Also corrected
the earlier backward-compatibility test's assertion: a pre-#4487 entry
reads as milestone: undefined, not null, and re-rendering it must not
introduce a milestone key at all.

Also ran npm run regen:derived: docs/features/broken-windows-ledger.md
(edited in an earlier commit) had never been propagated to its
generated docs/FEATURES.md projection, which is what was independently
failing tests/features-index-gate.test.cjs and, as a side effect of
staleness, tripping tests/fragment-single-edit-propagation.install.
test.cjs's second-source-surface check.

Manually verified via the compiled lib (500 fast-check iterations plus
direct legacy/new-entry roundtrip checks) before wiring the test file,
since this repo blocks local node --test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4487): materialize milestone via conditional spread, not undefined assignment

An object literal property set to `milestone: undefined` is still an
OWN property -- `'milestone' in entry` reads true regardless of the
assigned value, only JSON.stringify treats undefined specially. My
prior commit's own new backward-compat test asserted `'milestone' in
entry === false` for a pre-#4487 entry and failed on exactly this.
Switched to conditionally spreading the key in only when the source
object actually had it, so a genuinely absent milestone is not
materialized at all -- matching both the `in` check and JSON
serialization. Re-verified via the compiled lib (500 fast-check
roundtrip iterations, plus the specific in/undefined/JSON assertions
the failing test makes) before re-running gsd-test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4487): backfill changeset pr number to 4583

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 15:37:50 -04:00
Tom Boucher
37b965c0d1 enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code (#4553)
* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code

ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills
(src/init.cts) now selects between a canonical agents/<name>.md and a
token-minimized agents/<name>.compact.md sibling based on
workflow.compact_content, resolved in code (a real function call with a real
exit code) rather than a prose config-get gate — the same precedent stream 1's
spine/detail split established for a load-bearing seam, applied here because
this seam already runs through TypeScript instead of an eager @-include.

A missing compact sibling falls back to the canonical persona and discloses
the fallback in the served payload itself (a leading HTML-comment provenance
line), so the Done-when contract — compact when on, canonical when off, never
silent or empty — holds even for an agent nobody has compacted yet.

Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md),
each an independent, complete rewrite (not an extraction — nothing is "moved"
the way spine/detail moves text) that preserves frontmatter, every @-include,
every output-format contract, and every guardrail verbatim while cutting
restatement and verbose framing. Verified mechanically: every pair registers
(a canonical sibling exists), every compact file is strictly smaller, and the
full @-include set matches canonical's — including which references are
standalone eager-load lines versus inline prose mentions, since demoting one
to inline changes what the host actually substitutes.

Traced the install path before writing any code (.gsd/phase/.../40-design.md):
stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no
stem filtering under the default full profile, so the new .compact.md files
install for free with zero installer changes — matching issue #4407's stated
scope. A tiered agent profile that doesn't stage a compact sibling degrades
through the same fallback-with-provenance path already required for an
unauthored one, so no installer change is needed there either.

Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export
(deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are
reached by a generic code construction rather than a literal path in prose,
and checkReachability's markdown-search shape has nothing to find there).
Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new
"#4407 compact payload selection" describe block spawns gsd_run agent-skills
against real compact/canonical fixture pairs and asserts on the served
payload, which can only pass if the seam genuinely wires through.

Fixed a pre-existing test whose agents/*.md glob incidentally matched the new
.compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift
guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's
roster, both real, unrelated-to-content defects the new files' mere existence
surfaced.

Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files
under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token
benchmark baseline (npm run benchmark:compact-content-variants --write).

Closes #4407.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): apply orthogonal review findings from the compact-payload seam

Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath)
to collapse the duplicated read-and-empty-check shape between the compact
and canonical branches in cmdAgentSkills, and updated the adjacent comment
enumerating flat JSON extras to name agent_payload_variant alongside
source/degraded (added by the prior commit, comment left stale).

Security review and the Spec axis found no defects requiring a code change;
their non-blocking observations (a pre-existing, unmodified path-construction
pattern; the reasoned, documented substitution of a behavioral test for the
literal reachability check) are recorded in
.gsd/phase/enhance-4407-agent-skill-seam/60-review.json.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents

Root-caused via a real gsd-test run (93 failures) rather than guessing which
tests glob agents/ naively. Two classes of defect, both genuine:

1. Identity-roster confusion (11 files/areas): many tests and one production
   script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f
   => f.endsWith('.md'))`, which incidentally matched the new .compact.md
   variant siblings too — a compact file is a rendering of an EXISTING agent
   identity, not a new one. Fixed at the shared root
   (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests
   already consolidated on) and at each independent glob that didn't use it:
   agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix
   before checking XL/LARGE membership, so a compact file inherits its
   canonical sibling's tier instead of silently falling through to DEFAULT),
   agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual
   script, not just its test), codex-config.test.cjs (confirmed directly
   against generateCodexAgentToml that a compact role's derived sandbox_mode
   is byte-identical to its canonical sibling's before excluding it — not
   assumed), and copilot-install.test.cjs (two counts that legitimately DO
   need both files — an installed-file count and a full-conversion smoke test
   — fixed to expect 70, not stay pinned to 35).

   no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix:
   two compact files reproduce descriptive prose already allowlisted at their
   canonical file's line number; added matching entries at the compact files'
   own line numbers rather than excluding them from the scan (a genuine bare
   gsd-tools command-position bug in a compact file would be as real a defect
   as in canonical).

2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree
   run): six agents' compact renditions (gsd-debugger, gsd-executor,
   gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed
   the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction —
   confirmed structural, not a compaction-quality gap: each is dominated by
   content this phase's own rules require verbatim (the ~2.6 KB gsd_run
   bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in
   every agent that calls gsd_run, output-format contracts, guardrails).
   ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing
   spot in cmdAgentSkills's single-file synchronous read. Removed these 6
   compact files rather than ship an over-cap file or invent a multi-part
   read mechanism out of scope for this phase; recorded by name with the
   reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and
   50-test-matrix.md, per #4407's own "or explicitly recorded as not worth
   covering" allowance. Their canonical personas are served correctly today
   via the fallback-with-disclosed-provenance path this phase's own Done-when
   #2 already requires — 29 of 35 agents now have a compact variant.

Also fixes an unrelated, genuinely pre-existing defect this gsd-test run
surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own
DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in
this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md
| wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed:
two meaning-preserving trims in the <success_criteria> block (a repeated
parenthetical replaced with a same-exception reference; one redundant
qualifier dropped) bring it to 40,940 bytes.

Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant
benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6
now-orphaned roster rows removed alongside them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): make .compact.md-aware roster checks resilient to partial coverage

Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent
has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception),
breaking once 6 stems legitimately have none.

- tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was
  picking up the "### Compact Payload Variants" subsection's rows as
  phantom/uncounted entries in the primary/advanced/inventory-only
  classification this test validates — a compact row documents an existing
  agent's alternate rendition and never gets its own AGENTS.md heading, so it
  was never meant to participate in that classification. Excluded at the
  parser, not per-assertion.
- tests/copilot-install.test.cjs: the derived expected-file-list generator
  assumed every listAgentFiles() stem has a .compact.md source sibling;
  checks disk per stem now instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4407): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 12:38:59 -04:00
Tom Boucher
5823d2ec7a docs(#4484): correct native-plugin-install's install-time-config parity claim (#4579)
* docs(#4484): correct native-plugin-install's install-time-config parity claim

The doc claimed the plugin path and the npm installer "differ in
namespace and lifecycle only." False: the native plugin path
(claude plugin install, marketplace discovery, and the skills-dir
zero-friction load) materializes the repository tree directly and never
runs GSD's install engine, so install-time config baked into generated
artifact files at install time never applies there -- confirmed for
agent_tools (#4238/#4032, reproduced live in #4484: 35/35 files granted
via npm install, 0/35 via plugin install, even after
`claude plugin update`). model_overrides is the same architectural class
(install-time-only logic on the npm-install call tree, per
src/install-model-override-resolver.cts) but hedged, not claimed
confirmed, matching the issue's own hedging.

Reporter explicitly frames this as a docs-only fix: the code behavior
(zero install step on the plugin path) is presumably intentional design;
the bug is the doc's incorrect parity claim, not the missing
functionality. No code changed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4484): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 10:57:28 -04:00
Norman Yee
42c02a00c0 enhance(#3418): report real codebase drift instead of the whole repository (#4124)
* fix(#3418): write the codebase-drift baseline from code instead of agent prose

writeMappedCommit shipped correct and callerless, so no full map-codebase run ever wrote last_mapped_commit. The gate then read null and diffed HEAD against the empty tree, reporting every tracked file as newly added on every run.

Adds the stamp-codebase-map leaf verb and calls it from the map-codebase workflow and the execute-phase auto-remap path, replacing the prose instruction that asked the mapper agent to stamp its own output. An agent that concludes its work is already done skips a prose step silently, which is the failure the stamp exists to detect.

The gate now reports an absent or unresolvable baseline as skipped, with reason no-mapped-commit or unresolvable-mapped-commit, rather than as whole-repo drift. Files under .planning/ are excluded from the diff so the map's own commit does not read as seven new directories on the next run.

Emitted-Drift-Ack-Growth: map-codebase.md — adds the stamp_codebase_map step and its rationale, new workflow content this change requires

* test(#3418): cover the stamp writer and the absent-baseline gate

* docs(#3418): document how the drift baseline is written and skipped

* docs(#3418): note that a manual stamp reflows the map's whitespace

writeMappedCommit writes through platformWriteSync, which normalizes markdown whitespace on .md targets. Run in its workflow position the stamp lands on documents the mapper just wrote, so the normalization is folded into the same commit, but a hand-run stamp over an already-committed map reflows that map as a side effect. Reported on the issue thread.

* chore(#3418): add changeset fragment

Typed Changed to match the enhancement route the linked issue's label sets. The docs-required lint is satisfied by the ARCHITECTURE.md update already in this branch.

* fix(#3418): anchor the planning-artifact filter to the repo root

git diff --name-status prints repo-root-relative paths whatever the cwd, so computing the exclusion prefix against cwd yielded ".planning/" while git printed "sub/.planning/" and the filter silently matched nothing from a subdirectory.

* fix(#3418): derive the planning prefix from git, not from path arithmetic

Anchoring the exclusion prefix with path.relative() against `rev-parse --show-toplevel` broke on Windows, where os.tmpdir() hands back the 8.3 short form and git resolves the long one, so relative() produced a "../.." chain that matched nothing. `rev-parse --show-prefix` gives the cwd's root-relative prefix from the same producer as the diff paths, so the two sides cannot disagree.

* fix(#3418): take the planning lock around the codebase-map stamp

Stamping seven documents is seven frontmatter read-modify-writes, and two stampers can run at once: the full map-codebase run and the execute-phase auto-remap. Wrap the write loop in withPlanningLock, the same lock the other .planning/ writers take, so a concurrent pair cannot lose an update.

Also corrects the path-arithmetic comment, which read as if the Windows short-path hazard applied to the .planning half of the prefix. It applies to the rejected --show-toplevel alternative; both sides of the surviving relative() call are the same cwd string.

* fix(#3418): read HEAD and the map file list under the planning lock

The stamp resolved HEAD and listed the present codebase-map documents before it acquired the planning lock, so a stamper that then waited on the lock could write its now-stale sha over a newer one, or recreate a document deleted while it waited as a frontmatter-only stub. Both reads now happen inside the lock, matching the read-and-write-in-one-lock pattern config.cts and phase.cts already use. An empty --files value is refused as well instead of silently widening the stamp to all seven documents.

* fix(#3418): narrow the map stamp to the documents an update run refreshed

An "Update - only update specific documents" run reached the new stamp step with no --files narrowing, so the six documents the user did not select were stamped at HEAD and read as freshly mapped. The selection now threads through to --files, the same way the auto-remap path already does.

A bare --files (an unquoted empty shell variable drops the token) parsed to null, indistinguishable from an absent flag, so it skipped the empty-filter refusal and stamped all seven. Presence is now read off argv.

* fix(#3418): require the drift baseline to resolve to a commit, not any object

`git cat-file -t` exits 0 for a tree or blob sha and for a ref name, and `git diff <tree> HEAD` is valid, so an exit-code-only probe accepted a baseline that is not a commit and reported the resulting diff as real drift. Check the reported type instead of the exit code alone, which routes every non-commit stamp to the same `unresolvable-mapped-commit` skip.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-09 05:08:20 +00:00
Behruz Nassre Esfahani
3ad75a6d59 enhance(#4285): resolve context-monitor fire-points from .planning/config.json (#4366)
* enhance(#4285): resolve context-monitor fire-points from .planning/config.json

The monitor's WARNING (35%) and CRITICAL (25%) fire-points were module
constants, so the only way to tune them was editing gsd-context-monitor.js —
a file in the MANAGED hooks registry, whose body the next install re-stages,
silently discarding the edit. The alternative was turning the safety net off.

Both are now readable from the config block the hook already opens:
hooks.context_warning_threshold and hooks.context_critical_threshold. Absent
keys resolve to today's 35/25, so every existing project is byte-identical.

Resolution is total and never throws — this hook must not block the tool call
it rides in on. A value is usable only if Number.isFinite (type-strict, so the
string "30" and true are rejected) and inside the 0-100 domain of the
remaining_percentage it is compared against; anything else falls back to the
default. The PAIR falls back together: critical >= warning has no coherent
reading, and honouring one side silently picks which of the operator's two
numbers to discard. That also covers a single override contradicting the other
key's default.

config-set validates the domain per key so accept and honour agree, but
deliberately does not enforce the pair — it writes one key per call, so a
two-step retune is transiently inconsistent on disk and refusing it there
would block a legitimate configuration.

Registration follows the statusline.show_git precedent: schema manifest plus
src/config.cts validation, not config-defaults.manifest.json and not
buildNewProjectConfig — emitting 35/25 into every new project would pin the
defaults at creation time for a setting nobody has tuned.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2

* enhance(#4285): address Codex review — per-key fallback docs, discriminating tests

Codex full-PR review (gpt-6-astra, read-only) returned five findings. Each was
verified against source before acting; all five are real.

1. docs/CONFIGURATION.md described the wrong fallback. An out-of-domain value
   falls back PER KEY; both defaults apply only when the RESOLVED pair violates
   critical < warning. warning 150 with critical 30 resolves to 35/30, not
   35/25 — at remaining 28 that difference changes the severity emitted. The
   table now states the two rules in the order they compose, and
   docs/context-monitor.md gains the same worked example.

2. The inconsistent-pair test could not prove the CRITICAL side reverts: its
   pair was 20/25, and 25 is already the default, so an implementation that
   reset only `warning` passed it. A 45/50 pair — both halves away from their
   defaults — now pins each side with its own reading, and an equal 45/45 pair
   pins that the rule is strict (`<`, not `<=`).

3. The rejection table's rows could not tell rejection from acceptance: an
   accepted -5 pairs with the default critical 25, trips the pair check, and
   produces the same silence. Two rows now separate those: a below-domain
   critical must escalate remaining 20 to CRITICAL (proving -5 was rejected,
   not honoured), and an unusable critical beside a usable warning 45 must
   still fire WARNING at remaining 40 (proving per-key fallback rather than
   reset-both). The over-claiming comments are narrowed to what each row
   actually shows.

4. Scope, reproduced rather than assumed: config-set writes through
   planningDir(), so under GSD_WORKSTREAM it lands in
   .planning/workstreams/<name>/config.json while this hook reads only
   <cwd>/.planning/config.json. That is the pre-existing root-only scope
   hooks.context_warnings has always had, but this PR advertises the setter
   route, so both docs now say the keys are root-project settings.

5. Four other English docs still stated 35/25 as fixed: the REQ-CTX-02/03
   requirements fragment, ARCHITECTURE.md's hook table and threshold table,
   and INVENTORY.md's hook row. All now name them as defaults and point at the
   config keys; docs/FEATURES.md is regenerated from its fragment via
   scripts/gen-features.cjs --write, not hand-edited.

Four new mutations, each reverted after: resetting only the warning half on an
inconsistent pair (1 red), resetting both on any unusable key (1), dropping the
>= 0 bound (1), and accepting critical == warning (1). perf-317 is 116/0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2

* enhance(#4285): tighten claims after Codex round 2 — scoped paths, one more discriminator

Confirmation round found no runtime defect and confirmed the five round-1 fixes
landed. Four precision items, all real, all fixed here.

1. The scoped-write note named the wrong path for GSD_PROJECT. planningDir()
   composes three distinct shapes, confirmed by running config-set under each:
   .planning/<project>/config.json, .planning/workstreams/<ws>/config.json, and
   .planning/<project>/workstreams/<ws>/config.json. docs/context-monitor.md
   now tabulates all four cases instead of collapsing them into one.

2. The 45/50 silence row asserted empty stdout without pinning the exit code.
   runMonitorRaw turns a spawn failure, a non-zero exit or a timeout into empty
   stdout as well, so the row could have passed on a dead child. It asserts
   exitCode === 0 first now, like the equal-pair row already did.

3. The sibling row's message claimed it proved critical fell back to 25. It
   does not: coercing '30' to 30 yields WARNING at remaining 40 too, so the row
   pins the WARNING side surviving and nothing more. Message narrowed, and a
   new row reads the same config at remaining 28, where the two candidate
   resolutions diverge — rejected gives (45, 25) and WARNING, coerced gives
   (45, 30) and CRITICAL. Mutation-verified: swapping Number.isFinite for the
   coercing global reds it.

4. "Accept and honour must agree" was too absolute in the src/config.cts and
   tests/config.test.cjs comments. The agreement holds on the DOMAIN and per
   key: an accepted value can still lose to the hook's pair check at read time,
   and a scoped write never reaches the hook at all. Likewise a two-step retune
   only CAN be transiently inconsistent — 35/25 to 20/10 is valid throughout if
   critical moves first — so the docs now say what a setter-side pair check
   would actually cost: rejecting that intermediate write and forcing an order.

The same over-absolute phrasing is in b7d179c89's message, which is left as
written rather than rewriting history; this commit and the PR body carry the
precise claim.

perf-317 117/0, config 192/0, config-field-docs 47/0, features-index-gate 84/0,
lint:ci clean cold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2

* chore(#4285): add changeset

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2

* enhance(#4285): address review — planning-config rows, resolveThresholds properties

Two Minor findings from the maintainer review, no behaviour change.

Minor 1: gsd-core/references/planning-config.md's "Hook Fields" table gains
rows for hooks.context_warning_threshold and hooks.context_critical_threshold,
in that table's 5-column form, carrying the same per-key-fallback,
pair-reversion and root-config-scope claims docs/CONFIGURATION.md already
makes. hooks.workflow_guard's absence from that table is pre-existing and
out of scope here.

Minor 2: resolveThresholds() gets fast-check property coverage, which ADR 456
requires of a threshold/limit contract. Reaching it needed a require-time
seam: the resolver was previously observable only by spawning the hook, and a
subprocess per case cannot drive 200 runs — the same conclusion CONTEXT-INDEX
records for the ROADMAP Requirements parser. The stdin adapter therefore moves
into main() behind `require.main === module`, mirroring
gsd-cursor-subagent-start.js and gsd-statusline.js, and module.exports exposes
the resolver plus both default constants so a test asserts fallback against
the source of truth rather than a second copy of 35/25. Spawned behaviour is
unchanged: the 10s stdin timeout still arms per invocation (stdinTimeout is
now a module-scope let assigned in main(), still cleared by the end handler),
and the try/catch crash(ON_CRASH) path is untouched.

Seven properties: totality, ordering, exactness, togetherness, non-vacuity,
per-key fallback, non-object argument. Exactness is stated PER KEY — a mixed
result (one key honoured, one fallen back) is legal and is the documented
contract; the property falsified a per-pair phrasing of it in 4 runs.

Verified: cold lint:ci 0; perf-317 file 125/0; seven mutations killed and
restored, one of which (upper bound widened to 120) is invisible to the 17
hand-written cases and caught only by a property.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte

* enhance(#4285): close the Codex-found gap in the property coverage

Codex whole-PR review of round 3 returned no Blocker and no Major. Two items,
both in the tests added this round, both verified against source before acting.

Minor — the per-key fallback property was asymmetric: it required a usable
warning to survive an unusable critical, but never the reverse. A resolver
that reverted BOTH keys the moment warning was unusable passed all seven
properties. Reproduced exactly: that mutant answers 35/25 for
{warning: 150, critical: 30} where the resolver answers 35/30, and the file
stayed green at 125/0. The mirrored property closes it — with the mutant
re-applied it is now the single failing row, and it is the only row that
fails, so it is load-bearing rather than incidental.

Nit — the ordering property's comment credited it with catching a
half-honoured pair, which it does not: 45/50 "repaired" by resetting only
critical yields 45/25, perfectly ordered. That case belongs to togetherness.
The same comment claimed the behavioural rows sample an inconsistent pair at
exactly one point; stale — they cover 20/25, 45/50 and the 45/45 equality
boundary. Both claims corrected in place.

Verified: cold lint:ci 0; perf-317 file 126/0; the mutant above killed by the
new property alone and the hook restored byte-identical afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte

* enhance(#4285): name the installed-monitor prerequisite; close the negative-critical gap

Second Codex whole-PR pass, run because the base moved: the author's three
"Update branch" merges pulled ~26 upstream commits in, so the previously
reviewed diff sat on a base that no longer exists. No Blocker, no Major, two
Minor — both verified against source before acting.

Minor 1, and only reachable because of what the merge brought in: #2586
(03738824d) landed in that window and stops staging
hooks/gsd-context-monitor.js for Codex, since the metrics bridge it reads is
written only by hooks/gsd-statusline.js, which Codex never installs
(bin/install.js: "gsd-context-monitor.js is deliberately NOT copied for
Codex"). These two keys are read by that hook and nothing else, so on such a
runtime config-set stores and validates them and nothing consumes them — a
claim the docs this PR adds did not make. docs/context-monitor.md now carries
the explanation and both key tables carry a clause pointing at it; the FEATURES
and INVENTORY entries already link through to those two files, so they are not
edited again. The changeset says it too, because it is user-facing.

Accepting the keys on every runtime is kept deliberately: config is shared
across runtimes, so validation stays runtime-independent and the runtime
caveat lives in documentation rather than in the setter.

Minor 2: the per-key fallback property's junk generator had no negative arm,
though its mirror did — and that asymmetry hid a gap. A resolver reverting
BOTH keys whenever critical is negative answers 35/25 for {45, -5} where the
resolver answers 45/25, and it passed all 126 tests. With the negative arm
added it is the single failing row.

Verified: cold lint:ci 0; perf-317 126/0; both mutants above killed and the
hook restored byte-identical; 538/0 across the config, changeset, doc-parity
and emitted-attribution gates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte

* enhance(#4285): refuse the two dead threshold endpoints; resolve absent keys

Maintainer review round 2 raised two Minors and a nit.

Minor 1 — `hooks.context_warning_threshold: 0` was accepted and stored but can
never take effect: `critical < warning` must hold and both sides are clamped to
0-100, so nothing can sit below a warning of 0. Verifying it surfaced the MIRROR
case the review did not name: `critical: 100` is equally dead, since nothing can
sit above it. Both confirmed against the real resolver for partners {absent, 0,
50, 100}, with 0.001 and 99.999 honoured as controls.

`config-set` now refuses both, because storing a value the reader always
discards is the accept-then-discard shape this codebase refuses elsewhere. The
hook is unchanged and still total — it degrades to defaults rather than
throwing, so a project that already carries one of these on disk still loads.
The old "accepts the domain bounds 0 and 100" row asserted the misleading half
and is replaced by tables that make the asymmetry the point (0 is legal for
critical and illegal for warning; 100 is the reverse), plus a control row so
"refuse both endpoints outright" would not pass in its place.

Minor 2 — the keys are absent from config-defaults.manifest.json /
buildNewProjectConfig where the sibling `hooks.context_warnings` lives. Kept
that way: buildNewProjectConfig writes a hooks object into every NEW project's
config.json, which would freeze today's fire-points as an explicit per-project
override everywhere — the opposite of this PR's premise. But the underlying
complaint was real, so the actual symptom is fixed: `config-get` on an absent
key returned "Key not found" while the hook silently used 35/25. It now resolves
through SCHEMA_DEFAULTS. Restated rather than derived because CONFIG_DEFAULTS is
re-exported flattened and has no `hooks` member at runtime; the one resulting
copy of 35/25 outside the hook is pinned against the hook's exported constants
by a drift test (red-checked: moving the literal to 40 reds it).

Nit — PR-body counts unverifiable from the diff. Noted, no code change.

Codex round 3 then found a broken doc link (`context-monitor.md` resolved
inside gsd-core/references/, where it does not exist; the emitted tree's own
convention is `../../docs/...`) and a stale comment still describing the
manifest-derived approach I had backed out. Both fixed. It also corrected my
rationale on a point of fact: manifest entries alone would NOT have reached new
project configs, since buildNewProjectConfig builds its own literal — the
freezing argument applies to that function, not to the manifest. The comment now
says so rather than running the two together.

Verified: cold lint:ci 0; full suite 36,082 / 0 fail before these two fixes,
config + perf-317 321/0 after; drift pin red-checked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-09 04:31:14 +00:00
Lorenz Leslie Espinosa
b33df03726 enhance(#4089): add minimum-solution reasoning check (#4118)
* enhance(planning): add minimum-solution reasoning check

* chore: add changeset for planning guidance

* chore: bind changeset to PR 4118

* docs: document planning sufficiency check

* docs: distinguish planning sufficiency guidance

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-09 00:58:23 +00:00
Tom Boucher
a27cb6b2fa enhance(#4139): Phase 6 — the lazily-read remainder and the artifact templates (#4540)
* enhance(#4406): the lazily-read remainder and the artifact templates

ADR-4139 Decision 3, Phase 6 of the #4139 Compact Content epic. Covers stream 1b
(gsd-core/workflows/<name>/{modes,steps,templates}/*.md) and stream 4
(gsd-core/templates/**) with a variant-swap mechanism, confirmed with the user:
two independent, complete files per covered path (canonical + .compact.md
sibling), with the gate picking which one gets Read at the call site. This is
a different shape from Phase 5's spine+detail partition, and is safe here
specifically because these files are already reached only by a runtime Read —
a missed Read already means zero overlay content today, with or without
workflow.compact_content, so selecting between two independently-complete
files introduces no new failure mode (documented in
gsd-core/references/compact-content-gate.md's new "Streams 1b and 4" section).

Disposition, after inspecting every candidate rather than trusting a byte-size
threshold (same rigor Phase 5 applied to review.md):

- Stream 1b: 1 of 78 files compacted (help/modes/full.md, a user-facing
  reference doc emitted verbatim, not orchestrator instruction). The other 9
  size-threshold candidates are dominated by fail-closed guards, exact CLI
  invocations, or output-format contracts (AskUserQuestion blocks) — recorded
  not-worth-compacting, same reasoning as Phase 5's review.md.
- Stream 4: a ground-truth reachability audit replaced the initial size-only
  candidate list. Two files (summary.md, user-setup.md) got compact variants;
  a third (spec.md) was drafted, then dropped after discovering its only two
  call sites are eager @-includes, not a runtime Read — stream-1 material
  hiding under gsd-core/templates/, not stream-4's actual mechanism. summary.md
  itself has 3 eager call sites and only 1 genuine runtime-Read call site
  (execute-plan.md); only that one was wired, so the compact variant's savings
  apply to the sequential single-plan execution path only.
- Discovered while auditing reachability: 12 gsd-core/templates/** files with
  zero references anywhere in workflow/agent/command prose, compiled source,
  or tests — dead scaffolding predating this phase. Deleted in this same PR
  per this repo's no-defer policy, after re-verifying against a computed
  path.join(...) pattern (not just a plain-string search) that nearly caused
  two genuinely load-bearing templates (user-profile.md, dev-preferences.md)
  to be misclassified as dead.

New checker (tests/helpers/compact-content-variant.cjs): registration,
reachability, protected-content-preserved, size-smaller — replacing Phase
3/5's disjointness/completeness checks, which assume a partition rather than
two deliberately-overlapping documents. The reachability check's own
"unprefixed match" guard had a real bug (rejected the repo's own
`~/.claude/gsd-core/...` convention), caught by running it against the
already-wired help/modes/full.compact.md pair rather than only synthetic
fixtures — fixed to anchor on the nearest `gsd-core` path segment instead.

Template consumer parity (tests/compact-content-template-variant-parity.test.cjs):
proves each compact variant's `## File Template` fenced block — the actual
output-format contract a generated SUMMARY.md/USER-SETUP.md is parsed
against — is byte-identical to the canonical file, then runs the one real
deterministic consumer (gsd-core/bin/lib/coverage.cjs's classifyContent,
backing `gsd-tools uat classify-coverage`) against content built from that
shared contract.

Added a sibling benchmark script (scripts/benchmark-compact-content-variants.cjs)
rather than extending the existing spine/detail one — different data shape,
and the existing script's own contract deliberately isolates it from a
test-only helper's shape changing.

Emitted-drift acknowledgement: not needed. Every changed/added path in this
diff is hand-authored and present in the diff itself, so diffEmitted's
attribution loop resolves `via` to the path's own source before reaching the
ack-lookup branch (same reasoning Phase 5 verified for its own diff).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* enhance(#4406): address code-review findings on the variant-swap gate

- docs/CONFIGURATION.md and gsd-core/references/planning-config.md's
  workflow.compact_content entries described only the spine+detail mechanism
  (Phase 5) and were missing this phase's variant-swap mechanism and its
  benchmark:compact-content-variants script entirely — required since this
  PR's changeset is type Added (CLAUDE.md's "Missing Docs for Changesets"
  rule). Both now describe both mechanisms and which call sites are wired.
- Added the missing RED^-1/no-op fixture for checkProtectedContentPreserved:
  a canonical file with zero <!-- gsd:protected --> blocks must be a
  no-op, not a violation — the only branch of that function the existing
  fixtures didn't exercise.
- Collapsed findCompactFiles/findMarkdownFiles in
  tests/helpers/compact-content-variant.cjs into one findFilesWithSuffix
  helper — the two were identical recursive walks differing only in the
  extension predicate (minor Duplicated-Code finding).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4406): restore copilot-instructions.md, a false-positive dead-template classification

gsd-test caught this, not static analysis: 10 real failures in
tests/copilot-install.test.cjs, tests/installer-migration-install.integration.test.cjs,
and tests/repo-layout.test.cjs — all downstream of bin/install.js's Copilot install
path, which does
fs.readFileSync(path.join(targetDir, 'gsd-core', 'templates', 'copilot-instructions.md'))
after copying gsd-core/templates/** into the target project, then merges it into both
.github/copilot-instructions.md and (local installs) AGENTS.md. The reachability audit
that flagged this file as dead checked src/*.cts and gsd-core/bin/*.cjs but never the
repo-root bin/install.js — a separately maintained installer bundle outside the
src/-to-gsd-core/bin/lib/ compiled-output convention. The fs.existsSync guard around
that read degrades to a silent skip rather than a crash when the template is missing,
which is why this surfaced only once the real E2E install test ran, not from any
static check.

Re-verified the remaining 11 deleted filenames against bin/install.js specifically
(plain substring and quoted-filename search) before trusting that list — all 11 have
zero hits there, confirmed dead by the same standard this one file failed.

Regenerated the installer emitted-tree goldens (tests/fixtures/install-tree/*.json) to
reflect the restored file, and corrected the "Removed" changeset (jolly-lynx-sprint.md)
and the phase design doc from 12 to 11 deleted files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: execute-plan.md — call-site wiring for the summary.md and user-setup.md .compact.md variants
Emitted-Drift-Ack-Growth: help.md — call-site wiring for full.compact.md, same variant-resolution rule

* docs(#4406): backfill changeset PR numbers

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4406): resolve removed-but-needed lint findings on the dead-template deletion

CI's own full-test matrix (not gsd-test's matrix, which does not run this
check) caught 4 more false-positive dead-template classifications via
tests/removed-but-needed-lint.test.cjs / scripts/lint-removed-but-needed.cjs
— a literal, word-boundary basename check across .github/workflows/,
gsd-core/, and docs/ (excluding docs/adr/** and docs/research/**) for every
file a PR deletes. It has no semantic awareness, so a deleted template's
basename colliding with something else entirely still fires:

- claude-md.md: gsd-core/templates/README.md had a stale table row claiming
  /gsd-profile reads this template to generate CLAUDE.md. Verified false (no
  code reads it anywhere, same search that already covered bin/install.js) —
  fixed the row to *(inline)*, matching every other command-generated
  artifact in that table. File stays deleted.
- codebase/testing.md: collided with docs/guides/testing.md, an illustrative
  example row in docs-update.md's sample output table (an unrelated real
  generated-docs path). Swapped the example topic to "contributing" — the
  row is illustrative, any topic works. File stays deleted.
- codebase/architecture.md, codebase/stack.md: collided with docs/reference/
  planning-artifacts.md's directory listing of a user's own generated
  .planning/codebase/architecture.md and stack.md output — the same
  semantic mismatch already investigated and dismissed as unrelated earlier
  in this phase's audit, now caught by a gate instead of judgment. That
  listing repeats across 5 locale copies of the doc.
- continue-here.md: collided with the real .continue-here.md pause-work
  artifact, referenced across 15+ locale and workflow files.

For the last two, the lint's own error message offers "restore the file or
update every consumer in the same commit." Rewording 15+ files across
languages I cannot verify translation quality for, to shave 2 already-tiny
templates that were merely presumed dead, is disproportionate to this PR's
actual scope — restored codebase/architecture.md, codebase/stack.md, and
continue-here.md instead, and corrected docs/ARCHITECTURE.md's Templates
section accordingly.

Final confirmed-dead set: claude-md.md, codebase/concerns.md,
codebase/conventions.md, codebase/integrations.md, codebase/structure.md,
codebase/testing.md, debug-subagent-prompt.md, discovery.md — 8 files, down
from the original 12. Verified locally: GSD_REMOVED_BUT_NEEDED_BASE=next
node scripts/lint-removed-but-needed.cjs now passes clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: docs-update.md — swapped an illustrative example-table topic (testing -> contributing) to avoid a removed-but-needed basename collision with the deleted codebase/testing.md template; net +10 bytes

* fix(#4406): split codex-config.test.cjs to fix a genuine Windows CI timeout

Root cause of the `full test (windows-latest, 24, shard 2/3)` failure the
user asked to be actually fixed, not just re-run past: PR #4497 (landed
2026-09-07, one day before this PR's CI run) isolated
tests/codex-config.test.cjs into its own dedicated chunk because its
measured weight (17.87, ~45% of the post-cut Windows budget) made it unsafe
to share a chunk with any other file. That isolation was necessary but not
sufficient — even alone, with zero companion-file contention, the file's
real Windows execution time sits right at the 600s per-chunk ceiling. Two
independent CI runs on two unrelated PRs (this one and #4154) were both
killed within ~1.4s of the identical 600000ms mark — not random contention,
a deterministic near-miss the isolation fix couldn't address because it
never reduced the file's own cost, only removed the risk of a companion
file's cost stacking on top of it (which the PR #4497 comment explicitly
anticipated: "if a future profiling pass genuinely speeds up
codex-config.test.cjs itself, this isolation can be revisited").

The file itself explains why it's this heavy: 11,262 lines / 433 tests / 79
describe blocks, accumulated over dozens of bug-fix PRs (#2695, #2760,
#3245, #3285, #3346, #3426, #3427, #3562, #3566, #3582, #3808, and more),
several of which are explicitly documented as "folded" in from separate
files that were never actually split back out ("Verified non-duplicate
against both the pre-existing target and the other three folded sources").

Split into 4 files by top-level AST statement boundaries (never a naive
column-0 regex — an early attempt at that overcounted 79 apparent
"describe(" matches when only 21 are genuinely top-level; the rest are
nested inside a handful of large folded-in blocks, which a regex can't tell
apart from real top-level statements). Verified lossless twice: the split
script asserts byte-for-byte reconstruction of every source character, and
independently, total test()/describe() call counts match exactly between
the original file and the sum across all 4 new files (433/79 both sides).
Each new file carries the complete original shared header (imports/helpers)
for safety; per-file unused-import warnings from that duplication are
resolved via ESLint-precise alias renames (`{ foo: _foo }`, the standard
form for an intentionally-unused destructured binding — never a bare `{
_foo }`, which would destructure a different, nonexistent property).

No change needed to scripts/run-tests.cjs's ISOLATED_HEAVY_FILES or its
pinned test in tests/run-tests-harness.test.cjs: the file that keeps the
original name (tests/codex-config.test.cjs) is now only ~28% of the
original's size and safely isolated in its own chunk as before; the other
three new files re-enter normal weight-balanced packing, none individually
close to disproportionate. Confirmed no other file hardcodes the hardcoded
filename anywhere that would silently stop these tests from running (the
CI test-selection scripts determine scope algorithmically, not by literal
filename).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 14:31:04 -04:00
Tom Boucher
e03921c7d8 enhance(#4405): split the rest of the eager-window workflows worth splitting (#4536) 2026-09-08 00:17:22 -04:00
Dennis Alexis Valin Dittrich
18c899def5 enhance(#4209): optional external source reviewer lanes for /gsd:code-review (#4323)
* test(01-01): define reviewer-support trait contract

Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02):
validator rejects non-boolean values with an exact field path, accepts
missing/true/false, and the real code-review capability.json steps
must declare supportsReviewerLanes: true. Add loop-resolver projection
coverage proving the trait reaches activeHooks verbatim for a
provider-neutral synthetic step (not code-review-specific), and that
omitted/false values stay inert (no key on the active hook).

All 8 new assertions fail today: the validator has no such field, and
loop-resolver has nothing to project. RED before GREEN.

* feat(01-01): declare reviewer-capable steps

Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional
boolean opt-in trait, step-scoped (not capability-wide). Only a
literal true validates and projects; false/omitted stay inert (no
key on the projected active hook), and every non-boolean type fails
capability-validator.cjs with an exact field-path error.

Opt both existing code-review steps (execute:post, execute:wave:post)
into the trait in capabilities/code-review/capability.json. Project
the validated field through src/loop-resolver.cts into activeHooks
so a provider-neutral generic interpreter can read it without any
code-review-specific knowledge. Document the field in
docs/reference/capability-manifest.md and regenerate
gsd-core/bin/lib/capability-registry.cjs via the generator (never
hand-edited).

Makes all 8 RED assertions from the prior commit pass.

* test(01-02): define shared reviewer dispatch

- Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes:
  inert when the supportsReviewerLanes trait is off or nothing is selected,
  exactly-once plan/invoke per selected lane, duplicate-alias dedup, the
  bounded metadata-only source-review prompt (repo root, paths+baseSha,
  depth, four fixed prohibitions), and capability-neutral reuse via a
  second synthetic step context.
- RED: module under test (src/reviewer-step-dispatch.cts) does not exist
  yet, so require() fails and every assertion is unreached.

* feat(01-02): dispatch reviewers for opted-in steps

- Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps),
  ONE interpreter for a step's supportsReviewerLanes trait. Reuses
  resolveReviewerSelection for selection and resolveLanePlan for planning
  (both already-existing, pure building blocks); invocation is the one
  required, caller-injected seam (deps.invoke) since runLane needs
  OS-aware spawn plumbing this module does not own.
- trait !== true, or a selection resolving to zero lanes, dispatches
  nothing (zero plan/invoke calls). Each selected lane is planned and
  invoked exactly once, in the selector's deduped/sorted order.
- buildSourceReviewPrompt assembles a metadata-only bounded prompt
  (repo root, canonical paths + base SHA, depth, four fixed
  prohibitions) — never file contents — written once per dispatch and
  shared across every invoked lane.
- GREEN: tests/reviewer-step-dispatch.test.cjs now passes.

* test(01-02): define reviewer dispatch failures

- Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed
  matrix: an explicitly requested lane the selector could not resolve
  still lets the OTHER resolved lane run, but the aggregate result must
  never read as a clean success (and 'every explicit lane unavailable'
  must be distinguishable from the plain no-flags-passed inert case);
  request-level validation (path traversal, absolute paths outside
  repoRoot, empty/non-string paths, missing depth/base SHA) halts the
  whole dispatch before any lane is planned or invoked; a per-lane
  prompt-budget overflow hard-fails only that lane before invoke while
  its sibling still runs.
- RED: src/reviewer-step-dispatch.cts does not yet implement any of
  these guards, so 9 of the new assertions fail against the current
  (Task 1) implementation.

* fix(01-02): fail closed in reviewer dispatch

- src/reviewer-step-dispatch.cts: add the fail-closed guards the prior
  commit deliberately left out. An explicitly requested lane the
  selector could not resolve no longer lets the aggregate read as a
  clean success — lanes that DID resolve still run and keep their
  results (never narrow the requested set), but selection.errors now
  flips the aggregate ok to false, and 'every explicit lane
  unavailable' is now distinguishable (SELECTION_FAILED) from the
  plain no-flags-passed inert case (NO_LANES_SELECTED).
- Add request-level validation (validatePaths, depth/baseSha presence)
  that halts the WHOLE dispatch before any lane is planned or invoked:
  path traversal, absolute paths outside repoRoot, empty/non-string
  paths, and missing provenance are all rejected up front.
- Add per-lane prompt-budget enforcement (resolveBudget, mirroring
  gsd-tools.cjs's budgetFor convention including budget 0 = unbounded):
  a lane whose resolved budget the prompt exceeds hard-fails before
  invoke runs for it, without cancelling a sibling lane already
  planned.
- Document the supportsReviewerLanes trait and its dispatch-step
  interpreter in gsd-core/references/loop-hook-dispatch.md.
- GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass;
  no regressions in the review-lane/reviewer-selection/prompt-budget
  suites (356 passing).

* test(01-03): define optional source reviewer flow

RED: assert code-review.md dispatches roster-derived reviewer-lane flags
through a single review-lane dispatch-step call (DISP-01..05), that the
no-flag path stays byte-for-behavior unchanged (COMP-01), and that
external evidence reaching the internal reviewer prompt is marked
unverified (CONS-02). Also covers the CLI contract directly: no-op with
no explicit selection, and fail-closed on an explicit unknown lane
(SAFE-07) via real gsd-tools.cjs subprocess calls.

* feat(01-03): route optional source reviewers

GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches
canonical reviewer-lane flags against the merged first-party + installed
roster (never a hand-maintained list) and, only when at least one is
present, calls the shared reviewer-step interpreter exactly once with the
already-resolved repo root, file scope, depth, and base SHA. Its evidence
paths are appended to the internal reviewer prompt via
${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane
flag leaves the internal-only dispatch byte-for-behavior unchanged
(COMP-01).

Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane
dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI
route `dispatchReviewerLanes` wires through, but never implemented the
gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add
it to the existing review-lane router, reusing the same effort-aware plan
building and runner deps `plan`/`invoke` already use (factored into
buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard
the CLI's own `detected` set on whether an explicit flag was passed:
resolveReviewerSelection's no-explicit-selection fallback is "select every
detected reviewer" (the correct default for /gsd:review), and passing it
an unconditionally non-empty detected set would silently invoke the whole
roster on every no-flag code review, violating COMP-01.

* test(01-03): define external finding consolidation

RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as
untrusted input — independently re-verifies every claim against the actual
current source, resists a prompt-injection attempt embedded in evidence
text, and folds a verified claim into the existing Narrative Findings
section with no second REVIEW.md schema (CONS-01..03). Also assert
code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* feat(01-03): consolidate external review evidence

GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence>
as untrusted data, independently re-verifies every cited claim against the
actual current source before it can appear in REVIEW.md, and explicitly
resists prompt injection embedded in evidence text (never a command, no
matter what it claims to be). A verified claim folds into the existing
Narrative Findings section with (external: {slug}) provenance — one
REVIEW.md schema only, no separate external-findings section.
code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* fix(01-02): gitignore the reviewer-step-dispatch build artifact

01-02 added src/reviewer-step-dispatch.cts but never added its
npm run build:lib output to .gitignore, unlike every sibling
gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked
noise in git status.

* docs(01-04): publish user and command contract for reviewer-lane source review

- Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md
  and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on
  failure, findings independently consolidated into the single REVIEW.md
- Add the same contract to the docs/features/code-review-pipeline.md
  fragment and regenerate docs/FEATURES.md from it
- Preserve /gsd-review as the plan-review command; cross-reference it
  rather than duplicating the reviewer roster
- Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md
  drift owned by source already shipped in Plans 01-01/01-03 but never
  regenerated (npm run regen:derived had not been run in this worktree)

* docs(01-04): align architecture and agent ownership docs for reviewer-lane trait

- ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes)
  through the shared dispatchReviewerLanes interpreter to the existing
  review-lane plan/invoke machinery, ending at gsd-code-reviewer as the
  sole REVIEW.md consolidator
- AGENTS.md: document gsd-code-reviewer's full-context verification scope
  and its treatment of external reviewer evidence as unverified input
- No new diagram, abstraction, or config key; docs/CONFIGURATION.md is
  unchanged since the feature adds no setting or default

* fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact

Same gap as the earlier .gitignore fix: 01-02 added
src/reviewer-step-dispatch.cts but never added its generated
gsd-core/bin/lib/reviewer-step-dispatch.cjs output to
eslint.config.mjs's ignore list like every sibling generated file,
so tsc's emitted __importDefault CommonJS-interop var tripped
no-var.

* fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md

01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists
cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster
row in docs/INVENTORY.md — required by design, since a role sentence
cannot be generated — was never added.

* fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes

refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook
strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added
supportsReviewerLanes: true to that step and this fixture was not updated.

* chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md

Both files grew as a direct, intended consequence of wiring optional
reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes
step and the untrusted-evidence consolidation contract) — not
incidental drift.

Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209)
Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209)

* test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes

From internal code review: dispatched must be false when zero lanes
actually reached plan(), and a throwing plan()/invoke() for one lane
must not discard results already collected for a sibling lane —
matching the fail-closed pattern gsd-tools.cjs already uses for the
same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794).

Refs: gsd-core-dks.16, gsd-core-dks.17

* fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review

- WR-01: dispatched now tracks whether any lane actually reached
  plan(), not results.length — an unresolvable selected slug no
  longer reports dispatched:true.
- WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a
  throw for one lane can never discard results already collected
  for a sibling lane, matching the same guard gsd-tools.cjs already
  has around the identical resolveLanePlan call.
- IN-01: documents the intentional budget===0-is-unbounded
  convention (#2797) the caller already relies on.
- IN-02: review-lane dispatch-step no longer blocks indefinitely on
  an un-piped interactive TTY; fails closed to empty paths instead.

Refs: gsd-core-dks.16, gsd-core-dks.17

* docs(01-05): add changeset fragment for PR #17

* fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract

agents/gsd-code-reviewer.md's untrusted-evidence section and its
pinning regression test both quote injection phrases as the exact
attack they defend against/detect — same
DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing
allowlist entries, not an actual injection vector.

* test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip

From CodeRabbit review: WR-02's earlier fix only wrapped plan() —
writePromptFile()/deps.invoke() still ran unguarded, so a throw
there still aborted every later selected lane. Also covers the
dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit
lanes silently not running when no prior review and no phase-start
commit exist).

* fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved

Previously an explicit reviewer-lane request with no prior review and
no resolvable phase-start commit reached dispatch-step with an empty
--base-sha, which fails closed via missing_provenance — correct, but
silent about why explicitly requested lanes didn't run. Now skip
dispatch entirely in that case with a stderr warning naming the
actual cause.

* fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan()

WR-02's original fix only guarded plan() — a throw from
writePromptFile() or deps.invoke() still aborted the whole dispatch,
discarding results already collected for lanes processed earlier in
the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it.

* fix(01-05): WR-02b mock must throw only on the first writePromptFile() call

The committed mock threw unconditionally, so codex's retry also threw and
failed for the same reason as claude's — the test could not distinguish
'sibling still runs' from 'sibling also breaks'. Gate the throw to the
first call, matching WR-02/WR-02c's single-failure intent.

* fix(#4209): close review findings from adversarial + critical-code-reviewer pass

Two independent reviews (agy adversarial review, Opus critical-code-reviewer +
ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch
wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and
fixed here:

- dispatch-step's reducer silently swallowed whole-dispatch rejections
  (invalid paths, missing provenance, etc); it now checks parsed.ok/reason.
- spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the
  LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review;
  now shares the single compute_file_scope derivation.
- the external reviewer prompt had no actual review request or citation
  requirement, only prohibitions; added both.
- removed the supportsReviewerLanes trait plumbing (capability registry,
  validator, loop-resolver, docs, tests) — it was never consulted by the
  real dispatch path, which gates on explicit CLI flags instead.
- flag-resolution require() was a fragile cwd-relative literal that failed
  silently on non-vendored installs; now resolves via GSD_TOOLS's own
  directory and warns instead of swallowing failure.
- reducer didn't unwrap the @file: overflow protocol for large payloads.
- deduplicated resolveBudget/budgetFor into one resolveLaneBudget.
- lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a
  second dispatch can't overwrite prior evidence.
- validatePaths rejects control characters, closing a markdown-injection
  vector into the external prompt via crafted filenames.
- reworded the one line that tripped prompt-injection-scan.sh instead of
  allowlisting the whole production prompt file.
- fixed a stale docstring range and a dispatched-field ordering bug.
- added 3 integration tests executing the actual reducer against synthetic
  dispatch-step JSON, replacing markdown-substring-only assertions.

771/771 tests pass across every touched suite; tsc --noEmit clean.

* fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait

The maintainer's approval on issue #4209 explicitly redirected implementation
shape: reviewer-lane dispatch must be a reusable capability/step-dispatch
trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call
itself. My previous commit (e2558326) deleted that trait entirely after
finding it declared-but-never-consulted, which was backwards — the fix was to
wire it, not remove it.

Restores the trait (capability.json, generated registry, validator,
loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes
now resolves its own active hook via `gsd_run loop render-hooks` for the
configured workflow.code_review_point and only proceeds to CLI-flag matching
when supportsReviewerLanes reads true. Explicit flags no longer bypass the
trait; a matching flag with the trait false resolves zero slugs (proven by a
new integration test executing the real fence with both trait states).

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the
dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer
redirect requires the capability layer, not the workflow, own the opt-in
decision).

* fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point

Both an agy adversarial review and an Opus critical-code-reviewer pass
independently found the same gap in my previous commit (9b2c3773d): the trait
check I wired into code-review.md only protected code-review's OWN
invocation — gsd-tools.cjs's dispatch-step handler still hardcoded
`trait: true` unconditionally, so a second capability declaring
supportsReviewerLanes would get zero enforcement from the shared CLI unless
it correctly re-implemented the ~15-line render-hooks scrape itself. That is
exactly the "each workflow.md hand-wiring the call" the maintainer's redirect
said to eliminate.

Moves the trait check into dispatch-step itself: given --cap-id/--point, it
self-invokes `loop render-hooks <point>` (relocating the one subprocess
code-review.md used to spawn for this, not adding a new one) and derives the
real trait from that capId's active hook, rather than trusting a
caller-passed boolean. code-review.md now only passes
--cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or
gates on the trait itself — the ~20-line scrape it previously carried is
gone. Any other capability opts into the identical enforcement by declaring
the trait and passing the same two flags.

Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input
variable (they proved a bash branch honors a variable, not that the variable
reflects the real capability manifest) with three integration tests that
invoke the real dispatch-step CLI against the real first-party capability
registry: the real code-review trait resolves true, an unknown --cap-id
resolves false (trait_not_enabled, fail-closed), and omitting
--cap-id/--point entirely resolves false (no context means no opt-in).

Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check
(agy-F1 was incomplete), and delete the promptWritten per-lane coupling
flag — the prompt write is idempotent, so writing it once per lane instead
of gating on "did any lane write it yet" removes a latent bug where a
deps.plan override that ever varies promptPath per lane would silently skip
writing for a later lane.

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count
drops (the trait scrape moved into dispatch-step), but the file still grew
this session across multiple commits; acknowledging per the growth-tracking
convention.

* fix(#4209): remove per-run token waste from the shipped prompts

Runtime prompt content, not session tokens: two real, per-invocation token
costs in the code that ships.

1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of
   load_context step 5's ~180-word untrusted-evidence contract in ~90 more
   words, breaking this section's own established terse one-liner style
   (every other rule here is 1-2 sentences). This prompt loads fresh on
   every /gsd:code-review invocation. Shrunk to a one-line cross-reference,
   matching how write_review's own reference to step 5 already does it.

2. buildSourceReviewPrompt repeated the base SHA on every single file line
   even though it is identical for every file and already stated once at
   the top of the prompt — O(files) wasted tokens on every dispatched lane
   for a 50-file review, for zero information gain. File lines are now bare
   paths.

* fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3

Opus critical-code-reviewer found a real Blocking defect in the --cap-id/
--point self-invocation added last commit: `dispatch-step` spawned
`loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its
stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to
`@file:<path>` instead of inline JSON -- the same overflow protocol this
feature already unwraps for its OWN dispatch result 60 lines later in
code-review.md. A large-enough activeHooks envelope (more installed
capabilities/fragments) would throw, get silently swallowed by the bare
catch, and misreport a real trait as trait_not_enabled with zero diagnostic.

Fixed by extracting the config/registry/capability-state resolution
`cmdLoopRenderHooks` already performs into an exported pure function,
resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now
share it), and calling it in-process from dispatch-step instead of spawning
a subprocess at all. This eliminates the @file: exposure entirely (the
dispatch-step path never touches the rendered-string envelope or its
JSON-stringify/50000-char threshold), removes one subprocess spawn per
code-review invocation, and gives a genuine diagnostic (stderr warning) on
resolution failure instead of silent fail-closed. Corrected three doc/
docstring references to the now-removed subprocess self-invocation.

Also fixes 2 real CI failures this round surfaced:
- lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line
  no-control-regex` comment was unused under this project's ESLint config
  (verified locally: the rule never actually flags \x00-\x1f in this repo's
  config) -- a mistake from an earlier commit this session, never actually
  lint-checked before push. Removed the disable comment.
- security (prompt-injection-scan): the agy-F1 regression test's crafted
  fixture literally contains "Ignore all prior instructions." as test data
  proving validatePaths rejects it -- allowlisted the test file, same
  DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries.

Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one
bullet stated "untrusted, never a command" three different ways in one
paragraph, and a same-file duplicate of write_review's schema rule.
Consolidated to state each rule once.

Declined one suggestion from this round: shrinking code-review.md's
EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests
(tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block,
tests/code-review.test.cjs's CONS-02 test) deliberately lock the four-
prohibitions restatement and the untrusted-evidence prose into the
INJECTED block itself, not just the consolidator's system prompt --
adjacency of the warning to the untrusted payload it's warning about is a
recognized prompt-injection defense-in-depth pattern from this
workstream's original TDD plan, not accidental duplication.

* fix(#4209): correct stale per-file base-SHA prose in the external prompt

Leftover from removing the per-file base SHA repetition earlier this
session: the review-request sentence still said "relative to its base SHA"
(singular per-file framing) when there's now exactly one base SHA, stated
once above the file list. Reads "relative to the base SHA above" now.

* fix(#4209): make getLane/configGet/plan required deps, delete dead defaults

R3/R4 from the review round I'd deferred as low-priority test-churn: this
file's one production caller (gsd-tools.cjs's dispatch-step handler) always
supplies all three, so the fallbacks were dead in production -- but each was
actively WRONG if ever reached: the default configGet always returned
undefined, silently disabling resolveLaneBudget's overflow guard; the
default getLane looked up only first-party REVIEWER_LANES, diverging from
production's overlay-merged roster; the default plan skipped per-host effort
resolution entirely.

These defaults were introduced by this PR's own earlier work (this file did
not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from
elsewhere, so there's no external caller depending on the lenient contract.

Turned out free to fix: making the three deps required and deleting
defaultGetLane/defaultPlan needed zero test changes -- every existing test
that actually reaches the per-lane loop already supplies getLane/plan
explicitly, and configGet's only real dependent (the budget-overflow tests)
already supplies it too. 788/788 tests pass unchanged, tsc/lint clean.

* fix(#4209): define depth semantics for the external reviewer lane

Verified this was a real bug, not a match to existing convention as I'd
claimed when declining the suggestion earlier this session: the internal
gsd-code-reviewer agent's own system prompt carries a full <depth_levels>
block defining what quick/standard/deep mean and do (agents/gsd-code-
reviewer.md:68-99). The external reviewer lane has no access to that
persona at all -- it only ever sees buildSourceReviewPrompt's bounded text,
which sent the bare depth label with zero definition to a third-party CLI
with no other source of truth for what "standard" means.

Added depthMeaning(), condensed from the internal reviewer's own
<depth_levels> definitions so the two stay consistent, and interpolated it
into the review-request sentence. 150/150 tests pass, tsc/lint clean.

* fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation

CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the
roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a
SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide
whether to dispatch at all. This file's own documented rule (its
depth-resolution guard, stated explicitly a few hundred lines earlier) is
that a guard and the extraction it protects must run as one shell
control-flow decision, because markdown-fenced blocks do not share shell
state -- this step violated its own file's rule for the entire feature's
gating condition.

Merged the roster-resolution fence and the dispatch-decision fence into one
continuous bash block, removing the intervening prose that split them.
Fixed the stderr-based failure detection in the same edit (RQ-01: checking
whether stderr is non-empty misfires on any benign Node warning; now checks
the actual exit status of the roster-resolution command).

Verified by extracting the merged fence and executing it standalone, driving
both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a
real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty,
SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests
pass, tsc/lint clean.

* fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write

Batch of Required/Suggestion fixes from the Opus critical-code-reviewer +
writing-for-agents pass:

- CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch
  blocks, commented-out code) and deep (error propagation, state mutation
  consistency, circular dependencies) relative to the real <depth_levels>
  block, and had zero test coverage. Restored full accuracy and added tests
  that read the real agents/gsd-code-reviewer.md file directly, so drift
  between the two can't recur silently. Unrecognised depth now normalizes to
  standard's definition, matching that agent's own documented rule, instead
  of rendering an undefined bare label.

- RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt
  `paths` does, but weren't checked for control characters like paths were
  (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and
  applied it to all four fields at the same provenance-check boundary.
  runDir previously had zero validation at all.

- S1: deleted the dead `identity` parameter on `invoke` -- the one production
  caller already ignores it, no test read it by name.

- S2: hoisted the shared prompt write above the per-lane loop -- promptPath
  is derived from runDir alone (constant across lanes by construction), so
  writing it once is both correct and cheaper than the per-lane write R1
  introduced earlier this session. Discovered and fixed a real regression
  from the naive version of this hoist: an unguarded throw would have
  escaped dispatchReviewerLanes as an uncaught exception instead of a clean
  per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason,
  matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with
  a dedicated regression test.

- S3: moved `planned = true` past the budget-overflow gate, so `dispatched`
  only reports true once a lane has cleared BOTH plan and budget checks.

- S5: relayed gsd-code-reviewer.md's own "performance issues are out of
  scope unless also correctness issues" policy into the external-lane
  prompt, which previously had no such guidance and could return findings
  the internal reviewer's own contract excludes.

- RQ-05 (partial): shrunk this file's own header docstring's restatement of
  the trait-reuse architecture to a pointer at
  gsd-core/references/loop-hook-dispatch.md, the canonical home.

234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion

RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the
SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke`
already share. code-review.md's ~18-line inline `node -e` reimplementing
`loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in
gsd-tools.cjs) is now a single call to this subcommand -- the exact
violation code-review-flags.cjs's own header warns against ("this is the
canonical flag-parsing surface -- do not replicate inline bash parsing").

RQ-03: an empty --cap-id XOR --point now warns distinctly from the
legitimate no-context opt-out (both absent) -- a caller that named a
capability without its point was silently indistinguishable from a correct
opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only
ever fires when the config-get COMMAND ITSELF fails (config-get already
resolves the manifest's own schema default in the normal case), but that
failure was previously silent.

RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait
resolved inside dispatch-step" explanation was restated in full in 5
places across this session's own review cycles. Consolidated to ONE
canonical statement in gsd-core/references/loop-hook-dispatch.md; the other
4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md,
code-review.md's step-opening comment) now point at it instead.

W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two
inert cases when capability-validator.cjs already rejects non-boolean at
load -- restated as the two cases that actually reach this code. Removed a
"do not hand-roll trait resolution" prohibition whose target no longer
exists once the positive description precedes it.

W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing
block means proceed as normal") -- an absent optional block already means
proceed as normal without being told.

W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the
made-up compound "byte-for-behavior [un]changed" with the token this
session's own docs already coined for this concept (inert) and the word
that means what byte-for-behavior was reaching for (unchanged).

W-10: dispatch_reviewer_lanes had no completion criterion -- added one
sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set,
either populated or empty). This exact sentence would have caught the
cross-fence bug fixed two commits ago at authoring time.

Declined from this round, with reasoning: W-02/W-03 (trim the
untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) --
two tests deliberately lock this as intentional adjacency-based
prompt-injection defense-in-depth, not accidental duplication (see this
branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site
`trap ... EXIT`) -- would fire at the end of the CREATING fence, before
spawn_reviewer's agent ever reads the evidence files, given this file's own
documented fenced-block execution model; the existing named cross-reference
between creation and cleanup already satisfies the co-location concern
without introducing that regression.

853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex

Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed
for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get
fallback lived in an earlier, separate fence from the fence that consumes it
via --point, split only by prose (not a guard, per this step's own documented
rule). Merged into the single continuous fence and added a structural test
asserting exactly one bash fence in the step.

The new end-to-end regression test for this used --codex, which drives the
fence's real `review-lane dispatch-step` call and, with the codex binary
present on PATH, spawns the real external CLI — which then blocks on
interactive auth with no stdin (BL-01). Stubbed gsd_run for
`review-lane dispatch-step` only (captures argv instead of executing),
keeping the real config-get/explicit-from-argv calls the test is actually
about.

* fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment

Round-5 review (Opus) warning-tier findings:

- WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but
  a control-character injection attempt" — a caller distinguishing a config
  problem from a security event couldn't tell them apart. Split into
  MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid).
- WR-05: validatePaths' containment check was lexical only (path.resolve),
  so a symlink whose own path sits inside repoRoot could still point outside
  it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can
  legitimately name a file already deleted in a stale worktree), realpathing
  repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't
  false-positive-reject its own real children.
- WR-08: a comment in the per-lane loop still said a throwing writePromptFile()
  was caught there — stale since the prompt write was hoisted above the loop
  in an earlier round.

WR-03 (validate depth against the quick/standard/deep enum) was considered
and declined: this dispatcher is deliberately capability-neutral (see the
existing "synthetic step context" test, which passes a non-code-review depth
label on purpose to prove no code-review-specific special-casing exists).
WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and
WR-07 (reason omitted on the aggregate return) were verified against source
and are not bugs — see review notes.

* docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap

Round-5 review (Opus, BL-03) flagged that an early exit between
dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A
trap-based cleanup was considered and rejected: if a step genuinely runs as
a separate process, a trap set at creation time would fire at the end of
that SAME fence, deleting the directory before spawn_reviewer/commit_review
ever read it — worse than the leak it would fix.

review.md's own gather_context/cleanup pair for the identical resource class
(a run-scoped reviewer temp dir) already makes and documents this exact
trade-off: cleanup runs only on a documented success path, and a leftover
$TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording
that precedent here so this isn't re-raised as a live gap in a future review.

* fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path

reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes
paths: ['docs/spec.md'] as a synthetic, never-read path proving the
dispatcher has no code-review-specific special-casing. lint-docs-guard-
registration correctly flagged this as an unregistered docs/ path reference —
add the docs-guard-exempt marker and its pinned baseline entry, the same
pattern every other synthetic docs/ literal in this test suite already uses.

* fix(#4209): backfill changeset pr: field with the real upstream PR number

changeset-lint's fail_pr_field_drift caught the fragment still pointing at
the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this
branch is now also open against.

* docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam

trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every
decision in ADR-2782 (D1-D9) and every prior dated amendment governs the
`role: "reviewer"` capability body and its one consumer, /gsd:review. This
PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary
feature capability's `steps[]` entry, projected through loop-resolver.cts
and resolved in-process via resolveActiveHooksForPoint - is a different
capability axis (steps/gates/contributions) that the ADR's own scope note
explicitly places out of reach. Per docs/contributor-standards.md's
"Amending an accepted ADR", an in-place dated section is the established,
lighter-weight path for an addition that stays within the ADR's existing
decisions - used twice already in this same file - so this appends a third
dated entry documenting the new seam, its consumer, and why it reuses the
existing D1-D9-governed plan/invoke machinery rather than adding a second
one. No decision is reversed; no new Amends/Amended-by pair is needed since
the steps/gates/contributions axis already carries reciprocal links to
ADR-857 and ADR-894.

* fix(#4209): close two test-quality gaps trek-e's review found

Minor 1: validatePaths (a path-shape parser guarding the prompt-
injection/path-traversal trust boundary) had only example-based coverage,
violating ADR-456's rule that parsers/budget limits carry at least one
fast-check property test. Adds three: safe-segment paths are never
rejected, a single leading "../" always escapes the one-segment repoRoot,
and a control character anywhere is always rejected - one property per
rejection reason validatePaths owns.

Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only
ever exercised far below budget or at budget:0 (unbounded), never at the
exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds
three exact-boundary tests using the real estimateTokens/
buildSourceReviewPrompt the module calls internally, so the resolved
token count is exact rather than approximated: budget == estimate (must
pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must
pass).

Also extracts okPlan()'s fixture timeoutMs into a named constant -
local/no-adhoc-timeout-literal (#4446) landed on next after this branch
was authored and flagged the pre-existing literal on rebase; it is fixture
data for a synthetic plan object dispatchReviewerLanes never waits on, a
distinct class from tests/helpers/timeouts.cjs's real subprocess norms.

* fix(#4209): update docs-guard-registration baseline for the new ADR citation

reviewer-step-dispatch.test.cjs's new fast-check property tests cite
docs/adr/456-test-rigor-architecture.md in a justifying comment (never a
real read). lint-docs-guard-registration fingerprints every docs/ path
string an exempted test file mentions and fails on drift so a human
re-confirms the exemption still holds - re-confirmed, and the baseline is
updated to match.

* fix(#4209): point changeset pr: field at the fork PR for CI validation

changeset-lint's fail_pr_field_drift check compares the fragment's pr:
field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH),
not a fixed target. Rehearsing this branch on fork PR
davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's
pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but
fails here. Backfill to 4323 happens again, as the last commit, immediately
before the approved push to open-gsd#4323 - never leaving pr: 17 on the
branch that ships upstream.

* fix(#4209): reject promptChannel:none lanes from source-review dispatch

CodeRabbit found a real scope mismatch: coderabbit's lane declares
promptChannel: 'none' and reviews the working tree on its own terms,
fed nothing (review.md:367). Silently dispatching it through
dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope
buildSourceReviewPrompt promises and let the lane review whatever it
independently sees fit, violating this interpreter's own scoped,
metadata-only contract. Reject before plan()/invoke(), same as an
unresolved slug.

* fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file

CodeRabbit found the whole-file match on workflowContent would still
pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts
of this 1000+-line workflow, proving nothing about the actual evidence
block's contract. Line-filtered via splitLines (not a bare-\n regex
spanning readFileSync content) so this stays CRLF-portable and passes
local/no-unbounded-quantifier and local/no-crlf-fragile-split.

* fix(#4209): guard DISPATCH_JSON substitution and capture its stderr

CodeRabbit found the dispatch-step command substitution unguarded: a
non-zero exit could leave DISPATCH_JSON empty (or halt the step under
errexit with no warning), and the downstream reducer would only ever
report the generic unparseable_dispatch_output reason, discarding the
command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/
EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface
it in a warning on failure, and fall back to a parseable dispatch_
command_failed JSON stub so the reducer's existing reason-reporting
path still fires.

* docs(#4209): fix byte-for-behavior wording and missing colon, regenerate

CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the
established repo term for output-identical unchanged behavior) and a
missing colon after the bold "Optional external reviewer lanes (#4209)"
lead-in in docs/features/code-review-pipeline.md. Fixed in the two
hand-authored sources (commands/gsd/code-review.md, docs/features/
code-review-pipeline.md) and regenerated the two derived projections
(skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/
FEATURES.md via gen-features.cjs) so they stay in sync.

* fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI)

The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds
double-quoted JSON keys inside a single-quoted shell literal. That
extra quote density, inside an already quote-heavy ~8KB driver string,
passed bash -n and the full local suite on Linux but broke Windows
Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end
to end (#4209 round 5)` failed on two Windows CI shards with `bash -c:
unexpected EOF while looking for matching '''` — a Windows argv-to-
command-line re-quoting edge case, reproducible on rerun, not a flake.
Root-caused via gh api job logs plus a byte-identical local
reconstruction of the test's own driver script.

Fix: drop the fabricated stub. The downstream node -e reducer already
falls back to reason `unparseable_dispatch_output` on any JSON.parse
failure, so an empty/partial DISPATCH_JSON on command failure is still
handled correctly, with zero new quoting risk.

* revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI)

Two materially different mechanisms for the same CodeRabbit Nitpick
("Trivial | Quick win") both broke Windows Git-Bash reproducibly:
a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ...
matching '''") and, after removing that, a plain `head -1 "$VAR"`
inside a nested command substitution ("unexpected EOF ... matching
'"'"). Both passed bash -n and the full local suite on Linux every
time; both failed the SAME test deterministically on Windows CI. Two
attempts at the same class of fix (nested-quote construction near
this exact step) is the retry limit - reverting to the original,
already-shipped, Windows-verified unguarded form rather than
continuing to guess at a third quoting mechanism for a Trivial-
severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for
anyone attempting this again: the fix belongs outside this specific
markdown-fence-driver test harness (e.g., a real .sh helper script)
if it's worth doing at all.

* fix(#4209): backfill changeset pr: field to the real upstream PR before push

Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy
changeset-lint's PR-number check while rehearsing there; this is the
last commit before the approved push to the real upstream PR
(open-gsd/gsd-core#4323), so the field points at that PR number again.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 22:52:33 -04:00
Michel Moreira
733bec3ad1 enhance(#4261): report size-cap headroom on every run, with a reserved margin (#4418)
* enhance(#4261): report size-cap headroom on every run, with a reserved margin

The tier caps are red lines and none of them moves here. What was missing is
everything below the red line: a passing run said nothing, so a contributor
at 99.6% of a cap and one at 60% got identical feedback, and the density that
produces merge-time collisions was invisible to the people creating it.

Two levels, matching the shape execute-phase.md already carries by hand (a
hard ceiling plus a lower margin "so minor future edits don't re-trip the
gate") and which was until now the only capped file with one:

  1. a headroom census printed every run, green included, sorted
     least-headroom-first, and appended to the GitHub job summary
  2. a 95% reserved margin that names the files inside it and REPORTS
     rather than fails

The margin deliberately does not fail. A cap breach is a red line; a file at
96% is not broken, it is a file whose next contributor should extract before
adding. Failing there would create a second red line and force exactly the
+N bumps the policy forbids. Neither level asserts a count, so this adds no
snapshot to regenerate — the per-file size baseline was deleted by #2724 for
conflicting on 7 of 7 PRs that touched it.

Also deletes rather than refreshes the per-tier high-water comments in both
guard files. They were measured once and then diverged from the tree: the
LARGE line still claimed "gsd-executor 42,342 -> ~6.8 KB headroom" while the
real high-water sat at 99.6% of that cap, so the comment documenting the
margin was itself why nobody noticed the margin was gone.

As measured on next by the new census, the pressure has grown since the
issue was filed: gsd-plan-checker.md has 9 bytes of headroom, gsd-verifier.md
21, and plan-phase.md 14.

* chore(#4261): add changeset for the size-cap headroom census

* test(#4261): exercise reserved-margin boundaries

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 22:04:48 -04:00
Tom Boucher
93e141a006 enhance(#4139): Phase 4 — measure the window instead of asserting it (#4502)
* enhance(#4404): add offline token benchmark for compact-content splits

ADR-4139 Decision 2 requires the finite-attention justification for
workflow.compact_content to be measured, not asserted. `npm run
benchmark:compact-content` computes, per registered spine/detail split
discovered under gsd-core/workflows/, the token count with the split
active (spine alone) vs inactive (spine + all detail parts read back
in), using gpt-tokenizer (pinned exact devDependency — Anthropic
publishes no tokenizer for Claude 3+, so every output surface labels
this a PROXY-TOKENIZER comparison: the on/off delta is exact under one
tokenizer applied identically to both sides, the absolute counts are
not Claude's real ones).

Reporting-only by design and verified so: --check diffs the live
recompute against a committed baseline (tests/fixtures/compact-content-benchmark-baseline.json)
and prints drift, but never exits non-zero for a drifted or missing
baseline — the only thing allowed to fail this script is a genuine I/O
error reading a source .md file it's measuring. Not wired into lint:ci
or pretest.

Discovery is deliberately reimplemented rather than importing
tests/helpers/compact-content-split.cjs (Phase 3, #4403), keeping a
scripts/ reporting tool from depending on a test-only module.

tests/fixtures/deny-network.cjs preloads via NODE_OPTIONS=--require to
prove the benchmark makes no network call, monkeypatching http/https/
net/dns/fetch to throw rather than relying on sandboxing.

docs/CONFIGURATION.md documents the new benchmark against the
workflow.compact_content key to satisfy this repo's docs-required gate
for an Added-type changeset.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4404): address orthogonal review findings on the token benchmark

Standards axis found two hard violations against documented rules:
- CLAUDE.md's Generative Fix Divergence rule requires a parity assertion
  for shared discovery logic maintained in two places. Added a test
  comparing benchmark-compact-content.cjs's own discoverRegisteredSplits
  against tests/helpers/compact-content-split.cjs's version on the real
  repo tree, so the two can never silently drift apart.
- The changeset body closed its bold span with a period and continued
  as a second sentence, instead of the canonical
  `**<phrase>** — <explanation>.` shape CONTRIBUTING.md documents.

Spec axis found the "network disabled + identical output across two
runs" Done-when criterion was verified as two separate properties
(determinism tested without network denial, offline survival tested as
a single run) rather than as one combined property. Added a test that
runs the benchmark twice under the deny-network preload and asserts
byte-identical stdout.

Security axis found tests/fixtures/deny-network.cjs didn't patch
dns.promises (a separate binding from the callback dns API), tls.connect,
or http2.connect — inert today since nothing in the benchmark calls
them, but a silent gap in what the preload's own header claims to
guarantee. Patched all three.

CLAUDE.md's Property-Based Testing rule also requires a fast-check test
for budget-limit arithmetic; added one for computeAggregate's off/on
summation (true sum over N splits, never NaN/Infinity, never exceeds
100% when off >= on for every split).

Standards axis's remaining two findings (a Data Clumps observation on
the {offTokens, onTokens, reductionPct} triple, and mild duplication in
formatDriftReport's three line-formatters) are left as judgement calls:
introducing a named type for a 3-field local tuple, or a formatter
abstraction for three short lines, would be exactly the premature
abstraction CLAUDE.md's engineering guidance warns against for a script
this size.

All changes verified directly (parity logic, the fast-check property,
and the three newly-denied network surfaces actually throwing under the
preload) via node -e before committing; full npm run lint:ci passes
with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4404): backfill changeset pr number to 4502

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 19:19:03 -04:00
Tom Boucher
a0f8f956c4 enhance(#4139): Phase 3 — partition rules + the five checks (#4497)
* enhance(#4139): Phase 3 — partition rules + the five checks

ADR-4139 Decision 5, epic #4139 Phase 3. Issue #4403's own "Proposed behavior"
section lists four checks; the ADR's Decision 5 and its own phase table ("partition
rules + the five checks") list five — the same four plus "boundary moves are
declared, ongoing". Same issue-vs-ADR drift Phase 2 hit on the detail.md vs
detail/*.md layout: the ADR is the locked, reviewed document, so it wins. This PR
implements all five.

docs/PARTITION-RULES.md (new) is the partition-rules document: the partition rule
itself, the protected-content list and <!-- gsd:protected --> sentinel syntax
(relocated unchanged from gsd-core/references/compact-content-protected-content.md,
now deleted — it was never referenced by any runtime workflow Read, only by the
predecessor test as documentation, so nothing at runtime regresses, and removing it
from gsd-core/references/ also drops it from all 19 installed-project shipped-content
trees for a file nothing ever read), and the five checks explained for a human
reader. Referenced from a new CONTRIBUTING.md subsection under "Editing shipped
content".

tests/helpers/compact-content-split.cjs (new) is the shared mechanics: split
discovery (any gsd-core/workflows/<name>/detail/*.md paired with <name>.md — no
registry file, a pair is registered by existing on disk), line normalization
(carries forward Phase 2's bare-label-line isTrivial fix and the canonical
gsd_run-launcher-preamble exclusion), sentinel extraction, and a
Boundary-Move-Declared commit-trailer reader that is a direct structural port of
tests/helpers/emitted-runtime.cjs's Emitted-Drift-Ack-Hash/-Growth trailer reader
(ADR-3942) — same merge-base range, same fail-closed throw on an uncomputable range,
same dedupe/conflict rules.

tests/compact-content-partition-guard.test.cjs (new) is the actual guard, superseding
tests/plan-phase-compact-split.test.cjs (deleted — its per-pair checks are now the
general guard's job for plan-phase specifically). Checks 2 (disjointness) and 3
(registration + size cap) run unconditionally against every registered split. Checks
1 (completeness, fires once per split on the PR that introduces a new detail/ path),
4 (protected content — no trailer can ever excuse this one, unlike check 5) and 5
(boundary moves declared) are PR-diff-scoped against the resolved base ref and skip
cleanly when there's nothing to compare (a fresh clone, no PR in flight) — a
deliberate asymmetry from check 5's trailer reader, which must throw rather than
silently pass when ITS range is uncomputable, since that function is answering "did
this PR declare its moves" rather than "is there even a diff to look at". Each of the
five checks carries a RED (deliberately broken fixture) / GREEN (fixed) test pair,
built against synthetic temp files or real throwaway git repos, per this repo's rule
that a guard nobody has seen go red is not yet a guard. Building the real fixtures
caught and fixed one real bug before it shipped: check 4's line-presence test was
using the trivial-line-filtered normalizer, so a byte-identical spine falsely
reported its own protected code-fence line as "deleted" — fixed with a
non-filtering membership check.

Extending docs/INVENTORY.md's "Workflow Sub-Files" table for `detail` surfaced a
pre-existing, unrelated gap in the SAME area: gsd-core/workflows/<name>/templates/*.md
is a fourth workflow sub-file kind that already existed on disk and was already known
to lint-response-language-coverage.cjs's FRAGMENT_DIRS, but was invisible to
gen-inventory-manifest.cjs and undocumented in that table. Fixed alongside it, same
pattern, same PR, rather than deferred.

Also, mechanically required by the new fourth sub-file kind:
- scripts/lint-response-language-coverage.cjs: `detail` added to FRAGMENT_DIRS
  alongside modes/steps/templates — a detail/<part>.md inherits its parent's
  response_language coverage through the same per-file proof, not a parallel one.
- tests/workflow-size-budget.test.cjs: explicit regression test locking that
  detail/ files are governed solely by the hard, non-waivable NEW_FILE_CAP
  (tests/helpers/emitted-diff.cjs) and never by the XL/LARGE/DEFAULT spine tiers —
  true by construction (measureWorkflows/listWorkflowStems don't recurse), made
  explicit per the issue's own Done-when item rather than left true-by-omission.
- scripts/gen-inventory-manifest.cjs: `workflow_detail` and `workflow_templates`
  NESTED_FAMILIES entries; docs/INVENTORY-MANIFEST.json regenerated
  (plan-phase/detail/elaboration.md, discuss-phase/templates/*.md now tracked);
  docs/INVENTORY.md's table updated to four kinds.

Verified: `npm run lint:ci` clean with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4403): review findings + a real gsd-test failure in the new guard

Two orthogonal review passes (Standards + Spec, isolated sub-agents) plus a
separate security review ran against the prior commit. Fixed everything each
surfaced:

- Security (Low, path-traversal existence oracle): checkRegistration's
  dangling-reference check extracted detail-path-shaped substrings from spine
  PROSE via a regex that permits `.`/`/` freely, then joined them onto repoRoot
  and probed fs.existsSync with no containment check — a spine file containing
  `../../../etc/detail/passwd.md`-shaped text could make the guard test file
  existence outside the repo. Added a path.relative-based containment check
  before the fs.existsSync call; anything that resolves outside repoRoot is now
  reported as a dangling reference directly, never probed on disk.
- Standards (Boundary Coverage): the size-cap fixtures covered NEW_FILE_CAP and
  NEW_FILE_CAP-1 but not NEW_FILE_CAP+1 — added the third boundary-point case
  CLAUDE.md's TEST RULES require (limit-1/limit/limit+1).
- Standards (Property-Based Testing): extractProtectedBlocks (a sentinel
  parser) and the new parseBoundaryMoveTrailerValues (a declare/dedupe/conflict
  parser, bijective-shaped) had no fast-check property test. Added three: a
  render/parse bijectivity property for the trailer parser (mirroring the exact
  ADR-3942 sibling test's alphabet/idiom), a dedupe-is-idempotent property for
  the same parser, and a well-formed-sentinel-round-trips property for
  extractProtectedBlocks.

Then dispatched gsd-test on the resulting commit. It found a real bug the
reviews couldn't have caught (none of them can run inside gsd-test's sandbox):
checks 4/5's real-repo assertion failed against plan-phase's own split,
reporting DISK_PLANS/#3218-comment lines as "undeclared boundary moves" —
content Phase 2 (#4402) legitimately moved into detail/elaboration.md months
before this PR's Boundary-Move-Declared mechanism existed to require a
trailer for it. Root cause: `resolveBase()`'s own doc comment already documents
that no `origin/*` remote-tracking ref exists inside the gsd-test sandbox
container, and its fallback candidate (a bare `next` branch) can resolve to a
point in history that predates an already-merged, already-reviewed split —
making that split look "newly introduced" from the sandbox's vantage point.
Check 1 (completeness) already scopes itself correctly to only genuinely-new
detail paths (git diff status 'A'); checks 4 and 5 did not share that scoping,
so a stale base made them re-litigate a settled split retroactively. Fixed by
having checks 4/5 skip any split name check 1 already counted as newly-split —
their own premise ("did an EXISTING split shed/undeclare something") does not
apply to a split that is, from the resolved base's vantage point, brand new;
that is check 1's domain alone. Verified locally (25/25 tests pass via a
direct `node -e` require, since `node --test` is blocked in this repo) and via
re-reasoning through the exact real-repo scenario the gsd-test failure showed.

Also regenerated all 19 tests/fixtures/install-tree/*.json goldens — the
prior commit's deletion of gsd-core/references/compact-content-protected-content.md
was never reflected there, which is what golden-install-tree.test.cjs's other
19 failures in the same gsd-test run were.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4403): backfill changeset pr number to 4497

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4403): isolate codex-config.test.cjs into its own chunk, root-causing the Windows CI failure

PR #4497's "full test (windows-latest, 24, shard 2/3)" job failed: run-tests
killed chunk 3/8 at the 600s per-chunk backstop, with codex-config.test.cjs
(weight 17.87, by far the chunk's dominant cost) packed alongside 39 other
files. Traced, not assumed:

- scripts/run-tests.cjs's own timeout-headroom comment for the OUTER
  per-shard timeout documents that "adding one test file reshuffled 115 of
  268 unit files between shards" — shard/chunk composition is architecturally
  known to be unstable to single-file additions, which is exactly what this
  PR's own new tests/compact-content-partition-guard.test.cjs is.
- A second comment, dated 2026-09-06 (one day before this PR, PR #4428's own
  CI), already documents the SAME chunk hitting the SAME 600s backstop with
  the SAME file (codex-config.test.cjs, "a genuinely MEASURED weight of
  17.87 — not a stale-table miss") dominating it — the fix then was cutting
  the Windows per-chunk budget from 60 to 40. That cut clearly was not
  enough: two documented incidents in two days, at two different budget
  settings, both centered on one file that alone consumes ~45% of even the
  reduced Windows budget.
- tests/test-timings.json's own header confirms its source data
  (test-events-linux-node22/24.jsonl) is Linux-only, and run-tests.cjs's own
  chunk-timeout diagnostic already prints "real Windows cost runs ~2.2x the
  recorded figure" — the packer's weight-balancing is working off data that
  is both stale (table last regenerated 2026-08-07) and known to
  underestimate the platform where the failure occurs.

Given codex-config.test.cjs is disproportionately heavy AND every companion
sharing its chunk is decided by a packing algorithm already documented as
reshuffling unpredictably on any new file, tuning the shared budget a third
time only moves the marginal line to wherever the next new file happens to
land — it does not remove the gamble. Isolating codex-config.test.cjs into
its own dedicated single-file chunk, unconditionally and on every platform,
removes it at the source: the file never enters the pool packChunks balances,
so no other file's packing changes, and no future single-file addition
(mine or anyone else's) can silently reintroduce this exact failure by
landing in its chunk.

Extracted as a small pure function, partitionIsolatedFiles (mirroring this
file's existing pattern of pulling packing/analysis logic out of main() for
in-process unit coverage — see computeSweepProtectSet, analyzeChunkEvents),
with 6 new tests in tests/run-tests-harness.test.cjs covering basename
matching across path separators, near-miss non-matches, the empty-list case,
and the isolated-set contents.

Root cause is now closed rather than papered over with a retry: this failure
is a property of one specific heavy file's chunk placement, not something
that recurs randomly. If codex-config.test.cjs itself is ever genuinely sped
up, this isolation can be revisited — this is a packing-side mitigation for
a known file's cost, not a claim the cost is irreducible.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 16:44:53 -04:00
Tom Boucher
8b7a0b696b enhance(#4139): Phase 2 — one shared gate, one pilot split, one accuracy spot-check (#4471)
* enhance(#4402): split plan-phase into a spine + detail, add the shared compact-content gate

ADR-4139 Decisions 3-5, Phase 2 of the #4139 Compact Content epic. Pilot
split for plan-phase.md, the largest of the 58 eagerly-@-included workflow
files (98,290 bytes): the spine keeps every happy-path step, every
protected-content block (planner/checker prompt templates, quality gates,
the failing-direction few-shot example, the two ScheduleWakeup guardrail
paragraphs — each marked with a <!-- gsd:protected --> sentinel), and
condensed one-paragraph summaries of five rare/opt-in fallback paths
(planner and checker filesystem-hang recovery, phase-split recommendation,
source-audit gaps, the thinking-partner conditional, and plan bounce). The
full text of those five moves verbatim to gsd-core/workflows/plan-phase/detail.md
(9.9KB, well under the 32,768-byte NEW_FILE_CAP), read by the spine only
when workflow.compact_content is false (the default) — the exact same
resolution rule now stated once in the new shared
gsd-core/references/compact-content-gate.md, which every future split
references instead of restating.

Verified mechanically (tests/plan-phase-compact-split.test.cjs, scoped to
this one split — Phase 3/#4403 owns the generalized guard): the union of
spine + detail contains every non-trivial line the parent commit carried
(0 missing), no non-trivial line is duplicated between them (0 duplicated),
and every declared protected block is well-formed and non-empty. The spine
shrinks from 98,290 to 93,206 bytes (-5.2% of the eager-window cost this
epic exists to reduce); detail.md's 9,853 bytes are only ever paid by a
project that has NOT opted in.

Verified live, end to end, twice, against this actual repo (not a
synthetic fixture) — real gsd-planner and gsd-plan-checker subagent
spawns, real PLAN.md output:
- workflow.compact_content=false: planned a real disposable phase
  (a docs/how-to page for enabling the key itself); planner returned
  PLANNING COMPLETE, checker returned VERIFICATION PASSED, all fact-checks
  against real repo state confirmed.
- workflow.compact_content=true (detail.md never read): planned a second
  real disposable phase; planner returned PLANNING COMPLETE with
  frontmatter.validate and verify.plan-structure both clean, again fully
  grounded against real repo state. The five condensed fallback sections
  were independently re-read spine-only and confirmed sufficient to act on
  correctly without detail.md's elaboration.

Also drafts gsd-core/references/compact-content-protected-content.md — the
protected-content category list and <!-- gsd:protected --> sentinel syntax
ADR-4139 Decision 5 calls for, written to move to Phase 3 (#4403) unchanged
once it lands there.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4402): move detail.md into the ADR-4139-mandated detail/ subdirectory

Two independent review sub-agents (Standards and Spec axes of /code-review)
caught the same structural defect: ADR-4139 Decision 6 mandates
gsd-core/workflows/<name>/detail/*.md ("one or more parts... individually
skippable"), and this PR had shipped a flat plan-phase/detail.md instead,
copying issue #4402's own (inconsistent) restatement rather than the
locked ADR text. Fixed by git-mv to plan-phase/detail/elaboration.md and
updating every cross-reference (the spine's step 0.5 gate pointer, the
shared compact-content-gate.md's own resolution-rule wording, and the
completeness test's path constants).

Also, from the same review pass:
- docs/CONFIGURATION.md and gsd-core/references/planning-config.md's
  workflow.compact_content rows said "nothing branches on it yet" — no
  longer true now that plan-phase.md's spine does. Updated both to name
  plan-phase as the pilot and note the rest of the corpus is still pending.
- Regenerated all 19 tests/fixtures/install-tree/*.json golden fixtures
  (npm run gen:install-tree) — the three new shipped files were missing
  from the installer emitted-tree goldens.
- Found via a cache-busted `eslint . --max-warnings 0` (this repo's
  eslint --cache has produced false-greens before): the split test's
  `git show` call had a bare `timeout: 10000` literal, tripping
  local/no-adhoc-timeout-literal. Extracted to the existing GIT_TIMEOUT_MS
  constant from tests/helpers/timeouts.cjs instead of a second guessed
  copy of the same class of timeout.

Verified NOT needed, by tracing the actual mechanism rather than asserting
(tests/helpers/emitted-provenance.cjs's gsd-core-verbatim rule attributes
every gsd-core/{workflows,references}/** path to itself as an identity
source): an Emitted-Drift-Ack-Hash/-Growth trailer. Every changed/added
path in this diff is hand-authored and present in the diff itself, so
diffEmitted's attribution loop resolves `via` to the path's own source
before ever reaching the ack-lookup branch — there is no unattributed
delta to acknowledge. The spine also shrank (98,290 to 93,206 bytes), so
the growth ratchet has nothing to ack either.

Re-verified after these changes: the completeness/disjointness self-check
(0 missing, 0 duplicated) still holds against the relocated detail file,
and a full `npm run lint:ci` passes clean with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4402): restore literal content the pre-existing drift guards pin on

The first gsd-test run against this split (19 failures) surfaced real
regressions: several pre-existing structural guards pin the EXACT text
of the sections this split condensed, and paraphrasing broke them.

- tests/plan-phase-drift-guard.test.cjs expects the literal
  `DISK_PLANS=$(gsd_run query find-phase ...)` bash assignment inside
  plan-phase.md itself, not a prose description of the same check.
  Restored the exact line into both §9a and §11a's spine summaries.
- tests/thinking-partner.test.cjs expects plan-phase.md to literally
  offer "No, I'll decide" as the skip option. Restored that exact
  phrase into the condensed thinking-partner paragraph.
- Both restores would have duplicated the same text into
  plan-phase/detail/elaboration.md (which still carries the full
  elaboration). Removed the now-redundant restatements from the
  detail file instead of leaving them duplicated — the spine already
  computes DISK_PLANS before the detail elaboration is ever read, so
  the detail file references it rather than recomputing it.
- Re-running scripts/sync-runtime-launcher.cjs after that edit found
  the canonical gsd_run preamble had also become an unintentional
  spine/detail duplicate (both files call gsd_run and each is
  required, by runtime-launcher-parity's own contract, to carry its
  own copy). That's sanctioned duplication under a DIFFERENT
  contract, not lost/copy-pasted content, so
  tests/plan-phase-compact-split.test.cjs now excludes it from the
  disjointness check the same way it already excludes trivial
  fences/headings.
- Applied the adversarial-review finding on tests/plan-phase-compact-split.test.cjs's
  own isTrivial(): a blanket `line.length <= 15` cutoff silently
  swallowed real content (e.g. the 14-char `<quality_gate>`
  sentinel). Replaced it with a specific bare-label-line pattern
  (`Options:`, `Display banner:` etc.) — verified 0 missing / 0
  duplicated against the actual split, an improvement over both the
  original cutoff and a naive full removal (which produces
  false-positive "duplicates" on generic recurring labels).
- gsd-core/references/planning-config.md's own workflow.compact_content
  row used `/gsd-plan-phase` (hyphen). That file is Claude-facing
  source text (gsd-core/references/), which tests/slash-command-namespace.test.cjs
  requires in colon form; docs/CONFIGURATION.md's use of the hyphen
  form is correct as-is since docs/ is human-facing and outside that
  test's scanned directories. Fixed to `/gsd:plan-phase`.
- tests/plan-phase-compact-split.test.cjs's own `git show` of the
  parent commit failed inside the gsd-test sandbox ("detected dubious
  ownership") because the checkout is mounted under a UID the
  invoking user doesn't own. Scoped `-c safe.directory=<repo-root>`
  to that one git invocation rather than touching global git config.
- docs/INVENTORY.md still had one outstanding "detail.md part" wording
  fix from the earlier adversarial-review pass, staged now.

Re-verified locally against the exact assertions in all four affected
test files (all pass) before dispatching a fresh gsd-test run — no
change here should have broken any of the other 18 gates; `npm run
lint` is clean with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4402): restore the full marker enumeration to §9a's spine trigger line

The isolated Spec-axis review flagged that §9a's "Triggered when" line was
condensed to "Agent() returns but the return contains no recognized
marker" — dropping the literal `## PLANNING COMPLETE` / `## PHASE SPLIT
RECOMMENDED` / `## ⚠ Source Audit` / `## CHECKPOINT REACHED` /
`## PLANNING INCONCLUSIVE` enumeration, which is exactly the "machine-
parsed structural headings" category compact-content-protected-content.md
lists as protected. The load-bearing use of that same list (the
gsd_stall_watch call and the Handle Planner Return bullets a few lines
above) was never touched — only this one descriptive restatement was
genericized — but leaving any instance of a protected category
unsentineled is the silent erosion ADR-4139 Decision 4(c) warns
sufficiency isn't machine-checkable enough to catch on its own. Restored
the full enumeration into the spine.

That reintroduced an exact duplicate into plan-phase/detail/elaboration.md,
which still stated the same trigger sentence verbatim. Reworded the
detail file's version to reference the spine's trigger condition instead
of restating it, since the spine is now the single place that sentence
lives in full — mirroring the DISK_PLANS/"already computed above" pattern
from the previous commit.

Re-verified locally: completeness/disjointness (0 missing, 0 duplicated)
and all previously-fixed literal-content assertions still hold.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4402): backfill changeset pr number to 4471

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 12:18:38 -04:00
Tom Boucher
476394689a fix(#4254): pin sequential executor to the orchestrator's validated root (#4476)
* test(#4254): sequential executor root pin — failing-first regression + matrix

The new suite executes the shipped supplied-root-pin guard against real git
fixtures (drifted primary-checkout cwd halts before the write and the FATAL
names both roots; matching cwd permits it; unexpanded/empty pins halt;
normalization forms; submodule and sibling boundaries; metacharacter quoting;
drive-letter form gate) and locks the dispatch contract across execute-phase.md,
its sequential-root-pin step fragment, and worktree-path-safety.md. The #2772
per-plan serialization assertion retargets to the fragment that now carries
those rules (ADR-857 Phase 6 ceiling), plus the host-step wiring.

* fix(#4254): pin sequential executor to the orchestrator's validated root

Sequential-mode dispatch told the executor to self-derive PROJECT_ROOT from its
own cwd; every existing guard is worktree-mode-only or self-referential, so an
executor spawned with a drifted cwd committed onto the wrong checkout silently.

- worktree-path-safety.md step 0p: mode-agnostic supplied-root pin guard,
  composed by the orchestrator at build time with the literal $ORCHESTRATOR_WT
  (git-vs-git comparison on both sides — representation-safe on Windows, the
  #4296 lesson), fail-closed on empty/unexpanded pins, registered-submodule
  allowance, warn-and-proceed only when the dispatch carries no pin block.
- execute-phase.md sequential branch: build-time embed of the bound
  <project_root_pin> via the new execute-phase/steps/sequential-root-pin.md
  fragment (ADR-857 Phase 6 frozen ceiling — the host step cannot grow; the
  wave serialization rules move with the fragment, verbatim in substance) plus
  the per-write/commit pin instruction in <sequential_execution>. Worktree-mode
  dispatch untouched (its self-derived toplevel IS correct there).
- INVENTORY rows (5 locales) + INVENTORY-MANIFEST + install-tree goldens
  regenerated for the new fragment; changeset added.

* chore(#4254): backfill changeset PR number

* fix(#4254): accept backslash-separated Windows drive pins

CI on windows-latest showed every permit-path test failing with
"Actual root: <none>": pins composed from Node's path.join arrive in the
backslash drive form (C:\Users\RUNNER~1\...), which the guard's absolute-form
gate rejected before the cwd-side root was ever computed — a legitimate
matching pin could never pass. The gate now accepts either separator
([A-Za-z]:[\\/]); git -C resolves both forms (and 8.3 short names) to the
same canonical toplevel, so the git-vs-git comparison is unaffected. Form-gate
tests cover the emitted (C:/…) and produced (C:\…) spellings plus short names.

* fix(#4254): portable drive-form gate for MSYS bash

The bracket class [\\/] that accepted backslash drive pins parses
inconsistently on MSYS bash (the Windows CI leg still rejected C:\ pins —
every permit-path test red with "Actual root: <none>"). Replace it with
standard pattern escaping outside brackets: [A-Za-z]:/*|[A-Za-z]:\\* —
the escape form is version- and build-portable. Verified across all forms:
both drive spellings accepted; bare "C:", relative, empty, and unexpanded
rejected.

* fix(#4254): runtime-generated backslash comparator + self-describing FATAL

The Windows CI legs failed every #4254 permit-path row with
'Actual root: <none>' across two prior pattern spellings ([\\/] and \\*).
Stage misattribution: <none> appears whenever the FATAL fires BEFORE the
cwd-side capture assigns ACTUAL_ROOT — the absolute-form gate was what fired.

Mechanism: the test harness spawns bash -c <script> through the Windows
command-line boundary; that round-trip applies one extra shell-quoting pass
with double-quote semantics — a backslash written twice in the script text
arrives halved, while a lone backslash survives (the pin displays intact;
row 9's pure-bash gate independently showed the halved pattern rejecting
C:\ while C:/ still passed its surviving arm). On windows-latest every pin
carries backslashes (os.tmpdir() is the 8.3 short form C:\Users\RUNNER~1\...),
so the gate ate every pin before the actual root was ever computed.

Fix, robust by construction:
- the drive-form gate generates its backslash comparator at RUNTIME
  (BS=$(printf '\134'); match [A-Za-z]:"$BS"*) — the shipped guard now
  contains no doubled backslash anywhere, enforced by a regression
  assertion on the extracted guard text;
- the FATAL self-describes: Guard stage (pin-unbound / form-gate /
  actual-capture / pinned-capture / root-mismatch) plus a Diagnostic line
  carrying git's own stderr for capture failures and both compared values
  for mismatches — future platform failures name their stage in the log;
- row 9's hand-rolled duplicate case gate (transit-fragile copy, #4296
  Minor 1 duplication smell) is replaced by driving the SHIPPED guard and
  asserting the stage; rows 2/4 pin the new stage machinery.

Validated on darwin across drift/match/relative/unbound/empty/bare-drive/
forward-and-backslash drive forms, each also re-run under a simulated
Windows transit (every doubled backslash halved) with identical outcomes.

* fix(#4254): close the empty-comparator fail-open seam in the drive-form gate

Self-review of the runtime-generated backslash comparator: if printf's
octal escape ever returned empty, the drive arm [A-Za-z]:"$BS"* would
widen to drive-RELATIVE pins (C:foo) — the construction's one theoretical
fail-open path. Fail closed with a self-describing diagnostic instead of
trusting the shell's printf.

---------

Co-authored-by: sim <sim@local>
2026-09-07 10:54:30 -04:00
Tom Boucher
c4b6dbd486 fix(#4247): refuse update-plan-progress on a roadmap with no writable phase entry (#4468)
* test(#4247): failing-first regressions for checklist-form update-plan-progress

* fix(#4247): refuse update-plan-progress when the roadmap has no writable phase entry

* fix(#4247): single local source for the phase-heading anchor grammar

* docs(#4247): note the missing_phase_details refusal in cli-tools reference

* docs(#4247): backfill pr number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-07 02:20:47 -04:00
Brenden Smerbeck
e54d3aa159 enhance(#4401): register workflow.compact_content as a validated config key (#4441)
* feat(#4401): register workflow.compact_content as a validated config key

- Add compact_content: false to the nested workflow object in
  gsd-core/bin/shared/config-defaults.manifest.json
- Add 'workflow.compact_content': false to SCHEMA_DEFAULTS in src/config.cts
  so an absent key resolves to false via config-get --raw
- validKeys entry in config-schema.manifest.json already present

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(#4401): behavioral and boundary tests for workflow.compact_content

- 19 behavioral tests covering config-set/config-get round trip, invalid-shape
  rejection (banana, 42, empty string), the corrected null-unset semantics
  (#2046), absent-key resolution against config-defaults.manifest.json,
  config-new-project wiring, and doc-row shape assertions
- Drops the install-tree fixture-parity block (and its docstring item) that
  asserted gsd-core/references/compact-content-gate.md and
  gsd-core/workflows/compact/map-codebase.md fixture entries — those paths
  belong to #4402 and do not exist on this filtered branch

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(#4401): document workflow.compact_content in both config references

- One 4-cell row in docs/CONFIGURATION.md (workflow.* run)
- One 5-cell row under Workflow Fields in gsd-core/references/planning-config.md
- Both cross-reference ADR-4139

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(#4401): add changeset

- Added-type fragment, pr: 4401 (issue number; backfill to the real PR number
  is a required follow-up once the PR is opened, per D-08 and CHANGESET-PR-
  FIELD-DRIFT)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(#4401): backfill changeset pr field to #4441

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(#4401): derive workflow.compact_content default from CONFIG_DEFAULTS

SCHEMA_DEFAULTS['workflow.compact_content'] hardcoded the literal false
instead of deriving it from CONFIG_DEFAULTS the way 3 of its 8 sibling
entries do (smart_zone_tokens, pr_strict, inline_plan_threshold), leaving
a single-source-of-truth drift risk: a future manifest-only edit to the
default could silently diverge from this literal, only caught later by
the D-03 test if it ever happened to manifest.

Adds compact_content to CONFIG_DEFAULTS in src/config-loader.cts and
derives SCHEMA_DEFAULTS from it in src/config.cts, matching the majority
sibling pattern. Found during maintainer review (review-open-prs) of
this PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4401): map compact_content in config-field-docs NAMESPACE_MAP

The previous commit added compact_content to CONFIG_DEFAULTS in
src/config-loader.cts but missed the matching entry in
tests/config-field-docs.test.cjs's NAMESPACE_MAP, which maps flat
CONFIG_DEFAULTS keys to their namespaced doc form before checking
gsd-core/references/planning-config.md for a match. Without it, the
test looked for a bare `compact_content` doc reference instead of the
actual `workflow.compact_content` row, and failed:
"CONFIG_DEFAULTS keys missing from planning-config.md: compact_content".

Found by actually running gsd-test against the branch rather than
trusting the plausible-looking fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4401): register compact-content-4139 test in the docs-guard lane

tests/compact-content-4139.test.cjs's D-06 tests read docs/CONFIGURATION.md
directly (fs.readFileSync) to assert the workflow.compact_content doc row's
shape, which makes it a doc-reading test file under the #3753 docs-guard
lane. It was never added to scripts/docs-guard-registry.cjs's
DOCS_GUARD_TESTS map and carries no docs-guard-exempt marker, so
tests/ci-docs-guard-registry.test.cjs's registration lint correctly failed:
"compact-content-4139.test.cjs reads a docs/ path but is not registered in
the docs-guard lane and carries no docs-guard-exempt marker".

Registers it with ['docs/CONFIGURATION.md'] (the only real docs/-prefixed
path it reads; gsd-core/references/planning-config.md is outside this
registry's docs/ scope, matching the sibling config-field-docs.test.cjs
entry's existing convention).

Found by actually running gsd-test against the branch — this gap predates
the maintainer's config-loader.cts fix and was already present in the
original PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: sim <sim@local>
2026-09-06 19:52:59 -04:00
Tom Boucher
2cf119f57e fix(#4217): reconcile artifacts before classifying an abnormally-ended executor (#4442)
* fix(#4217): reconcile artifacts before classifying abnormal ends

* test(#4217): pin the completion-reconciliation contract

* chore(#4217): regen derived inventory and install-tree fixtures

* test(#4217): follow the #4003 anchoring pins into the reconciliation fragment

Emitted-Drift-Ack-Growth: execute-phase.md — the runtime-neutral completion-reconciliation pointer, the two Codex wait-rule bindings, and the step-7 reconcile-first gate net +33 bytes over the extracted fallback block (#4217)

* chore(#4217): add changeset fragment

* chore(#4217): backfill PR number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-09-06 18:57:38 -04:00
Michel Moreira
19b66c3ec8 fix(#4218): stop the orchestrator steering an executor that is still working (#4391)
* fix(#4218): stop the orchestrator steering an executor that is still working

An executor with recent RED/GREEN/REFACTOR commits and passing verification had
not yet written its SUMMARY because it was finishing closeout. The parent saw no
local OS test/build process, inferred an "idle tail", and sent "Finalize
immediately" into a working child; in CLI runs the same inference interrupted an
executor before GREEN, leaving a RED commit and an uncommitted edit.

The stall block said only "if no completion signal, no SUMMARY.md, and no
expected-branch commits appear for N minutes" — it never said what to do when
commits DO exist and only the SUMMARY is outstanding, never defined the
threshold as a period without progress rather than a total runtime, and never
ruled out a process listing as an idleness signal. Four rules close that:

- the threshold measures a period WITHOUT MEANINGFUL PROGRESS, from the last
  sign of progress, not from dispatch — a long verification tail is not a stall;
- commits + missing SUMMARY + recent activity resolves to KEEP WAITING, with
  steering, interrupting and re-dispatching each named and forbidden;
- urgency/finalization messages ("Finalize immediately" and family) are
  forbidden outright — they arrive mid-verification and truncate a correct run.
  The existing user-facing pause is the only sanctioned stop, and `kill and
  retry` is a clean restart, not a nudge;
- the absence of a local OS test/build process is NOT idleness: a native
  subagent runs in the runtime's own session, and an executor between two tool
  calls shows no process at all. Progress is judged only by the signals this
  workflow names.

Five prose-contract assertions in tests/execute-phase-wave.test.cjs, all red on
next.

* fix(#4218): extract the progress policy to a step fragment

CI's #1168 gate caught it: execute-phase.md sits 77 bytes under a frozen 93600
ceiling and the four rules added ~2.3 KB. "Extract, not bump" is the repo's
stated remedy, and this workflow already carries policy detail that way.

execute-phase/steps/executor-progress-policy.md owns the policy. The
worktree-recovery arm moved with it — `kill and switch to inline execution`
qualifies the stop this policy governs, so it belongs beside the rule about when
stopping is sanctioned at all, not stranded in the host. The #3212 recovery
OPTIONS stay in the host, where tests/config.test.cjs pins them.

The host keeps what must be read before the orchestrator acts: the verdict, the
threshold definition, and a pointer that fires before any message is sent to the
child. execute-phase.md is now 93475 bytes — 48 SMALLER than next.

* chore: add changeset for #4218

* chore(#4218): regenerate the inventory manifest for the new step fragment

docs/INVENTORY-MANIFEST.json is the authoritative per-file list behind
INVENTORY.md's `<workflow>/steps/*.md` row, so a new fragment has to appear
there or gen-inventory-manifest --check reds the lint-tests lane.

* chore(#4218): restore the issue ref on the allow-test-rule marker

ADR-456 requires a #NNN on a new exemption; the block rewrite that moved the
policy into the fragment dropped it.

* chore(#4218): regenerate the install-tree fixtures for the new step fragment

The fragment ships with the workflow, so every runtime's golden install tree
gains one path — gen:install-tree is the generator that owns those fixtures.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-06 17:51:34 -04:00
Tom Boucher
38e4ce5f62 fix(#4186): anchored status vocabulary, record-session arg guard, recount pin (#4381)
* fix(#4186): anchored status vocabulary, record-session arg guard, recount pin

Three defects from #4186:

1. normalizeStateStatus ran a first-match-wins SUBSTRING chain over the
   free-prose body Status field, so prose merely mentioning a status word
   was silently rewritten to a credible wrong token (a .planning/ path in
   Italian prose -> status: planning; verifica -> verifying; completezza ->
   completed). Recognition is now an ANCHORED whole-field match against a
   declared vocabulary (STATUS_EXACT_TOKENS + STATUS_ANCHORED_PATTERNS,
   state-document.cts) — case/whitespace-tolerant, branch-order artifacts
   preserved (Planning complete -> planning; Phase complete — ready for
   verification -> verifying). The recorded lenient fallback (#3873 row 26)
   stands: unrecognized prose passes through verbatim. Read-side consumers
   (W011, statusline) ride the same function.

2. The progress recount skew (stray *-SUMMARY.md inflating
   completed_plans) is already dead on next via #1988/PR #2016
   (countMatchedSummaries pairs summaries to plans) — verified live and
   pinned with regression rows composed against the #4129/#4359 ratchet.

3. state record-session with no args executed and wrote STATE.md; it now
   errors like state update (stopped-at or resume-file required), handler-
   side so SDK callers are covered too. Four tests pinning the bare-call
   write are updated to the new contract.

* fix(#4186): update status pins to the anchored vocabulary contract

Bench round 1 follow-ups:

- Legacy bare 'Milestone complete' kept as reader-side vocabulary
  (ADR-2207 removed the writers, not recognition of legacy files).
- state.test pins updated: 'Paused at Plan 3' and round-trip
  'Executing Plan 5' were pins of the substring guessing itself —
  the round-trip now uses the real handler form 'Executing Phase 5'.
- record-session no-op/no-fields tests repurposed to the usage-error
  contract (CLI + SDK-level ExitError), byte-unchanged assertions kept.
- statusline tests repinned: vocabulary values collapse to keywords;
  narratives render the documented first-word fallback instead of a
  guessed token. Hook doc comment updated to match.
- docs-guard exempt baseline: state.test.cjs now cites docs/CLI-TOOLS.md.
- docs/CLI-TOOLS.md: record-session signature notes the required flag.

* fix(#4186): repair a dangling sentence in the schema docstring

* test(#4186): bound the completed_plans scan regex (#2128 class)

* chore(#4186): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-06 17:08:24 -04:00
Tom Boucher
b7917882bb fix(#4398): render the pending-todo bullet link repo-relative (#4416)
* test(#4384): failing-first regression rows for the macOS long-base todo-cap failure

The 240-char pending-todo bullet cap must be deterministic w.r.t. where the
repo is checked out. Deterministic long-base-path fixtures (a single 110-char
segment, no real macOS dependency) reproduce next's own macos shard 3/3
failure (run 34038716700) on every OS: with an absolute link the bullet
exceeds the cap and the documented needs-first truncation drops the
'Needs <solution>' clause. Rows cover the determinism property (byte-identical
bullets under short and long bases), the CLI surface, relative-path stability,
legacy no-projectRoot behavior, drop-order preservation, and adversarial
edges (outside-root, path===root, non-string path).

* fix(#4384): render the pending-todo bullet link repo-relative

renderPendingTodosMarkdown gains an optional projectRoot; when given and the
todo's path is absolute, the bullet's [todo file](…) target becomes
toPosixPath(path.relative(projectRoot, path)) — the idiom already used for
project_exists. cmdInitTodos passes cwd.

The JSON todos[].path field stays absolute (#2376). Only the rendered display
link changes: embedding the machine-variable absolute base let macOS's
/private/var/folders/… temp paths consume the 240-char budget and drop the
'Needs' clause on long-path machines only — next's own macos-latest shard 3/3
went red on exactly this (run 34038716700), Linux's short /tmp passed. The
240-char whole-bullet cap and the needs→title→area drop order are unchanged;
this matches PR #4384's own canonical example, docs, and unit tests, which all
show repo-relative links. Docs updated at all three surfaces that describe the
bullet (COMMANDS.md, templates/state.md, reference/state-md.md — the last was
still pre-#4384 'count and reference' prose).

Fixes the macOS regression introduced by #4384; next is red on its own CI.

* test(#4384): fix substring false positive in the outside-root regression row

The ../-form relative link legitimately contains the absolute path as a
substring, so !line.includes(absolutePath) fired on correct output (caught by
the first remote verify run, linux-node24 44018/44019). Assert the property
itself instead: extract the link target and require it to be non-absolute and
not equal to the absolute path.

* chore(#4398): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-06 13:23:26 -04:00
Tom Boucher
aad04f4e9a docs(#4400): ADR-4139 — the compact-content seam (#4410)
* docs(#4400): ADR-4139 — the compact-content seam

Phase 0 of epic #4139. Locks the design before any code lands.

#4139's stated mechanism cannot reach the stream it exists for: 58 of 72
shipped commands deliver their whole workflow file through an eager
@-include, which the host expands before any project config is in context.
An in-content gate is evaluated after those bytes are already paid.

The ADR declines the obvious fix (convert the 58 execution_context blocks
to runtime Reads) because that removes the host guarantee for every user,
not only opted-in ones — a global install shares one skill tree, so an
@-include cannot be conditional. It instead keeps every @-include exactly
where it is and splits what sits behind them: the canonical path becomes a
runnable spine, elaborations move to a sibling detail file read at runtime.
A missed Read then degrades to "runs correctly with fewer tokens", never to
"runs with no instructions".

Also records: the rename to workflow.compact_content, the re-pitch onto
ADR-1610's context-rot argument rather than the cost argument ADR-1610
discounts, partition-not-duplication (which dissolves the dual-maintenance
cost the Feature Review called disqualifying), the guard-scope and
NEW_FILE_CAP mapping for the new subtree, and an argued reconciliation of
the acceptance criteria this design does not meet literally.

Corrects ADR-3646 §Context: it cites #3647 as open; #3647 closed
2026-09-01 as a duplicate of #3606. ADR-3646's Decision is unaffected —
it explicitly disclaimed any dependence on #3647's state. The residual
prose-dispatch variance named in #3647's own closure thread is unresolved,
and this ADR routes around it rather than assuming it away.

Refs #4139
Closes #4400

* docs(#4400): fold the two orthogonal review findings into ADR-4139

Code review (isolated context) and security review (isolated context) both
returned findings. Fixed here rather than carried.

Critical, from code review: the ADR repeated earlier research's claim that
discuss-phase, manager and pause-work all reach a workflow by runtime Read.
manager and pause-work carry plain eager @-includes and are inside the 58,
not outside. discuss-phase is the only precedent, and it is one file. The
Open Questions section is corrected with it.

The NEW_FILE_CAP mapping was wrong in a way that changes the layout. It
lives at tests/helpers/emitted-diff.cjs:96, not in workflow-size-budget,
and its own doc comment records that it is a hard cap, not ack-able, and
NOT tier-exemptible -- the pre-#2724 test-file version was. So a single
detail.md holding plan-phase.md's elaborations is blocked outright with no
exemption path. Detail content is now one or more parts under
workflows/<name>/detail/, each below the cap, named by the spine in the
dispatch-table shape discuss-phase.md already uses.

commit-files-pathspec is in scope and earlier research called it
irrelevant. Per CONTRIBUTING.md:1164-1170 it sweeps every .md under
gsd-core/workflows/ for unscoped commit-seam invocations. Added to the
guard table.

Two byte figures were inherited rather than measured, against this ADR's
own evidence note. Templates and agents re-measured; the table now carries
the method and the exact numbers.

From security review: "a spine that has shed a protected-content marker
fails" never defined what a marker was, leaving the strongest check in the
set resting on a prose-category judgment. Protection is now a literal
greppable sentinel in the existing gsd: comment namespace, and the guard
rule has no discretion in it. Also added: an explicit statement that the
detail path is never user- or project-supplied and cannot be shadowed by a
project-local file, and an exact-version pin commitment for gpt-tokenizer.

Code review also found a real hole in the central fail-safe argument:
spine sufficiency is verified once at split time and never again, so
load-bearing procedural text carrying no sentinel could later drift into a
detail part with every check green. Sufficiency is not machine-decidable,
so a fifth ongoing check is added -- a spine that loses lines which
reappear in its parts fails unless the PR declares the boundary move. The
ADR now states plainly that this is authoring discipline with a forced
checkpoint, not a structural invariant, and that the partition relocates
the Feature Review's cost rather than fully eliminating it.

Refs #4139
Refs #4400

---------

Co-authored-by: sim <sim@local>
2026-09-06 12:24:33 -04:00
Tom Boucher
fd4aac5670 fix(#4192): honor explicit model pins on the claude runtime (#4396)
* fix(#4192): honor explicit model pins on the claude runtime

Two documented model-configuration contracts did not hold on the claude
runtime (confirmed-bug scope from the issue triage):

Finding 1 — model_profile_overrides.claude.<tier> was inert. Step 3 of
resolveModelInternal gated runtime-aware tier resolution on
configRuntime !== 'claude', so the key's only reader was never consulted,
while workflows/settings-advanced.md writes it for claude-runtime users.
A new step 4.5 resolves ONLY the user's override entry (never the builtin
claude tier map, so unpinned installs keep resolving aliases). An
override value that maps to a current tier alias collapses to that alias
(byte-equivalent, the #2041 protection); anything else — a pinned older
generation, a bare alias repoint, a non-Anthropic id — resolves verbatim.
It sits after the resolve_model_ids:'omit' gate so an explicit project
omit still wins (#2297) and before the alias return so
resolve_model_ids:true cannot re-materialize the pin to the latest id.

Finding 2 — fully-qualified claude-* ids in model_overrides were
warn-dropped to tier resolution (mapClaudeOverrideForRuntime unmappable
branch, #2041), while the docs promise any fully-qualified model id is
valid. The unmappable branch now passes the pin through verbatim with a
warn-once breadcrumb (text describes the pass-through). Dropping it
silently unpinned the operator's explicit choice — the exact 'profile
can misrepresent what actually runs' defect of #4192. Mappable ids and
non-claude values behave exactly as before; resolveModelForTier shares
the mapping; the tier honesty signal is unchanged (raw ids still report
'unknown'); the model_policy path is untouched.

Docs updated to the agreed contract (CONFIGURATION.md false 'Claude
example' corrected; how-to + shipped reference document the pin
semantics, the fable alias, and the tier-override composition).

* test(#4192): pin explicit model pin resolution on the claude runtime

28 failing-first rows across the resolver seam and the resolve-model CLI:
pinned-generation fidelity (tier override + per-agent verbatim pins,
object form, explicit runtime), unpinned controls byte-stable (no
override, other runtime/tier, inherit, project omit, precedence),
adversarial rows (prototype-chain keys, malformed values, warn-once
dedupe, 64-char stderr cap), and behavioral AC1/AC2 rows through
runGsdTools. The stale #2041 fall-through assertions now pin the
pass-through contract; mappable-id collapse assertions unchanged.

* chore(#4192): add changeset fragment

* chore(#4192): backfill PR number in changeset fragment

---------

Co-authored-by: ZCode <zcode@localhost>
2026-09-06 10:17:50 -04:00
Tom Boucher
b7406b293f enhance(#2618): render pending todos as one bounded bullet per todo (#4384) 2026-09-06 08:06:39 -04:00
Tom Boucher
03738824de enhance(#2586): stop installing Codex context-monitor hooks without metrics (#4367) 2026-09-06 05:46:49 -04:00
Tom Boucher
7bb366e836 fix(#4130): --context flag for check decision-coverage-plan + parseDecisions quadratic-backtracking hardening (#4374)
* test(#4130): failing-first regressions for --context flag + parseDecisions hardening

Block A (flag): check decision-coverage-plan --context <path> must route
identically to the positional form; flag wins over positional context;
valueless --context falls through to the #2770 fail-closed caller error;
verify keeps its positional surface (flag is plan-only). RED on base:
the flag token lands in the args[2] phase slot (false uncovered) or the
args[3] context slot (silent CONTEXT.md-missing skip).

Block B (hardening): regex-lattice asserts pin the atomic-ID wrapper
(?=(X))\1 and the em-dash first-separator narrowing [^*—–]*[—–] plus the
no-adjacent-overlap property; a differential property compares the module
against a frozen copy of the pre-hardening grammars (reference validated
against the base build: 60k generated lines, 0 mismatches); 40k cliff
shapes assert correct outcomes with no wall-time asserts (repo rule).

A12: partitionPredicateArgs keeps one parser behind parsePredicateFlags.

* fix(#4130): --context flag for check decision-coverage-plan + quadratic-backtracking hardening in parseDecisions

(A) check decision-coverage-plan --context <path> — sibling convention
(check predicate, #2008): --flag value pairs parsed by the new shared
partitionPredicateArgs (parsePredicateFlags reimplemented as its flags
half — one parser, cannot diverge), the flag winning over a same-purpose
positional, positionals kept (no sibling deprecates them; the plan-phase
workflow caller passes positionals), valueless --context falls through
to the #2770 fail-closed caller error. Repair of the routing accident
where --context landed in the args[2] phase slot (false uncovered) or
the literal token in the args[3] context slot (silent green skip).

(B) parseDecisions regex seam hardened, byte-identical on all legal
inputs: the three bullet grammars consume the ID atomically via the
(?=(X))\1 lookahead emulation (kills the tail/[^:*]* O(n^2) re-split,
~1.1s @ 40k), and the em-dash first separator narrows [^*]*[—–] to
[^*—–]*[—–] (kills the dash-position O(n^2) retry, ~1.7s @ 40k). Group
indices unchanged (handlers untouched). Pinned by regex-lattice tests,
a differential fast-check property vs the frozen pre-hardening grammars,
and 40k cliff/legal-shape outcome tests (no wall-time asserts per repo
rule — no deterministic engine step counter exists in Node).

* docs+test(#4130): document --context invocation; harden lattice test tooling

- docs/CONFIGURATION.md Decision Coverage Gates: new 'Invoking the plan
  gate directly' block documenting both the positional and --context
  forms, flag precedence, and the valueless-flag fail-closed semantics
  (same place the gate's behavior is documented; sibling check predicate
  documents its flags the same way).
- Two changeset fragments per the maintainer brief (Added: flag; Fixed:
  hardening), PR numbers to be backfilled.
- tests/decisions.test.cjs review fixes: readRegExpTemplate template
  escaping (bare ')' SyntaxError), range-aware lattice checker with
  backreference skip and template unescape, honest A1 contract, lint
  escape warning.

* fix(#4130): valueless --context fails closed per #2770; A8 isolates flag-vs-positional context

Suite-caught fixes from the first verify run:
- cmdDecisionCoveragePlan now refuses a flag-shaped token as the
  positional context path: a bare valueless --context stays a positional
  (sibling parser semantics, unchanged) but reading it as a PATH would
  turn a caller mistake into a silent 'CONTEXT.md missing' green skip —
  exactly what #2770's fail-closed law forbids. Now falls through to
  the missing-context-argument error, as documented.
- A8 test compares decoy-positional+flag against flag-with-phase (phase
  held constant) so the row isolates WHICH context was read; the old
  form compared against a no-phase invocation that could never match.

* chore(#4130): backfill PR number in changeset fragments (PR #4374)

---------

Co-authored-by: sim <sim@local>
2026-09-06 02:55:17 -04:00
Tom Boucher
6adf3098ac fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans (#4364)
* test(#4145): regression rows for hash-matching prefix-less pristine baselines

RED skeleton: src/pristine-baseline.cts exports findPristineByHash as a
null-returning stub so the new rows fail behaviorally, not at require time.
Failing-first rows: verifier resolution (no_baseline must drop to 0 when an
exact-hash orphan exists), findPristineByHash unit row, and the two
saveLocalPatches relocation rows. Negative-space rows pin today's behavior:
missing baselines still report ok_no_baseline, mismatching orphans are never
adopted or deleted, canonical precedence and the #3657 drift posture are
untouched.

* fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans

Both pristine readers joined the manifest-keyed path strictly, so a snapshot
stored without the gsd-core/ prefix (an earlier release's writer) was reported
as ok_no_baseline by the verifier and pushed into regeneration by
saveLocalPatches — where incoming-release candidates can never satisfy the
recorded outgoing hash, leaving the correct baseline permanently unconsumed.

- src/pristine-baseline.cts (new, ADR-457): shared findPristineByHash —
  deterministic sorted scan of gsd-pristine/, exact sha-256 equality with the
  recorded pristine_hashes entry (the same authority the #3657 drift guard
  trusts), symlink-skipping, canonical path excluded via skipRel.
- verify-reapply-patches.cjs verifyFile(): on canonical miss with a recorded
  hash, adopt byte-identical content found anywhere under gsd-pristine/ before
  reporting OK_NO_BASELINE. Drift posture (#3657), canonical precedence, and
  the frozen REASON/report shapes are untouched; the verifier stays read-only.
- install.js saveLocalPatches(): preserve-check rescue — relocate a
  hash-matching orphan to the canonical path (copy, hash-verify, then remove
  the orphan) so the state self-heals on the next update instead of repeating
  forever. Honest accounting: new non-overlapping rescued counter.
- Workflow doc: one-sentence note on hash-based snapshot resolution.
- Derived ripples: INVENTORY-MANIFEST.json regen, eslint ignore + .gitignore
  entries for the compiled artifact, seedFixture mkdir fix in the new rows.

Emitted-Drift-Ack-Growth: reapply-patches.md — one-sentence note on hash-based pristine snapshot resolution (#4145)

* fix(#4145): review follow-up — orphan scan never consumes a canonical path

Adversarial review finding: with two modified files sharing byte-identical
outgoing content, recoverOrphanedPristine could adopt the OTHER file's
canonical pristine as its rescue source — relocating it (copy + delete at
its home path) and ping-ponging the single baseline between the two files
across updates. findPristineByHash's skip parameter now accepts a Set, and
saveLocalPatches passes the normalized manifest keys so every canonical
path is excluded; only genuine non-canonical orphans are eligible for
removal (no strict-join reader ever consults those). Adds the
canonical-theft regression row, a Set-skip unit assertion, and tightens the
workflow doc sentence the same pass flagged as overstated.

* fix(#4145): INVENTORY roster row + symlink-fixture correction

Two leftovers from the ab17b7a1e5 bench run, both root-caused:
- docs/INVENTORY.md roster row for cli_modules/pristine-baseline.cjs
  (#3762 gate: every manifest entry carries a row).
- The findPristineByHash symlink unit fixture placed its symlink target
  INSIDE the scanned root, so the walk legitimately matched the real target
  file. The implementation skips the symlink itself; the fixture now keeps
  the target outside the scanned tree so the assertion tests what it claims.

* changeset(#4145): fixed fragment for pristine baseline hash resolution

---------

Co-authored-by: gsd-agent <agent@gsd.local>
2026-09-06 02:04:18 -04:00
Tom Boucher
06eba5fdb0 fix(#4130): parse phase-prefixed decision IDs (D4-01) (#4357)
* test(#4130): failing-first regression for phase-prefixed decision IDs

Add the #4130 matrix: D4-01/D12-01 across all three bullet forms, tags,
discretion, wrapped lead-ins, gate-level plan/verify end-to-end rows, and
parity properties (well-formed digit-prefixed ids parse to their exact id;
a non-digit injected into the prefix fails loud). Update the #2347
non-D-prefix fixture from D5-NN (now a legal grammar) to DEC-NN, and
graduate the representative d5-prefix corpus fixture from could-not-parse
to parsed-but-uncovered.

All new rows are RED against origin/next; they go green with the parser
fix in the next commit.

* fix(#4130): parse phase-prefixed decision IDs (D4-01)

The three declaration grammars, the parse-miss guard, the #3939 join
regexes, and the token evidence all anchored on the literal 'D-' (or
'**D-'), so an ID carrying a digit-run phase prefix between the leading
letter and the hyphen matched nothing — while the #2347 shape detector
correctly called those bullets decision-shaped, collapsing the whole
CONTEXT.md to could-not-parse with 0 extracted instead of a coverage
verdict.

Derive the extractor ID grammar from one shared DECISION_ID_SOURCE
('D[0-9]*-' + the existing alnum tail, full id captured), widen the
guard/join anchors to ID_ATTEMPT_SOURCE (bare 'D-' or a digit-initial
prefix run, so a typo'd 'D4x-01' fails loud while letter-initial prose
like 'Deferred-until' stays none-present), and align the bare-token
evidence. Both gates and the gap-checker share the parser, so all three
surfaces read phase-prefixed decisions now; the gate messages name the
accepted forms including the phase-prefixed one.

* docs(#4130): document the phase-prefixed decision identifier form

The canonical CONTEXT.md reference said decisions carry 'a sequential
D-NN identifier' with no mention of the optional phase-number prefix the
parser now accepts (D4-01) or the alphanumeric tail it always accepted
(D-INFRA-01). Name both in the Decision identifier format section, EN
and ja-JP.

* chore(#4130): changeset

* chore(#4130): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 21:05:19 -04:00
Tom Boucher
f9f72cb54c enhance(#3777): opt-in concurrent per-plan planners in chunked mode (#4346)
* test(#3777): add failing-first coverage for concurrent per-plan planner dispatch

Extracts and executes the real bash blocks this PR is about to add to
plan-phase.md and chunked-planning-mode.md (CHUNKED_PARALLEL resolution and
the BATCH_PLAN_IDS dedup guard), plus config-set/config-get coverage for the
new planning.chunked_parallel key. Expected RED against the current shipped
workflow text — the extraction anchors do not exist yet.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(#3777): dispatch chunked mode's per-plan planners concurrently within a Wave

Adds opt-in planning.chunked_parallel (default false, byte-identical to the
existing serial loop). When true and the runtime's negotiated dispatch
capacity (dispatch-capacity, #3673) is greater than 1, chunked planning's
per-plan Tasks that share one outline Wave are issued together instead of
one at a time; a later Wave still waits for the current one to be verified
on disk and committed. A host with no declared maxConcurrency (most
non-Claude runtimes today) stays serial regardless of the setting.

Resolution and the Plan-ID dedup guard live in chunked-planning-mode.md
itself (gated on the section's own CHUNKED_MODE skip-check) rather than in
plan-phase.md, so a non-chunked run pays no extra gsd_run calls.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#3777): repoint extraction at chunked-planning-mode.md after the move

CHUNKED_PARALLEL resolution moved out of plan-phase.md into
chunked-planning-mode.md itself (see the preceding commit); update the
test's extraction path and header comment to match.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3777): relocate the canonical runtime-launcher preamble before its first use

The CHUNKED_PARALLEL resolution block's two gsd_run calls landed earlier in
the file than the sole existing preamble (in the commit step), which
tests/runtime-launcher-parity.test.cjs's (B) check requires to precede every
gsd_run call in the file. Move the preamble (not duplicate it) to the top of
the resolution block; the commit step's fenced block now just calls
gsd_run directly.

Caught by the GREEN checkpoint gsd-test run before push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3777): strip the canonical preamble from the extracted resolution block

The CHUNKED_PARALLEL resolution fence now carries the relocated
runtime-launcher preamble as its first line (previous commit). Extracting
the whole fence and running it after the test's own gsd_run stub let the
embedded preamble's own resolver logic `unset -f gsd_run` and exit 1 before
reaching the resolution logic, since no real gsd-tools.cjs exists in the
temp script dir — every test calling runChunkedParallelResolution() failed.

Strip the preamble (sourced from gsd-core/workflows/_runtime-launcher.snippet.sh,
the same file scripts/sync-runtime-launcher.cjs treats as canonical) before
splicing in the stub, so this suite tests only the resolution logic it is
actually about.

Caught by the post-rebase gsd-test run before push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3777): add the How-To page the phase gate requires

Enablement is 2 commands (config-set, then --chunked), which this repo's
own doc-quadrant gate flags as how-to-owed: a reference table cannot carry
a sequence. Covers enablement, the dispatch-capacity gate's honest
"most runtimes today: no effect" case, and the two accepted trade-offs.

An earlier reasoning pass (recorded in .gsd/phase/.../70-docs.json before
this commit) had incorrectly claimed #3034 shipped with no equivalent
how-to page, as precedent for skipping one here. That claim was false —
docs/how-to/enable-parallel-reviewer-lanes.md exists and is indexed. The
phase gate caught the omission before merge; corrected here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3777): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:57:17 -04:00
Tom Boucher
1db726ebbf feat(#3806): canonize the Review Dispositions Ledger contract (#4345)
* test(#3806): add parity tests for the Review Dispositions Ledger contract

Failing-first: asserts references/planner-reviews.md, workflows/plan-phase.md,
and agents/gsd-plan-checker.md agree on a single canonical "Review Dispositions
Ledger" heading, its round-scoping, L##@{sha} anchor format, and append-only
supersession rule. These fail until the canon and its two references are added.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(#3806): canonize the Review Dispositions Ledger contract

Promote the existing planner-reviews.md Step 4 return-payload tables
(Review Feedback Addressed/Deferred) into a canonical `## Review
Dispositions Ledger` PLAN.md section, stated once in planner-reviews.md
and referenced (not restated) from plan-phase.md's
<review_incorporation_contract> and gsd-plan-checker.md's Review
Incorporation dimension. Adds round-scoping (`### Round {N} —
{REVIEWS_sha}`), a `L##@{sha}` line-anchor format so a REVIEWS.md
reference survives the file being rewritten each round, and an
append-only supersession rule. Scoped to part 1 only per the
maintainer's approved-feature verdict — the deterministic lint/check
verb (part 2) is explicitly deferred to a follow-up.

Also: ADR-3806 recording the decision, a docs/features/ fragment
(FEATURES.md is generated), and a changeset fragment.

Closes #3806

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): fenced-example count bug and lint findings from review

- tests/plan-review-convergence.test.cjs: the "heading exactly once"
  test counted the canonical heading text globally, so it also matched
  the illustrative fenced-code example in planner-reviews.md that shows
  the same heading as sample content, always failing 2 !== 1. Rewritten
  as a bounded line scanner that skips fenced blocks (found by an
  isolated adversarial review pass). Also bounded an unbounded regex
  quantifier over readFileSync content flagged by
  local/no-unbounded-quantifier.
- docs/features/review-dispositions-ledger.md: match house fragment
  style (bold-lead paragraphs, not #### headings) per the Standards-axis
  review; regenerated docs/FEATURES.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): fit reference-cite fix within size hard caps; ack growth

Trims the plan-phase.md / gsd-plan-checker.md reference-cite text to a
single short clause pointing at gsd-core/references/planner-reviews.md
(also fixes the bare `references/planner-reviews.md` cite the #3576
shipped-reference-cites gate rejects), bringing both files back under
their SIZE hard caps and the plan-phase.md phase6 shrink-only baseline.
Both files still grow slightly versus origin/next, acknowledged below
per ADR-2719's emitted-drift-ack contract.

Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): correct malformed Emitted-Drift-Ack-Growth trailer block

The previous commit's two Emitted-Drift-Ack-Growth trailers were
separated from the Co-Authored-By trailer by a blank line, so git's
own trailer parser (which tests/helpers/emitted-runtime.cjs reads via
`%(trailers:key=...)`) only recognized the last contiguous block
(Co-Authored-By) and treated the Ack-Growth lines as ordinary body
text — invisible to the emitted-attribution gate, not malformed data.
Restating them here as one contiguous trailer block, git log over the
PR range aggregates trailers from every commit, so this is additive.
Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): isolate the ack-trailer paragraph as its own trailer block

Git's trailer parser requires the trailer paragraph to be the message's
final paragraph, preceded by a blank line, and to contain nothing but
trailer-shaped lines. The prior commit's blank line before the trailer
lines was missing, which folded the leading Emitted-Drift-Ack-Growth
lines into an ordinary prose paragraph.

Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3806): backfill PR #4345 into changeset and ADR

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:50:27 -04:00
Tom Boucher
c20675cc4d fix(#3819): widen executor's pre-commit guard beyond worktree mode (#4343)
* fix(#3819): widen executor's pre-commit guard beyond worktree mode

The pre-commit protected-branch assertion in the executor agent (#2924)
only fired inside a Claude Code worktree and matched a hardcoded
five-name branch list. It never ran in an ordinary checkout and never
covered this repo's own default branch ("next"), so gsd-executor could
commit planning-repo documents directly onto a shared checkout's
default branch with no PR ever created.

Widen the guard to run in every isolation mode, and resolve the
protected branch via the repository's actual default branch (with the
existing five-name list retained as a fallback when the resolver
itself cannot be invoked) plus any configured git.protected_branches.
Add a git.allow_default_branch_commits escape hatch for projects that
intentionally execute on their default branch. Also point the
separate <final_commit> commit helper back at the same guard, so it
cannot be sidestepped by that path.

Emitted-Drift-Ack-Growth: gsd-executor.md — widened pre-commit protected-branch guard (#3819); tightened comments to stay under the size cap.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3819): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:35:42 -04:00
Tom Boucher
4e1c449281 enh(#3811): add hooks.commit_types config surface to gsd-validate-commit (#4340)
* enh(#3811): add hooks.commit_types config surface to gsd-validate-commit

Extends the opt-in Conventional Commits hook with a hooks.commit_types
config array that adds project-specific types to the 10 built-ins
without replacing them. Configured values pass a safe-token filter
before reaching the compiled regex, so a config entry can never alter
the pattern's structure. The regex alternation, the human-readable
error text, and a new typed valid_types JSON field all derive from one
list instead of the two hand-synced copies this replaces.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3811): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:35:09 -04:00
Tom Boucher
17e163f15c docs(#4333): document the ADR Amends/Amended-by convention (#4334)
* docs(#4333): document the ADR Amends/Amended-by convention

Two patterns for amending an accepted ADR are established practice —
an in-place `## Amendment (YYYY-MM-DD)` section, and a separate ADR
that declares `Amends` with a reciprocal `Amended by` back-link — but
only the first was ever written down. #4030 shows the cost: a
contributor concluded no ADR owned a contract that ADR-857 already
covers, because nothing said the second pattern (used by ADR-1244 and
ADR-2782 to extend ADR-857 itself) existed.

Document both patterns in docs/contributor-standards.md, note the
Amends/Amended-by reciprocity rule in docs/adr/README.md alongside the
existing Supersedes/Subsumes rule (and that it isn't yet gated by
scripts/gen-adr-index.cjs the way those are), and point CONTRIBUTING.md's
new-ADR process at the amendment path for revisiting an existing one.

Closes #4333

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4333): fix imprecise Amends/Amended-by precedent citations

Orthogonal review caught two inaccuracies: PR #1643 doesn't match the
in-place dated-section pattern (it rewrites the original Decision text
rather than appending an untouched dated section), and ADR-1244's
relationship to ADR-857 is prose ("extended by"), not the structured
Amends/Amended-by header field. ADR-2782 is the verified precedent for
the structured field pair — its one Amends field names four targets
(857, 894, 1016, 1244), all four carrying the reciprocal back-link.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 16:45:57 -04:00
Dennis Alexis Valin Dittrich
1017898cb9 fix(#3771): make remediation examples non-binding and surface revision conflicts (#3916)
* fix(#3771): separate the binding property from the advisory remediation

Checker findings fused "what property failed" with "how to fix it" into a
single `fix_hint` and never marked which half binds. The checker rendered
every hint under a "must fix" heading, the orchestrators injected the issues
verbatim and ordered targeted updates, and the shared revision references
mapped each hint to a prescriptive strategy — so a contract-following planner
applied a hint literally even when a smaller mechanism satisfied the same
property, or when the hint contradicted a locked decision. There was no
channel to report that conflict, and every attempt burned a revision
iteration.

Checker side: every issue now carries a binding `required_property` (the
invariant that failed) plus its evidence and severity, and `fix_hint` is
labelled non-binding wherever it appears — including the human-facing blocker
rendering, so "must fix" unambiguously names the property and never the
example.

Planner side: revision re-checks locked decisions, capability guidance and
existing plan constraints before editing; satisfying a blocker through a
smaller valid alternative counts as addressing it; and a hint that conflicts
with any of those returns `REVISION_CONFLICT` carrying the conflict and the
alternatives considered. Orchestrators route that to user choice or the
configured plan-review convergence loop without consuming retry budget.

Also applied to the UI-spec revision loop and the gap-plan hint, and the
generic pattern's stray `suggested_fix` field name is reconciled to the
plan-checker's `fix_hint`.

Nothing legitimately binding is weakened: blockers still block, severity
still gates, iteration caps and stall escalation still fire, and required
task fields and decision coverage still hold.

Refs #3771

* test(#3771): pin the binding/advisory split across the revision chain

Locks the separation at every link that carries it: the checker's issue
schema and blocker rendering, the planner's constraint re-check and
REVISION_CONFLICT return, the generic pattern's reconciled field names, and
each orchestrator's conflict routing without retry-budget consumption. Also
pins what must not have been weakened — blockers, severity gating, iteration
caps and stall escalation.

Red against the pre-fix prose: 32 of 34 assertions fail (the 2 that pass are
the preservation checks, correctly).

Refs #3771

* chore(#3771): add changeset fragment for the remediation-binding fix

* chore(#3771): acknowledge the remediation-binding growth

Five runtime-loaded files grow: the two checkers carry the binding/advisory
split where the model reads it (a `required_property` on every dimension
example, since a schema the examples contradict teaches the examples), and
the three orchestrators carry the REVISION_CONFLICT route, which has to live
with the `iteration_count`/`revision_count` state it declines to spend.

Deletes tests/emitted-drift-acks/3172-stated-failing-direction.json: it is
fully spent on next and still owned plan-phase.md, so it walls off a key it
can no longer clear (#3078). Its removal is the documented remedy for the
duplicate-key collision, not drive-by cleanup.

* fix(#3771): close the review gaps in the conflict contract

Adversarial review (Codex, Antigravity) found four real defects in the first
pass, each confirmed against the source before acting:

- The UI checker's structured return still ordered `Fix: {exact fix required}`
  and "list each BLOCK dimension with exact fix required". The dimension
  examples had been marked non-binding but the rendering the researcher
  actually reads had not — the same omission this issue is about.
- `ui-phase` and the canonical `revision-loop` flow incremented their counter
  BEFORE dispatching the reviser, so "do NOT increment on REVISION_CONFLICT"
  was unreachable prose: the iteration was already spent. The increment now
  sits on the return path in both.
- The conflict gate offered "accept as-is", which is an early exit from a
  still-failing blocker — a weakening the brief explicitly forbids. The three
  options are now adopt an alternative / override the constraint / amend the
  constraint; every one resolves the conflict. Accepting an unaddressed blocker
  remains available only at the unchanged iteration-cap escalation.
- The convergence route was declarative: nothing in
  plan-review-convergence.md could receive a conflict. plan-phase now records
  it in REVIEWS.md — the channel that loop already consumes — convergence
  refuses to declare convergence over an open entry, and routing back into a
  run convergence itself started is explicitly excluded as a cycle. `quick` has
  no REVIEWS.md and no phase, so its convergence branch was dead prose and is
  deleted in favour of asking the user.

Also reconciles the last two drifted field names (`finding`, `affected_field`)
to the plan-checker schema, and repairs a silent no-op: the few-shot
`required_property` insertion never applied because those lines are
blockquoted, and the test's own block filter was anchored on indentation only,
so a vacuous loop passed over zero blocks. Both are fixed and the filter now
asserts it found blocks.

Refs #3771

* chore(#3771): extend the growth acknowledgment for the review round

plan-review-convergence.md joins the list: the conflict route needed a
receiving end, and it lands on the seam that loop already reads (REVIEWS.md)
rather than a new mechanism. The plan-phase, ui-phase and gsd-ui-checker
entries gain the second-pass reasoning — an executable convergence branch, the
increment moved onto the return path, and the structured return that still
ordered an exact fix.

* fix(#3771): make the conflict route bounded, ordered, and owned

Round-2 adversarial review found five more defects, each confirmed in the
source before acting:

- The convergence gate sat AFTER `gsd_run state planned-phase` and the success
  banner, so a run could write and announce convergence over an unresolved
  conflict. OPEN_CONFLICTS is now read from REVIEWS.md and is part of the
  converged CONDITION, evaluated before any write.
- plan-phase's cycle-exclusion ("unless this run was invoked by convergence")
  was not a question the orchestrator can answer at runtime. plan-phase now
  never invokes convergence at all — it records the conflict when a phase
  REVIEWS.md exists and resolves it with the user in-place, which removes the
  cycle instead of describing it.
- Closure had no owner. plan-phase writes the row, so plan-phase strikes it
  resolved; convergence only reads. An open row is a live blocker, never a
  stale artifact.
- Declining to increment the counter removed the only bound on the conflict
  path: an agent returning the same conflict forever would loop unattended. A
  conflict naming the same `required_property` twice in a row is now a stall
  and escalates through the existing gate.
- verify-work's gap-plan revision loop hands `<revision_context>` to
  gsd-planner and so inherits the whole contract, but stated none of it and
  could not handle the conflict return. It is now covered like the others, and
  is in the test's orchestrator table.

Refs #3771

* chore(#3771): acknowledge the round-2 growth

verify-work.md joins the list — the flow the second review found missed — and
the plan-phase, ui-phase and plan-review-convergence entries gain the
round-2 reasoning: the gate moved ahead of the state write, the convergence
hand-off replaced with a runtime-checkable record-and-resolve, and the
recurrence bound that replaces the counter the conflict path stopped spending.

* fix(#3771): make the convergence gate countable and stop the conflict fall-through

Third adversarial round (Antigravity) found three defects:

- The OPEN_CONFLICTS pipeline had no `grep -v '~~'` despite its own comment
  claiming one, and `grep -c '^| '` also counts a markdown table's header and
  separator rows — every resolved conflict would have read as open and
  convergence would have deadlocked instead of converging. plan-phase now
  records each conflict as a `- [ ]` checklist line and flips it to `- [x]`, so
  the gate is an exact fixed-string match with no table parsing.
- "then continue below" fell through to the checker spawn, so a SECOND
  REVISION_CONFLICT would have been handed to the checker as though it were a
  revised plan. plan-phase, quick and verify-work now re-evaluate the return
  from the top of the conflict handler; ui-phase already looped back.
- revision-loop.md still described plan-phase routing a conflict to the
  convergence loop instead of asking — the behaviour round 2 removed. Recording
  is now stated as being in addition to asking, never instead of it.

Refs #3771

* chore(#3771): bring the changeset in line with what shipped

Two review rounds widened the change after the fragment was written:
verify-work's gap-plan revision and the convergence loop are covered, two
more drifted field names are reconciled, and the conflict path carries an
explicit recurrence bound.

* fix(#3771): declare and emit the REVISION_CONFLICT marker

check:contract-drift on CI caught what local lint never reached: four
workflows dispatch on `## REVISION_CONFLICT`, but no agent declared or emitted
it — an orphan consumer, matching a marker nothing produces. The shared
reference (planner-revision.md Step 7b) described the return; the agent
definitions did not carry it.

gsd-planner and gsd-ui-researcher now emit the marker in-fence alongside their
other return markers, and both registry rows in agent-contracts.md declare it.
gsd-planner's Consumed by gains the two workflows that dispatch on it and were
missing from the row.

The gate is right: a return contract belongs where the agent is defined, not
only in a reference the agent happens to load.

Refs #3771

* chore(#3771): acknowledge the return-marker growth

gsd-planner.md and gsd-ui-researcher.md each gain the REVISION_CONFLICT
marker that check:contract-drift requires them to emit.

* fix(#3771): hoist the shared conflict protocol out of the workflows

Two CI failures, both correct gates:

- tests/few-shot-calibration.test.cjs pins the plan-checker calibration file
  at exactly 4 examples (2 positive, 2 negative). The example added in the
  first pass broke that balance — and described PLANNER behaviour in the
  CHECKER's calibration set, which is the wrong surface for it. Removed; the
  smaller-alternative rule is already normative in gsd-plan-checker.md and
  planner-revision.md, and pinned by the regression suite.
- tests/phase6-capstone-conformance.test.cjs (ADR-857 phase 6, #1168) requires
  plan-phase.md to stay BELOW its pre-phase-6 baseline of 94519 bytes. The
  inline conflict block pushed it to 94988.

The fix for the second is the one that should have been made first: the
record/resolve/close protocol and the recurrence bound were identical in four
workflows, and revision-loop.md — which plan-phase already @-imports — is what
a shared contract is for. The protocol now lives there once; plan-phase states
only its bindings (which counter, which artifact, which next step) and points
at it. plan-phase.md: 94988 -> 92739, under the ratchet with headroom, and the
four-way duplication is gone.

quick, ui-phase and verify-work do not import the reference, so they keep their
inline statements. The suite asserts each rule against what the runtime
actually loads for that orchestrator, not against the file in isolation.

Refs #3771

* docs(#3771): state the shared-protocol relationship accurately

Three of the four revision-bearing workflows do not @-import revision-loop.md,
so 'follows it verbatim' overstated the coupling. Only plan-phase defers; the
others restate the rules inline and this section is the authority they must
agree with.

* refactor(#3771): name the authority instead of restating it four times

Self-review finding: the same ~700-byte reviser paragraph was inlined in four
prompts while gsd-planner already loads planner-revision.md whenever
<revision_context> is present (agents/gsd-planner.md:555, :588) — a fifth copy
of the same contract. Four places to edit in lockstep is precisely the drift
class this PR exists to fix; the generic pattern calling the field
suggested_fix while the checker emitted fix_hint is what that looks like after
a year.

Each prompt now carries only the load-bearing clauses and names the authority
it summarises. Deliberately NOT reduced to a bare pointer: these are LLM
prompts, and a contract stated only in a file the reader is supposed to fetch
is the failure mode of this very bug. The saving is modest (~50 bytes each) —
the point is the named source of truth, not the bytes.

Suite: dropped seven assertions that pinned heading text and bold-lead
phrasing a reword would break without changing what the runtime is told; their
neighbours already pin the same contract by content. 519 -> 499 lines, 61
tests. Red gate against origin/next: 56 of 61 fail.

Refs #3771

* fix(#3771): sanitize agent-authored conflict text and bound total conflicts

Cross-AI review (agy/Gemini) raised two security findings. I confirmed the
first by running the gate's own awk against a crafted file and reproduced it
exactly; the second is a real hole in my bound.

1. FAIL-OPEN in the convergence gate. The conflict text is agent-authored and
   is appended to REVIEWS.md, which the gate scans with an awk that stops at
   the next '## ' heading. One line of agent text beginning '## ' ends that
   scan early, so conflicts below it are never counted and convergence declares
   success over a live blocker. Measured: 3 open conflicts, awk returned 2.

   Fixed at the write boundary, which is the trust boundary: every field has
   newlines and tabs collapsed to spaces and a leading '#', '-', '|' or fence
   stripped, so one conflict is exactly one line. Both producing agents now
   declare their fields single-line plain text, and the reader states the
   invariant it depends on so a later edit cannot silently break it. Verified:
   3 open + 1 resolved now counts 3; missing file and absent section count 0.

2. The recurrence bound was 'same required_property twice in a row', which an
   agent alternating property names never trips, leaving the un-incremented
   conflict path unbounded. Now bounded twice: the repeat rule catches the
   common case, and the THIRD conflict return of a loop escalates whatever
   property it names. A conflict still never consumes a revision iteration;
   this cap is separate from and additional to the revision cap.

Rejected from the same review: deleting 'a planner that reaches
required_property by a smaller or different mechanism has addressed the issue
in full' from the CHECKER prompt as misplaced. It is load-bearing exactly
there. A checker that does not know a different mechanism counts will re-flag
the issue on re-check, which is the revision loop that never terminates. The
argument offered for deleting it, that the checker evaluates the new state
independently, describes the failure mode.

Refs #3771

* fix(#3771): fail closed on an unverifiable convergence gate

Second cross-AI pass (agy, this time with the full files rather than the diff)
found two more, both real:

1. The gate read REVIEWS_FILE with `2>/dev/null || echo 0`, so an unreadable or
   empty path counted as ZERO open conflicts and converged. That path is
   resolved a few lines earlier by a pre-existing unquoted
   `ls ${phase_dir}/${padded_phase}-REVIEWS.md` (line 346, not touched by this
   PR), which yields an empty string rather than an error when the path
   contains a space. Unverifiable is not clean: the gate now tests -z and -r
   first and BLOCKS. Verified both branches.

   The unquoted ls itself is left alone deliberately — it predates this change
   and belongs to the reviews lookup, not the conflict gate. Fixing it at my
   own boundary removes its effect on this gate without widening scope.

2. REVIEWS.md is writable by the review agent, which could flip a `- [ ]` to
   `- [x]` or delete the section and forge the state of a blocking gate. The
   section now declares a single writer: /gsd:plan-phase appends and closes,
   every other agent leaves it byte-for-byte alone, readers read.

Also trimmed a clause that explained the increment ordering by reference to
what the file said before this PR. Commit history is not instruction, and
these files are prompts.

Rejected: the claim that quick's conflict gate deadlocks autonomous pipelines
by asking the user. Its existing max-iteration escalation in the same file
already asks the user the same way; this adds no new interaction class.
Noted but out of scope: the per-dimension YAML example blocks and the shim
boilerplate duplicated across agent prompts both predate this change.

Refs #3771

* fix(#3771): count conflicts by line shape, not by section

CodeRabbit review on the rehearsal PR. Five findings, all valid, all applied.

The best one is a deletion. The convergence gate scanned between
'## Plan-Revision Conflicts' and the next '## ' heading, and that scan stops at
the FIRST heading it meets — so one stray '## ' line hid every conflict beneath
it and returned 0, converging over a live blocker. Reproduced: section-scan 0,
shape-scan 1. Sanitizing at the write boundary does not cover a hand-edited,
legacy, or foreign-written REVIEWS.md, so the reader needed its own guarantee.

It now matches the conflict line SHAPE anywhere in the file:

  grep -c '^- \[ \] .*required_property:'

No section bookkeeping, nothing a heading can truncate, and it composes with the
writer's sanitization (which strips a leading '-' from agent text, so agent prose
cannot forge the shape). Verified: injected heading -> 1, all resolved -> 0.

The other four:

- Both checkers told the author never to emit a contradictory fix_hint, then
  offered an escape hatch that put the forbidden route in the hint anyway. They
  now name NO route in that case and state only that the property conflicts with
  the constraint. A hint carrying a forbidden route is applied by anyone who
  trusts hints.
- The REVISION_CONFLICT marker description in gsd-planner.md was narrower than
  planner-revision.md: it covered a contradictory hint but not an unreachable
  required_property. A planner reading only the agent file would have burned
  retry budget on the case the reference routes to a conflict.
- The few-shot calibration examples used uppercase BLOCKER/INFO while the schema
  defines blocker/warning/info. Pre-existing, but it is the same schema-vs-example
  disagreement this PR exists to end, and the file was already being edited.
- verify-work's re-entry instruction existed but sat after the Bounded clause, so
  the paragraph read "re-spawn ... stop re-spawning ... after re-spawning". The
  re-entry now immediately follows the re-spawn, and states that only a
  non-conflict return may reach the checker or increment iteration_count.

Refs #3771

* fix(#3771): resolve the contradictory scope_sanity severity examples

Sixth CodeRabbit finding, posted outside the diff range and missed on my first
read — I had claimed all findings were addressed after reading only the five
inline comments. This one was in the review body.

agents/gsd-plan-checker.md carried TWO scope_sanity examples with identical
metrics (5 tasks, 12 files) and OPPOSITE severities: warning in Dimension 5,
blocker in <examples>. Line 872 states "2-3 tasks/plan good, 4 warning, 5+
blocker" and the severity table lists warning as "Scope 4 tasks (borderline)",
so the warning example contradicted both.

ADR-2629 Decision 5's "over budget is a WARNING, never a blocker" does not
excuse it: that rule governs the smart-zone TOKEN estimate (the estimate-check
verb, lines 299-306), which is a different axis from task count. Verified in
source before touching it.

The contradiction is pre-existing but this PR made it binding and visible:
severity is now declared part of the binding payload, and both examples were
given the same required_property, so they now disagree on the severity of an
identical finding about an identical property.

Deviating from the proposed correction, which was warning -> blocker: that
would duplicate the <examples> entry outright (same tasks, files, severity).
The Dimension 5 example is instead made a genuine 4-task borderline warning, so
the file keeps one worked example per severity and the thresholds, the severity
table and both examples finally agree.

Refs #3771

* fix(#3771): stop laundering a grep error into zero open conflicts

Seventh CodeRabbit finding — from a SECOND review round my own CR-4 push
triggered, which I had not looked for. This one is a regression I introduced
while fixing the previous fail-open.

CR-4 replaced the truncatable section scan with:

  OPEN_CONFLICTS=$(grep -c '^- \[ \] .*required_property:' "$REVIEWS_FILE" || true)

`|| true` masks every grep failure. grep exits 1 for "no matches" (a legitimate
zero) but 2 for a read error, and `|| true` turns both into an empty capture
that `${OPEN_CONFLICTS:-0}` renders as 0. If REVIEWS.md is removed or becomes
unreadable between the -r check and the scan, the gate reports no conflicts and
convergence proceeds. Proven: unreadable file -> captured empty -> 0.

The status is now inspected, and only exit 1 counts as zero; anything else
blocks.

My first attempt at this fix was itself wrong and my own harness caught it: I
wrote `if ! grep ...; then grep_status=$?`, but `!` inverts the status, so `$?`
in that branch is 0 and every failure reads as success — the clean-file case
printed "BLOCKED (grep exit 0)". The status must be read in the ELSE branch of a
non-negated `if`, which is what CodeRabbit proposed. Both traps are now pinned
by tests.

Verified end to end: all resolved -> 0, no conflicts at all -> 0, injected
heading -> 1, unreadable file -> BLOCKED with grep exit 2.

Refs #3771

* test(#3771): execute the conflict gate instead of reading it

CodeRabbit round three: 0 actionable, 1 nitpick — "these assertions inspect
Markdown source only; they do not prove that grep status 1 produces zero
conflicts or that a scan error exits before convergence." Rated Trivial. It is
the most valuable finding of the three rounds.

This gate has been wrong three times: a section scan a heading could truncate, a
`|| true` that laundered grep's error status into zero, and an `if !` whose `$?`
reported the negation rather than the command. Every one of those passed the
text assertions that existed at the time. I proved each fix by hand in a shell,
and none of that proof lived in the suite.

The gate is one self-contained fenced block, so the test now extracts it from
the workflow — located by content, not line number — writes it to a script and
RUNS it against fixtures: two open plus one resolved counts 2; no matches counts
0 and does not fail; a conflict below an injected `## ` heading still counts; an
unreadable path and an empty path both BLOCK with a non-zero status and no zero
count on stdout.

Non-vacuity proven by mutation rather than asserted. Reverting the gate to each
of its three historical broken forms reds the suite:

  section-scan awk  -> 7 failures (5 in the gate cases)
  || true           -> 4 failures (3 in the gate cases)
  if ! (negated $?) -> 4 failures (3 in the gate cases)
  restored          -> 69 pass, 0 fail

The prose assertions stay: they are the right instrument for a prompt. This
covers the one part of the change that is real shell an orchestrator executes.

Refs #3771

* test(#3771): route the gate harness through the shared test helpers

ESLint's project rules caught three violations in the new harness: an unbounded
execFileSync (DEFECT.UNBOUNDED-SUBPROCESS — an unbounded spawn is an indefinite
hang, and on macOS CI that is how a stuck run stops reporting instead of failing)
and two raw fs.rmSync calls, which skip the Windows-EBUSY retry budget that
helpers.cleanup carries.

Now uses createTempDir/cleanup from tests/helpers.cjs and passes an explicit
30s timeout. Suppressing the rules was available and would have been the wrong
call: both exist because of real CI failure modes on platforms I am not testing
on.

* chore(#3771): backfill the changeset PR number

The pr: field is drift-checked against the PR event payload, so it cannot be
written before the PR exists. Set to 3916.

* fix(#3771): close revision conflict persistence gaps

Use the authoritative review artifact, keep conflict and normal retry paths disjoint, and enforce one writer-reader grammar so malformed state fails closed.

Emitted-Drift-Ack-Growth: diagnose-issues.md — #3771 marks the gap-plan remediation hint non-binding while keeping root_cause authoritative
Emitted-Drift-Ack-Growth: gsd-plan-checker.md — #3771 separates binding required_property evidence from advisory fix_hint examples across the checker contract
Emitted-Drift-Ack-Growth: gsd-planner.md — #3771 declares the REVISION_CONFLICT return used when remediation contradicts governing constraints
Emitted-Drift-Ack-Growth: gsd-ui-checker.md — #3771 applies the same binding-property and advisory-hint split to UI review findings
Emitted-Drift-Ack-Growth: gsd-ui-researcher.md — #3771 defines the UI revision producer's structured REVISION_CONFLICT return
Emitted-Drift-Ack-Growth: plan-phase.md — #3771 routes and persists bounded revision conflicts before spending the normal retry budget
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #3771 adds the fail-closed owned-block parser and prevents convergence over open conflicts
Emitted-Drift-Ack-Growth: review.md — #3771 emits and preserves the canonical writer-owned conflict block across review regeneration
Emitted-Drift-Ack-Growth: ui-phase.md — #3771 routes UI revision conflicts to resolution before consuming revision_count
Emitted-Drift-Ack-Growth: verify-work.md — #3771 gives gap-plan revision the same bounded conflict route before iteration_count

* test(#3916): guard rebases against schema drift

Load the current-base progressive-disclosure examples so every integrated issue remains bound by required_property after branch reconciliation.

* test(#3916): skip the extracted-gate suite's bash spawns on win32

Third review round's sole survivor: runConflictGate()/withReviews() spawn
bash against a Node-native temp path built by createTempDir(), which is
backslash-separated on the Windows CI lane and not a path Git Bash is
guaranteed to accept (DEFECT.WINDOWS-TEST-PORTABILITY, matching the
observed CI failure at revision-remediation-binding.test.cjs:844,
ENOENT on a path Windows read as a directory separator). No eslint rule
catches it since the call has neither a chmod nor a `bash -c` form.

Guards the four call sites with the repo's existing skipOnWin32
convention (describe/test `{ skip: IS_WINDOWS }`) rather than
normalizing the harness path to forward slashes, which would defeat the
one test whose purpose is proving the production gate does NOT rewrite
a literal backslash in a POSIX filename.

* fix(#3916): backfill changeset pr field to the fork validation PR number

* fix(#3771): forbid silently accepting an open plan-revision conflict at max-cycles escalation

The max-cycles escalation prompt only surfaced HIGH_COUNT and ACTIONABLE_COUNT; an open
plan-revision conflict (OPEN_CONFLICTS > 0) was never disclosed there, and "Proceed anyway"
could exit successfully over it — exactly the failure mode this PR exists to close (a success
banner over an unresolved conflict nobody resolved). Blockers still block: withhold "Proceed
anyway" and route to Manual review whenever a conflict is open.

* fix(#3771): do not hard-block REVISION_CONFLICT persistence when no REVIEWS.md exists yet

A phase's first-ever revision cycle can return REVISION_CONFLICT before any REVIEWS.md has been
written — REVIEWS_PATH is then legitimately empty, not a corrupt or deleted file. The persistence
gate's own accompanying prose already says the record channel applies 'when REVIEWS_FILE is
non-empty', but the bash condition never checked that, so it hard-blocked every conflict on a
brand-new phase regardless of whether persistence was even expected to run. Require a non-empty
REVIEWS_FILE before treating a missing file as an error.

* chore(#3771): raise the plan-phase.md ADR-857 host-loop ceiling to 96700

The frozen pre-phase-6 ceiling (94519) collided on rebase: this PR's own
REVISION_CONFLICT persistence/routing gate is core planner control flow, not an
un-extracted optional feature, and landed alongside an unrelated, already-merged
same-file growth (the #4.6 context-drift pre-check) already on next. Same
rationale #1298 already established for execute-phase.md's ceiling.

* chore(#3916): backfill changeset pr field to the upstream PR number

* fix(#3771): make the writer-side REVISION_CONFLICT sanitize step real shell

The Conflict Return record channel sanitized agent-authored fields via a
prose instruction ("Sanitize each agent-authored field before appending")
for the orchestrator LLM to apply by hand, while the reader-side gate in
plan-review-convergence.md parses the same slot with real, executed awk.
Flagged Minor across two review rounds (round 4, round 6) since no code
performed the sanitize anywhere.

plan-phase.md's Conflict Return step now runs a real bash gate: sanitize
each field (collapse newline/tab to space, strip a leading #/-/|/fence),
build the one-line record, skip the append if an identical line already
exists (idempotent), insert before the writer-owned end delimiter, and
fail closed if that delimiter is missing rather than silently dropping
the conflict.

tests/revision-remediation-binding.test.cjs extracts and RUNS the new
fence (matching how the reader gate is already tested), composing it
with the existing reader gate: hostile-field fast-check fuzzing, a
repeated-conflict idempotency check, and a missing-delimiter fail-closed
check that the file is left byte-for-byte unchanged on failure.

Emitted-Drift-Ack-Growth: plan-phase.md — #3916 turns the writer-side
REVISION_CONFLICT sanitize+insert step into real, executed shell instead
of a prose instruction, matching the reader gate's existing rigor

* fix(#3916): backfill changeset pr field to the fork validation PR number

Fork CI's changeset-lint reads the real PR number from its own event
payload; the fragment still carried the upstream number from the last
sync, so the DEFECT.CHANGESET-PR-FIELD-DRIFT check failed on this fork
PR. Re-backfill to the upstream number before the final push.

* fix(#3771): close the awk -v forgery and same-session close gaps agy found

Adversarial review (gemini-3.8-flash-high via the internal agy review
lane) on the full PR found two BLOCKERs against the just-added
writer-side conflict gate:

1. `awk -v line="${LINE}"` decodes escape sequences in its argument, so
   a literal two-character `\n` in agent-authored text became a real
   newline inside awk, splitting the appended record across two
   physical lines. `tr` only strips actual control bytes, so it never
   saw this — it defeated the exact forgery the gate exists to
   prevent, both the reader's zero-count and the writer's own
   idempotency check. Fixed by passing LINE/END through awk's
   ENVIRON, which is not escape-decoded.

2. A conflict resolved and re-spawned within the same plan-phase
   session was never flipped from `- [ ]` to `- [x]` — the record
   channel bullet said "plan-phase closes it," but no step did. Only
   a *separate* `--reviews` re-entry (line ~622, still prose-only)
   closes conflicts; the in-session resolve path left them open
   forever, permanently blocking convergence. Fixed by carrying the
   just-written line in `PENDING_CONFLICT` and closing it in the
   `Otherwise` branch before the checker re-spawns.

Also fixed a MAJOR: docs/COMMANDS.md described the `--max-cycles`
escalation gate as uniformly offering "proceed or review manually,"
but the code (this PR's own change) withholds "Proceed anyway"
specifically when a plan-revision conflict is open — only manual
review is offered in that case. Docs now say so.

Not applied: the reviewer's `\r` truncated to plain tr from a MINOR
that also asked for temp-file permission preservation across `mktemp`.
Applying `chmod --reference` is not portable to macOS/BSD `chmod`, so
this is left as a documented low-severity tradeoff — the temp file now
sits alongside REVIEWS.md (same filesystem, atomic `mv`), which was
the same finding's more substantive half. Also not applied: a
suggested `gsd_run review record-conflict` CLI subcommand to
deduplicate the two `awk` blocks — a new command plus wiring is out of
scope for a review-remediation fix.

tests/revision-remediation-binding.test.cjs adds regression coverage
for both BLOCKERs: a literal-backslash-n hostile field composed with
the reader gate, and a close-gate extraction that verifies the flip to
`[x]`, the reader's count dropping to 0, and a fail-closed path when
the pending line is missing.

Emitted-Drift-Ack-Growth: plan-phase.md — #3916 fixes an awk -v escape-
decoding forgery and adds the missing same-session conflict-close step
an adversarial review found in the writer-side gate

* chore(#3916): backfill changeset pr field to the upstream PR number

Fork-validation CI needed pr: 1 to pass its own changeset-lint; restore
pr: 3916 before this push reaches open-gsd/gsd-core.

* fix(#3771): trim plan-phase.md prose back under the XL byte cap

Merging origin/next's unrelated growth pushed plan-phase.md 473 bytes
past the workflow-size-budget XL cap and the ADR-857 phase-6 baseline,
both tripped by CI after review approval. Removed an unpinned inert
bash comment and tightened connective prose in three REVISION_CONFLICT
bullets; no executable shell or test-pinned substring changed.

* chore(rehearsal): pin changeset pr field to fork rehearsal PR #25

Scratch-only commit for the rehearsal branch's own CI. Will not be
carried onto the branch backing upstream #3916 — that keeps pr: 3916.

* fix(#3771): address CodeRabbit findings on the REVISION_CONFLICT protocol

Fork rehearsal PR #25's first CodeRabbit pass surfaced 7 findings against
the already-approved #3916 diff; each verified against current code
before fixing (none hallucinated):

- plan-phase.md: writer-side awk gates now strip a trailing \r before
  comparing lines, matching the reader gate (plan-review-convergence.md)
  -- a CRLF REVIEWS.md previously made both writer gates fail closed.
- plan-phase.md: the close-fence's REVIEWS_FILE/PENDING_CONFLICT/
  CONFLICT_RESOLUTION were read without ever being (re)defined in that
  fence -- shell state does not survive across separate fenced blocks
  (same convention already documented in review.md). Added the explicit
  recompute/set instruction.
- revision-loop.md: previous_conflict_property was never reset after a
  normal (non-conflict) revision, so a later, unrelated conflict on the
  same property could be misread as a repeat and escalate prematurely.
- gsd-plan-checker.md / few-shot-examples/plan-checker.md: two example
  required_property strings were unconditionally binding in a way their
  own dimension's rules aren't (no-analog RESEARCH.md fallback; tasks
  that create no functions), now scoped to match.
- quick/steps/plan-checker-loop.md: added the same disjoint
  "Otherwise (not REVISION_CONFLICT)" branch plan-phase.md already had,
  closing an ambiguity between the conflict and non-conflict return paths.
- revision-remediation-binding.test.cjs: the REVIEWS_PATH init-order
  assertion used indexOf() without checking for -1, so it would pass
  vacuously if either anchor were renamed away.

Also restores an "Export the row's CONFLICT_*" instruction I had cut in
the prior byte-budget trim -- checked non-pinned by tests, but it was the
only text telling the agent to set those vars before the awk block reads
them via ENVIRON.

Net growth from these fixes required reclaiming bytes elsewhere in
plan-phase.md (verified against every pinned substring in
revision-remediation-binding.test.cjs) to stay under the XL tier's
hard 98304-byte cap; final size 98245 bytes.

* fix(#3771): resync the #4079 shrink-only mirror to the current PRE_PHASE6 line

tests/plan-phase-background-wait-wakeup.test.cjs (landed on next via an
unrelated #4079 PR, merged in by this branch's next-sync) mirrored
plan-phase.md's phase6 shrink-only ceiling as a hardcoded local constant
(94519) rather than reading tests/phase6-capstone-conformance.test.cjs's
PRE_PHASE6 value. That value has since been legitimately raised twice
during this PR's own review (94519 -> 96700 -> 98300) to accommodate the
REVISION_CONFLICT persistence/routing gate. The two branches' independent
histories left the mirror stale post-merge -- not a textual git conflict,
but the same class of thing. Resynced to 98300.

* fix(#3771): address round-2 CodeRabbit findings on the conflict gates

CodeRabbit's re-review of the previous remediation commit found two real
issues in what it had already flagged:

- Both writer-side awk CRLF fixes used \`sub(/\r$/, "")\` directly on \`\$0\`,
  which mutates it in place -- \`{ print }\` then emitted the CR-stripped
  copy for every passed-through line, silently rewriting an unrelated
  CRLF REVIEWS.md to LF on any insert or close. Now compares against a
  separate \`cur\` copy and prints the original, untouched \`\$0\`.
- The close-fence's "recompute REVIEWS_FILE/PENDING_CONFLICT" prose
  implied in-fence derivation, but the fence has no such code and the
  test harness (\`runCloseGate\`) deliberately supplies all three as
  pre-set env vars -- matching how the open fence's "Export the row's
  CONFLICT_*" instruction already works. Reworded to "export ... in the
  same invocation", matching that established, test-verified pattern
  instead of promising logic that isn't there.

Added a regression test proving the CRLF fix no longer touches
passthrough lines (red against the mutate-in-place version, green now).

* fix(#3771): use a CRLF-safe check in the new passthrough regression test

local/no-crlf-fragile-split forbids splitting readFileSync content on a
literal \n (Windows git-autocrlf checkouts yield \r\n). My CRLF
passthrough-preservation test from the previous commit did exactly that
to inspect the first line. Replaced with a direct startsWith() check
against the known CRLF-terminated header, which needs no split.

* test(#3771): assert the record itself is inserted in the CRLF passthrough test

CodeRabbit nitpick (round 3): the passthrough-preservation test checked
gate status and the pre-existing line's CRLF ending, but never asserted
the new REVISION_CONFLICT record was actually written.

* fix(#3771): address agy/gemini-3.8-flash-high adversarial review findings

Full-PR adversarial review (internal /gsd-review antigravity lane,
gemini-3.8-flash-high) surfaced 9 findings; each verified against current
code before fixing (none hallucinated):

HIGH:
- quick-batch/steps/plan-checker-loop.md never received the
  required_property/fix_hint binding language or REVISION_CONFLICT
  handling this PR added everywhere else -- a genuinely unmigrated
  producing context. Migrated to match quick/steps/plan-checker-loop.md,
  and added it to the ORCHESTRATORS consistency battery in
  revision-remediation-binding.test.cjs so future drift is caught
  automatically.
- The close-fence's PENDING_CONFLICT was an agent-supplied env var that
  had to exactly reconstruct a five-field sanitized line across a
  multi-minute subagent dispatch -- fragile, and a scalar var also meant
  a second simultaneous conflict silently dropped the first on overwrite.
  Redesigned to match the open conflict by CONFLICT_DIMENSION/
  CONFLICT_PLAN identity instead: the agent re-supplies two short,
  already-tracked identifiers rather than reconstructing the full
  sanitized text, and each conflict resolves independently regardless of
  how many are open. Updated the test harness's runCloseGate contract to
  match, and added a two-open-conflicts regression test.

MEDIUM:
- plan-phase.md's `--reviews` replanning path told the reader to "flip
  the matching line to [x]" in prose only, with no executable path to
  it -- pointed it at the same close gate used in step 12.
- plan-review-convergence.md's reader-gate awk tolerated a blank line
  before the opening delimiter but not before the heading that follows
  it; a formatter or LLM writer inserting one would hard-abort
  convergence on an otherwise well-formed REVIEWS.md. Added the same
  tolerance already granted above it, with a regression test.

LOW:
- Clarified that the escalation destination for a stalled conflict is
  the same iteration/revision-count cap gate already defined in each of
  quick, quick-batch, ui-phase, and verify-work, rather than an
  undefined "stall" concept.
- Clarified "twice in a row" means no successful revision intervened,
  matching revision-loop.md's now-explicit previous_conflict_property
  reset.
- Fixed gsd-ui-researcher.md's stale rationale text, copied verbatim
  from planner-revision.md: ui-phase presents the conflict table
  directly to the user, it does not persist to a shared file scanned by
  heading.

Net growth again required reclaiming bytes in plan-phase.md (verified
against every pinned substring in revision-remediation-binding.test.cjs)
to stay under the XL tier's hard 98304-byte cap; removed a now-dead
PENDING_CONFLICT assignment in the process. Final size 98258 bytes.

* fix(#3771): scope row 48's quick/steps guard away from plan-checker-loop.md

tests/gsd-quick-batch-quick-regression.test.cjs's row 48 (#3676) flagged
this branch's quick-batch/steps/plan-checker-loop.md migration (the agy
HIGH finding) as a violation, because it also edits
quick/steps/plan-checker-loop.md for the same underlying #3771 protocol
fix.

Verified against git history before scoping: 2f64e6230 (#3676's own
landing commit) CREATED quick-batch/steps/plan-checker-loop.md as a new,
independent 119-line file, never a call-site into quick/'s copy. Row
48's "shared primitives, never edits the ordinary quick command" premise
was never about this specific file -- it was always meant to carry its
own per-flow copy of whatever revision-loop contract applies, same as
ui-phase.md/verify-work.md throughout this PR. This is the same
false-positive class the row's own comments already document scoping
away twice (#3730, #2529 round 40); excluded plan-checker-loop.md from
its touched-quick-steps check with the same evidence trail.

* chore(#3771): point changeset pr field at upstream PR 3916

---------

Co-authored-by: davdittrich <davdittrich@gmail.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 15:16:38 -04:00
Michel Moreira
86b745b48b fix(#4270): forward Codex spawn model routing (#4281)
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 14:20:26 -04:00
Tom Boucher
7ff196c505 fix(#4096): honor --dry-run in todo complete and write completion keys inside the frontmatter fence (#4325)
* fix(#4096): honor --dry-run in todo complete and upsert completion keys inside the frontmatter fence

* review(#4096): tighten todo complete flag rejection to any dash-prefixed token

* chore(#4096): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 13:49:58 -04:00
Tom Boucher
2e1ede6d99 fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline (#4318)
* test(#4093): regression matrix for advance-plan zero-labeled-fields decline

* fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline

* refactor(#4093): collapse IIFE to a plain block (review finding)

* docs(#4093): document the advance-plan recovery decline + changeset

* chore(#4093): backfill PR number in changeset

* fix(#4093): budget lint-compiled-artifact-sync's tsc compile as a compile, not a probe

---------

Co-authored-by: sim <sim@local>
2026-09-05 10:46:37 -04:00
Behruz Nassre Esfahani
5d804dd287 fix(#3709): clear the context-monitor warn sentinel on PreCompact (#3808)
* fix(#3709): clear the context-monitor warn sentinel on PreCompact

The monitor's per-session warn sentinel survived a compaction, so once the
first CRITICAL of a session had fired, `lastLevel` stayed pinned at 'critical'
for the rest of the run. The hook was already wired to PreCompact (#772), but
read the event only at the very END, and solely to pick an output envelope.

Two documented behaviours died as a result:

  - "First warning always fires immediately" — the first warning of the
    post-compaction cycle was debounced instead.
  - "Severity escalation (WARNING -> CRITICAL) bypasses debounce" — computed as
    `lastLevel === 'warning'`, which can never be true again, so every later
    CRITICAL waited out the full five-tool-use debounce, exactly when an
    immediate warning matters most.

`criticalRecorded` was equally sticky: a session that compacted and later truly
ran out kept a /gsd:resume-work breadcrumb (#1974) describing the earlier
near-miss rather than the exhaustion that ended the run.

Reproduced first, with the issue's own literal repro, including the detail that
the compaction consumed a debounce slot (callsSinceWarn 0 -> 1).

The reset runs BEFORE the metrics read, deliberately: a post-compaction reading
is healthy again, so the ENOENT / stale / above-threshold branches would all
exit first and never reach it. Returning early also stops the compaction from
eating a slot of the cycle it was meant to restart. The event name is now read
once through a shared `readEventName()` helper, so this reset and the #2289
output allowlist cannot drift on what counts as "no event name".

Seven rows against a real sequence (the defect is state carried ACROSS calls, so
they need their own driver — the existing helpers delete the sentinel after each
invocation). Reverting the reset turns SIX of them red; the seventh is the
non-vacuity row asserting a NON-compaction event must not clear the sentinel,
which correctly passes either way.

AC4 initially passed with and without the fix — asserting `criticalRecorded ===
true` is vacuous when the seeded stale sentinel already carries it. It now seeds
a `staleProbe` marker that can only survive if the sentinel survives, so its
absence is what proves the state was rebuilt.

hooks/dist/ is gitignored and regenerated by build:hooks, so no committed dist
copy needs syncing.

Verified: `npm run lint:ci` exit 0; acceptance criteria 1-6 driven end-to-end
against the real hook.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3709): reset ahead of the config gate, and pin the placement itself

Codex review of the #3709 fix, before opening the PR. Three findings, all in
this change's own new code.

1. `context_warnings: false` prevented the reset. The config early-exit sits
   ABOVE where the reset was placed, so a session that disabled warnings,
   compacted, then re-enabled them mid-session resurrected the stale sentinel
   and the original bug with it. Config is re-read per invocation, so that
   sequence is supported rather than hypothetical. The reset now runs ahead of
   the config gate: clearing the sentinel is CLEANUP, not a warning — state that
   must not outlive a compaction should not outlive it merely because warnings
   are switched off right now. It cannot emit anything from there, so the
   disabled contract is untouched.

2. Nothing pinned the "before the metrics read" placement. Every row wrote a
   fresh metrics file, so the reset could have been moved below the metrics
   read, the stale check, or the healthy-threshold exit with all seven rows
   still green — while a REAL PreCompact, which carries no fresh metrics and
   follows a recovery to healthy usage, silently kept its sentinel. Three rows
   now pin it: no metrics file at all, usage recovered to healthy, and warnings
   disabled. Each catches a distinct wrong placement — moving the reset below
   the config check reds the third; below the metrics read reds all three.

3. The absent-sentinel row proved nothing. `assert.doesNotThrow` was vacuous
   because the driver caught every child exit, so a hook that exited 1 on the
   ENOENT unlink would still have passed. The driver now returns the exit code
   and the row asserts it is 0.

Also corrected the `readEventName` comment: it said the event is "read once",
which is not literally true — there are two call sites. The point is one
DEFINITION of what counts as an event name, so the reset and the #2289
allowlist cannot drift; the comment now says that.

Verified: 60 rows in tests/perf-317-context-monitor-fs.test.cjs, 0 fail, with
both placement mutations driven to red and reverted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3709): backfill changeset pr number

The fragment shipped with the documented `pr: 0` placeholder, which the
changeset lint treats as always-silent, because the PR number does not exist
until the PR is opened. Backfilled to 3808 now that it does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3709): a compaction clears the stale reading too, not just the state

Review round 1. Major 1 was right and it mattered: clearing only the sentinel
traded a warning that never fires for one that fires when it must not.

The statusline bridge still holds the PRE-compaction reading, and STALE_SECONDS
is 60, so for up to a minute it still reads fresh and still says the context is
exhausted. With the sentinel gone, firstWarn is true, so the next PostToolUse
emitted a spurious CONTEXT CRITICAL immediately after the compaction that FREED
the context — and flipped criticalRecorded, spawning a false context-exhaustion
breadcrumb. That is the same breadcrumb inaccuracy #3709 exists to fix, re-entered
from the other side. Reproduced before fixing, exactly as the review described.

A compaction now invalidates the warning state AND the reading that produced it.
Removing the bridge loses nothing: the statusline owns that file and rewrites it
on every render, and its absence is already the "no reading yet" state a fresh
session starts in, which exits silently.

Two things my own verification caught while fixing it:

  - The first attempt did NOTHING. metricsPath was declared below the PreCompact
    block, so referencing it hit the temporal dead zone, threw, and the outer
    catch swallowed it into a silent exit 0. The probe still printed "silent",
    which looked like success but was the old debounce. metricsPath is now
    hoisted beside warnPath.

  - The new Major 1 row was VACUOUS. The driver's `metrics: false` DELETES the
    bridge, but the defect is a bridge that is still there and still reads fresh,
    so the row passed on the ENOENT early-exit rather than on the fix. Only the
    sentinel-only mutation exposed it. The driver grew a `metrics: 'keep'` mode
    that leaves the stale file in place; both Major 1 rows now red under that
    mutation.

Also from the review:
  - Minor 1 — the compaction-abort path is now stated in the source rather than
    left silent, including why a conditional reset (SessionStart source "compact")
    is out of scope for this fix.
  - Minor 2 — docs/context-monitor.md completed: PreCompact wiring and the early
    return under How It Works, a table of all three things the reset clears, the
    breadcrumb guard, the warnings-disabled interaction, and the never-block
    property under Safety.
  - Minor 3 — changeset trimmed from ~1,400 chars of implementation narration to
    the user-visible change.
  - Nit 1 — a failed unlink (Windows EPERM/EBUSY) no longer leaves the bug
    silently intact: the file is neutralised in place instead, with a shape safe
    for each (an empty sentinel, a timestamp-0 bridge).
  - Nit 2 — reviewer-process narration removed from shipped test source. The
    remaining "Codex" mentions are pre-existing and name the RUNTIME.
  - Nit 3 — the debounce-slot row now asserts the observable consequence (the
    first post-compaction warning fires) rather than repeating AC1's assertion.
  - Nit 4 — the file docblock now lists the folded-in blocks and asks the next
    contributor to extend it.

Verified: `npm run lint:ci` exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3709): the unlink-failure fallback truncates to empty, matching deletion

The fallback wrote well-formed neutral values, and neither was equivalent
to the deletion it stood in for: '{}' parses, so firstWarn was false and
the first post-compaction warning was debounced — AC2 undone on exactly
the path the fallback exists for — and '{"timestamp":0}' was never stale
(the guard is `metrics.timestamp && ...`), so the flow reached emit with
remaining === undefined and injected a literal 'Usage at undefined%'.
Truncating to '' makes JSON.parse throw on both reads: the sentinel read
keeps firstWarn true, the bridge read falls to the outer catch and exits
0 silently (review of #3808, Blocker 1).

The branch is now executed for real: an EPERM is injected into the
child's fs.unlinkSync via --require preload — method monkeypatching,
never chmod 0o000, which root bypasses under Docker/CI (Blocker 2). Both
rows proved failing-first against the neutral-value fallback. The
boundary trios at WARNING=35 / CRITICAL=25 are completed on the emit
path with 34, 26, and 24 (Major 3); 36/35/25 were already pinned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3709): the truncation fallback refuses to follow a planted symlink

The per-session files live in a shared sticky tmpdir, where an unlink
failing EPERM is exactly what another user's planted file produces — and
a planted SYMLINK would make the fallback's plain truncating write empty
out its TARGET, weaponising the hook against any file its own user can
write. Open with O_WRONLY|O_TRUNC|O_NOFOLLOW instead: a symlink fails
ELOOP into the same give-up arm. On Windows the constant is absent and
'|| 0' keeps the fallback alive there, where the held-handle case it
exists for occurs and temp dirs are per-user. Found by Codex review;
the new row proved failing-first against the writeFileSync fallback
(victim file truncated to zero bytes).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3709): refuse non-regular files everywhere, not only where O_NOFOLLOW exists

Codex round 2: '|| 0' removed the no-follow protection exactly where it
cannot be expressed as an open flag — Windows, whose tmpdir is NOT
guaranteed per-user (TEMP/TMP overrides, system-temp fallback). An
lstat isFile() guard now rejects symlinks and every other non-regular
shape on all platforms before the truncating open; O_NOFOLLOW stays, as
the lstat->open substitution-race backstop where the platform has it.
The symlink row additionally asserts the planted link SURVIVES the call,
so a preload match that stops engaging can no longer pass the row
vacuously off a successful unlink.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(#3709): tolerate the Windows give-up, still outlaw neutral values

Both windows-latest CI lanes fail the two EPERM rows deterministically:
the runners hold freshly written files with a share mode that allows
DELETE (every real-unlink row passes) but refuses a truncating
write-open, so the fallback's give-up arm engages — which is the
fallback working as designed, not the defect the rows exist to catch.
The rows are now platform-aware: POSIX still requires exact truncation
and the behavioural follow-ons; Windows accepts truncated-or-untouched
but still rejects the Blocker-1 regression class (a parseable neutral
value is never legal anywhere), with the follow-ons gated on the
truncation actually landing. Also corrects the hook comment: libuv
defines O_NOFOLLOW as 0 on Windows — a no-op, not an absent constant.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3709): a compaction watermark closes the window bridge deletion only narrowed

Round-3 Major 1: the statusline is an uncoordinated process that
re-writes the bridge on every render, so a render landing between the
PreCompact clear and the compaction's completion re-created the
PRE-compaction reading under a CURRENT timestamp — past STALE_SECONDS,
into a spurious post-compaction CRITICAL and a false exhaustion
breadcrumb: the exact failure the deletion was added to prevent.
PreCompact now also writes claude-ctx-<id>-compacted.json ({at}) and the
metrics read drops any reading not STRICTLY newer than it — which also
covers unstamped/zero timestamps once a compaction happened. Written
unlink-then-O_EXCL so a planted file or symlink is never followed;
failure degrades to the old narrowing. Docs and changeset now describe
the watermark instead of overclaiming for the deletion.

Round-3 Major 2: DEBOUNCE_CALLS and STALE_SECONDS get their trios — the
gate increments BEFORE comparing, so seeds 3/4/5 pin 4-debounced,
5-emits, 6-emits; ages 59/60/61 pin the strict >. The child's clock is
pinned via a --require preload (a wall-clock boundary row would flip on
one second of startup delay). timestamp-0's falsy bypass is pinned
directly as characterized behaviour. Mutation-proven: dropping
O_NOFOLLOW, <= for <, and >= for > each red exactly one row.

Minors: the symlink row's comment now names the lstat guard it actually
pins, and a preload-blinded-lstat row drives the O_NOFOLLOW substitution
-race backstop for real (3); absence assertions use warnRaw so a
corrupt leftover cannot pass as deleted (4); the Windows give-up is an
explicit t.skip, never a silent if (5); readEventName is total via
String(), keeping #2289's side-effects-always-run contract for
malformed event names, with a row (6); the PreCompact rationale lives
once in docs/context-monitor.md with the code keeping only line-level
constraints (9); the changeset is release-note-sized (10).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3709): the grace window covers the compaction's duration, not just its start

Codex on the first watermark cut: the watermark stamps the compaction's
START, so a statusline render one second later — still mid-compaction,
still the old reading — passed 'strictly newer' and re-fired the false
CRITICAL. Readings inside COMPACT_GRACE_SECONDS (60) past the watermark
are now dropped: the window covers the compaction's own duration, a
healthy reading dropped there behaves identically to an accepted one
(it exits above-threshold anyway), and a genuine exhaustion warning is
delayed at most one window after a compact. A watermark stamped ahead
of the reader's clock is ignored — a clock step backwards or a stray
file must degrade to plain staleness, never mute the monitor
indefinitely. Both proven failing-first.

readEventName is strict about TYPE, not coerced: String() rendered
['PreCompact'] as 'PreCompact' and would run the reset off a malformed
payload. typeof: every non-string is 'no event' — silent, side effects
intact — with rows for the number, hostile-object, and array-wrapped
cases. The lstat-claim preload arm now writes an engagement marker the
substitution-race row asserts on, so a match string that silently stops
matching can no longer let the row pass off the real lstat guard.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: retrigger CI — the previous wave was cancelled by an Actions outage, zero job failures

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3709): drive the compaction rows on the clock, not on a future stamp

Round-4 review raised three majors, all in the test scaffolding around the
fix rather than in the fix itself.

Major 2 (taken first — it is the cheapest and it unblocks Minor 6): call()
passed process.env to the child unmodified, so two rows depended on ambient
GEMINI_API_KEY. The preserved Gemini fallback is `eventName === "" &&
!!process.env.GEMINI_API_KEY`, and readEventName returns "" for every
malformed name, so with the key set the malformed-event row's `stdout === ''`
assertion failed outright — reproduced by running it under GEMINI_API_KEY=x.
call() now takes an explicit env, the way the sibling runMonitor helper in
this file always has, and both rows pin the variable unset. (The array row
survived an ambient key only because its reading was debounced — incidental,
not independence, so it is pinned too.)

Major 1: the AC2/AC3 rows drove the hook with a bridge stamped 62 seconds in
the FUTURE — a shape hooks/gsd-statusline.js cannot produce, since it always
stamps Math.floor(Date.now()/1000) on the same clock. They proved "the
sentinel was cleared" while their assertion messages claimed the documented
immediate-warning behaviour, which is gated behind the grace window and went
unexercised. Both rows now run the real sequence on the clock-pinning preload
this PR already added for the STALE trio: PreCompact at a fixed instant, then
a normally-stamped render one second past the window. Verified non-vacuous —
stubbing the sentinel unlink reds both.

Major 3: COMPACT_GRACE_SECONDS, the one constant this PR introduces, was the
only threshold without a limit-1/limit/limit+1 trio, in a PR that adds full
trios for four pre-existing ones. The seeded offsets were +0, +1 and +61; the
boundary itself (+60) and limit-1 (+59) were untested. Added, driven by
advancing the reader's clock rather than post-dating the reading, so the
reading is never ahead of the reader and only the grace gate can drop it.
Verified against three mutations — `>` to `>=`, the constant to 59, and the
constant to 61 — each of which reds exactly one row of the trio.

No production code changed. Verified: 85/85 in this file, lint:ci exit 0,
and the two Minor-6 rows now pass under GEMINI_API_KEY=x as well as unset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG

* fix(#3709): harden the watermark read and pin the thresholds it introduces

Codex review of the full PR found four majors. All reproduced here against
the real hook before fixing.

MAJOR — the watermark was write-hardened but read-untrusted. PreCompact
already refuses to follow or overwrite a planted object (unlink-then-O_EXCL),
but the read was a bare readFileSync, so anything the write side gave up on
was followed by every later invocation. In a shared sticky os.tmpdir() that
is a mute primitive — a planted recent watermark suppresses monitoring — and
a symlink to a FIFO stalls a synchronous read. Measured against the
pre-hardening file: a symlink to a planted watermark WAS honored and muted
the monitor. The read now uses the same lstat + O_NOFOLLOW pair the sentinel
path uses, plus a size bound; symlink, directory and oversized cases are all
refused, with a plain-file control proving watermarks still work.

MAJOR — the `now + 5` skew tolerance was an unnamed, untested threshold. It
is now WATERMARK_SKEW_SECONDS with a +4/+5/+6 trio, verified against two
mutations (`<=` to `<`, and the constant to 6), each of which reds one row.
This is the same class as round 4's Major 3, one layer up.

MAJOR — the malformed-event row shared one session across both subcases, so
the hostile-object iteration's `assert.ok(s.warn())` passed off the sentinel
the `42` iteration left behind. A regression throwing before the bookkeeping
would have kept it green — vacuous for exactly the subcase it exists for.
Fresh session per subcase, with an explicit no-sentinel precondition.

MAJOR — the stale-reading row's non-vacuity is an artifact of call()'s future
stamp: with a production stamp the watermark suppresses the same reading, so
the row cannot isolate bridge deletion. The two guards genuinely overlap
inside the window, so no end-to-end row can separate them; the comment now
says so and points at the direct pin (s.metrics() === null) instead of
claiming an isolation it does not have.

Docs corrected where measurement contradicted them: the window NARROWS the
race rather than covering the compaction's duration, and the delay is not
bounded by the window alone — first recovery is watermark+61s with no skew
but watermark+66s at the accepted +5s skew. Aborted compactions are muted
the same way. The truncation fallback is documented as best-effort, which is
what the code and the Windows rows already do.

Verified: 89/89 in this file, lint:ci exit 0, symlink/directory/oversize all
refused where the pre-hardening file honored them, both new trios
mutation-checked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG

* fix(#3709): move the PR's two new exits onto the declared-policy vocabulary

#3911 / ADR-3889 migrated this hook off raw process.exit() while this PR was
in review, replacing every exit with hooks/lib/hook-exit.js's allow(), which
forces each call site to name its crash policy. The PreCompact reset and the
watermark gate are added by THIS PR, so they did not exist to be migrated and
came through the merge as the only two raw exits left in the file — caught by
the new local/require-registered-exit rule. Both are ALLOW: a compaction is
never blocked by this hook, which is the policy the rest of the file declares.

Caught only in CI, not locally: `npm run lint` runs eslint with --cache, and
the cached entry for this file predated the new rule, so a warm local cache
reported clean. Re-verified with the cache cleared.

allow() terminates rather than throwing, which matters for the watermark call
site because it sits inside a try/catch — a throwing helper would unwind into
that catch and silently drop the grace-window mute. Verified behaviourally,
not by reading: the grace trio, the skew trio and the non-regular-file rows
all still pass.

Verified: lint:ci exit 0 with a cold eslint cache, full suite exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG

* fix(#3709): harden the routine sentinel writes and the read beside them

Round 7 ruled that the three routine debounce-accounting writes to the warn
sentinel must match the three writes this PR already hardened: leaving the
fourth unhardened beside them is the asymmetry that invites the defect back.
They now go through one writeSentinel() helper using the compaction
watermark's own unlink-then-O_EXCL shape, rather than a second policy — the
unlink removes any existing object, and O_EXCL then refuses to create through
one, so a write can only land on a fresh regular file this process made.

The routine READ beside them was the last bare readFileSync on warnPath, and
the same rationale applies to it verbatim; the watermark's read was hardened in
round 4 for exactly this reason. Same lstat + O_NOFOLLOW + size bound. Its
scope is stated in the test rather than overclaimed: lstat establishes that the
sentinel is a plain regular file, not that it is trustworthy, so a cross-owner
regular file at the predictable path is still read and is left as a disclosed
pre-existing residual.

Also fixes an accept-direction regression this PR introduced and six rounds of
review missed. readEventName collapsed an ABSENT event name and a MALFORMED one
onto the same '', and the preserved Gemini fallback keys off eventName === "",
so with GEMINI_API_KEY set a malformed payload began emitting an AfterTool
envelope. At the merge-base, data.hook_event_name.trim() threw on a truthy
non-string after the side effects and nothing was ever emitted. Measured
base-vs-head with a fresh sentinel per run: 42, ['PreCompact'] and {} all went
silent -> EMITS, while an absent name and 'PostToolUse' were unchanged.
readEventName now returns '' only for an absent name and null for a
present-but-non-string one; both call sites compare for equality only, so every
well-formed payload behaves identically.

Five new rows, each proven fail-first with the mutations attributed separately:
reverting the writes reds the write-through and non-regular rows, reverting the
read reds the mute and non-regular rows, and reverting the absent/malformed
split reds the Gemini row. The changeset's "behaves like a fresh session" is
narrowed to name the 60-second suppression window and the best-effort reset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018FUAVz49BghqxoJgwt7EW9

* test(#3709): pin both 4096-byte read bounds at their boundaries

Round 8 asked for limit-1/limit/limit+1 coverage on the size bound the
round-7 sentinel read-hardening introduced (gsd-context-monitor.js:335).
The existing refusal row pads to 8192 -- a full 4096 bytes clear of the
fence -- so `>` vs `>=`, or an off-by-one in the constant itself, was
invisible to it.

Covers the sibling bound too. The identical check guards the round-4
WATERMARK read at :278 and its refusal row pads to 8192 in exactly the
same way; the review's own rationale (this file already holds
WATERMARK_SKEW_SECONDS to a boundary trio, so an uncovered bound is the
odd one out) applies to it unchanged. That half is a class sweep of a
pre-existing bound and is test-only -- say the word and it comes out
without touching the rest.

Both trios assert on observable hook output rather than an internal
error. Sentinel: an honored {callsSinceWarn:1,lastLevel:'warning'} keeps
the debounce arm taken at remaining=30, so nothing is emitted, while a
refused one falls back to first-warn defaults and emits. Watermark:
honored mutes (stdout empty), refused leaves the warning. The 4097 row is
the non-vacuity control for the two accept rows. Payloads are sized by
measurement, with Buffer.byteLength asserted to equal the target, not by
arithmetic on an assumed prefix width.

Proven fail-first in both directions, with the hook restored after:
`> 4096` -> `>= 4096` reds both trios (94/96); `> 4096` -> `> 4097` reds
both trios (94/96); restored, 96/96. Under both mutations only the two
new rows fail -- the pre-existing 8192-padded rows stay green, which is
the review's fencepost claim demonstrated rather than assumed.

* fix(#3709): correct the changeset's mute-window claim and a superseded comment

Both from the pre-push Codex pass on the full PR.

The changeset said readings are "suppressed for up to 60 seconds after a
compaction starts". That is false at the accepted skew boundary, and this
repo's own docs/context-monitor.md already carried the accurate figure:
first recovery is watermark+61s with no skew and watermark+66s for a
watermark at the +5s skew limit. Measured independently at +64 silent,
+65 silent, +66 warning. The changeset now states the window plus the
accepted skew, matching the doc rather than contradicting it.

A comment in the malformed-event row still described readEventName as
returning "" for every malformed name. Round 7 superseded that: a
present-but-non-string name returns null and only an ABSENT one returns
"", so a malformed payload can no longer reach the Gemini fallback at
all. Marked as historical and corrected. The GEMINI_API_KEY pin stays --
the row is about readEventName's typing, not the fallback, and an ambient
key would still change what it measures.

Codex's three Major findings are not taken, on attribution rather than
logic; the reasoning is in the PR reply. In short: the watermark does not
exist at the merge-base at all (0 occurrences), so "base emits, HEAD
mutes" compares a new feature against its absence rather than showing a
regression; and the base sentinel read is a bare readFileSync, which
blocks on a planted FIFO exactly as the hardened read would, so the
TOCTOU stall is not introduced here. The underlying limits -- watermark
provenance, and lstat->open races on a non-symlink substitution -- are
real, pre-existing, and already offered to the maintainer as follow-ups.

* fix(#3709): read both sentinels through one hardened helper; state the two limits precisely

Round 9's Major, with a correction to its premise, and both Minors.

The review names "watermark read/write helpers this PR adds" that a call
site at :238-250 duplicates inline. There are no such helpers: this PR
adds readEventName and writeSentinel, the latter a write-side primitive a
read cannot call, and :238-248 is base code the diff never touched. What
IS duplicated is the hardened READ. The watermark read (round 4) and the
warnPath read (round 7) are the same ten lines twice -- lstat, isFile and
a 4096-byte bound, O_RDONLY|O_NOFOLLOW, readSync, close -- differing only
in the path variable and the error string, and that is two copies to keep
in step by hand. Now one function, readSentinel(target), beside
writeSentinel. Refusal throws; both callers already wrapped the read in a
try/catch that degrades to "no file", so behaviour is unchanged by
construction.

Proven rather than assumed: with the helper replaced by a bare
readFileSync in a complete scratch tree, exactly the five hardened-read
rows in tests/perf-317-context-monitor-fs.test.cjs go red -- round 7's
symlinked and non-regular sentinel and its size bound, round 4's
non-regular watermark, round 8's watermark size bound -- so the helper
carries both call sites' guarantees and the rows pin it. 96/96 with the
helper in place.

Minor, drop vs delay: the grace-window comment said "dropped" on one line
and "delayed" three lines later, and docs/context-monitor.md said
"delayed". A genuine exhaustion reading inside the window is skipped, not
queued: its warning and its #1974 breadcrumb both fire on the next reading
after the window, so both are delayed when a later reading comes and lost
when none does -- a session ending inside the window records neither.
Comment and docs now say exactly that, and that the loss is accepted over
trusting a reading that may be the pre-compaction value under a fresh
timestamp.

Minor, ordering: the PreCompact unlink and the debounce
writeSentinel(warnPath) are two writers with nothing serialising them; a
debounce invocation that read pre-compaction state and lands its write
after the unlink would resurrect the sentinel the reset removes. The hook
relies on the host dispatching a session's hooks one at a time, which
Claude Code does and the other runtimes are assumed to. Stated at the
reset as an assumption, with the lock-file alternative named and not
taken.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

* fix(#3709): write the compaction watermark through writeSentinel

Review of #3808, round 10. The PreCompact watermark write was the block
writeSentinel was lifted from in round 7, and it kept its own inline copy
of unlink-then-O_EXCL a few lines below the helper. Round 9 flagged that
write-side duplication; the round-9 reply misread it as the read side and
unified only the reads. The write now calls the helper too, so the hook
holds one copy of the hardened write, not two.

Behaviour is unchanged: same unlink-then-O_EXCL sequence, same flags,
same best-effort outer catch. The one difference is that writeSentinel
closes the descriptor in a finally, where the inline copy leaked it if
writeSync threw before closeSync.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R7bQLXAKubb4EFLtCPiLiL

* fix(#3709): read the statusline bridge through the same hardening as the sentinels

Review of #3808, round 11. `metricsPath` is built one line from `warnPath` and
`watermarkPath` — same tmpdir, same predictable `claude-ctx-{sessionId}` shape,
same threat model this PR documents at length for its siblings — and it is the
only one of the three read on EVERY invocation. It was also the only one still
reached by a bare `readFileSync`, so the symlink-follow and the symlink-to-FIFO
stall that rounds 4 and 7 closed on the other two stayed reachable here, on the
file's highest-traffic path. It now goes through `readSentinel` like the rest.

The 4096-byte bound is ample for it: the statusline writes four fixed fields
(`gsd-statusline.js`), about 140 bytes with a UUID session id, so no legitimate
bridge approaches it. A refusal lands in the same rethrow an unreadable or
malformed bridge already did.

The comment introducing `readSentinel` claimed the warn sentinel was "the one
bare readFileSync". Read as scoped to `warnPath` that was true, but it reads as
a claim about the file and it is not one — the bridge kept its own until this
round. Corrected rather than left to mislead the next reader.

Round 11 Minor: `readSentinel` discarded `fs.readSync`'s return value and
assumed the buffer was full, so a file truncated between the `lstat` and the
read left a zero-filled tail. It now refuses a short read. Stated plainly
because it was measured: this guard has NO observable behavioural delta —
deleting it leaves the new row green, because the NUL tail makes `JSON.parse`
throw one line later and both paths degrade to "no sentinel". It is a
consistency fix in a function whose purpose is refusing to trust what it read,
and the test comment says exactly that rather than implying coverage it lacks.

Five rows added: the bridge refusing a planted symlink (with an attacker-chosen
reading that WOULD warn if followed, so silence is proof), a non-regular bridge,
an oversized bridge, the shrink path end to end, and the direction that matters
most — a healthy bridge still warns, so the hardening is not a mute. Proven by
mutation: reverting the bridge to `readFileSync` reddens two rows. The shrink
injection carries an engagement marker for the same reason the lstat-claim one
does, learned the same way: the hook rewrites the sentinel later in the
invocation, so a size check afterwards passes whether the truncation landed or
not.

An independent full-PR pass on this round added two more, both taken:

`writeSentinel` discarded `fs.writeSync`'s return value, and a short write is
permitted by the syscall — so a truncated sentinel could reach disk and every
later read would reject it, silently losing the debounce accounting or the
watermark this write exists to record. It now loops until the payload is
written, as Node's own `writeFileSync` does, with an explicit no-progress guard.
Pinned by a row that injects a one-byte first write; reverting the loop reddens
it.

The directory row's comment claimed it pinned the `lstat` isFile() check. It
does not — measured: deleting that condition leaves the row green, because
reading a directory fails on its own a line later. The comment now says the row
pins the outcome, and names the symlink row as the one that pins isFile().

DISCLOSED, NOT FIXED HERE — a session id long enough to push the derived
filenames past NAME_MAX. The bridge is `claude-ctx-{id}.json`; the sentinel and
watermark add longer suffixes, so on a 255-byte limit the watermark stops fitting
at a 230-character id and the sentinel at 233. Measured base-vs-HEAD at 233+:
base is SILENT, HEAD emits the warning, because the bare `writeFileSync` base
used threw ENAMETOOLONG out of the warning path while `writeSentinel` degrades
best-effort and lets the warning through. That is an accept-direction delta and
it is in the delivering direction — base swallowed a warning the user should
have seen, which is this issue's own failure class. The underlying limit is a
property of the per-session filename scheme, shared by two files that predate
this PR, and bounding session ids belongs to whatever writes them
(`gsd-statusline.js`), not to the sentinel logic. Happy to fold a length guard
in here if you would rather have it in this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 06:29:06 -04:00
Dennis Alexis Valin Dittrich
5869febb16 enhance(#4155): invalidate verification results when covered inputs change (#4290)
* enhance(#4155): invalidate verification results when covered inputs change

readVerificationStatus() now recomputes a deterministic sha256 fingerprint
over a VERIFICATION.md's declared covered_files (phase PLAN/SUMMARY,
requirements, implementation files in the verified change set) and returns
stale on any mismatch, fail-closed when a covered file is missing,
unreadable, or escapes the project root. Legacy reports with no fingerprint
metadata keep the prior SUMMARY-mtime staleness check unchanged.

The verifier computes covered_digest via the new verification.fingerprint
CLI command rather than by hand, since a digest is deterministic math, not
an LLM-estimated value.

* chore(#4155): backfill fork PR number in changeset

* fix(#4155): trim gsd-verifier.md fingerprint instructions to fit LARGE tier byte cap

* fix(#4155): address CodeRabbit findings on fingerprint fail-closed behavior

Partial fingerprint metadata (one of covered_files/covered_digest present,
the other missing or malformed) now fails closed to stale instead of
silently downgrading to the legacy mtime-only check. computeCoveredDigest
also canonicalizes with realpathSync before re-confining, so an in-root
symlink whose target escapes the project root can no longer produce a
matching digest. gsd-verifier.md restores the completeness requirement and
checklist item trimmed by the earlier size-budget fix, within the LARGE
tier byte cap.

* chore(#4155): acknowledge gsd-verifier.md growth for the #4155 fingerprint instructions

Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the covered-input fingerprint instructions and frontmatter fields the #4155 verification staleness mechanism requires; trimmed to stay within the LARGE tier byte cap

* fix(#4155): address gemini adversarial review findings

computeCoveredDigest now threads the caller-supplied opts.fs seam through
its confinement and read paths instead of always using raw node:fs — a
caller like planning-inspect.cts's containmentEnforcingVerificationFs (GAP
2, #2790 follow-up) was silently bypassed for covered-input reads. The
project-root anchor itself still canonicalizes through real fs (it is a
trusted value the caller derived, not attacker-influenced covered-input
data); only per-file candidate reads go through the injected seam.

Covered-file paths are now canonicalized (./ prefixes, redundant slashes,
internal .. segments) before becoming dedup/sort/hash keys or confinement
subjects — closes both a spurious-stale false positive (two spellings of
the same file hashing differently) and a confinement gap (an internal ..
segment that doesn't start the string).

gsd-verifier.md now states covered-file paths are project-root-relative,
not phaseDir-relative, closing an ambiguity that would have made a real
verifier agent's first fingerprint invocation fail closed.

defaultFsImpl's methods now late-bind through fs.<method> rather than
capturing function references at module load — the earlier direct-capture
form was invisible to existing tests' t.mock.method(fs, 'statSync', ...)
seams, a real regression caught by the full suite (not the reviewer).

* fix(#4155): catch a plan/summary added to the phase dir after verification but never declared

The content digest only recomputes hashes for paths the verifier actually
declared in covered_files — it had no way to notice a plan or summary
added to the phase directory after verification if that new file was
never declared, silently regressing behind the legacy mtime check it
replaces (which scans the live directory, not a declared list).

findUncoveredCurrentArtifact re-scans the live phase directory for every
current *-PLAN.md/*-SUMMARY.md and requires each to be represented in
covered_files, closing that gap; a directory scan failure fails closed to
stale rather than silently skipping the check.

CONTEXT.md's Verification Module entry corrected to describe the
fingerprint path's stricter fail-closed FS-error contract (routes to
stale) instead of the module's original degrade-to-safe one (missing /
not-stale), which only the legacy path still keeps.

* refactor(#4155): extract canonicalizeCoveredFiles, add real nested-project e2e test

computeCoveredDigest and cmdVerificationFingerprint each normalized/deduped/
sorted covered_files independently — one shared helper now backs both
(gemini review's ponytail-lens finding).

Adds one CLI-to-readVerificationStatus test against a genuine
.planning/phases/NN-x/ project with an implementation file outside
.planning/ entirely, closing the review finding that prior #4155 unit
fixtures put phaseDir directly under an ownerless tmpdir (findProjectRoot
falls back to phaseDir itself there) and never exercised real multi-level
path resolution.

* fix(#4155): route computeCoveredDigest through real fs, fail closed on unreadable plans/

Two independent review rounds (opus critical-reviewer + opus ponytail +
agy, run twice) found two instances of the same fail-open class:

- computeCoveredDigest's per-file reads routed through the caller's
  injected fsImpl. planning-inspect.cts passes a `.planning/`-confined
  containment fs into readVerificationStatus's opts.fs, so any covered
  implementation file outside `.planning/` (mandatory per the issue)
  made the confinement wrapper throw, which was caught and turned into
  a stale digest -- reporting every fingerprinted phase permanently
  stale via `planning.inspect`, regardless of actual drift. Per-file
  reads now always use real node:fs, matching the pre-existing
  treatment of root canonicalization; the realRel-vs-realRoot check is
  the real confinement boundary for this data and needs no seam.

- allCurrentArtifactsCovered's try/catch never fired (scanPhasePlans
  reports readdir failures via a `scope` field, it never throws), so
  an unreadable nested plans/ dir was silently treated as "zero
  artifacts, all covered" instead of failing closed. Now branches on
  scope !== SCOPE.COMPLETE.

Also, per ponytail's second-round findings: reverted an unwarranted
FINGERPRINT_VERSION bump and digest length-prefix from the first fix
(no v1 digest has ever existed -- the feature is unreleased -- and the
prefix closed a collision that grants no capability beyond what a
writer of covered_files already has more cheaply); removed a
verifier-facing escape-hatch instruction whose own example was a case
that should trigger staleness, not bypass it; corrected CONTEXT.md
references to the renamed allCurrentArtifactsCovered and a stale
"unconditional" rescan claim; simplified the isStale derivation,
removed dead FsLike members, and tightened test coverage.

Regression tests for both fail-open bugs are included and were each
confirmed to fail against the pre-fix code before the fix landed.

full test suite: 2558/2560 pass, 2 skipped, 0 fail

* fix(#4155): trim gsd-verifier.md under the LARGE size cap

Fork CI caught what my local runs missed: the superseded/nested-plans
instruction added earlier pushed gsd-verifier.md to 49299 bytes,
147 over the LARGE tier's 49152-byte hard cap
(tests/agent-size-budget.test.cjs). Tightened the #4155 instruction's
wording and dropped a redundant inline comment tag; no content lost.

* chore(#4155): point changeset at the upstream PR number

pr: 19 was the fork PR opened for internal review-lane CI; now that
open-gsd/gsd-core#4290 exists, the changeset field must match it per
CONTRIBUTING.md's release-notes convention.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 05:42:52 -04:00
Behruz Nassre Esfahani
5ad9a36f35 fix(#4255): resolve reviewer-lane effort from the lane, not from gsd-plan-checker (#4275)
`review-lane plan` resolved every cross-AI reviewer lane's reasoning effort by
spawning `query resolve-execution gsd-plan-checker --host <slug>`. The agent id
was a hardcoded literal, so `--host` chose only the argv RENDERING while the
LEVEL always came from the installed plan-checker's frontmatter — `low` under
every shipped model profile. Every prompt-fed lane therefore ran at a fast
structural verifier's effort, and because the rendered argument is a CLI config
override it silently beat the effort the operator had configured for that CLI.
At `low` a large source-grounded prompt makes a model end its turn with no final
message, so the lane came back empty and its stub read as a crash.

Effort is a property of the review, so the lane declares it. Two new fields on
ReviewerLane — `effortConfigKey` (`review.effort.<slug>`) and `defaultEffort` —
carried through each capability manifest and the generated registry, set on the
three lanes with an argv effort channel and null on the other nine. A new pure
`resolveLaneEffort()` resolves config key -> lane default -> nothing, where
"nothing" emits no effort argument at all and the reviewer CLI's own
configuration decides; `inherit` selects that path explicitly and an
unrecognized level falls back to the lane default rather than being forwarded to
a CLI that would reject it. The host's negotiated effortSurface still gates the
rendering, so ADR-1239/#2481's trust boundary holds on this path too. Resolving
in-process also removes up to twelve subprocess spawns per review.

The empty-output stub now names the effort the lane ran at and distinguishes a
clean exit from a timeout kill, a non-zero exit, and a process that never ran —
`status` is null for both a timeout and a signal, so those were indistinguishable
before. The hint is hedged: a clean empty exit is most often a model stopping
short, but it is also consistent with a CLI writing its output elsewhere.

Also: the capability validator now knows both fields, rejects a malformed key or
an out-of-vocabulary default, and rejects a default declared without a config
key (a level the operator could never override). An existing end-to-end row in
tests/effort-surface-axis.test.cjs asserted the old coupling; it now configures
the lane's own key and pins the decoupling in the same real spawn, with the
agent execution tier set to a level that must not appear.

Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort; leaving the new key undocumented there is the same invisibility that made the plan-checker coupling survive this long.

Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort, so leaving the new key undocumented there is the same invisibility that let the plan-checker coupling survive.

Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 05:25:44 -04:00