260 Commits

Author SHA1 Message Date
Jakub Zych
6cfa0c55d2 refactor: drop 12 runtimes, keep Claude, Codex, OpenCode, Cursor, ZCode, Antigravity
Removes kilo, kimi, kimi-code, copilot, windsurf, augment, trae, qwen, hermes,
cline, codebuddy and pi end to end: capability descriptors, installer branches
and converters (bin/install.js 14.9k -> 11.2k lines), TypeScript converters,
hook surfaces and runtime homes, review lanes qwen/kimi-code, the two pi
migrations, Kimi payload normalization in the hook guards, dead hostBehaviors
vocabulary, launcher home probes, fixtures, runtime-specific tests and the
prose that presented them as supported.

Installer output for the six kept runtimes is byte-identical to before the
prune. The Kimi tool-vocabulary tests in workflow-guard, read-guard and
read-injection-scanner are left in place pending a decision.
2026-10-06 20:02:40 +02:00
Jakub Zych
a9a7a328e6 refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD
across contents and paths, upstream package/repo coordinates -> @golem15/msd-core
and golem15com/msd-core. Deep links into upstream history, sibling upstream
packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is.

Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line,
package/plugin identity, regenerated lockfile, install-tree fixtures, derived
registries and benchmark baseline; migration checksum baseline re-locked
(MSD keeps its own install state, so no install had applied the old sums);
sort-order and regex-escaped expectations in tests adjusted.
2026-10-06 01:47:40 +02:00
0xdhx
b90eef28e8 enhance(#3829): report code review severity counts and record a per-finding disposition (#3861)
* enhance(#3829): report code review severity counts and record a per-finding disposition

`code_review_gate` extracted `status:` from REVIEW.md's frontmatter and discarded
the `critical`/`warning`/`info`/`total` values sitting in the same range, so its
output was byte-identical for a review with one `info` finding and a review with a
Critical. Nothing anywhere recorded what happened to a finding: no file under
`gsd-core/workflows/` branches on `issues_found`, and `gsd-verifier.md` has zero
references to REVIEW.md. A phase therefore reached `phase.complete` with Criticals
standing and no trace they had been seen.

Both halves were approved on the issue; the gate stays advisory.

A — severity surfacing, in `execute-phase.md`. The gate states the breakdown it
already parsed, accepting `blocker:` as the documented tier-equivalent of
`critical:`. The breakdown is shown only when all four counts are numeric
(`REVIEW_COUNTS_OK`); otherwise the countless message stands, because gating on
the total alone still emits `6 findings —  critical` for a review carrying a total
and nothing else.

Frontmatter is extracted by an `awk` that emits only when it saw the CLOSING
delimiter, after stripping CR. A `sed` range re-opens on a body `---` and runs to
EOF: first-match protects a key the frontmatter always carries, but not an
optional one, so a review with no `findings:` block and a body `total:` line would
have reported the body's number. An unterminated block would leak the whole body
the same way.

Every read is guarded and `|| true`-terminated. This step is advisory, and under
`set -e`/`pipefail` a non-matching `grep` exits 1 — an assignment whose command
substitution fails would take the step down with it. A REVIEW.md that is missing,
a directory, or unreadable now leaves the counts empty and execution continues.

B — per-finding disposition, in a new lazily-read step file,
`gsd-core/workflows/execute-phase/steps/code-review-disposition.md`, referenced
from the gate in the established plain read-and-execute form. One row per finding
ID, defaulting to `open`, and:

- `fixed`/`skipped` are reconciled from REVIEW-FIX.md, whose section headings are
  matched WHOLE — a prefix match let `## Fixed Issues Verification` classify every
  finding beneath it as fixed — and only when the fix report names the SAME
  finding. Finding ids are reused across re-reviews, so matching on the id alone
  let a stale fix report declare a brand-new CR-01 already fixed.
- headings inside fenced blocks are ignored; a quoted example is not a finding.
- an id listed under both sections resolves by first occurrence, not row order.
- a recorded disposition is preserved together with the reason in its Source cell,
  escaped pipes included, and a hand-mangled row missing its trailing pipe still
  keeps its decision.
- a decided finding the current review no longer reports is CARRIED and marked;
  `--auto` rewrites REVIEW.md each iteration, so this is routine, and dropping the
  row would erase the record that it was seen. An untriaged `open` row for a
  vanished finding is not carried. A review reporting nothing still reconciles an
  existing ledger rather than freezing it.
- a run that changes no disposition rewrites nothing, so a re-executed phase does
  not produce a docs commit whose only delta is a timestamp.

The record is a sibling artifact, not a section inside REVIEW.md: `--auto`'s
re-review loop rewrites REVIEW.md every iteration, so a ledger kept inside it
would not survive the next pass, and REVIEW.md has a single writer that this step
is not.

B lives in an extracted step file because `execute-phase.md` was 91,493 bytes
against a 98,304 hard cap the size-budget test calls a red line, and because
`scanWiredKinds` caps a call site's dispatch-coverage region at 6000 characters —
an inline version pushed the `kind == "gate"` paragraph out of that window, which
silently drops `gate` from the covered set and fails
`gen-capability-registry --check` while pointing at the capability rather than at
the prose that displaced it. Extraction is what that size test's own message
prescribes, and it leaves the file at 93,854 bytes.

The tests execute the shipped script rather than modelling it. Three adversarial
review rounds each refuted "the mirror is faithful", and mutation testing agreed:
with a hand-written model, deleting the carried-row logic from the shipped file
turned nothing red. The suite now extracts the embedded script — undoing exactly
the four shell double-quote escapes — and runs it, so all ten mutations of its
behaviour are caught.

* chore(#3829): set changeset fragment pr to 3861

* fix(#3829): keep execute-phase.md under both size ceilings and propagate the launcher probe

The first push failed `full test (macos-latest, 24, shard 3/3)`. Two things it
caught that the CI-selected scope for this diff does not run, and that I
therefore did not run either:

1. `execute-phase.md` is governed by TWO ceilings, not one. The XL hard cap in
   `tests/workflow-size-budget.test.cjs` (98304) was satisfied at 95179, but the
   frozen ADR-857 pre-phase-6 ceiling in `tests/claude-orchestration.test.cjs`
   (93600) was not. The whole budget from base is 2107 bytes, which the inline
   reporting half alone did not fit. That half now lives in the extracted step
   file alongside the disposition half, and the parent carries only the paragraph
   that reads and executes it — 91529 bytes, 36 over base.

2. The step file calls `gsd_run`, so it owes the hermes runtime-home probe that
   `tests/runtime-launcher-parity.test.cjs` (E) requires of every workflow file
   that does. Propagated with `node scripts/sync-runtime-launcher.cjs`, the
   remedy that test names.

Verified with the FULL unit suite this time rather than the scoped selection —
14 shards, 0 failures — plus `npm run lint:ci`, and a re-run of the ten mutations
of the shipped disposition script, all still caught.

* fix(#3829): stop the disposition step instructing the agent to execute itself

Blocker 1 and Minor 7 of the round-1 review are one defect. The step file
carried a copy of execute-phase.md's pointer paragraph, so it named its own
path as something to "read and execute" — unbounded self-recursion at runtime
— and that copy is also the duplicated paragraph, sitting immediately above
the full instruction it duplicates.

Removing the copy resolves both. execute-phase.md remains the only surface
that points here, which is what it always intended.

Two structural tests guard it. Both are red against the pre-fix file: no
behavioural test could see either defect, because they execute the node
script through the process seam and so never read the prose that tells the
agent what to load.

* fix(#3829): re-derive the ledger paths in the block that uses them

Blocker 2. The disposition block reads REVIEW_FILE, DISPOSITION_FILE and
PADDED, all derived in the step's FIRST shell block. Each fenced block is
dispatched as its own shell, so all three are empty by the time the second
block runs: the ledger write lands on a bare `-REVIEW-DISPOSITION.md` path
and the review read finds nothing. The step then reports success having
produced no artifact — the feature's central acceptance criterion, silently
unmet, with no error to notice.

The tell was already in the file: the gsd_run shim preamble is re-emitted in
the second block for exactly this reason. These three paths belong beside it,
and now are.

The guard test asserts the general property rather than the instance — every
block derives what it reads, inheriting only the step's declared inputs
(PHASE_DIR, PHASE_NUMBER) — so a third block added later cannot reintroduce
it. Red against the pre-fix file.

* test(#3829): assert the counts mirror against the shipped shell, and execute its guards

Major 4, with Minor 6 and part of Minor 9.

The disposition builder stopped being a mirror three rounds ago, and the
reason given then was that a hand model of a shell-embedded script drifts
while the tests stay green. parseGateCounts kept its mirror anyway. That
argument does not stop applying at the boundary between the step's two shell
blocks, so the mirror now loses its authority: it is asserted against the
shipped awk and greps, run under `set -euo pipefail` in a real shell, across
every fixture it is exercised on.

Negative-controlled in both directions. Dropping `blocker:` from the mirror
alone fails the parity test; replacing the shipped awk with the leaky
`sed -n '/^---$/,/^---$/p'` range fails it on the unterminated-frontmatter
fixture. Divergence in either half is now red, which is what the finding asks
for. Skipped on win32, where there is no bash to compare against.

Minor 6: the zero-count edge is covered — `0` is numeric, so a zero-finding
review reports `0 findings — 0 critical, …` rather than falling back to the
countless form. A guard written against truthiness would have failed here
silently, and now cannot.

Minor 9, partially: running the block makes its advisory guards behavioural,
so the four `src.includes()` assertions that stood in for them are retired —
a missing and an unreadable REVIEW.md are now proven not to abort under
`set -e`, rather than asserted to contain a string. The remaining docs-parity
assertions are kept deliberately; see the PR discussion.

* test(#3829): add the render/re-parse fixed-point property for the ledger

Major 3. RULESET.TESTS.property-based-testing asks for at least one fc
property on a parsing/transformation contract, and the ledger is one with a
fixed point stated in its own prose: re-running the gate preserves every
disposition except `open`, and rewrites nothing when nothing changed.

Two properties, both driving the SHIPPED script rather than a model of it:

  idempotency — a second run reports `unchanged` and leaves the file
                byte-identical. Without it, the timestamp alone dirties the
                tree on every phase re-run.
  round-trip  — a hand-recorded decision AND the reason beside it survive
                render -> re-parse -> render, escaped pipes included. The
                Source cell is where a human writes why something was
                deferred, so losing it loses the only thing that instruction
                asks for.

Negative-controlled per property: disabling the unchanged-check fails the
first and only the first; discarding the carried source cell fails the second
and only the second.

numRuns is 40 rather than the shared 200 because each case spawns the shipped
script twice through the process seam. The seed stays pinned, so a failure
still reproduces; the deviation is stated in the file header rather than made
silently.

* fix(#3829): state a stale fix-report match instead of dropping it silently

Minor 5, plus the finding-id census this round owes.

Exact-title coupling stays — ids are reused across re-reviews, so a stale
REVIEW-FIX.md must not mark a brand-new CR-01 as already fixed. What changes
is the silence. A row that stays `open` because the report named a different
finding under the same id is indistinguishable, to any reader, from a row
that stays open because no report mentioned it. The gate now names the ids it
could not reconcile, on both report paths, and stays advisory throughout.

The census (RV4, self-found — the review did not ask for this). The script
enumerates finding-id prefixes in three places: the heading matcher, the
ledger re-parser, and the severity map's keys. The DOMAIN those enumerate is
owned elsewhere — gsd-code-reviewer.md's body template and its
Label-equivalence paragraph — so it can acquire a member without this script
changing.

Reached: CR, BL, WR, IN — 4 of 4, all present. Not reached: none today. What
follows if that changes is the payload: an unlisted prefix is not mis-tiered,
it is INVISIBLE — the finding never enters the order list and gets no row at
all, so the artifact silently under-reports the review it is meant to record.
Adding a prefix to two of the three copies fails the same way, and additionally
drops carried rows on the next run.

Two guards rather than a rewrite: hoisting the alternation into one constant
means rebuilding three regexes inside a double-quoted shell string, which is
the exact class of edit that produced both of this round's blockers. The
guards make the drift loud instead, and are negative-controlled against each
of the two ways it can happen.

* docs(#3829): keep the feature reference descriptive, not instructional

Minor 8 — a Diataxis mode mix. "Set `deferred` by hand and put the reason in
the Source cell" is a how-to instruction sitting in a reference doc. The
information belongs there (a reader needs to know the field exists and what
preserves it); the imperative does not.

Rewritten to describe the field instead: `deferred` is the one disposition the
gate never writes, and the reason recorded beside it survives re-runs. The
same pass records Minor 5's new behaviour, since the reference described the
title coupling but not what happens when it misses.

The imperative form is kept where it belongs — inside the ledger the gate
renders, which is where a reader meets the field and the only place an
instruction has an audience.

docs/FEATURES.md regenerated from it; `gen-features.cjs --check` is green.

* fix(#3829): close six defects found by reviewing this round's own fixes

None of these came from the maintainer's review. They came from adversarially
reviewing the five commits above before pushing them, and two are worse than
anything the round was opened to fix.

1. A foreign fence marker swapped an example for a finding. The heading scanner
   toggled fenced/not-fenced on ANY fence marker, so a ~~~ line inside a ```
   example closed the fence and the example's real close reopened one. Driven:
   a review quoting ~~~ inside a fenced example produced a ledger recording
   CR-77, the illustration, and omitting CR-01, the actual finding. A
   confidently-written artifact wrong in both directions at once. The open
   marker's character and length are now remembered, and a fence closes only on
   the same character at least as long, per CommonMark.

2. The disposition block had no status gate at all. The prose above it says it
   runs only when the review reports issues — but block 1 computes
   REVIEW_STATUS, emits nothing, and its shell is discarded, so no later block
   could act on that condition even in principle. A prose gate on a value
   nothing downstream can see is not a gate, and a clean re-review would rewrite
   a ledger it was never meant to touch. Re-derived in block 2's own shell.

3. A numeric breakdown could still be internally false. `total: 0` beside
   `critical: 1` is four valid numbers rendering `0 findings — 1 critical, …`.
   Numeric was necessary and not sufficient; an inconsistent breakdown is now
   withheld for the same reason a partial one is.

4. The carried-marker strip ate hand-written prose. It removed a trailing
   `(not in the current review)` unboundedly and unconditionally, so a deferral
   reason that merely ENDED in that phrase lost it — the one field a human
   writes into this artifact. Now bounded to one occurrence, and only on rows
   the marker can legitimately be on. The no-growth property it exists for is
   re-pinned.

5. parseGateCounts diverged from the shipped pipeline in two ways no fixture
   reached. The shipped reads are `cut -d: -f2 | tr -d ' '`: `tr` removes
   INTERNAL spaces (`1 0` -> `10`) where `.trim()` keeps them, and `cut` takes
   only the second colon-field where a tail capture keeps the rest. The mirror
   models the pipeline now, and both counterexamples are fixtures — a parity
   assertion that agrees only on well-formed input asserts very little.

6. The prefix census guards were both partly vacuous. The drift guard read the
   two regex alternations and not the severity map, so a set could agree in both
   regexes while mis-tiering in the map. The domain guard scanned only `### XX-01:`
   headings — and BL appears in no heading at all, only in the Label-equivalence
   prose, so the guard passed purely because BL happened to be hard-coded and
   would have missed the next prose-defined prefix exactly as it missed BL. Both
   widened; the domain the guard now sees is BL, CR, IN, WR.

Each fix fails a named test on reversion and none fires on the ordinary path.
The property generator now deliberately produces the reserved suffix from (4),
which a generator drawn only from innocuous characters could never reach.

Also corrected: the previous commit's account of the empty-path failure. The
script did not write a bare `-REVIEW-DISPOSITION.md`; it threw on reading the
empty review path and the trailing `|| echo` swallowed it as a non-blocking
skip. Same silent outcome, different mechanism, and the comment said the wrong
one.

* fix(#3829): the tests now run what bash runs — and six fixes to the fixes

A second adversarial pass over the previous commit. It found a regression that
commit introduced, and the reason it slipped through is the finding worth
keeping.

THE FIDELITY GAP. Every test here extracts the embedded script as TEXT and
runs it. Bash does not: it expands the double-quoted `node -e "..."` argument
first, so a backtick inside it is COMMAND SUBSTITUTION. The previous commit put
one in a code comment. Bash duly ran it, failed with `+: command not found`,
and handed Node a script two bytes shorter than the one 122 green tests were
exercising. No behavioural test could see this, because none of them ever
asked bash what it would actually pass. One now does, and it is the general
guard: it catches an unescaped backtick, an unescaped $, and any other
expansion the extractor cannot model.

Then, in the shipped step:

- A padded count silently disabled the sum check. `$((08 + …))` fails on base
  inference; it does not abort — the expansion sits in an `if` condition, where
  set -e does not fire — so the check simply never ran and an inconsistent
  breakdown passed with a stray diagnostic as its only trace. `10#` on every
  operand.
- The status guard made the script's own reconciliation unreachable. A clean
  review with an EXISTING ledger must still be reconciled — decided rows
  carried, stale `open` rows dropped — or the ledger freezes showing findings
  as open that the review no longer reports. The guard now skips only when
  there is nothing to reconcile.
- The carried marker is no longer stripped at parse time at all. Bounding the
  strip still ate a carried row's human-written reason. No-growth is a property
  of the RENDER, so it is enforced there: a marker already present is not
  appended again. Nothing is stripped, nothing doubles.
- Fence openers are bounded to three leading spaces, per CommonMark.
- parseGateCounts matched `[ \t]` where the shipped grep uses `[[:space:]]`,
  which covers form feed and vertical tab. Third counterexample of the same
  class, and a fixture.
- The census drift guard checked only one direction, so a tier for a prefix the
  regexes never admit stayed green as dead code that reads as coverage.

TWO OF MY OWN TESTS WERE VACUOUS, and the controls are what said so. The
leading-zero test asserted exit 0 and a consistent verdict — both true before
the fix. The clean-review test drove the node script directly, which never
executes the shell guard at all: it passed unchanged with the guard made
unconditional. Both are rewritten to test the layer the defect lives on, and
both now fail when their fix is reverted.

Every fix in this commit fails a named test on reversion, each mutation
verified to have applied before its verdict was read.

* fix(#3829): the carried marker can no longer outlive the carry

A third adversarial pass. Its most important finding is a defect the SECOND
pass talked me into, which is worth recording as plainly as the fix.

THE MARKER BECAME A LIE. Pass 2 objected that bounding the carried-marker strip
still altered a human-written reason, and proposed storing the cell verbatim
instead. That objection was a preference, not a defect — its own driven output
showed exactly one marker, which is correct — and adopting it created a real
one: once the generated marker is stored it can never leave, so a carried
finding that REAPPEARS in a later review still renders "not in the current
review". The ledger then contradicts its own contents. Driven both runs.

The strip is back, bounded to one occurrence and unconditional. The residual
ambiguity is irreducible — a reason ending in exactly that phrase is
indistinguishable from the marker — and it costs nothing real: on a carried row
the render puts the phrase straight back, and on a current row the phrase was
self-contradictory to begin with. The unbounded quantifier is what had to go,
not the strip. The property now states that contract rather than asserting a
verbatim survival the code deliberately does not provide.

Also:

- An ABSENT REVIEW.md abandoned the ledger it was meant to reconcile. The guard
  proceeds when a ledger exists, then the script read the review unconditionally,
  threw, and the trailing fallback swallowed it — the freeze the reconciliation
  path exists to prevent, reached through the door the guard opened.
- Counts are length-bounded as well as digit-only. Bash integers wrap at 2^64,
  so a 20-digit count arrived at the sum as 0 and an inconsistent breakdown
  passed.
- A closing fence must carry only whitespace after its marker; a line with an
  info string is an opener's shape and ended the fence early.
- parseGateCounts matched [ \t\n\v\f\r] where the shipped grep uses
  [[:space:]], which under this UTF-8 locale matches EM SPACE. `\s` is the
  faithful model. Fourth counterexample of that class, and a fixture.
- The agent-domain scan required [A-Z]{2,}, so a one-letter prefix like `C-01`
  — explicit and parseable, not prose — was invisible to it.

AND THE FIDELITY GUARD PAID FOR ITSELF INSIDE ONE SESSION: writing this round's
first draft I put backticks around a token in a code comment again, in the very
commit whose subject is that mistake. The probe failed, named it, and no test
of behaviour could have. Two of my own tests also had to be rewritten: one
asserted things true before its fix, and one drove the node script directly
where the defect lived in the shell.

383 pass across the touched files and the two size ceilings; ten lint gates
green; every fix fails a named test on reversion, each mutation verified to have
applied before its verdict was read.

* docs(#3829): the Source reason is preserved, but not verbatim — say so

Found by claim-auditing the response comment before posting it, which is the
one place this would have been caught: the doc and the code were written in
different commits and only a reader holding both notices they disagree.

The feature reference said the hand-written reason is "preserved verbatim
across re-runs". It is not, and deliberately so — a reason ending in the
literal phrase "(not in the current review)" loses that trailing phrase,
because it is indistinguishable from the carried marker the gate appends.

The exception is stated rather than dropped, with the reason it is the better
trade: storing the marker instead means it never leaves, and a carried finding
that later reappears goes on claiming it is absent from the very review that
reports it. A ledger wrong about its own contents beats losing a duplicated
phrase, but only if the doc admits which one it chose.

FEATURES.md regenerated; gen-features --check and lint:docs green.

* fix(#3829): the gate now emits the counts it computes (B1a/B1b)

Block 1 computed REVIEW_STATUS and the four counts and printed none of
them, then the prose below asked the agent to display four of them. The
shell exits at the closing fence and the agent sees only stdout, so those
values were unobtainable: REQ-REVIEW-08 was unreachable in every shipped
path and the fence was decorative.

The rule was already stated one block down -- "a prose-only gate on a
value no later block can see is not a gate" -- and applied only to block
2. It now governs the block that is this step's primary deliverable.

Both arms emit, and the status gate is mechanical rather than prose:
a clean/skipped/absent review prints nothing, an inconsistent or partial
breakdown prints the countless form, and the full breakdown prints
otherwise. Driven against the review's own case (critical: 1, warning: 9,
info: 8, total: 18) with no appended emitter:

  Code review: 18 findings - 1 critical, 9 warning, 8 info.
  Consider running: /gsd:code-review 1 --fix

* test(#3829): the counts harness stops manufacturing the output it asserts on (B2)

runShippedGateCounts extracted the shipped fence and then APPENDED its own
printf of the six internal variables before running it. Every counts
assertion was green against a script that existed only inside the test
process: the shipped fence emitted nothing, the tested fence emitted six
lines because the test added them. That is why B1a shipped past a suite
that looks like it covers exactly that surface -- the green was
structurally incapable of turning red for it.

The emitter now lives in the fence, so the harness reads the fence's own
stdout and synthesizes nothing. Parity with the mirror moved up a level
with it: renderGateMessage() renders both arms from the mirror's parsed
counts and the assertion compares the WHOLE emitted message, so a drift
in any parsed value changes the string or the arm it selects. Asserting
on the observable is strictly stronger than asserting on five
intermediates, and it can express what the old probe could not -- an
absent review now reports NOTHING, which is a different fact from
reporting a countless review.

A fifth src.includes() assertion converted with it (round 1 retired
four). It pinned the PROSE stating the countless condition, so it went
red when the emitter moved into the fence while the behaviour it named
was untouched -- the pin arguing for its own conversion.

Negative control: reverting the shipped echo now turns 16 tests red.
Before this commit the same reversion turned zero red, which is the
finding.

* fix(#3829): the disposition column is an enum, not any lowercase token (B3)

ADR-227 requires a trust boundary to validate semantic SHAPE and to coerce
a failure to the contract's safe default. The ledger is a trust boundary by
construction -- the rendered instruction tells a human to hand-edit it --
and the prior-row parser captured column 3 as ([a-z]+), checked against
nothing.

One transposed character was enough. `| CR-01 | critical | opne | - |` is
not the literal 'open', so it beat the default, was excluded from the
`open:` headline count, and was carried forward forever. The ledger then
reported the phase fully triaged off a typo.

The asymmetry is what made this a correctness bug rather than a style
point: a typo OUTSIDE [a-z] ('Deferred') already failed to match, lost the
decision and reset the row to open -- safe. A typo INSIDE [a-z] was unsafe.
The parser failed open in the one direction that matters. A row that fails
the enum now yields no prior entry and the row falls back to 'open', by the
same path the capital-D case already took.

The property test could not have caught this: DECIDED is drawn from the
vocabulary, so no property built on it can present an out-of-vocabulary
token. Added JUNK, the arbitrary for the complement, deliberately
lowercase so it stays inside the old capture's own character set -- the
unsafe half is the token that LOOKS like a decision and is not. The new
property also asserts the headline count agrees with the row it renders,
which is the half the defect actually reported wrongly.

Negative control: the new property fails against the ([a-z]+) capture and
passes against the enum.

* fix(#3829): a finding the heading parser cannot match is surfaced, not dropped (B4)

Two independent parsers produce two numbers one paragraph apart -- the
counts from REVIEW.md's frontmatter, the rows from `### <ID>:` heading
matches against a closed CR|BL|WR|IN alternation -- and nothing reconciled
them. A finding the alternation could not reach contributed no row, no note
and no diagnostic, and the ledger then declared `open: 3 of 3` over a set
strictly smaller than the console line had reported one paragraph earlier.
Two findings recorded nowhere, and neither artifact said so.

The PR's own argument for the closed alternation -- that an unlisted prefix
produces no row rather than a MIS-CLASSIFIED one -- is the wrong trade under
this repo's fail-safe rule. A dropped finding is demoted below every finding
that parsed, and an unparseable finding is precisely the one a human most
needs to see.

Block 2 now derives the frontmatter total (anchored inside the findings:
mapping, digit-and-length-bounded like block 1's) and hands it to the
script, which reconciles it against the CURRENT review's matched findings --
order.length, never rows.length, which also counts carried rows and would
either understate the shortfall or invent one. Surfaced exactly as the
stale fix-report case already is: a non-blocking `unparsed: N` key plus the
console line, both naming the two numbers so the claim is checkable.

  Code review disposition recorded: 3 of 3 finding(s) open (2 finding(s)
  recorded NOWHERE: the review reports 5, but only 3 matched the expected
  heading shape `### <CR|BL|WR|IN>-NN: <title>`)

The key is emitted only when there IS a shortfall, so an ordinary ledger
gains no noise key and the unchanged-run check is unaffected.

Four tests, including three negative controls the round owed itself: a
clean review gains no key, an absent/non-numeric total reconciles nothing
rather than fabricating a shortfall, and a total SMALLER than the row count
cannot render `unparsed: -1`. Reversion control: dropping the key turns the
first red.

* fix(#3829): pass --raw to the commit_docs config-get (#3763)

Not from the review -- from a gate the base range added after it. #3763
lands `tests/config-get-raw-guard.test.cjs`, and this branch was its sole
offender: a config-get command substitution without --raw feeds
JSON.stringify output into a bash string comparison, where it silently
never matches for string values. The consumer here is exactly that:

  if [ "$COMMIT_DOCS" = "true" ]

Every other shipped call site in the tree already passes --raw
(spike.md, fast.md, new-milestone.md, sketch-wrap-up.md, ...), so this is
sibling convention, not a new posture.

Worth recording because the two readings are both correct and they
disagree: round 2's review cleared this exact line under ADR-3409 as "the
safe member of that family", since `query config-get <key>` with no --pick
exits 1 on absence and the fallback arm is reachable. That is still true --
--raw does not change it. The base then moved and added a gate that reads
the same line for a different property.

* fix(#3829): scope the count reads to the findings: mapping, not just the frontmatter (m1)

`^[[:space:]]*total:` matches any indented key anywhere in the block, so a
top-level key later named `total:`, `info:` or `critical:` was picked up
ahead of the nested one. The block's own extensive comment is about scoping
the FRONTMATTER, and the scoping stopped one level short of the mapping the
values actually belong to. `status:` was never exposed -- it is anchored to
column 0 because it IS top-level.

The reads now run over the `findings:` block alone, selected by awk and cut
at the next column-0 key. Block 2's REVIEW_TOTAL derivation (added with B4)
already used that filter; this brings block 1 to it, so the two agree by
construction rather than by coincidence.

The mirror models the same scoping, and two fixtures drive it: a top-level
`total: 999` ahead of a nested `total: 1`, and top-level `critical:`/`info:`
ahead of theirs. Reversion control: unanchoring the shipped reads turns them
red.

* fix(#3829): severity comes from the section heading, not just the id prefix (M3)

gsd-code-reviewer.md emits findings under '## Critical Issues' /
'## Warnings' / '## Info', and that heading is the reviewer's own statement
of a finding's severity. The walker already visits every line -- the
fix-report path tracks '## ' sections -- so the signal was in hand and
discarded in favour of the id prefix alone.

A reviewer who mis-numbers a Critical as WR-04 while filing it under
'## Critical Issues' produced a row reading 'warning'. The ledger's Severity
column is the whole basis for triaging it, and it then disagreed both with
the review it summarizes and with the frontmatter count line block 1 prints
from findings.critical.

Section first, prefix as fallback: a finding under no recognized section --
a review that does not use the documented headings, and every row carried
from an earlier review -- keeps the prefix mapping, BL- included. Sections
are matched WHOLE, exactly as the fix-report sections are, so
'## Critical Issues Verification' does not re-tier what sits under it, and
a heading inside a fenced example does not govern.

Five tests: both mis-numbering directions, the prefix fallback across all
four prefixes, the lookalike heading, and the fenced-example case.
Reversion control: prefix-only turns the first two red.

Sixth src.includes() assertion converted with it -- it pinned the exact
source LINE of the enumeration loop, so it went red when that loop was
reformatted while the property it names was strictly widened. It now
asserts the property: every finding id, in order, once each.

* fix(#3829): an untriaged row is carried too, not silently deleted (M1)

The carry-forward kept a prior row only when its disposition was not
'open', so an untriaged row for a finding the current review no longer
reports was dropped entirely. Combined with the reconciliation gap that
left EVERY row open, a re-review deleted the whole ledger.

The re-review loop rewrites REVIEW.md on every iteration, so REVIEW.md does
not retain it either: run 1 records CR-01 open, the re-review renumbers it
to CR-02, run 2's ledger contains neither. That is #3829's complaint
verbatim -- "no trace of what happened to them" -- reproduced by the
artifact built to prevent it. The old justification, "nothing was decided
about it", is exactly the state #3829 says must leave a trace.

Every prior row is now carried, and the carried marker is what keeps it
honest: the row does not claim the finding is live, it records that it was
seen and never triaged. Two costs, stated rather than discovered: a
renumbered finding shows twice until the old row is triaged, and a carried
untriaged row persists until decided. Both are bounded by the phase's own
findings, both are legible from the marker, and both beat a silent delete.

Five tests updated -- they encoded the dropped-untriaged behaviour as the
contract -- plus one new test for the renumbering case M1 names. Reversion
control: restoring the guard turns six red.

Two self-inflicted defects caught while writing this, both by probes round
1 built:

  - Four unescaped backticks in a comment inside the double-quoted node -e
    argument, which bash ran as command substitution. The extractor-parity
    probe fired ("--auto: command not found"). Third time that trap has
    been sprung in this PR, third time the probe caught it.
  - The reworded ledger footer contained the literal carried-marker phrase,
    and the marker-accumulation assertion counts it across the whole file,
    so a doc line read as a second marker. The assertion was right.

* test(#3829): cover the count-length threshold at limit-1, limit and limit+1 (M2)

The guard is `?????????*` -- nine or more characters -- so the limit is
8 digits accepted, 9 rejected. The only cases were 'x', single digits and a
20-digit value, none of which pins the boundary. RULESET.TESTS
boundary-coverage is a hard rule here and it was unmet.

All three points asserted, with the sum kept consistent at each so the
LENGTH rule is what decides the verdict rather than the sum check
incidentally agreeing.

Reversion control is the off-by-one M2 names: dropping one `?` moves the
limit to 7 digits, which no test could previously notice, and now turns
this one red.

* feat(#3829): wire the disposition ledger into the fix path (B1c/B1d)

REQ-REVIEW-09 was unreachable in every shipped path. execute-phase.md's
code_review_gate invokes review with neither --fix nor --auto, so
<NN>-REVIEW-FIX.md cannot exist when the gate runs and every row it writes
is `open` by construction. The operator then runs /gsd:code-review N --fix
by hand -- the very suggestion the step prints -- which writes REVIEW-FIX.md
and never touched the ledger. A phase with 23 findings, all fixed, ended at
`open: 23 / total: 23`: the artifact that exists to distinguish a triaged
finding from a forgotten one asserted that 23 triaged findings were
forgotten. Worse than recording nothing, because it looks authoritative and
is inverted.

Taking remedy (i), not (ii). Narrowing the docs to say the ledger reflects
the previous phase execution is a legitimate choice, but it ships a feature
whose central artifact is inert and then documents the inertness.

ONE ADAPTATION, because the prescribed site does not exist. The review says
to wire code-review.md's --fix/--auto path. code-review.md is not the writer
(gsd-code-fixer writes the report, code-review-fix.md commits it), and more
decisively it has no point that is AFTER the report exists: it delegates
through code-review/steps/dispatch-fix.md, which calls
Workflow(code-review-fix.md) and then exits the workflow. There is nothing
downstream of that call to wire to.

The site is code-review-fix.md, immediately after commit_fix_report. That
is where the report is on disk and committed, it is the canonical
implementation for all fix logic by dispatch-fix.md's own statement, and it
additionally covers a direct invocation of that workflow -- which a wiring
in code-review.md would have missed.

The same step, not a second copy: it consumes PHASE_DIR and PHASE_NUMBER,
both already parsed from the init JSON, and it is idempotent, so a phase
that reaches the gate and then a fix run ends with one ledger reflecting
both rather than two competing ones.

Driven end to end: the gate writes `open: 2 of 2`, the fix path reconciles
to fixed/skipped and `open: 0`. Two tests -- one pins the wiring and its
ordering relative to commit_fix_report and present_results, one drives the
two call sites in sequence. Reversion control: removing the step turns the
first red; the second covers the reconciliation the wiring makes reachable
rather than the wiring itself.

* fix(#3829): a reflowed fix-report title is the same title (m2)

The stale-fix-report guard compared titles with trim() equality. The strict
instinct is right -- ids are reused across re-reviews, so a stale
REVIEW-FIX.md must not mark a brand-new CR-01 as already fixed -- but
gsd-code-fixer.md writes '### {finding_id}: {title}' under no contract that
the title is copied byte-for-byte from REVIEW.md. A fixer that reflows a
long title produced a spurious mismatch note, left a genuinely-fixed row
'open', and told the reader the report named a different finding. That
false-positive mode was acknowledged nowhere.

Whitespace is normalized, and only whitespace: a wrapped title is the same
title, and it is the one divergence that carries no information. Case
changes and truncation stay strict on purpose -- they are the shapes a
genuinely DIFFERENT finding takes, and widening to them would trade a
visible false positive for the silent false negative the strict match
exists to prevent. The residual is now stated in the step rather than left
to be rediscovered.

The note's wording changed with it. It asserted the report "names a
different finding"; both causes reach that branch and the step cannot tell
them apart, so it now reports the observation -- "titles its finding
differently from the review ... a stale report, or a re-titled one" --
rather than a conclusion it has not earned.

Three tests: the reflow case reconciles cleanly, the re-cased case still
reports, and the stale case still reports with the new wording. Reversion
control: restoring the strict comparison turns the reflow test red.

Seventh src.includes() converted -- it pinned the comparison EXPRESSION, so
it went red when the comparison gained normalization while the property it
names was unchanged.

* docs(#3829): describe the flow that ships, not the one implied (m3)

Both reference pages said "/gsd-code-review <N> --fix records fixed and
skipped, which the gate reconciles from REVIEW-FIX.md" -- true in the
abstract, materially misleading in practice, because no shipped path
performed that reconciliation. With B1c/B1d wired it is now real, and the
pages say WHERE it happens rather than leaving a reader to assume the
in-phase gate does it: the gate runs before any fix report exists and
writes all-open, and --fix is what records what happened.

The round's other behaviour changes land here too, since a reference page
that lags the artifact is worse than none:

  - the disposition column is a closed vocabulary, and a value outside it
    falls back to open rather than being treated as a decision
  - severity comes from the section heading when the review uses one, and
    from the ID prefix otherwise
  - an unparsed shortfall is stated rather than dropped
  - titles are compared ignoring whitespace, so a reflowed title still
    reconciles, and a mismatch is reported as an observation rather than as
    a claim that the report is stale
  - EVERY row is carried now, triaged or not, with the cost of the
    renumbered-finding double-entry stated rather than left to be found

docs/FEATURES.md regenerated from the fragment; lint:generated-sync and
lint:docs both exit 0.

* chore(#3829): migrate the emitted-drift ack from a fragment to commit trailers

ADR-3942 landed on next in #3954: the acknowledgment is a git commit
trailer now, and tests/emitted-drift-acks/ no longer exists.

Worth noting for anyone reading the rebase: this did NOT surface as the
modify/delete conflict the migration guidance predicts. This branch ADDED
its fragment rather than modifying an existing one, and the base deleted
only the files that were already there, so the replay was clean and the
fragment survived silently into a directory that no longer exists. Quieter
than a conflict, and worse -- the gate is what catches it, not git.

Two Growth keys rather than the fragment's one: round 2 wired the ledger
into code-review-fix.md, so that file grew too. Both key on the bare
filename, per the Growth namespace.

Emitted-Drift-Ack-Growth: code-review-fix.md — #3829 review round 2, blocker 1c/1d: REQ-REVIEW-09 was unreachable in every shipped path because the in-phase gate runs before any REVIEW-FIX.md exists, so every ledger row it wrote was open and nothing ever reconciled them. This file gains one step, record_disposition, that reads and executes the same lazily-read step after commit_fix_report. It is the only point in the fix flow that is after the report is on disk: code-review.md delegates here through steps/dispatch-fix.md and exits, so it has no such point at all. Growth is one step of prose, no logic is duplicated, and the step is idempotent so the two call sites converge on one ledger.

* chore(#3829): the changeset describes the round's behaviour, not round 1's

It renders into CHANGELOG, so it carries the same misleading implication
minor 3 was about: "the gate ... reconciling fixed/skipped from
REVIEW-FIX.md" reads as though the in-phase gate does it, when the gate
runs before any fix report exists. Says where it happens, and picks up the
round's other user-visible changes -- carried untriaged rows, section-based
severity, the disposition vocabulary, and the unparsed shortfall.

* fix(#3829): a dotted phase number no longer aborts the step

Found by this round's own adversarial review, in its MISSED section: no
finding asked about it, and it is the most serious thing in the round after
the two blockers.

Both callers explicitly accept a dotted phase -- code-review.md:60 and
code-review-fix.md:36 both validate ^[0-9]+(\.[0-9]+)?$ and name "03.1" in
their own error text -- and both fences reconstructed the path with
`printf "%02d" "${PHASE_NUMBER}"`, which cannot format one. Driven with
PHASE_NUMBER=3.1: bash prints `invalid number` and exits 1, and under
`set -euo pipefail` that aborts the step on its FIRST line. An advisory
gate that promises never to block took the phase's entire review report
down with it, and the newly wired fix-path call site inherited the same
defect.

Pad the integer part and carry the sub-number verbatim, so 3.1 -> 03.1 and
3 -> 03, with both arms falling back to the raw value rather than aborting.
Driven: 3.1 now reads 03.1-REVIEW.md and writes
03.1-REVIEW-DISPOSITION.md; the integer path is unchanged.

Two other findings from the same review, both about claims rather than code:

MINOR 2's TEST WAS MIS-NAMED, and the reviewer was right to refute the
claim. It called itself the "reflowed" case while substituting triple
spaces, which is not a reflow. Driven: a genuinely WRAPPED heading is still
not reconciled, because a `###` heading is one line by definition and the
continuation is a separate paragraph. Not widened -- absorbing whatever
follows a heading into the title would swallow arbitrary prose and make the
stale-report check meaningless, and the kept failure mode is the safe one
(a visible mismatch note, never a wrong "fixed"). The test is renamed to
what it covers and the bound is now pinned by its own test.

THE SHELL-SHARING GUARD DID NOT GUARD. Negative-controlling it -- rather
than reading it -- showed that deleting block 2's real REVIEW_FILE
derivation left it GREEN, on the exact defect it was written for. Block 2
prefixes its `node -e` with `REVIEW_FILE="${REVIEW_FILE}" ...` to put the
values in the child's environment, and the detector counted that
self-referential pass-through as a derivation. Pass-throughs are now
excluded, and the control fires. Pre-existing, not introduced here: the
original column-0 anchor matched that same line.

Also worth recording: my first attempt at that control silently patched
nothing and reported clean. Same lesson this PR already learned once.

* fix(#3829): validate the phase number before formatting it, and make the shell guard executable

Three findings from the round review's continuation pass, all confirmed by
driving them.

1. MY OWN DOTTED-PHASE FIX WAS WRONG on the fallback path. `printf "%02d"
   abc` writes `00` to stdout BEFORE it fails, so
   `$(printf ... || printf %s ...)` CONCATENATES the two: `abc` became
   `00abc`, empty became `00`, and a legitimate `08.1` became `0008.1`
   because bash reads the leading zero as octal. An unset PHASE_NUMBER also
   aborted under `set -u` -- in the step that promises never to abort.

   Validate, then format: never format and fall back on failure. Driven
   across every edge the review named -- 3.1 -> 03.1, 3 -> 03, 08.1 -> 08.1,
   09 -> 09, 1.2.3 -> 01.2.3, and abc / empty / -1 / unset carried verbatim
   with exit 0.

2. THE SHELL-SHARING GUARD STILL DID NOT GUARD. Excluding pass-throughs was
   not enough: a structural predicate recognises assignment TOKENS, never
   assignments that derive a usable value, so `REVIEW_FILE=`,
   `REVIEW_FILE=$REVIEW_FILE` and a commented-out assignment all evaded it.
   No regex closes that class.

   The authority moves to execution -- the third time this PR has learned
   that lesson. The real second fence now runs in a fresh shell with nothing
   but the step's two declared inputs and must write the ledger at the
   correct derived path. All four mutations are caught: empty assignment,
   self-reference, commented-out, and deletion. The textual check stays as a
   cheap fast-fail and is labelled as one.

3. THE TITLE-BOUND CORRECTION HAD NOT REACHED THE DOCS. The step comment and
   both docs pages still said a reflowed title reconciles, contradicting the
   bound pinned one commit earlier. Superseded prose left standing reads as
   current to anyone arriving cold, so all three surfaces are rewritten
   rather than annotated, and FEATURES.md regenerated.

Also hoisted `HAS_BASH` to the file's other top-level constants. `const` is
in the temporal dead zone until its declaration runs, and a
`{ skip: !HAS_BASH }` option object is evaluated eagerly, so a bash-gated
test added above the old mid-file declaration threw a ReferenceError that
aborted its whole describe and CANCELLED its siblings -- while the summary
line still read `fail 0`. It caught three separate additions in this round
before I stopped moving tests and moved the constant.

* fix(#3829): refuse an out-of-shape phase number instead of carrying it into a path

Self-found while writing the prompt for the next review pass, which is the
honest provenance: I asked the reviewer whether a path traversal was
reachable through PHASE_NUMBER, then checked before dispatching.

It was, and I had introduced it. The previous commit's fallback carried an
unusable phase number VERBATIM, and PHASE_NUMBER is interpolated into a file
path:

  PHASE_NUMBER='../../etc/passwd'
  -> REVIEW_FILE=/tmp/phase/../../etc/passwd-REVIEW.md

The `printf "%02d"` it replaced had at least mangled that to `00`. A fix
that makes a path more reachable than the bug it replaced is a regression,
whatever it does for the case it was written for.

Both callers already validate ^[0-9]+(\.[0-9]+)?$ (code-review.md:60,
code-review-fix.md:36), so this is defense in depth rather than a live
exploit -- but the step has two call sites now and should not take either
caller's word for its own inputs. It validates the WHOLE value and, on
failure, builds no path at all: PADDED is empty and each fence refuses by
name rather than coercing. Block 1 declines to report counts read from a
path made out of the bad value; block 2 declines to write, which also keeps
it clear of the bare-name ledger defect round 1 closed.

Driven across the shape boundary: 3.1 / 3 / 08.1 / 09 accepted; abc, empty,
unset, 1.2.3, -1, 3., .1, +1, "3 1" and ../../etc/passwd all refused with
exit 0 and a named diagnostic. Reversion control: restoring carry-verbatim
turns the traversal test red.

* fix(#3829): bound the phase number's length, and make the shell guard prove derivation

Third adversarial pass. Two of its three refutations were already closed by
the previous commit (the ../escape and 1/../../escape traversals, and the
unset-input abort); these two were not.

1. A 54-DIGIT PHASE NUMBER WRAPPED SILENTLY. The validator accepted any
   all-digit value, so `$((10#$_int))` overflowed 64 bits and PADDED became
   `-7908320945662590977`. Length-bounded now at 8 digits, exactly as the
   counts already are and for the identical reason -- and the counts guard
   sitting twenty lines away is why this one is embarrassing rather than
   subtle. Driven at the boundary: 8 digits accepted, 9 rejected.

   The bare `${PHASE_NUMBER}` in the suggestion line is hardened to
   `${PHASE_NUMBER:-}` while here. The empty-PADDED guard makes it
   unreachable today, but it is one refactor away from an unbound-variable
   abort under `set -u`, in the step that promises not to abort.

2. THE EXECUTED SHELL GUARD PROVED THE FENCE WORKS, NOT THAT IT DERIVES.
   A single-phase probe is satisfied by a hardcode, and the review
   demonstrated exactly that: replacing the derivation with
   `case ... in 1) PADDED=01 ;; 7) PADDED=07 ;; *) PADDED=07 ;; esac`
   breaks every real phase and passed the entire suite. It now runs two
   distinct phases, 7 and 3.1 -- a hardcode cannot satisfy both, and the
   dotted one additionally pins the integer-part split.

   The claim "given only the declared inputs" was also overstated: the test
   spreads `...process.env` (it needs PATH and HOME). The DERIVED names are
   now explicitly deleted from that environment, so the claim is true rather
   than merely intended.

Also rewrote a comment that had become false: it pinned a describe to the
end of the file because of the HAS_BASH temporal-dead-zone constraint, which
the hoist removed. Superseded prose left standing reads as current to
anyone arriving cold.

The changeset's "the gate stays advisory and never blocks" is now verified
rather than asserted: both fences exit 0 under an unset PHASE_NUMBER and a
traversal-shaped one.

* fix(#3829): validate both inputs, refuse before building a path, and never write through a symlink

Fourth adversarial pass. Four findings, all confirmed by driving them.

1. PHASE_DIR WAS NOT VALIDATED AT ALL. Unset, both fences died with
   `PHASE_DIR: unbound variable` under `set -u` -- the same class as
   PHASE_NUMBER, which I had just spent two commits fixing while its sibling
   input sat one line away. The step declares two inputs; it now validates
   two.

2. THE LENGTH BOUND WAS ON THE WRONG THING. The nine-character glob applied
   to the WHOLE value rather than the integer part, so it falsely rejected
   `12345678.1` (a legal 8-digit phase) while accepting `1.123456`. Each
   component is bounded on its own now; the sub-number is bounded too, since
   it is likewise interpolated into a filename.

3. REJECTED VALUES STILL HAD PATHS BUILT FROM THEM. The refusal guard sat
   AFTER the assignments, so an unusable input still assembled
   `${PHASE_DIR}/-REVIEW.md` and stat'ed it before refusing. The guard is
   now the first thing after validation, and both fences construct paths
   from validated locals rather than from the raw environment.

4. THE LEDGER WRITE FOLLOWED SYMLINKS. From the review's MISSED section, and
   the sharpest thing in it: `fs.writeFileSync` follows a symlink, so a
   pre-existing symlink at the ledger path replaced the contents of whatever
   it pointed at -- outside the phase directory, with the link left intact
   so nothing looked wrong. Driven, and the target's contents were gone.
   This PR introduces the artifact, so it owns the check: an existing ledger
   that is not a regular file is not a ledger, and the advisory gate says so
   and steps over.

The executed shell guard now draws its phases AT RUN TIME. Fixed fixtures
cannot establish derivation -- the review defeated the one-phase version
with a hardcode, then defeated the two-phase version by adding one more arm
to the same case. Any finite sample loses that race. A phase picked per run
cannot be enumerated in advance; the drawn values print in every assertion
message so a failure stays reproducible. Control: the review's three-value
hardcode now fails on three consecutive runs.

Eighth src.includes() converted -- it pinned the literal `${PHASE_DIR}`
interpolation and went red when construction moved to a validated local,
while "writes a REVIEW-DISPOSITION sibling" was untouched. It now asserts
that property, and that REVIEW.md is not written.

The changeset's "stays advisory and never blocks" is verified rather than
asserted: 8 of 8 hostile-input cases across both fences exit 0 -- both
inputs unset, PHASE_DIR unset, a traversal-shaped phase, and a missing
phase directory.

* fix(#3829): check the ledger path before reading it, and pin the write-safety behaviour

Fifth adversarial pass, and the last one this round. Three fixes, three
disclosed residuals.

FIXED

1. A FIFO AT THE LEDGER PATH BLOCKED FOREVER. readFileSync on a FIFO never
   returns, so the step documented as "advisory, never blocks" blocked
   indefinitely -- the literal counterexample to its own headline claim. The
   non-regular-file check ran after that read.

2. THE UNCHANGED-RUN FAST PATH BYPASSED THE CHECK. A symlink whose target
   already matched the rendered ledger read through the link, reported
   `unchanged`, and never reached the refusal.

   Both fixed by the same move: the check is now the FIRST thing the script
   does, before any read or write of that path. Ordering was the defect, not
   the predicate.

3. THE COMMIT TEST FOLLOWED THE LINK the script had just refused. `[ -f ]`
   resolves symlinks, so the guard and its consumer disagreed about the same
   path and the helper could still be handed one. `[ ! -L ]` added.

   And the behaviour shipped with NO regression control -- I hand-drove it
   last commit and did not pin it, which the review caught by grepping for
   the words. Five tests now: symlink, symlink-with-matching-target, FIFO,
   directory, and an ordinary ledger as the negative control so the refusal
   is not a blanket one. mkfifo goes through the process seam like every
   other spawn here.

DISCLOSED, NOT FIXED -- these are stated in the step rather than carried
silently:

- TOCTOU between the lstat and the write. Node exposes no portable
  O_NOFOLLOW write, and an attacker who can write into the phase directory
  mid-run already has what the check would protect. It narrows a real
  accident; it is not a security boundary and the docs claim none.
- A hard link passes isFile() by construction.
- The REVIEW.md and REVIEW-FIX.md reads still resolve symlinks. They are
  reads of files the operator owns, in their own phase directory.

Also narrowed a comment that overclaimed. The randomized guard's domain is
FINITE -- 88 integer and 792 dotted values -- so a mutation enumerating all
880 passes forever, and Math.random() is unseeded, so "reproducible" means
only that the drawn values are printed on failure. Raising the bar is what
it buys; proving derivation is not, and nothing short of reading the fence
is. The previous comment claimed otherwise and was refuted.

* test(#3829): make the write-safety controls portable to the Windows lane

CI caught what neither the local suite nor five adversarial review passes
could: every one of those ran on Linux.

The FIFO test gated on `mkfifo`'s exit code. On the Windows lane mkfifo
EXISTS and exits 0 while producing something that is not a FIFO, so the
guard passed, the test ran against an ordinary path, the ledger wrote
normally, and the assertion failed for a reason unrelated to the behaviour
under test. It now gates on `lstatSync().isFIFO()` -- what was actually
created, not what the command claimed. Control: with the shipped guard
disabled the test still goes red on Linux, where the FIFO is real.

The two symlink tests are skipped on win32, following this repo's existing
convention for symlink-planting tests (tests/settings-jsonc.test.cjs:389
skips the same class; tests/unreachable-guard-drift.test.cjs:726 records the
reason -- symlink creation requires elevated privileges on Windows CI). The
privilege happened to be available on the lane this round, which is exactly
why the convention is not "try it and see".

* fix(#3829): a bare `|` in a deferral reason is prose, not a parse failure

Review round 3, the one blocker. The Source cell is the one field this ledger asks a
human to hand-edit, and "waiting on team A | team B to align" is an ordinary thing to
type there. The prior-row capture admitted a pipe only when escaped, so a bare one
failed the WHOLE line: prior.get() was undefined, the row fell through to `open` with
an empty Source, and the console line read "1 of 1 finding(s) open" — a Critical a
human explicitly deferred, with a documented reason, rendered indistinguishable from
one never triaged, and the reason gone. The exact ambiguity #3829 exists to remove,
reachable by one missing backslash.

The Source cell is the LAST column, so it is now captured through to the end of the
line, less an optional trailing pipe; a bare `|` inside it is prose. The render
escapes a bare pipe on the next write so the table stays a table, and the escaped
form re-parses to itself, so the second run reports `unchanged` — the fixed point
holds. The ledger's own instruction line says so instead of asking the human to
escape.

Why the property never caught it: SOURCE_CELL only ever appended a PRE-ESCAPED pipe,
so the arbitrary built to stress this cell could not reach the one input that broke
it. It now also emits a bare pipe, and the round-trip expectation is the escaped
form of what the human wrote. A fixed regression case drives the reviewer's exact
input through two runs and asserts the decision, the reason, the headline count and
convergence. Negative-controlled: both new tests fail against the previous capture.

The src.includes() pin on the old capture text is retired for the behavioural case —
it was pinning the defect.

* fix(#3829): escape every bare pipe in one write, whatever precedes it

Round 3, found by the adversarial pass over the round's own fix rather than by the
review. The first escape used /(^|[^\\])\|/g, which CONSUMES the character before the
pipe: adjacent bare pipes were escaped one per run (A||B -> A\||B -> A\|\|B, a third
run to converge, breaking the advertised second-run fixed point), and an escaped
backslash before a pipe (A\\|B) hid the pipe behind the wrong parity and left it bare
in the rendered table. The property generator emits at most one bare pipe, which is
the one case the old form got right, so no property reached either.

Scan as pairs instead: an escaped pair (backslash + anything) is kept verbatim and only
a pipe outside one is escaped. One write, then a fixed point. Regression case drives
`A||B and C\\|D` through two runs; it fails against the previous escape.

* fix(#3829): the script leaves by return, so an explicit exit cannot drop its verdict line

Round 3, from the adversarial pass over the round's own fix. The embedded node script
printed its verdict and then called process.exit(0) -- on the 'unchanged' branch only;
the 'recorded' branch fell off the end. Node's "A note on process I/O" documents
process.stdout writes to pipes and sockets as asynchronous on POSIX, and process.exit()
as forcing exit before pending asynchronous stdout writes complete -- so on a POSIX
lane the caller can see exit 0 with no verdict line. This is a hardening against that
documented hazard, not a reproduced defect: the reviewer's empty-second-run stdout,
which first pointed here, turned out to be its own sandbox -- a bare console.log child
printed nothing there either -- and that attribution is withdrawn.

The script now runs inside main() and leaves by return on all four early-exit paths,
so the event loop drains stdout before the process ends. Same exit status either way,
and the || echo fallback is unaffected. A structural test pins the absence of the call
(comment-stripped; dotted, bracketed and whitespace-split spellings). The empty-review
docs-parity pin that asserted the literal process.exit(0) line is retired -- the
round-2 describe drives that property behaviourally.

Also widens the property generator: SOURCE_CELL now reaches adjacent pipes and a
backslash of either parity before a pipe, the two shapes the first render escape got
wrong while passing every input the generator could then produce -- checked against
an independent parity-walk oracle rather than a copy of the render's own scan.

* fix(#3829): record what an --auto iteration fixed, instead of reporting it open

Round 5's major. `record_disposition` runs once, after the whole capped-at-3
`--auto` loop converges — but this workflow keeps ONE final version of REVIEW.md
and REVIEW-FIX.md rather than per-iteration copies, and deletes the .iterN.md
backups on convergence. A finding fixed in iteration 1 was therefore absent from
the final review (it was fixed, so the re-review stopped reporting it) AND from
the final fix report (overwritten by the last iteration), so the row fell back to
the gate's `open` and rendered `open ... (not in the current review)` — the same
bytes a finding that vanished for an unrelated reason produces. That is the one
distinction #3829 exists to make, undone by the artifact built to make it.

The precise site was the two-arm `applied` construction: for an id the current
review does not report, `sameTitle(undefined, h.title)` is false and
`title.has(id)` is false too, so the entry entered NEITHER `applied` NOR
`staleFix`. It was dropped in silence.

Four changes, one defect:

- A third arm. When the review does not report an id at all there is no title to
  disagree with, so this is not the stale-report case — it is what a finding
  looks like once it has been acted on. Record it. The id-reuse hazard stays
  closed by the arm below it: when the review DOES report the id, a title
  mismatch still goes to `staleFix` and is never applied, so a renumbered
  finding cannot inherit an earlier iteration's `fixed`.
- Rows for decided ids the review no longer reports, carried and marked. A
  decision the ledger cannot render is a decision lost — the same silent drop
  the carry-forward loop already refuses for prior rows, one source over.
- The .iterN.md fix-report backups are read alongside the final report, newest
  first, so the most recent statement about an id wins — the precedence a
  duplicate id already gets within one report.
- The shell guard proceeds on a fix report, not only on an existing ledger. A
  direct `/gsd-code-review N --auto` writes no gate ledger, and a converged loop
  leaves `status: clean`, so a fully successful multi-iteration run recorded
  nothing at all.

And the backups now go in `cleanup_iteration_backups`, after the ledger has read
them. #3190's rule is untouched — spent scratch on convergence, retained on
degradation — only the timing moved; deleting them inside the loop erased every
early fix before anything read it. `CONVERGED` does not survive the loop's shell
and is re-derived from the final review's status, which is exactly how the loop
sets it; anything but a proven-clean review retains.

Seven new regression tests plus an ordering test, all eight reversion-controlled
against pre-fix code — every one fires. One is the negative control that matters:
a reused id whose title differs must stay `open`, never inherit `fixed`.

Residual, stated: an id appearing only in an iteration fix report takes its
severity from the id prefix rather than a section heading, because `sectionSev`
is built from the current review. That is the documented fallback for carried
rows, not a new gap.

* test(#3829): pin the two PADDED derivations against a silent desync

Round 5's minor 1. Each fenced block runs in a fresh shell and must derive what
it reads, so the PADDED derivation — the traversal fence between an
attacker-influenceable phase number and a file path, plus the per-component
length bound — is duplicated verbatim. Both copies were independently tested and
nothing asserted they stay in step, which is the shared-parallel-surface shape
CLAUDE.md requires a parity test for, on security-relevant validation logic
rather than incidental repetition.

Compared line by line rather than through a normalizing rewrite: a normalizer
has to be told what may differ, and whatever it is told to tolerate stops being
asserted. Exactly one line may differ — each block refuses by its own name — and
the test names both forms. It also asserts the slice is substantial, since a
parity test over an empty slice passes vacuously.

Control: dropping one `?` from block 2's length bound, which moves that copy's
limit to 7 digits while block 1 keeps 8, turns it red. That is the exact silent
divergence the finding describes.

One correction to the finding's own statement, since it is worth recording: the
cited lines are :324 and ~:480, which are node-script lines; the derivations are
at :48-83 and :211-246. And they are 35-of-36 identical rather than
byte-identical — the refusal message differs, deliberately.

* docs(#3829): state the PHASE_DIR trust boundary instead of carrying it

Round 5's minor 2 asked that the assumption behind PHASE_DIR's validation be
confirmed rather than silently carried forward at the two new call sites. It is
confirmed, and the comment that stood here was wrong about it: "PHASE_DIR is the
step's other declared input and gets the same treatment" describes something the
code does not do.

Both inputs have the SAME provenance — each caller binds them from
`gsd_run query init.phase-op` (code-review-fix.md:7,17; execute-phase.md the
same) — so neither is raw user input and neither is more trusted. The asymmetry
is not about trust. It is that only one of them has a shape: PHASE_NUMBER
carries a documented contract, `^[0-9]+(\.[0-9]+)?$`, asserted by both callers,
so a value outside it is provably wrong and is refused. PHASE_DIR's contract is
"a filesystem path", which admits `..`, absolute and relative forms and
symlinked parents alike; no predicate separates a legitimate planning directory
from an illegitimate one, so a shape check would reject working setups while
proving nothing.

So the emptiness check is adopted as what it actually is — the guard against
`PHASE_DIR: unbound variable` aborting a step that promises never to block — and
the shape check is declined, with the reason written where the next reader meets
it rather than left to be re-derived.

The residual is restated in place rather than left in a PR comment: PHASE_DIR
may itself be a symlink and the ledger is then written through it, outside the
phase directory, deterministically. Left alone deliberately — the write goes
where the caller pointed. Not a security boundary, and nothing here claims one.

* docs(#3829): record why HAS_BASH is a platform assumption, not a probe

Round 5's minor 3 is DECLINED, and the reason is the repo's own contract rather
than a judgement call — written at the constant so the next reader does not
"fix" it and re-enable what the rule exists to prevent.

The gap is real and confirmed: 22 tests carry `{ skip: !HAS_BASH }`, so block
1's bash severity-reporting path has no Windows-lane coverage. But
`local/no-unguarded-nonportable-exec`
(eslint-rules/no-unguarded-nonportable-exec.cjs, DEFECT.WINDOWS-TEST-PORTABILITY)
REQUIRES this guard around `sh -c` / `bash -c` in tests, and its own remedy text
names `if (process.platform !== 'win32')` as the sanctioned form, because these
constructs fail under Windows Git Bash. So the constant is the repo's answer to
this question, not an oversight in this PR.

Swapping it for a runtime `bash` probe would light 22 tests up on a lane the
rule has already determined they cannot pass — trading a legible, rule-encoded
skip for a red matrix. Reversing that is the rule's decision; a change here
belongs with a change there.

* docs(#3829): describe how --auto's iterations reach the disposition ledger

The reconciliation section described the `--fix` path accurately and said
nothing about `--auto`, which is where round 5's major lived. It now states
that the loop overwrites its fix report each pass, that the re-review drops a
finding once it is fixed, that the gate therefore reads the per-iteration
backups newest-first, and that the backups are removed after the ledger has read
them rather than before. It also states the converged-with-no-ledger case: a fix
report on disk is reason enough to record.

FEATURES.md regenerated (176 features / 21 groups). Changeset extended to name
the shipped behaviour rather than only the `--fix` half.

* fix(#3829): clear lint-workflow-shellcheck, a gate the base range added

Not from the review. The rebase onto `next` brought in `lint-workflow-shellcheck`
(#4109), whose baseline was generated before this PR's new step file existed — so
that file's findings are new by construction and `lint:ci` exited 1 on the
rebased head before this round touched anything. The last green CI run predates
the gate. Caught locally rather than by a red push.

Three fixes and one baseline entry, split by whether the finding is real:

- STRUCTURAL (not ShellCheck, not baselineable): the guard's
  `for _f in "…${PADDED}-REVIEW-FIX.iter"*.md` is the bare `for x in $VAR` shape
  that word-splits differently under bash and zsh. Wrapped in
  `$(printf '%s' "$PADDED")`, the linter's own prescribed remedy.

- SC2097/SC2098, and this one was a genuine latent bug rather than a lint nit:
  `FIX_REPORT_FILE="${_pd}/${PADDED}-REVIEW-FIX.md"` sat in the same env-prefix
  list that sets `PADDED`, so its `${PADDED}` expanded the OUTER variable, not
  the one two entries earlier. Both happen to hold the same value here, which is
  exactly why it would have kept being wrong quietly. Built before the command
  now.

- SC2317 ×3 is baselined, not fixed. It fires on
  `return 0 2>/dev/null || exit 0` — the deliberate idiom that lets a fence
  refuse whether it is sourced or executed — and the verdict is a false
  positive: the `exit 0` is reached precisely in the executed case. Rewriting a
  dual-mode refusal to satisfy a wrong unreachability claim trades a real
  behaviour for a clean report. Baseline 207 -> 210.

`lint:ci` exits 0. 173 tests pass across the two touched files.

* fix(#3829): a reused finding id no longer inherits the old finding's decision

Found by this round's own adversarial review, which drove it rather than
reasoned about it — and it refuted the arm I had named as my strongest
suspicion, so it is recorded as a correction, not a discovery.

Finding ids are reused across re-reviews: the --auto loop renumbers. `row()`
inherited a prior decision on an id MATCH ALONE, with nothing checking it was
the same finding. Driven: a prior `CR-01 fixed` row against a review reporting a
brand-new CR-01 rendered the NEW finding `fixed`. A false decision in the
artifact whose entire purpose is telling triaged from forgotten — the same
failure mode round 4's blocker was, reached by the other door.

I had argued this was closed by the stale-report arm. It is not: that arm guards
the FIX-REPORT path only. The PRIOR-LEDGER path had no title check at all.

- The ledger now records each finding's title, in the FRONTMATTER rather than a
  fifth table column: the Source cell is the field a human hand-edits and the one
  that must escape pipes, and a second free-text column doubles that surface for
  no reader benefit.
- A prior decision is inherited only when the recorded title still matches. An
  ABSENT prior title inherits, deliberately — a ledger written before titles were
  recorded carries none, and refusing there would reset every decision in it,
  which is the loss this guard exists to prevent, caused by the guard.
- A decision whose id has been reused is PRESERVED under a `superseded:` key
  rather than dropped. The review's driven refutation was precisely that the
  mismatch was surfaced while the decision was lost. It cannot keep a row — the
  id is taken, and two rows under one id is an ambiguity, not a record — so it is
  carried in the frontmatter, re-emitted every run, deduped by id+title, and
  named on the console.
- And an iteration-derived decision now cites the report it actually came from.
  The Source cell hard-coded the unsuffixed `<NN>-REVIEW-FIX.md`, so a decision
  read out of an iteration backup cited a file that may not exist. A citation the
  reader cannot follow is worse than none. Also the review's finding.

Five new tests. Four fail against the pre-fix step; the fifth — that a ledger
with no recorded title still inherits — is a BACK-COMPAT guard and passes both
ways by construction. It is not a reversion control and is not counted as one.

* fix(#3829): follow the cleanup move through, and stop miscalling a converged run

Three loose ends the earlier cleanup relocation left, two of them found by the
round's own review and one by the suite.

**#3190's own test still pinned the old placement.** T6 asserted the `.iterN.md`
removal lives inside `auto_iteration_loop` — exactly what moving it broke. Its
SEMANTICS are unchanged and still asserted: removed on convergence, retained on
degradation, creation intact. What it now pins additionally is the ordering that
forced the move — the ledger reads the backups BEFORE they are removed — and that
the loop no longer removes what it just wrote. Rewritten rather than deleted: the
assertion was superseded, the guarantee was not.

**`CONVERGED` had become a decoy.** With the removal gone from the loop, the flag
was set in two places and read in none. Deleted, and the prose that still said
"the loop sets it" rewritten to what is true: the loop breaks on exactly one
condition, a clean re-review, which leaves REVIEW.md at `status: clean` — and that
is what `cleanup_iteration_backups` re-derives from.

**A converged final iteration reported the opposite of what happened.** The
post-loop message keyed on the iteration COUNTER alone, so a run that converged ON
iteration 3 exited with `ITERATION == MAX_ITERATIONS` and printed "Reached maximum
iterations. Remaining issues documented in REVIEW-FIX.md" over a run in which
every finding was fixed. Convergence is re-derived from the review the loop left
behind — the same signal the cleanup step reads, so the two cannot disagree.

* docs(#3829): retract two claims this round made and could not support

Both were caught by the round's own adversarial review, both were driven, and
both would have reached the maintainer. Recording the retraction where the claim
was made, rather than only in a PR comment.

**The env-prefix "latent bug" does not exist.** An earlier commit in this round
claimed that `FIX_REPORT_FILE="${_pd}/${PADDED}-REVIEW-FIX.md"`, sitting in the
same `node -e` env-prefix list that sets `PADDED`, expanded the OUTER variable
rather than the one two entries earlier — reading ShellCheck's SC2097/SC2098 as
a defect report. Driven in bash and in dash: assignments in one prefix list take
effect left to right, and the later entry DOES see the earlier one. The warning
is a false positive here. The split is kept, but for readability only; the
comment no longer describes it as a fix.

**The HAS_BASH decline rested on a rule that does not govern these call sites.**
It cited `local/no-unguarded-nonportable-exec` as REQUIRING the
`process.platform !== 'win32'` guard. Checked, and wrong on both halves: the rule
fires only on a file that also chmods an exec bit with an octal literal, and this
file has none — so it never runs here — while
`eslint-rules/lib/platform-guard.cjs` accepts four guard shapes plus
`os.platform()`, not one. A constraint that exists is not a constraint that
applies, and I did not check which.

The decline stands on narrower and honest grounds: whether these fences PASS on
the Windows lane is UNVERIFIED. What evidence there is points at divergence
rather than absence — the rule's subject line is that `bash -c` constructs "fail
on Windows Git Bash", and this PR already measured `mkfifo` existing on that
runner, exiting 0, and creating no FIFO. So a probe would not be a clean win; it
would light 22 tests on a lane whose shell semantics are known to differ and
unknown in detail. That is a measurement to make deliberately, not a change to
make in passing. The gap is real and is now stated as a gap.

* fix(#3829): close four defects the review drove out of the first title fix

The round's own adversarial review re-ran against the reworked tree and refuted
two more claims. Every item below is its finding, verified before acting.

**An iteration-only decision recorded no title, so the reuse guard leaked.**
`applied` stored `{d, src}` and the row took its title from the current review —
which does not report the finding at all. The row shipped with no title, and the
next review reusing that id hit the title-ABSENT back-compat exception and
inherited the old `fixed`. The exact defect the title machinery exists to close,
surviving through the hole opened for legacy ledgers. `applied` now carries the
title it was decided under.

**A changed decision was dropped in favour of the obsolete one.** The dedupe was
a has()-guard, so re-superseding a finding whose decision had since changed left
the older record standing. It now replaces.

**Re-spaced titles double-recorded.** The dedupe keyed on the raw title while
`sameTitle()` collapses whitespace; the key now agrees with the comparison.

**And the frontmatter was not valid YAML.** `title: Parser: loses data` is
rejected outright by a real reader, and the `superseded:` line format was not
YAML at all. Values are emitted as JSON scalars — YAML 1.2 is a JSON superset —
and superseded records are properly nested. Round-tripped through js-yaml in the
tests.

One more, self-inflicted while fixing the above: the parse registered each
carried superseded record TWICE, once at `- id:` under an empty-title key and
again at `title:`. Records doubled on every run. They are collected during the
walk and registered once, complete.

**T6 was vacuous.** The review flipped `= "clean"` to `!=` in the cleanup and the
rewritten T6 still passed — it greps for `FINAL_STATUS`, `rm` and "retained"
occurring somewhere, never wiring them to a branch. T6b now EXECUTES the fence in
both directions against real files. It fails on that exact mutation.

**And a converged final iteration printed two success messages** — the loop's
break already reported it. This branch now stays silent and exists only to
withhold the degradation warning.

Three CI gates the base range brought in, all tripped by this round's own text:

- `/gsd-code-review` in a comment — runtime workflow artifacts take the colon
  form. Now `/gsd:code-review`.
- The preamble-ordering parity test: my PHASE_DIR comment wrote the literal
  `gsd_run` before the shim preamble. Reworded.
- Prompt-stuffing: the file passed 50K. I trimmed 5.8K of my own commentary
  first; even removing every added comment leaves the added CODE over the line,
  and the file entered this round at 44,523 — 89% of the budget. Added to
  SIZE_ONLY_WORKFLOWS with the same reasoning the two existing entries carry, and
  the same acknowledgement: splitting is the real fix.

* test(#3829): extract the cleanup fence without an ad-hoc markdown regex

T6b's helper used `/```bash\n([\s\S]*?)\n```/`, which trips two of the repo's own
rules: `local/no-adhoc-markdown-parsing` (use the sectionizer, not a hand-rolled
fence regex) and `local/no-crlf-fragile-split` (a bare `\n` against readFileSync
content is wrong under Windows autocrlf).

Line-scanned now, CRLF-normalized first — the same shape `bashFences()` in
tests/code-review-pipeline-regression.test.cjs already uses, which solved this
first. `npm run lint` is clean and T6b still fails on the inverted-branch
mutation it exists to catch.

* fix(#3829): withdraw the superseded-decision store; keep the identity guard

Three adversarial passes over this round each found real defects, and passes 2
and 3 were entirely inside the `superseded:` block added in pass 1 — a second
identity scheme, keyed on (id, title), living beside the row store keyed on id.
Pass 3 refuted it on three separate counts: a legacy title that merely looked
like JSON lost its quotes and fabricated a record; a finding that was deferred,
superseded, then returned and fixed left an active row and an obsolete
superseded record standing together, reporting `unchanged` forever; and my own
test for the replacement path never passed the earlier ledger in, so it guarded
nothing.

The construct had no terminal state. It is withdrawn.

**What survives is the safety property.** The ledger records each finding's
title, and a recorded decision is carried forward only while the id still names
the same finding. That is what stops a renumbered `CR-01` inheriting an earlier
`CR-01`'s `fixed` — a false decision in the artifact whose purpose is telling
triaged from forgotten, and the same class as round 4's blocker.

**What is given up, and it is disclosed rather than hidden.** On a detected
reuse the earlier decision loses its row. The drop is reported on the console
naming the id and what had been decided, the previous ledger is committed so the
row remains in git, and docs/features/code-review-pipeline.md states the
limitation.

Two defects from pass 3 are fixed rather than deleted, because they are in the
guard and not the store:

- **Known-empty and NOT-KNOWN were conflated.** `### CR-01:` yields an empty
  title; that is a title. While it emitted no `title:` key it read back as a
  pre-format ledger and inherited across a reused id — the same leak, three
  passes running. Emitted whenever the title is known, empty included; a carried
  row no source knows stays absent, which is the legacy-compatible read.
  Underneath it was a falsy fallback: `(act && act.t) || priorTitle.get(id)`
  discards `''`. Now a typeof check.
- **JSON.parse ran on legacy values.** A pre-format ledger whose bare title was
  written `"quoted"` was parsed and lost its quotes, so the decision stopped
  matching. The frontmatter now declares `titles: json` and the parse is gated on
  it; a ledger without the marker keeps its scalars.

One defect from pass 3 is NOT mine and is not fixed here: a converged run prints
a success message from the loop break AND another from `present_results`. Both
predate this round. My earlier claim that "the duplicate is gone" was true only
of the pair I introduced; the pre-existing pair stands, and widening this round
into `present_results` is not warranted.

188 tests pass. The three new tests fire against the pre-simplification step.
`lint:ci` exits 0. The step file is 55,590 chars, down from a 62,220 peak.

* docs(#3829): stop the ledger promising a preservation it no longer makes

Fourth review pass. No machinery defects this time — both findings are claims in
text this step SHIPS, which is the class this whole stack exists to prevent.

**The rendered ledger still said "Re-running the gate preserves every row and
every disposition."** That was true until the same round gave the step an
intentional drop for a reused finding id, and then it was false in the artifact's
own user-facing footer. It now states what the step does, including the one
exception, where a reader actually meets it.

**And the console asserted "the previous ledger is in git."** Committing the
ledger is gated on `commit_docs`, and a failed commit is swallowed — so under
`commit_docs=false` the overwritten decision may exist nowhere. The note reports
the drop and stops there; asserting a recovery path that may not be there is the
same overclaim in a smaller font.

Two residuals from the same pass are DECLINED and documented rather than fixed,
because both would need the second identity scheme just withdrawn:

- A pre-titles ledger carries no titles, so its decisions inherit on the id
  alone. Refusing there resets every decision in every existing ledger, which is
  the loss the guard exists to prevent.
- Two genuinely distinct findings sharing both an id and a title are
  indistinguishable to an (id, title) key.

The pass also refuted the `titles: json` marker on a ledger written by
`b86ea6065^`, which emitted JSON titles before the marker existed. Declined:
that revision is an intermediate commit on this unpushed branch and has never
been released. The PR's published head writes no titles at all, so a real ledger
is either pre-titles (unmarked, bare — handled) or written by the shipped version
(marked). The unmarked-JSON state cannot reach a user.

Test pinned, and it fails against the pre-correction step.

* docs(#3829): fix four wrong citations and one false size justification

All four came out of a claim-audit of this round's own response comment — an
audit of the text, not the code, which is where the remaining errors were.

- **The caller citation was wrong.** The in-code note said both inputs bind from
  `gsd_run query init.phase-op`. `execute-phase.md:85` uses `init.execute-phase`;
  only `code-review-fix.md:22` uses `init.phase-op`. The substantive point is
  unchanged — both are orchestrator-derived, neither is raw user input — but the
  citation was not checked.
- **A leftover "the prior row is in git."** Removed from the console note last
  commit, left standing in the comment two lines above it.
- **The docs still carried the promise the ledger had just dropped.** The
  rendered footer was corrected; the same sentence in
  `docs/features/code-review-pipeline.md` was not.
- **The SIZE_ONLY_WORKFLOWS justification was false.** It claimed the added CODE
  alone exceeded the threshold. Removing every round-added comment leaves 47,148
  chars against a 50,000 limit, so the file CAN fit — the claim was wrong, and an
  exemption defended on a wrong premise is worse than no exemption.

So the entry is re-justified on what is actually true, and earned first: another
**10,188 chars** of this round's own commentary are cut (62,220 → 52,032, from a
44,523 baseline that was already 89% of the budget). Fitting under is possible
only by stripping essentially all remaining explanation from logic three review
passes found defects in. That is the wrong trade in a file whose house style is
heavy in-fence documentation, and the entry says so rather than implying the
file had no choice.

One measurement corrected while checking: the Windows-lane skip count is **37**,
not the 22 the review cited nor the 26 I first counted. Twenty-two and 26 count
`{ skip: !HAS_BASH }` CALL SITES; a skip on a `describe` cancels its subtests.
Forced the constant false and counted what actually skips.

236 tests pass. `lint:ci` exits 0.

* fix(#3829): the drop report is conditional, and two published claims were not

A fifth adversarial pass, run against the two commits that went out AFTER the
fourth pass and were never reviewed, refuted three claims this round published.

1. The drop is NOT reported unconditionally. `row()` reports only a RECORDED
   decision (`was.d !== 'open'`); a prior row still at `open` is replaced in
   silence. The behaviour is right — `open` records no decision to lose — but
   the shipped ledger legend and BOTH feature docs asserted the report happens
   every time. Text corrected in all three places, which is the same defect
   class this round already corrected once for the preservation promise.

2. The test guarding that console wording was VACUOUS: it ran with no prior
   ledger, so no reuse occurred and its `is in git` assertion could not have
   failed however the console was worded. Driven through a real drop now, with
   the drop asserted as a precondition. A new test covers the `open` arm and
   fails on the pre-fix legend.

3. The HAS_BASH gap is now MEASURED rather than assumed, on native Windows with
   Git Bash 5.2.37 / MINGW64 first on PATH, node v25.2.1:

       HAS_BASH left alone:  179 tests, 127 pass,  0 fail, 52 skipped
       HAS_BASH forced true: 179 tests, 140 pass, 24 fail, 15 skipped

   So 37 of the skips are this guard's, confirming the count the round
   published — and unskipping is NOT a clean win: 24 fail, clustered on
   `bash -c` quoting and spawn failures, exactly the divergence the eslint
   rule's subject line names. The guard stays; it now documents a measured gap.
   The stale "the count is 22" comment is gone.

4. The size-exemption justification was wrong a second time. The overshoot is
   ~2.3K normalized chars, not "essentially all remaining explanation": the
   round's committed peak was 59,246 chars (not 62,220, which was never
   committed), and it entered at 44,466 chars, not 44,523 — both earlier
   figures mixed bytes into a character measurement. Rewritten to the numbers
   the scanner actually produces.

Also: the shipped comment said both callers validate the phase shape without
naming that they validate PADDED_PHASE, not the raw PHASE_NUMBER this step is
handed.

* fix(#3829): renumber this PR's two REQs, which #3661 took while the branch sat

The rebase onto current `next` surfaced a REQ-number collision, not a text
conflict. #3661 landed `REQ-REVIEW-08` (`workflow.code_review_point`) on
`docs/features/code-review-pipeline.md` while this branch also claimed 08 and
09 for severity surfacing and the per-finding disposition. Two different
requirements under one identifier is the kind of thing that reads as correct in
both diffs and is wrong in the merged tree.

Base numbering wins, because it shipped: `REQ-REVIEW-08` stays #3661's. This
PR's two become **REQ-REVIEW-09** (severity surfacing) and **REQ-REVIEW-10**
(per-finding disposition). Swept the whole tree rather than the conflict hunk —
two references sat in files git merged cleanly and never flagged:

- `gsd-core/workflows/code-review-fix.md:450`, the prose stating why
  `record_disposition` is the step's only reachable call site.
- `tests/code-review-pipeline-regression.test.cjs:1782`, the comment on the
  test that pins that call site.

`docs/FEATURES.md` is regenerated from the fragment rather than hand-edited;
`node scripts/gen-features.cjs --check` is green (178 features, 21 groups) and
`lint:generated-sync` exits 0.

Two things stated rather than quietly carried. The `Emitted-Drift-Ack-Growth`
trailer on the round-2 commit still reads `REQ-REVIEW-09` for what is now
REQ-REVIEW-10 — it is a historical acknowledgment of that commit's growth, and
its purpose is unaffected, so it is left rather than rewritten across 52
replayed commits. And `docs/INVENTORY-MANIFEST.json` appeared stale immediately
after the replay, reporting two missing `cli_modules/` entries; that was the
lane's pre-rebase build output, not manifest drift. Rebuilding in the replayed
lane and re-checking shows it in sync and unmodified. Regenerating before the
build would have committed the deletion of two base-added entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): reach the title round-trip with a generator that can break it

Round 6's only finding. The round-5 title tracking introduced a fresh parser
(the `titles: json` / `  - id:` / `    title:` frontmatter walk) and a fresh
bijective contract (`JSON.stringify(oneLine(t))` out, `/^    title: (.*)$/`
plus `JSON.parse` back in), and `tests/code-review-disposition.property.test.cjs`
was untouched since round 4 with no reference to `title` at all. Every heading
the generator built was `'### <id>: finding number <i>'` — never a colon, a
quote, a backslash, or the empty string.

You were right that this is the round-3 shape again, and I would rather
demonstrate that than assert it. Two mutations to the shipped step, each a
plausible edit rather than a contrived one:

  A. render `titles: raw` instead of `titles: json`, so the re-parser never
     JSON.parses and stores the quoted scalar as the title;
  B. `yv = (t) => oneLine(t)` — the bare scalar, no JSON at all.

    mutation A — new generator: FAIL     old generator: pass (3/3)
    mutation B — new generator: FAIL     old generator: pass (3/3)

Both ship past the pre-round suite. The gap was reachable, not theoretical.

What changed:

- `TITLE`, a new arbitrary drawn from the class the render's own comments say
  the escaping is for — `:` (why `yv()` exists), `"` and `\` (what
  stringify/parse must round-trip), the empty string (the known-empty vs
  not-known distinction the render draws explicitly) — plus scalars that MIMIC
  the ledger's own frontmatter grammar (`findings:`, `titles: json`, a nested
  `    title: ` line, `  - id: CR-99`), unicode, surrounding whitespace, and one
  title long enough to outrun a scanner assuming short scalars.
- `FINDINGS` now carries a title per id, so all four properties run the cycle
  over the title contract instead of over a constant. `IDS` keeps the old
  id-only shape it is built from.
- A fourth property asserting the round trip in the two places it is observable:
  the stored scalar must `JSON.parse` back to the trimmed heading title, and a
  hand-recorded decision must survive the next run.

The second half is the one that matters, and its construction is the point.
The decision is made by EDITING THE RENDERED LEDGER IN PLACE, never by writing
a bare row the way the existing properties do. A bare row carries no
frontmatter, so `priorTitle` is empty, `sameFinding()` returns true through its
`!priorTitle.has(id)` back-compat arm, and the title contract is never
consulted — the property would pass over a completely broken round-trip. Both
mutations above go green against the bare-row form. That collapse is why the
property is written this way, and the comment says so in place.

So the assertion is the consequence, not the JSON: a lossy round-trip does not
corrupt a title, it makes `sameFinding()` false and resets a human's `deferred`
to `open` with the reason gone — this PR's own founding failure mode, reached
through the field the round-5 work added.

BOUND, stated rather than quietly omitted: the generator emits no CR or LF. A
`###` heading is one line by definition, so a newline is not an input the
heading parser can be handed; `oneLine()` guards the value's other producers,
not this one.

Two things found while writing it, both corrected here rather than left:

- `runOnce` now returns stdout. The reuse report is a CONSOLE note, not a
  ledger key, so my first draft's `assert.doesNotMatch(ledger, /^reused:/m)`
  was vacuously true forever — a test that cannot fail.
- `expectedTitle` is a TRIM, not a `\s+` collapse. Collapsing is `sameTitle`'s
  COMPARISON rule; `oneLine()` is the STORAGE rule and preserves internal
  whitespace. The collapse form fails on an internal tab against entirely
  correct code, which is how a test gets weakened instead of believed the first
  time it goes red.

The file header claimed "two properties" while three were running; it now
states four, one line each.

239 tests pass across the four pipeline files, 0 skipped. `lint:ci` exits 0
(`lint-workflow-shellcheck`: 203 baseline findings, 0 new).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): the prefix census guard now says four sites, because round 5 added one

Self-found, from re-deriving the round-1 finding-id census this round rather
than carrying the round-1 verdict forward.

The census guard's comment says the prefix set is "written out three times —
the heading matcher, the ledger re-parser, and (by its keys) the severity map".
That was true when it was written. Round 5's title tracking added a fourth
copy: the frontmatter `- id: ((?:CR|BL|WR|IN)-\d+)` matcher that rebuilds
`priorTitle`.

The guard itself did not fall behind, and the reason is worth keeping visible:
`idAlternations()` scans the extracted script by PATTERN rather than walking a
fixed list of sites, so the new alternation was absorbed with no edit. Verified
by running the extractor at this head — three alternations found, one distinct
set, severity map keys `CR,BL,WR` with `IN` on the documented `info` default,
0 domain members not reached.

Only the prose fell behind. Corrected, with the pattern-scan rationale stated
in place so the next reader does not helpfully convert it into the hand-listed
enumeration it deliberately is not — which would be exactly the defect this
guard exists to catch, in the guard.

Census discharge for this round: re-derived at the rebased head over the
extracted shipped script, 3 enumeration sites reached, 0 not reached; the
domain (the prefixes `gsd-code-reviewer.md` can emit, walked across both its
heading template and its prose Label-equivalence paragraph) is unchanged since
round 1 at 4 of 4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): catch a duplicate REQ id in a fragment, since nothing did

Not from your review — this is the test the round owed itself, and I would
rather say why than let it look like scope creep.

The renumber commit earlier in this round has no reversion control without it.
I reverted that fix to check, and the first attempt LOOKED controlled: reverting
only the fragment turned `gen-features --check` red. That is the generated-sync
gate noticing the projection went stale, not anything noticing the collision.
Reverting CONSISTENTLY — fragment plus a regenerated `docs/FEATURES.md` — is
silent:

    gen-features --check   rc=0
    lint:ci                rc=0
    pipeline suite         rc=0

with two `REQ-REVIEW-08` entries standing in one requirement list. Nothing in
the repo reads REQ ids at all, so there was no second place for it to be caught.

The failure this guards is a MERGE, not an edit, which is why review does not
see it: two PRs open at once each append "the next" REQ number to the same list,
and whichever lands second is rebased onto a list that already used it. git
merges them as different lines of one file and reports nothing. Neither PR's
diff shows a collision — each is correct against the tree it was written on.
That is exactly how #3661 and this PR both ended up claiming REQ-REVIEW-08.

Scope, stated because it is the part that could be wrong: the check is WITHIN a
fragment, never across the corpus. Two different features legitimately both
carry `REQ-REVIEW-01..07` — the cross-AI review feature and the code-review
pipeline — so corpus-wide uniqueness would be false on the committed tree and
would have to be weakened the day it first ran. A requirement list belongs to
its feature; that is the scope of the identifier.

It lives in `describe('the committed docs/features/ corpus')` because it is an
invariant over the committed corpus, which is that block's stated job, and it
pins no count — the file's own header rules out counts as shared mutable cells
that every feature PR would have to edit.

Control: green on the committed tree (no fragment carries a duplicate today);
red on the restored collision, naming the file and the id. 85 tests pass in
this file.

Happy to drop this if you would rather the round stayed inside the review's
four corners — but then the renumber ships uncontrolled, and I would rather put
that choice in front of you than make it quietly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): finish the census comment correction, which stopped one line short

Found by this round's own pre-push adversarial review, which refuted the claim
the previous commit made about itself.

`7f019d985` said the census comment correction was complete. It corrected one
site and left two, both in the helper block twelve lines above the test it
belongs to:

- `severityMapKeys`' header still read "The THIRD copy: the severity map's
  keys". With three alternations the map is the FOURTH copy, and has been since
  round 5.
- `idAlternations`' header said "adding a prefix to only two of them is silent",
  written when there were two alternations and never updated to three.

This is the defect the original correction was ABOUT, committed inside the
correction: a fragment of prose carries no supersession marker, so a reader
landing on line 2810 gets the dead count stated as current fact, and the fixed
comment eighty lines down does not reach them. Fixing one surface and leaving
its neighbour is not a partial fix, it is the same fix not done.

The region is now consistent end to end, and both headers say the thing that
actually matters — the scan is by PATTERN, not a fixed list of sites, which is
why round 5's new matcher needed no edit here and why converting it to an
enumeration would reintroduce exactly the drift it guards.

WHILE HERE, a disclosure that was narrower than the truth. `7a6680e8f` said the
`Emitted-Drift-Ack-Growth` trailer still names REQ-REVIEW-09 for what is now
REQ-REVIEW-10, and left it deliberately rather than rewrite 52 replayed
commits. That is right, but it is not the whole set: the message BODIES of
`c94106568` ("wire the disposition ledger into the fix path") and `06282f668`
("migrate the emitted-drift ack") both state "REQ-REVIEW-09 was unreachable in
every shipped path", meaning the disposition requirement, which is now
REQ-REVIEW-10.

Same decision, stated at its real size: three historical references, not one.
They are commit history rather than living documentation — git is the record of
what was believed when — and rewriting the branch to correct a number in a
message would cost every review round its correspondence to the commits it
reviewed. The TREE carries no stale reference; `docs/`, the workflows and the
tests all read REQ-REVIEW-09 for severity surfacing and REQ-REVIEW-10 for the
disposition.

Regression file: 181 tests pass, 0 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): the third stale count, and a disclosure that over-counted itself

Both found by re-running this round's pre-push review after the last fix. It
refuted the commit that claimed the region was consistent — for the second time
in a row — and it was right again.

**The third site.** `:2917` said "And the third copy, which is not an
alternation" and `:2919` said "Without this, both regexes can gain a prefix".
Written when there were two alternations; there are three, so the map is the
fourth copy and it is three regexes that can drift.

Worth saying how it survived two passes, because the mechanism is the point and
it is the same one this PR keeps re-learning. Both earlier passes VERIFIED with
a grep built from the strings I had just fixed — `THIRD copy`, case-sensitive,
plus a handful of phrasings I expected. `the third copy` in lowercase matched
none of them, and `both regexes` was not a phrasing I thought to look for. A
grep returns what you already thought of; that is not a verification of prose,
it is a re-statement of your own assumption. The region is now checked by
reading it end to end, and all four count statements agree: three alternations
(heading matcher, ledger row re-parser, frontmatter `- id:` matcher), with the
severity map as the fourth copy.

**And the disclosure over-counted.** The previous commit widened the historical
REQ-REVIEW-09 references from one to three. Three is wrong. There are TWO
underlying statements:

  - `c94106568`'s message body, and
  - the `Emitted-Drift-Ack-Growth` trailer on `06282f668`.

I counted `06282f668` twice — once as "the trailer" and once as "a body" — when
its only mention IS that trailer (`git show -s --format=%B 06282f668 |
grep -c REQ-REVIEW-09` outside the trailer line: 0). Over-counting is the safe
direction and it is still a wrong number in a message, which is the thing this
round has been correcting all along.

The decision is unchanged: both are commit history rather than living
documentation, and rewriting the branch to fix a number in a message would cost
every review round its correspondence to the commits it reviewed. The TREE
carries no stale reference — 08 is #3661's `workflow.code_review_point`, 09 is
severity surfacing, 10 is the per-finding disposition.

Comment-only in one test file; no assertion, regex or extracted-script
expectation moved. Regression file: 181 tests pass, 0 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* fix(#3829): join the disposition-step dispatch so REQ-LANG-04 inheritance is provable

`lint-response-language-coverage` (#2529, which landed on `next` after this PR was
approved) reported `execute-phase/steps/code-review-disposition.md` as having no
response-language coverage. The step does inherit it: `execute-phase.md` imports
`references/execute-phase-response-language.md` and dispatches the step with
`Read and execute`. The dispatch stub wrapped, leaving the verb at the end of one
line and the path at the start of the next, and `namesFragmentAsEntryPoint` matches
within a single line — so a genuine inheritance was unprovable to the linter.

Rejoining the verb and the path restores it: `namesFragmentAsEntryPoint` goes
false -> true and the lint reports `OK (165 workflows covered)`. Only line breaks
move — the word stream is identical to the previous revision, and the file is
unchanged at 93,390 bytes, so no growth acknowledgment is owed.

This takes the third coverage form the lint documents — inheritance — rather than
the inline directive the CI message names first. Where inheritance is provable the
lint's own comments say a second copy "buys no coverage and adds a sentence that
can drift", and the step file already sits over the prompt-stuffing threshold.

Swept all 76 fragments in the catalog: this is the only one whose parent's previous
line ends with a dispatch verb. The 17 others that are mentioned without a provable
entry point are table-routed or bare prose references carrying no dispatch verb at
all, and correctly hold the pinned inline directive instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JgX6QQmygeZnQqbc3o8RNC

* chore(#3829): regenerate derived artifacts after rebase onto next

The rebase onto current `next` conflicted on the 19 install-tree goldens and
`docs/FEATURES.md`. Those are generated, so the conflicts were resolved
arbitrarily and the generators re-run (`npm run regen:derived`) rather than
hand-merged — a clean textual merge of a generated file attests the merge, never
the content.

Reconciled per artifact against the base's own committed copy rather than against
the pre-regen tree, because the pre-regen tree is the arbitrary resolution:

  - all 19 `tests/fixtures/install-tree/*.json` now differ from
    `upstream/next` by exactly one key,
    `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`;
  - `docs/FEATURES.md` differs by exactly REQ-REVIEW-09/10 and this PR's own
    reference section;
  - `docs/INVENTORY-MANIFEST.json` differs by exactly the same one step file,
    and needed no regeneration to get there.

Nothing the base added was dropped by the arbitrary resolution: the restored
entries (the `gsd-core/agents/` and `gsd-core/commands/gsd/` families, the
compact templates, the `detail/elaboration.md` files, `gsd-secret-read-guard.js`)
are all base-owned and came back through the generator, which is what the
resolve-arbitrarily-then-regenerate discipline is for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* fix(#3829): repair the rebase's conflict resolution in the regression suite

The rebase onto current `next` hit one add/add conflict in this file: #4209's
external-reviewer-evidence describe and this PR's #3829 block were added at the
same insertion point. Resolving it by keeping both sides was correct in
substance and wrong in mechanics — the conflict boundary cuts through two open
blocks that the SHARED trailing `  });\n});` closes, so each side carries +2
unbalanced braces on its own and concatenating them left the file with 683 `{`
against 680 `}`.

`node --check` fails outright, so the whole file deregistered rather than
failing a test — 188 tests silently stopped existing. Rebuilt the region as a
real three-way merge (ancestor a262ad6b6, ours upstream/next, theirs 77ee739c3)
and closed the first side explicitly before the second begins.

Both feature blocks are present exactly once, braces balance 683/683, and the
file runs 188/188 locally. The sibling markdown file resolved the same way is
unaffected and was checked rather than assumed: prose has no block structure to
unbalance, and it differs from the base by 112 added lines with zero removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* test(#3829): inject the unreadable-review failure in a way root cannot bypass

Round 9 finding 1. `runShippedGateCounts({ mode: 0o000 })` does not simulate an
unreadable review under root: root bypasses POSIX read permission bits, so the
fence's `[ -f "$REVIEW_FILE" ] && [ -r "$REVIEW_FILE" ]` guard stays true, the
fixture is read, and the assertion sees a real breakdown where it expects
silence. Reproduced as reported — `node:24-slim`, euid 0:

    not ok 18 - an unreadable REVIEW.md leaves the counts empty and does not abort
    actual: 'Code review: 4 findings — 1 critical, 2 warning, 1 info.\n...'

**The prescribed remedy does not reach this site, so this adapts it rather than
applying it.** Stubbing `fs.readFileSync` to throw EACCES is the right fix where
the read happens in-process; here the read is performed by a spawned `bash`, so
node's `fs` is not on the code path and the stub would change nothing.

What the guard actually has is two legs, and only `-r` is defeated by root:

  - the `-r` leg keeps the mode-bit fixture and declares the lanes it cannot
    bind on (`win32`, `euid 0`), which is exactly what
    tests/plan-review-convergence.test.cjs:2326 does for its own shell-side
    `-r` arm — the repo's existing precedent for this shape;
  - the `-f` leg is new and root-immune: a DIRECTORY at the review path fails
    `-f` for every euid, reaching the same non-reporting arm with the same
    observable. It binds on the bench lane where the first test is skipped.

Skipping the first without adding the second would have traded a false failure
for lost coverage on the only lane that found this.

Reversion control, run as root: reverting this commit fails exactly
`an unreadable REVIEW.md leaves the counts empty and does not abort` and its
enclosing describe `#3861 round 1 — the counts mirror is asserted against the
shipped shell`, with no other change to the failure set. **That answers the
round's open question** — the review flagged the describe as possibly a second
root cause; it is the first one's rollup, and there is no second.

Local (euid 1000): 189/189, both tests run.
Root: 179 pass / 1 skip, the skip naming its reason, the directory test running.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* chore(#3829): refresh the compact-content benchmark baseline for the gate's edit

Self-found in this round; not raised in review. The base range landed #4139's
compact-content benchmark, whose committed baseline records per-workflow token
counts. This PR replaces a 12-line bash block in `execute-phase.md`'s
code_review_gate with a 5-line dispatch paragraph, which moves that workflow's
measured counts by 17 tokens — so the baseline the base just added drifts
against a tree it was measured before.

    DRIFT: split "execute-phase": off 25631 -> 25614 (-17), on 23380 -> 23363 (-17)
    DRIFT: aggregate: off 106923 -> 106906, on 90275 -> 90258

Refreshed with the remedy the script itself names
(`node scripts/benchmark-compact-content.cjs --write`). The regenerated diff
touches only the `execute-phase` entry and the aggregate — every other
workflow's numbers are byte-identical, which is the reconcile this PR's edit
predicts.

Attributed rather than assumed: `tests/benchmark-compact-content.test.cjs` is
27/27 at `upstream/next` with no PR content, and was 26/27 on this head. So the
drift is this PR's, not base noise — and it is invisible to a diff-scoped sweep,
because the PR never touches the baseline file and the base range is what
created it. The sibling `benchmark:compact-content-variants` was checked in the
same pass and reports up to date, so this is the only one of the pair affected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* test(#3829): cover an EMPTY REVIEW.md, a case the body claimed and no test reached

Found by this round's own adversarial audit of the PR body, not by review. The body
has said since round 1 that the suite covers "a REVIEW.md that is missing, empty,
or a directory". Two of those three were true. The empty one was not.

Every `reviewText: ''` call in this file also passes `writeReview: false`, which
makes the file MISSING, not empty — so the arm the body named had no test at all.
They are genuinely different paths through the shipped fence: a missing file never
gets past `[ -f ]`, while an empty one passes both `[ -f ]` and `[ -r ]` and is
actually opened and read.

Probed the shipped fence directly against a real empty file before asserting
anything: exit 0, empty stdout. So the behaviour was already correct and only the
coverage claim was false — which is the same "documented as covered, not covered"
shape this PR exists to make visible in the review gate, found in its own body.

The explanatory comment is deliberately precise about WHY the scan yields nothing,
because the plausible reading is wrong and a later reader would inherit it: it is
not the `NR==1{if($0!="---") exit}` guard. A zero-byte file gives awk no record, so
that action never runs (NR stays 0); the output is empty because `closed` is never
set. The comment also states what the test does not prove on its own — its
observable is identical to the missing-file case, so "the file was read" rests on
the harness and the fence, not on the assertions.

189 -> 190 tests in this file, all passing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* chore(#3829): regenerate derived artifacts after rebase onto next

The rebase onto current `next` conflicted in the 19 install-tree goldens,
`docs/INVENTORY-MANIFEST.json` and the compact-content benchmark baseline.
Those are generated, so they were resolved arbitrarily and regenerated
with their own producers rather than hand-merged: every golden now differs
from `next`'s committed copy by exactly the one PR-owned entry
(`gsd-core/workflows/execute-phase/steps/code-review-disposition.md`), the
manifest by the same entry, and the benchmark baseline by the
`execute-phase` split plus the aggregate.

* chore(#3829): regenerate the platform-conformance tier for the added property test

`next` gained the conformance-tier classifier (#4591) and its CI gate after
this branch was cut. The branch adds `tests/code-review-disposition.property.test.cjs`,
so the generated tier list was one file short (546 != 547). Regenerated
with `node scripts/gen-platform-conformance-tier.cjs --write`; the macOS
tier (`--target macos --check`) already matched.

* chore(#3829): regenerate the two conflicted generated artifacts after the rebase

`next` moved 6 commits past the previous base and conflicted in exactly two
files, both generated:

- `tests/fixtures/compact-content-benchmark-baseline.json` — #4208 (`4cc2a466b`)
  and #4619 (`db4d8a9ba`) both moved the measured token counts, and this branch
  moves the `execute-phase` split too.
- `scripts/lib/platform-conformance-tier.generated.cjs` — #4641 (`4d65c248e`)
  made test-conformance the sole Windows selector and narrowed the tier to
  28.5%, and #4568/#4619 re-ran it after.

Both were resolved arbitrarily during the replay and then regenerated with
their own producers rather than hand-merged, per the generated-artifact rule:
`node scripts/benchmark-compact-content.cjs --write` and
`npm run regen:derived` (which runs `gen-platform-conformance-tier.cjs
--write` for both the default and the macOS target).

Reconciled against `next`'s own committed copies rather than the pre-regen
tree:

- the benchmark baseline differs from `next` by exactly the `execute-phase`
  split (`offTokens` 25827 -> 25810, `onTokens` 23576 -> 23559 — the 17-token
  delta this PR's step-file extraction has carried since round 5) plus the
  `aggregate` that sums it;
- the conformance tier differs from `next` by exactly one added entry,
  `tests/code-review-fix-pipeline-regression.test.cjs`. Under the narrowed
  28.5% selector that is the file the classifier now picks from this PR's
  test set. The arbitrary resolution had carried 282 stale lines computed
  under the pre-#4641 selector (`--numstat` on this commit: 3 insertions,
  282 deletions), and regeneration collapsed them.

The full derived sweep was run, not just the two named producers: all 19
install-tree goldens, `docs/INVENTORY-MANIFEST.json`, `docs/FEATURES.md`,
the macOS conformance tier and the exit-code registries regenerated
byte-identical, so nothing else drifted under the new base.

* fix(#3829): accept N-segment phase ids, and bound length per component

The base range added `scanMarkdownSingleSegmentPhaseRegex` (#4568,
`a2331c01f`), which refuses the single-optional-segment phase regex on
phase-carrying markdown lines under three roots — `gsd-core/workflows/`,
`gsd-core/references/` and `agents/` (`lint-phase-id-drift.cjs:301`,
`:373-387`). It flagged two lines in this step file. Chasing the flag
turned up two real defects behind it, so this commit is those rather than
the comment edit the flag literally asked for.

## Defect 1 — the step refused ids both its callers accept

The comments asserted that both callers validate `^[0-9]+(\.[0-9]+)?$`.
#4568 had widened those two call sites to `^[0-9]+(\.[0-9]+)*$`, so the
prose was stale. Correcting only the prose would have shipped a comment
promising N-segment support over code that refused it, because the step
carried a third `case` arm:

    *.*.*)        _ok=0 ;;   # more than one dot: not the documented shape

`23.1.2` took the refusal arm, `PADDED` came back empty, and the step
printed `Code review reporting skipped (unusable phase number ...)` and
wrote **no ledger** — for a phase id both of its callers accept.

It degraded loudly rather than silently; there is a diagnostic on stdout.
The traversal-fence test asserted that refusal as *correct*, listing
`1.2.3` among the values that must be rejected, so an arity bound and a
shape bound sat folded into one `case` arm with a test pinning the pair.

The arity arm is gone. Deleting it alone would have left the step
**wider** than its callers in one direction — `1..2` has an empty
segment, which `^[0-9]+(\.[0-9]+)*$` refuses and the retired arm had been
masking — so a third arm replaces it:

    *..*)         _ok=0 ;;   # EMPTY SEGMENT

## Defect 2 — the length bound was not per-component, though its comment said so

Removing the arity arm made a second defect reachable. The bound read:

    case "$_pn" in *.*) case "${_pn#*.}" in ?????????*) _ok=0 ;; esac ;; esac

`${_pn#*.}` is the whole tail after the first dot — one component only
while an id has at most two. With N-segment ids accepted, that form
rejects `1.1234567.1`, whose every component is a legal 7 digits, purely
because the tail measures 9 characters. The comment directly above it has
read **"LENGTH-BOUND EACH COMPONENT SEPARATELY"** since before this PR,
and had itself named the composite bound as "too strict" — the same
mistake, surviving one level up.

Both fences now walk the segments and bound each:

    _rest="$_pn"
    while [ -n "$_rest" ]; do
      case "$_rest" in
        *.*) _seg="${_rest%%.*}"; _rest="${_rest#*.}" ;;
        *)   _seg="$_rest";       _rest="" ;;
      esac
      case "$_seg" in ?????????*) _ok=0 ;; esac
    done

The `$((10#...))` overflow guard is preserved, per component: bash
integers wrap at 2^64, so an unbounded integer segment would silently
become a negative padded phase.

**The remaining divergence from the callers is a CLASS, not a list:** any
id carrying a component of nine or more characters is caller-accepted and
fence-refused — `123456789`, `1.999999999`, `1.123456789.1`,
`123456789.1`, `1.1.123456789` and so on. That narrowing is deliberate
and is the overflow guard. An earlier draft named two examples as though
they were exhaustive; that wording is withdrawn.

## What is NOT claimed

- The canonical grammar is `PHASE_NUMBER_TOKEN_SOURCE` in
  `src/phase-id.cts:65`, `\d+[A-Z]?(?:\.\d+)*`, added by **#2128**
  (`09be501eb`, 2026-07-10). An earlier draft dated it to #865
  (2026-06-08); that was the first commit to touch the *file*, not the
  one that added the constant, and it is withdrawn.
- #4568 gave the six shell sites **segment-count** parity with that
  grammar, not textual parity: the canonical source permits an optional
  `[A-Z]`, and the shell literals remain digit-only. Driven: this step
  and both callers all refuse `23A.1`, so they agree with each other and
  are jointly narrower than `src/phase-id.cts`. That is a question about
  the six sites rather than about this step, and it is not touched here.
- **#4619 does not produce N-segment ids.** It only transforms an
  already-supplied `{phase_number}` so `$((10#...))` does not abort on
  one. An earlier draft cited it as the producer; that is withdrawn, and
  is stated rather than silently swapped so a reader can see it was
  corrected.
- That the folded `case` arm is *why* nothing caught this is an
  observation about the test's shape, not an established cause.

## Tests

- `an N-SEGMENT phase number reports counts, exactly as its callers
  accept it` — drives `23.1.2` **and** `1.2.3.4`: the retired guard was
  arity-shaped, so a bound merely moved from two dots to three would pass
  a three-segment-only test.
- `the length bound is PER COMPONENT, not over the whole tail after the
  first dot` — drives an 8-char and a 9-char **middle** segment, the
  position the old form got wrong.
- `the fence agrees with its callers across a probed set spanning both
  boundaries` — example-based, and says so: a finite probe cannot prove
  congruence over an infinite language, and one review pass demonstrated
  that by injecting a `2) _ok=0` arm this test still passed. It is a
  regression pin over the values that actually broke.
- The traversal list loses `1.2.3` (legal at this base) and gains `1..2`
  and `1.2.` — the malformed-dot cases the arity guard had masked.

**Negative controls, re-measured against reconstructed fences:**

    fence state              N-seg   probed   per-comp
    fully pre-fix            FAIL    FAIL     FAIL
    shape-fix only           PASS    FAIL     FAIL
    this tree                PASS    PASS     PASS

Two tests, not one, catch Defect 2: the caller-agreement probe includes
`1.1234567.1`, so the whole-tail bound breaks it too. An earlier draft
claimed the per-component test failed alone — that table was written
before `1.1234567.1` was added to the probe and was not re-measured
afterwards. It is corrected here from a fresh run. The N-segment test
correctly does not fire on Defect 2; it predates the bound work and is
insensitive to it.

Found by this round's own adversarial review passes.

* docs(#3829): name this step's two dispatchers correctly, in code as well as in the PR body

`code-review.md` is not a call site of this step. It carries an identical
`^[0-9]+(\.[0-9]+)*$` validator, which is why it kept getting cited as one, but
it never dispatches `code-review-disposition.md`. The two dispatchers are
`execute-phase.md` (`code_review_gate`) and `code-review-fix.md`
(`record_disposition`) -- and only the second validates anything.

This round corrected that in the PR body and simultaneously wrote the old
conflation into the shipped comments, so the file asserted at line 24 what the
body denied in public, and contradicted its own line 48. Four false assertions,
each duplicated because the fenced block is emitted twice:

- "Both callers explicitly accept ... (code-review.md:63, code-review-fix.md:39)"
- "#4568 widened both of this step's callers" -- it widened the one that
  validates; the other has no validator to widen
- "Both callers already validate ... (code-review.md:63, code-review-fix.md:39)"
- "a SHAPE (..., asserted by both callers)" -- asserted by one

The same conflation had propagated into three comments in
`code-review-pipeline-regression.test.cjs`; corrected there too.

Adds the one fact that follows from naming the dispatchers correctly and that
nothing else in the tree records: `execute-phase.md` applies NO shape gate, so
this fence is not mirroring an upstream guarantee -- it IS the guarantee. A
later reader who believes the caller validates will "simplify" it away.

Comments only. The executable shell is byte-identical to 03adf9474 (verified by
stripping comment lines and diffing). 200/200 regression + property, 47/47
prompt-injection security scan, eslint and lint:generated-sync clean.

The wider [A-Z]-axis divergence between these sites and the canonical
`PHASE_NUMBER_TOKEN_SOURCE` is tracked separately as #4660 and deliberately not
restated here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xdw628PpLYveWfJu7kDyvZ

* chore(#3829): regenerate the two conflicted generated artifacts after the rebase onto next

Both conflicted during the replay onto `eb49ff98d` and were resolved arbitrarily, then
regenerated with their own producers (`npm run regen:derived`,
`node scripts/benchmark-compact-content.cjs --write`) rather than hand-merged. Reconciled
against next's committed copies: the conformance tier differs by the one entry this PR
adds, the benchmark baseline by the `execute-phase` split (the same 17-token delta this
PR's step-file extraction has carried since round 5) plus the aggregate that sums it. The
rest of the derived sweep — 19 install-tree goldens, INVENTORY-MANIFEST, FEATURES, the
macOS tier, the exit-code registries — regenerated byte-identical.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* fix(#3829): a carried row keeps the severity the ledger recorded, instead of re-inferring it from the prefix

The ledger always wrote a severity for every row (table cell and frontmatter key) and
nothing read either back: the prior-row regex skipped the cell as `[^|]*`, the frontmatter
walk collected only titles, and a carried row was rebuilt through sev() from the id prefix,
because sectionSev holds only the findings the CURRENT review reports. So a WR-04 the
reviewer filed under `## Critical Issues` was recorded critical, a human deferred it, and the
next run -- the review no longer reporting it -- silently re-recorded it warning. The one
artifact whose purpose is remembering a finding's severity lost it on the second run, in the
unsafe direction (round 11, reproduced by executing the shipped script twice).

Both persisted copies are now read back, enum-validated (ADR-227, as the disposition column
already is): the table cell first, the frontmatter `severity:` as the fallback for a
hand-mangled cell. Severity precedence is the current review's SECTION, then the RECORDED
value, then the id PREFIX, and the recorded value is inherited only while the id still names
the same finding -- the identity rule the disposition already obeys -- so a reused id starts
from its own review. sev() moves below sameFinding() because it now depends on it.

Tests: a new describe drives the reviewer's exact case (WR-04 under `## Critical Issues`,
deferred by hand, dropped by the next review -> stays critical) plus five controls: recorded
outranks prefix under no recognized section; the current section still outranks recorded; a
REUSED id does not inherit; a mangled cell falls back to the frontmatter and a mangled pair
to the prefix; a bare pre-severity row still infers. A fast-check property assigns each
finding a section independent of its prefix, carries every row through an empty review, and
asserts the section severity survives and the third run reports unchanged.

Negative control, measured against the pre-fix step: the two carry tests, the mangled-cell
test and the property fail; the three precedence/back-compat controls pass at both ends, as
they pin behaviour that predates the fix. Every prior carried-row test used CR-01/IN-01,
whose prefix already matched, so the lossy path had returned the right answer by coincidence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* fix(#3829): a malformed REVIEW.md is reported as unparsed, not passed over as clean

A REVIEW.md with three criticals and an unterminated frontmatter yielded REVIEW_STATUS='',
and the counting arm then printed nothing -- byte-identical to a clean review. The guard
that scopes the frontmatter scan was right to yield no values from an unterminated block; the
reporting arm was wrong to treat 'no status' as 'no review'. Block 2 said `status: none`
rather than `clean`, which is why a careful reader could still separate them (round 11,
Minor).

Both fences now record whether the file was actually READ, separately from what it yielded.
A read file with no parseable status -- unterminated frontmatter, no frontmatter, no
`status:` key, a zero-byte file -- prints `Code review status unparsed: ...` with no
breakdown (there is none to trust) and no --fix suggestion (nothing proves there are
findings). Absent, directory and unreadable stay silent: nothing was read, so nothing is
described. Block 2's skip line names the same distinction, `status: unparsed` vs `none`.

The counts mirror follows the shell: a mirror is always handed a text, so its empty-status
arm is the unparsed one, and the existing 'unterminated frontmatter' and 'no frontmatter at
all' parity fixtures now bind the new message on both sides. The EMPTY-file test from round 9
changes its assertion deliberately: its observable is no longer identical to the missing-file
case, which is the point. Five new tests drive the arm, its three shapes, the three shapes
that stay silent, and block 2's wording.

Negative control, against the previous step: the unterminated, no-status, no-frontmatter and
empty-file tests fail, both parity fixtures fail (the mirror moved and the shell had not),
and block 2's `unparsed` assertion fails; the stays-silent controls pass at both ends.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* docs(#3829): name the unlocked read-modify-write and reference #3780 rather than solving it

The ledger is rendered whole from a prior read with nothing serializing two writers, and
this step has two dispatchers plus an invited hand-edit, so the window is real. It is the
shape #3780 reported for WINDOWS.md under parallel executors, which #4681 closed with a
cross-process lock in src/broken-windows.cts. Not taken here, deliberately: the step is a
shell-embedded script with no dependency on the compiled tree, and adopting the lock module
is its own change. Stated at the write site and as a residual in the feature doc; no lost
update has been reproduced (round 11, Minor).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* docs(#3829): cross-reference the two "review disposition" ledgers in both directions

ADR-3806 canonizes a `## Review Dispositions Ledger` section inside PLAN.md for reviews-mode
planning: append-only per round, over REVIEWS.md findings. This PR's
`<NN>-REVIEW-DISPOSITION.md` is a sibling file beside REVIEW.md for the code-review pipeline,
rewritten idempotently with rows carried. Adjacent names, opposite durability rules, and
neither document mentioned the other -- the round-11 review checked the ADR gate against
3806, cleared it, and flagged exactly that mis-read hazard.

An in-place dated amendment section on ADR-3806 (contributor-standards "Amending an accepted
ADR", pattern 1) and a paragraph in the pipeline feature doc, each naming the other and the
axis on which they differ. docs/FEATURES.md regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* chore(#3829): regenerate the compact-content benchmark baseline after the rebase onto next

`next` moved three commits while the round was in flight and the baseline conflicted again;
resolved arbitrarily during the replay and regenerated with its own producer. It differs from
next's copy by the `execute-phase` split this PR has carried since round 5, plus the aggregate.
The rest of the derived sweep regenerated byte-identical.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* fix(#3829): accept letter-variant phase ids, matching the canonical grammar #4744 widened the dispatchers to

#4744 (#4660) landed on `next` while this round was in flight: it widened the six
shell/markdown phase-number mirrors -- `code-review-fix.md:39`, this step's validating
dispatcher, among them -- to the canonical grammar's letter axis (`12A`, `3A`, `23A.1.2`),
and added a `lint-phase-id-drift` ratchet that flags any digit-only mirror left in the
workflow tree. Rebased onto that base, this step was the one it flagged (two fences, two
sites): `12A` was refused by name and wrote no ledger, for a phase id its own dispatcher
now accepts -- the round-10 class ("the step refused phase ids its validating dispatcher
accepts") re-opened by the base. Found by running the base range's modified gates against
the rebased tree, not by the review.

Both fences now admit an uppercase letter in the character class and pin WHERE it may sit --
only as the last character of the integer part, at most once -- so `23a`, `A23`, `2A3`,
`23AB` and `23.1A` stay refused. The per-component length bound is on the DIGITS (the letter
is one character the `$((10#...))` overflow guard has no stake in, so `12345678A` is within
it exactly as `12345678` is), and the letter is carried verbatim after the padded digits,
`3A` -> `03A`, as `src/phase-id.cts` pads it. The two fences stay line-identical except for
their refusal message (the parity test holds), and every comment literal of the old shape
reads the canonical one.

Tests: a new fixture drives `12A`, `3A`, `23A.1.2` and `12345678A` through the shipped
fence to the padded path; the traversal-fence list gains the five wrong placements; the
caller-agreement probe's regex gains the letter axis with both-direction cases, and
`123456789A` joins the deliberate over-bound narrowing. Negative control against the
pre-widening step: the letter-variant test, the caller-agreement probe and the base's
`scanMarkdownLetterlessPhaseMirror` gate all fail; the traversal-fence test passes at both
ends (the five new placements were already refused, by the narrower class).

Also re-anchors the fence and test comments' `code-review-fix.md` / `code-review.md` citations by
content (the validator, not a line number): the line numbers had drifted by one against the rebased
base, and drift again on every rebase.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* chore(#3829): regenerate the two conflicted generated artifacts after the rebase onto next

Both files conflicted during the replay and were resolved arbitrarily rather than
hand-merged, then regenerated with their own producers — `npm run regen:derived`
and `node scripts/benchmark-compact-content.cjs --write`.

Reconciled against next's own committed copies rather than against the pre-regen
tree, because a clean textual merge of a pinned-number file attests the merge and
not the numbers:

- `tests/fixtures/compact-content-benchmark-baseline.json` differs from next by
  exactly the `execute-phase` split (offTokens 26264 -> 26200, onTokens 24013 ->
  23949) and the `aggregate` that sums it. That is the step-file extraction this
  PR has carried since round 5, re-measured against the new base; no other entry
  moved.
- `scripts/lib/platform-conformance-tier.generated.cjs` differs from next by
  exactly one added entry, `tests/code-review-fix-pipeline-regression.test.cjs` —
  the file the classifier picks from this PR's test set.

The full derived sweep was run, not just the two named producers: all 19
install-tree goldens, docs/FEATURES.md, docs/INVENTORY-MANIFEST.json, the macOS
tier and the exit-code registries regenerate byte-identical under the new base,
so nothing else drifted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): stop block 2 computing an unparsed shortfall from a self-contradicting findings block

Round 12, Minor. Confirmed, and the premise is slightly stronger than stated: block
2 does not merely skip block 1's `critical + warning + info == total` cross-check —
it never derived the three severity counts at all, so the check's inputs were
absent. It bounded `total` for digits and length only and handed it to the
`unparsed:` reconciliation.

So a REVIEW.md whose `findings:` block disagrees with itself (`total: 10` beside
`critical: 1, warning: 1, info: 1`) made block 1 print the countless form —
breakdown suppressed as untrustworthy — while block 2 still computed a shortfall
from that same untrusted number. Two trust models for one field, one fence apart,
with the weaker one downstream. It fails in the safe direction, which is why the
review did not raise it as a blocker; it is still a real inconsistency.

Block 2 now derives `critical`/`blocker`, `warning` and `info` through the same
`findings:`-anchored filter block 1 uses, with the same digit-and-length bound and
the same `10#` on every operand, and blanks `total` when the three disagree with it.

**Adapted, not applied verbatim — and the divergence is the point.** The finding
says to re-apply block 1's cross-check. Block 1's gate is `REVIEW_COUNTS_OK`, which
demands all four counts be numeric, because block 1 DISPLAYS all four and
`6 findings —  critical` is the half-filled line that rule exists to prevent. Block
2 displays none of them; it uses `total` alone, against the number of headings the
row parser matched. Applied verbatim, the all-four rule blanks a perfectly usable
`total: 5` on a review carrying no severity keys and SILENTLY DROPS an `unparsed:`
shortfall this step reports correctly today — trading a safe-direction over-report
for a silent under-report, which is the wrong way round and is the exact failure
class the `unparsed:` key was added to close. Only the CONTRADICTION ports: absent
counts are not a disagreement, because there is nothing to disagree with.

Driven against the shipped fence, not a mirror:

  consistent 1+1+1=3          -> total 3       CONTRADICTION total:10 -> withheld
  blocker: alternation        -> total 2       counts absent          -> total 5 (kept)
  leading zeros 01+01+01=03   -> total 03      one count absent       -> total 5 (kept)
  no findings block           -> withheld      non-numeric count      -> total 5 (kept)

Six regression tests drive the second markdown fence end to end through `runHook`
under bash, asserting on the rendered ledger. Negative controls fire in OPPOSITE
directions, which is what pins the narrowing rather than only the fix:

- revert the fence fix          -> `a contradicting findings block yields no
                                   shortfall` and the `blocker:` twin go red
- apply the VERBATIM all-four   -> `a total with NO severity keys still reconciles`
  prescription instead             and the partial/non-numeric case go red

Restored tree: 214/214.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* docs(#3829): state the one input that produces no unparsed shortfall

The reference page said the shortfall "is stated" whenever `total:` exceeds the
parsed headings. After the round-12 fix that is conditional, and a doc asserting
the unconditional form describes behaviour the step no longer has.

Names the boundary in both directions, because the narrowing is the part a reader
would otherwise get wrong: a `findings:` block whose three severities are all
present, numeric and do not sum to `total` produces no key — the same input on
which the console line already withholds the breakdown — while counts that are
merely absent, partial or non-numeric are not a disagreement and still reconcile
from `total` alone.

`docs/FEATURES.md` regenerated; `gen-features --check` green (182 features, 21
groups).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): close two holes the round's own adversarial pass found in its first attempt

Neither is from the maintainer's review. Both were found by the pre-push adversarial
pass over this round's own claims, which refuted them by execution.

**1. A malformed severity could suppress a real shortfall.** The sibling frontmatter
reads use `cut -d: -f2 | tr -d ' '`, and `tr -d` deletes INTERNAL spaces, so
`critical: 1 0` arrives as the perfectly numeric `10`. That is long-standing in those
reads — its mirror is pinned as a fixture from round 1 — and it was INERT in block 2
until this round made that block read the severities at all. At that point a repaired
number could satisfy the new sum test and suppress an `unparsed:` shortfall that is
genuinely owed. Driven, pre-fix: `critical: 1 0 / warning: 0 / info: 0 / total: 5`
against three parsed headings emitted no `unparsed:` key where `unparsed: 2` was
correct.

The three severity reads now trim the ends only, so an internal space survives into
the digit check and fails it — `_sum_ok=0`, nothing is suppressed. Fail-safe in the
only direction that matters: when the frontmatter is malformed the step declines to
suppress rather than trusting a repaired number.

**Scope, stated:** only the SUPPRESSION inputs are strict. `REVIEW_TOTAL`'s own read
still uses `tr -d ' '`, unchanged and identical to block 1's — narrowing it would
change the `unparsed:` computation itself, which is pre-existing behaviour and wider
than this round. So `total: 1 0` is still read as `10` by both blocks, as before.

**2. The repointed #4748 gate could not see a later rebinding.** Its derivation slices
stop AT the first anchored assignment, so inserting the canonical lookup and then
overriding it with `REVIEW_FILE="${_pd}/WRONG-REVIEW.md"` left every assertion green —
the slice pins a line, not the path the fence actually consumes. A new test pins the
whole file instead: the only `REVIEW_FILE=` bindings permitted are the canonical
lookup (exactly twice, once per fence, each being a fresh shell) and the identity
pass-through that hands it to the embedded node script as an env prefix.

Negative controls, both the adversarial pass's own mutations, against the restored
tree at 215/215 and 164/164:
- restore `tr -d ' '` on the severity reads -> the internal-space test reds
- insert the WRONG-REVIEW override after the lookup -> the rebinding test reds

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): one parser for every count block 2 reads, and widen the rebinding guard

A second adversarial pass, run against the first pass's own fixes, refuted three of
them by execution. Fixes to review findings are the class most likely to carry a new
defect, which is why that pass exists; all three were real.

**1. `cut -d: -f2` takes the SECOND FIELD, not the scalar.** So `critical: 1: junk`
arrived as the perfectly numeric `1`, and `1+0+0 != 5` was read as a contradiction
that SUPPRESSED a shortfall genuinely owed. The previous fix trimmed the ends but
still cut at the wrong place, so it closed the internal-space shape and left this one
open. `-f2-` keeps everything after the first colon; the malformed scalar stays
malformed and `total: 5` still reconciles.

**2. The deliberate asymmetry was wrong, and it FABRICATED.** The previous fix parsed
the severities strictly and left `total` lenient, on the reasoning that narrowing
`total` was out of scope. Driven: `critical: 5 0` with `total: 1 0` repaired only the
total to `10`, rejected the severity, skipped the contradiction check, and invented
`unparsed: 7` against three parsed headings. Both uniform policies behave sanely —
strict rejects the malformed total, lenient detects `50 != 10`. A field is either
trustworthy or it is not; parsing one leniently and its sibling strictly is the shape
that fabricates. Every count this block reads now goes through one parser.

**Scope, restated because it moved:** the previous commit said `REVIEW_TOTAL`'s read
was deliberately unchanged. That is no longer true and the reasoning behind it did not
survive contact — the asymmetry it protected is what produced the fabrication. Block
1's reads are still untouched; its own all-four gate runs over consistently-parsed
values, so it has no equivalent split.

**3. The rebinding guard missed an indented or exported assignment.** `^REVIEW_FILE=`
let both `  REVIEW_FILE=...` and `export REVIEW_FILE=...` through, and each executes
exactly like a bare one. The predicate now absorbs leading whitespace and an optional
`export` before the accept-list decides.

Driven after the fix, against three parsed headings:

  critical: '1: junk'  total: 5      -> total 5 kept, unparsed: 2 reported
  critical: '5 0'      total: '1 0'  -> total rejected, no unparsed key
  critical: '1 0'      total: 5      -> total 5 kept (unchanged)
  consistent / contradiction / blocker / leading-zero / absent — all unchanged

Negative controls, each the adversarial pass's own mutation, against 381/381:
- `-f2-` back to `-f2`            -> the second-colon test reds
- `total` back to lenient `tr -d` -> the fabricated-shortfall test reds
- an INDENTED rebinding           -> the rebinding guard reds
- an `export` rebinding           -> the rebinding guard reds

Residual, disclosed: a duplicate `critical:`/`blocker:` key is still resolved by
`grep -m1` taking the first match. Duplicate keys are invalid YAML and the same
first-match rule is long-standing in the sibling reads; detecting them is a wider
change than this round.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): one parser for the whole step, and prove a contradiction from a partial sum

A third adversarial pass, run against the second pass's fixes. Three findings, plus
one this round's own negative control caught afterwards.

**1. An absent severity still bounds the sum from below.** Counts are non-negative,
so a missing one can only ADD: when the severities that ARE present already sum to
MORE than `total`, the block disagrees with itself whatever the absent value is.
Requiring all three before comparing missed that — driven: `critical: 4`,
`warning: 4`, no `info:`, `total: 5` reconciled against a total the present counts had
already refuted. The comparison is two-armed now: EQUALITY when all three are known, a
LOWER BOUND when they are not. An UNDERshoot stays reconcilable, because that is
exactly what the absent count explains.

**2. Block 1 now uses the same parser, so the console and the ledger cannot
contradict each other.** Tightening block 2 first left the two fences disagreeing about
the same bytes. Driven: `critical: 1 0` with `total: 1 0` repairs to 10 and 10, which
SUM — so block 1 reported `10 findings — 10 critical, 0 warning, 0 info.` from a
`findings:` block containing no such numbers, while the ledger recorded three rows and
no shortfall. Block 1's reads move to `cut -d: -f2-` plus an end-trim; both fences now
take the countless arm on that input. The counts mirror moves with them — its whole job
is modelling the shipped pipeline, and it modelled the retired one.

This is wider than the review's finding and I want that visible: the finding was about
block 2 alone. But a disclosed divergence between a console line and a ledger is the
confusion this PR exists to remove, so it is fixed rather than documented.

**3. `REVIEW_FILE+=-wrong` executes and was missed.** The rebinding guard matched only
`=`; `+=` appends (driven: `REVIEW_FILE=good; REVIEW_FILE+=-wrong` prints `good-wrong`).

**4. My first negative control for (2) was VACUOUS, and that is the reason for the new
`BLOCK 1 withholds a breakdown built from REPAIRED counts` test.** Reverting block 1's
parser left the suite green: on every fixture that existed both parsers landed on the
same arm, so parity could not see the difference. A SELF-CONSISTENT repaired breakdown
separates them, and the test pins it directly rather than through parity.

**Correction to the previous commit's claim.** It said moving `total` to a strict read
"changes nothing for a well-formed review". That is false: `total:\t5\t` is valid YAML
(`yaml.parse` returns 5) which the old `tr -d ' '` rejected and the new trim accepts.
The change is an improvement, not a no-op, and the claim was the wrong shape.

Also from the third pass's MISSED: the fixtures exercised malformed `critical` and
`total` only, so they did not pin the four-field symmetry the fix claims. Every field
now gets every malformed shape.

Negative controls, against the restored tree at 383/383 (220 in the pipeline file):
- revert block 1's parser        -> the repaired-counts test reds (was vacuous; now fires)
- revert the overshoot arm       -> the absent-severity contradiction test reds
- a `+=` rebinding               -> the rebinding guard reds

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): finish the mirror update, and give status the same parser as the counts

A fourth adversarial pass. The important finding is that the PREVIOUS commit's mirror
update was HALF APPLIED, and the suite could not see it.

**1. The counts mirror was still on the retired parser.** That file carries TWO
helpers: `first`, used for `status:`, and `firstIn`, used for ALL FOUR counts. The
previous commit updated `first` and the comment above it, and left `firstIn` on
`split(':')[1].replace(/ /g,'')` — so the shipped block had moved to `-f2-` + end-trim
and the mirror had not, while the parity assertion stayed green.

It stayed green for the same reason this round's earlier negative control was vacuous:
on every fixture that existed, both parsers reach the COUNTLESS arm, so the rendered
message is identical and parity cannot see the divergence. Two fixtures now separate
them — `a self-consistent repaired breakdown` (10 == 10+0+0, so the retired parser
renders a full breakdown from a `findings:` block containing no such numbers) and
`tab-separated counts` (valid YAML the retired `tr -d ' '` made non-numeric). Reverting
`firstIn` reds both.

This is the same shape this PR's round-3 reply already recorded about itself: a fix
verified with a grep built from the strings just fixed. The region is checked by
reading it end to end now.

**2. `status:` kept the retired parser after the counts moved off it, and it is the
read where truncation costs most.** `cut -d: -f2` turned the valid YAML scalar
`status: clean:junk` into the bare `clean`, so an unusable status took the CLEAN arm
and suppressed BOTH the console report and the ledger. Driven, both parsers side by
side. The whole scalar matches no arm now, so the step reports. One parser for every
scalar this step reads, in both fences.

**Two claim corrections, no code change:**

- The previous commit implied block 1's console output was preserved for every
  well-formed review. It is not: `critical:\t1` is valid YAML that the retired
  `tr -d ' '` left non-numeric (countless form) and the trim now reads (full
  breakdown). That is an improvement, and the claim was the wrong shape. The
  `tab-separated counts` fixture pins it.
- "Block 1 and block 2 can no longer contradict each other about counts" was
  overstated. It is true of the PARSER, which is what changed. They can still differ
  when the body carries MORE findings than `total:` declares: the reconciliation
  reports a shortfall only, and the excess direction is deliberately clamped so a
  review under-declaring its own total cannot render `unparsed: -1` — pinned by the
  round-2 test `a total SMALLER than the rows is not reported as a negative shortfall`.
  That is pre-existing and out of this round's scope; stating it rather than widening
  scope again.

Negative controls, against the restored tree at 223/223 (387 across both files):
- revert the mirror's `firstIn`  -> both new parity fixtures red
- revert the `status:` parser    -> the clean-arm suppression test reds

`npm run lint:ci` exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): pin the reads to LC_ALL=C, so the parser cannot depend on the machine

A fifth adversarial pass. It refuted the claim that the shipped reads and their JS
mirror are equivalent, and the counterexample is a locale.

**The POSIX character classes are locale-defined, and glibc's C.UTF-8 disagrees with
both C and en_US.UTF-8.** Driven, same sed, same input, three locales:

    LC_ALL=C          clean<U+2003>  ->  clean<U+2003>   (kept)
    LC_ALL=C.UTF-8    clean<U+2003>  ->  clean           (trimmed)
    LC_ALL=en_US.UTF-8 clean<U+2003> ->  clean<U+2003>   (kept)

C.UTF-8 classifies U+2003 — and U+1680, U+2000-U+200A, U+205F, U+3000 — as BOTH
[[:space:]] and [[:blank:]]. So `status: clean<U+2003>` trimmed to the bare `clean`,
took the CLEAN arm, and silently suppressed both the console report and the ledger —
but only on machines whose locale said so. That is the same suppression the previous
commit fixed for `clean:junk`, with a machine-dependent trigger instead of a parse one.

`[[:blank:]]` is NOT the fix — it is locale-defined too, and C.UTF-8 puts U+2003 in it
as well. Nor is `[ \t]`: POSIX bracket expressions provide no escape at all, so `\t` there is a
backslash and a `t`. GNU sed's default reading of it as TAB is an extension, and the
same GNU sed asked for conformance shows the other reading on this host:

    sed -E          's/[ \t]+$//'  draft  ->  draft
    sed --posix -E  's/[ \t]+$//'  draft  ->  draf

A POSIX-conforming sed is therefore expected to truncate `status: draft` to `draf`.
That expectation is derived from POSIX plus the `--posix` demonstration above; it was
NOT driven against a BSD/macOS sed, because this host has none. The portable
fix is to pin the locale: under C the class is exactly {space, tab, NL, VT, FF, CR},
which is precisely what the mirror already spells out literally. The two now agree by
construction rather than by coincidence of the machine.

All 20 read sites (10 `grep`, 10 `sed`, both fences) are pinned. The mirror's four
anchors move from JS `\s` to the same literal class, closing the divergence in the
other direction — `\s` matches a U+2003 indent that the pinned `grep` does not.

**A second gap, found by this round's own control rather than by the reviewer.** The
mirror has two helpers, and the previous commit proved `firstIn` (the counts) was
pinned by a fixture. `first` (the status) was NOT: reverting it left the suite green.
Every pre-existing status fixture left both parsers on the SAME arm — `issues:found`
truncates to `issues`, which is no more `clean` than `issues:found` is — so the status
mirror could drift unseen, exactly as `firstIn` had. The fixture that separates them
is one where truncation FLIPS the arm: `status: clean:junk`.

That is the third time this round a mirror edit was invisible to the fixtures that
existed, and the question that finds it every time is: what input actually separates
the two versions?

**Two prose corrections in the step**, which had gone stale rather than wrong-headed:
the block-2 comment still said "the sibling reads use `tr -d`" after they had all been
moved off it, and the mirror's class comment claimed an equivalence it did not yet have.

Negative controls, each driven against the committed tree:
- drop LC_ALL=C from the shipped seds  -> 5 red, incl. both locale-invariance tests
- revert the `first` status mirror     -> `a status whose truncation would flip the arm` reds
- revert the mirror anchors to `\s`    -> `locale-invariant on a unicode-space indented count key` reds

The locale-invariance tests are the durable guard: the parity fixtures only run under
whatever locale the suite inherits, so they can catch this only on a machine that
already has the bug. These drive the same input under both locales and assert the
shipped fence does not care.

238/238 across the two pipeline files. `npm run lint:ci` exits 0, with
lint-workflow-shellcheck reporting 212 pre-existing findings and 0 new.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): pin the awk selectors too, and assert the invariant instead of claiming it

A sixth adversarial pass, and it refuted the previous commit's central claim. That commit
pinned all 10 `grep` and all 10 `sed` reads to LC_ALL=C and then said the parser no longer
depends on the machine. It does: the `findings:` MAPPING SELECTOR is an `awk`, and both
copies of it were left unpinned.

`awk '/^findings:[[:space:]]*$/{f=1; next} f&&/^[^[:space:]]/{exit} f'` resolves its two
character classes through the ambient locale exactly as grep's and sed's did. Driven, on
`findings:<U+2003>`:

    fence 1, LC_ALL=C        Code review found issues.
    fence 1, LC_ALL=C.UTF-8  Code review: 1 findings — 1 critical, 0 warning, 0 info.
    fence 2, LC_ALL=C        TOTAL=''
    fence 2, LC_ALL=C.UTF-8  TOTAL=1

Under C.UTF-8 the opener matched and the mapping opened; under C it did not. The same
review rendered a breakdown on one machine and the countless message on another, with
every grep and sed already pinned.

**Why it was missed is the more useful part.** The previous commit's census counted
`grep` and `sed` sites and reported zero unpinned — because it SEARCHED FOR THE TOOLS IT
HAD JUST EDITED rather than for the tools that were there. That is the same shape as this
round's other three misses: a check built from the thing just changed cannot see what the
change forgot. So this commit does not just add the fourth and fifth pins; it replaces the
claim with an assertion the next edit cannot fool:

  `every locale-sensitive tool in the step is pinned to LC_ALL=C` walks the step file and
  fails on ANY unpinned `grep`/`sed`/`awk`, naming line and call. `cut -d: -f2-` and
  `tr -d '\r'` stay exempt, and the exemption is principled rather than residual: neither
  resolves a character class or a collation — one splits on a single ASCII byte, the other
  deletes one literal byte.

The mirror's block boundary moves to the same literal classes, for the same reason the
anchors did last commit — `/^findings:\s*$/` and `/^\S/` model neither pinned side.

**A correction to the previous commit's message, made in place.** It asserted that BSD sed
reads `[ \t]` as a literal backslash and `t`, stated as driven fact. It was not driven —
this host has no BSD sed. The claim is now stated as what it is: POSIX bracket expressions
provide no escape, GNU's TAB reading is an extension, and GNU sed asked for conformance
demonstrates the other reading here (`sed --posix -E 's/[ \t]+$//'` turns `draft` into
`draf`). The conclusion is unchanged; the evidence class was overstated.

Also measured while establishing that LC_ALL=C is safe for non-ASCII, and worth recording
because it makes the pin a strict improvement rather than a wash: on a REVIEW.md carrying a
single invalid UTF-8 byte, GNU grep under C.UTF-8 reports `binary file matches` and emits
nothing, blanking EVERY read; under C the value parses and is rejected on its merits. Valid
UTF-8 is untouched either way — the C space class is entirely bytes < 0x80, which no UTF-8
multibyte sequence contains, so the trim cannot split a character.

Negative controls, each driven and restored:
- unpin the four `awk` selectors  -> 3 red, incl. the new invariant test naming both lines
- revert the mirror block boundary to `\s`/`\S` -> both `findings:` opener tests red

241/241 across the two pipeline files. `npm run lint:ci` exits 0 — after it caught a real
defect in the new test itself: `split('\n')` on readFileSync content is banned here
(DEFECT.WINDOWS-CRLF-TEST-PORTABILITY), and it now uses `splitLines()`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): make the pin invariant see calls that are not piped

The invariant added in the previous commit keyed on `| grep|sed|awk`, which is true of
every call in the step today and is exactly the wrong thing to rely on. A guard written
around the shapes that happen to exist cannot see the shape a later edit introduces —
`awk '...' < "$f"` or `$(grep ...)` would have walked straight past it, which is the same
property that let the two awk selectors sit unpinned through a commit claiming the parser
was locale-independent.

It now blanks the PINNED calls and treats anything still naming one of the three tools on
a non-comment line as an offender, so the check is "every call is pinned" rather than
"every piped call is pinned". Comment lines stay exempt: the step's prose names unpinned
forms while explaining why they were retired.

Driven both ways against the invariant alone:
- inject a NON-PIPED unpinned `awk '...' < "$REVIEW_FILE"` -> reds (the old form did not)
- unpin the four piped `awk` selectors                     -> still reds (no regression)
- unmodified tree                                          -> green

241/241 across the two pipeline files, `npm run lint:ci` exits 0. No shipped behaviour
changes in this commit; it only widens what the test can see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): close the invariant's path-qualified hole, and state the limit it keeps

A seventh adversarial pass. It confirmed the shipped fix — census clean at 24 pinned calls,
whole-fence output identical under C and C.UTF-8 on every input built to separate them, and
both fences byte-identical on a well-formed review carrying `café 東京` and `naïve résumé` —
and then refuted the claim I made about the GUARD, not the feature.

`/usr/bin/awk '/[[:space:]]/{exit}' < "$REVIEW_FILE"` passed the invariant. The preceding-
character class shielded any match preceded by `/` or `.`, so a path-qualified call was
invisible. A path-qualified call is still a call; the class no longer shields either. The
substrings that motivated the exclusion are unaffected — `parsed`, `passed` and `awkward`
have a word character on one side or the other, so the boundary still rejects them.

**The rest of that finding is disclosed rather than fixed, deliberately.** `$AWK "$f"` and a
command name computed inside the embedded `node -e` block also evade the check, and they are
not closable by this mechanism: it scans text, not shell or JavaScript command structure. The
same pass that found them showed that widening the regex further only trades those false
negatives for false positives on quoted strings and awk program text. So the test now STATES
its boundary instead of implying it has none — the previous comment claimed a census "a future
edit cannot fool", which was exactly the kind of overclaim this round has been correcting.
What it catches is enumerated there, driven, along with which direction its false answers go.

Driven against the invariant alone, each mutation reporting its own substitution count:
- `/usr/bin/awk ... < "$f"` (the exact evasion) -> reds; before this commit it did NOT
- `env awk ... < "$f"`                          -> reds
- `LC_ALL=C.UTF-8 awk ... < "$f"` (wrong pin)   -> reds
- unmodified tree                               -> green

241/241 across the two pipeline files, `npm run lint:ci` exits 0. No shipped behaviour changes
in this commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): stop the guard's comment promising a closed list of what it misses

An eighth adversarial pass. It confirmed the delta was test-only and that both fences are
byte-identical across it — exit, stdout, stderr and rendered ledger — and then refuted the
comment again, with two more evasions: a command name fragmented in shell (`a''wk`), and an
executable command substitution on a physical line starting with `#` inside a multiline
quoted argument, which the comment exemption skips.

Both are real. Neither is the point. Three passes running have each found one more evasion of
a TEXT scan, which is the actual finding: **the list cannot be closed.** A comment that
enumerates residuals is false the moment someone is cleverer than the enumeration, and fixing
it by appending the newest example just resets the clock.

So the comment no longer claims an inventory. It says what this is — a regression guard against
the accident that has now happened twice in this round, a read added or edited without its pin
in a file where every other read has one — and what it is not: a proof. The examples are marked
as illustrations. The operative instruction is the one that survives any future evasion: treat
anything it reports as real, and never treat its silence as proof a new read is pinned.

No logic changed; the guard catches exactly what it caught before. 241/241 across the two
pipeline files.

`npm run lint:ci` exits 0 — after it caught this commit twice over, which is worth recording
because both were in prose I had just written to be careful:
- the previous message's example path `docs/grep.md` read as a genuine docs reference from this
  file, and lint-docs-guard-registration demanded a baseline entry for a path that does not
  exist. The illustration is now `bin/grep-wrapper`, and the reason is stated inline.
- an earlier commit's `split('\n')` on readFileSync content tripped the CRLF-portability rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): the guard's two directions are not symmetric, so stop saying they are

A ninth adversarial pass. It confirmed the delta before it was comment-only with the guard's
logic byte-identical, and that all six advertised forms are still caught — and then found the
one absolute the rewrite left behind.

The comment said "treat anything this guard reports as real" three lines above admitting the
guard produces loud false positives. Driven: `echo ok # grep is discussed` is reported, and
labelled `(unpinned)`, though it is prose and no unpinned read exists. Both sentences were
mine, in the same comment, written in the same edit that was supposed to remove overclaiming.

The instruction is now the accurate one, which is that the two directions are NOT symmetric:
a REPORT is cheap to adjudicate — read the line, a trailing comment or a path is obvious — while
SILENCE proves nothing, because the known evasions are silent and so is any evasion nobody has
thought of yet. Investigate every report; never read silence as proof a new read is pinned.

Comment-only. Guard logic untouched, `gsd-core/` byte-identical to cac648aab — three consecutive
passes have now confirmed the shipped behaviour unchanged. 241/241 across the two pipeline files,
`npm run lint:ci` exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): a report needs its context to adjudicate, not just its own line

A tenth adversarial pass, and the last one this round runs. It confirmed `gsd-core/` is the
SAME TREE OBJECT as at cac648aab (105b5ed9b) with the guard's caught and silent sets unchanged,
and refuted one more sentence of the same comment.

"A report is cheap to adjudicate (read the line)" is false, driven: the identical reported
physical line `  grep` is a COMMAND after `:` and an ARGUMENT after `printf '%s\n' \`. The guard
reports both, and the reported line alone does not distinguish them — the preceding line is what
settles it. The comment now says so, with that counterexample in it.

This is the fourth consecutive pass to find a defect in this one comment and none in the shipped
code, which is itself the result worth recording: the shipped fix has been frozen since cac648aab
and confirmed byte-identical by four passes, while the prose describing a best-effort text scanner
took four attempts to stop overclaiming. Writing an accurate description of what a heuristic does
NOT do turns out to be harder than the heuristic.

Comment-only; guard logic untouched. 241/241 across the two pipeline files, `npm run lint:ci`
exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): re-home the lookup's phase-id gate beside the step, not in a block upstream removed

#4781 (#4628) removed #4748's letter-axis work from `tests/nsegment-phase-grammar.test.cjs`,
including the `#4748 — the REVIEW.md lookup` describe block. This PR had four tests living in
that block, because that is where the gate was when #3829 moved the lookup out of
`execute-phase.md` and into the lazily-read step file.

Those four tests assert properties of THIS PR's step file, not of #4748's sites. Rebasing onto
the removal would have deleted them silently — the branch would still be green, with its own
coverage quietly gone. They move here instead, unchanged in substance, beside the step they
guard: an unrelated upstream revert can no longer take this PR's coverage with it.

One assertion did NOT come along. The old block also checked that `execute-phase.md`'s init
parse list names `padded_phase`; #4781 removed that field from the list, and the assertion is a
property of #4748's site rather than of this step. Carrying it here would only have pinned
someone else's revert to this PR.

The step never depended on that field in the first place — it computes PADDED itself, validating
PHASE_NUMBER for shape and traversal and padding the digit run through `10#` while carrying an
optional letter verbatim. That self-containment is why the removal costs this PR nothing but the
tests' address.

5 pass, 0 fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* chore(#3829): regenerate the derived artifacts the round-12 rebase invalidated

The base moved from 092d9256b to 003d982c8 (six merges) while this round was in flight, so the
branch was rebased and the derived state had to be re-derived rather than hand-merged.

`platform-conformance-tier.generated.cjs` regains `tests/configured-entrypoint-validation.test.cjs`,
added upstream by #4249. Resolving the conflict hunk-by-hunk in favour of this branch had dropped
that entry; `lint:generated-sync` caught it, which is what that check is for.

`compact-content-benchmark-baseline.json` carries the measured values at the new base rather than
this branch's stale pair: split "execute-phase" off 26200 -> 26242, on 23949 -> 23952
(8.59% -> 8.73%); aggregate off 108243 -> 108285, on 91595 -> 91598 (15.38% -> 15.41%). The
benchmark reports drift and exits 0 either way, so a stale baseline does not announce itself here
— it announces itself in CI.

**The growth acknowledgment is back, and the reason is worth stating.** Against the previous base
this PR left `execute-phase.md` a net -103 bytes: the extraction removed more than the dispatch
paragraph added, so the file ended up smaller than the base's copy and the
`Emitted-Drift-Ack-Growth` trailer became false and was dropped. #4781 then rewrote that file
upstream, and against the new base the same extraction nets +36 bytes (93421 -> 93457) — which is
the figure this PR originally reported at round 2. The size delta was never a property of this
change alone; it is a property of this change against whichever base it sits on, and it has now
been both signs in one round. The trailer was restored then. Round 14 rebased onto `029acd915`, where #4830's
re-land moved `execute-phase.md` again and the same extraction is -103 once more
(93564 -> 93461), so the trailer is false a second time and this commit no longer
carries it. Third sign flip, same reason each time.

`lint:generated-sync` and `benchmark-compact-content --check` both clean afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* chore(#3829): regenerate derived artifacts after rebase onto next

The rebase onto `c9a5cc3e1` conflicted in the 19 install-tree goldens.
Those are generated, so they were resolved arbitrarily and regenerated
with their own producers (`npm run regen:derived`, then
`benchmark-compact-content.cjs --write`) rather than hand-merged — a clean
textual merge of a generated file attests the merge, never the content.

Reconciled per artifact against the base's own committed copy rather than
against the pre-regen tree, because the pre-regen tree is the arbitrary
resolution:

- every install-tree golden now differs from `c9a5cc3e1`'s copy by exactly
  one entry, `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`,
  which is this PR's own new step file;
- the compact-content benchmark baseline by the `execute-phase` entry
  (offTokens 26184 -> 26167) and the aggregate that sums it.

`docs/INVENTORY-MANIFEST.json` was in the at-risk set but regenerated
byte-identical, so it carries no change here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RJPv9LNaPV3Cd3NCpfFKGb

* chore(#3829): regenerate derived artifacts after rebase onto 029acd915

The rebase onto current `next` conflicted in two generated files — the
compact-content benchmark baseline and the platform conformance tier. Both
were resolved arbitrarily and regenerated with their own producers
(`npm run regen:derived`, then `benchmark-compact-content.cjs --write`)
rather than hand-merged: a clean textual merge of a generated file attests
the merge, never the content.

Reconciled per artifact against the base's own committed copy rather than
against the pre-regen tree, because the pre-regen tree is the arbitrary
resolution. Every differing key belongs to a file this PR actually touches:

- `tests/fixtures/install-tree/claude.json` differs from `029acd915`'s copy
  by exactly one entry,
  `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`, this
  PR's own new step file;
- `scripts/lib/platform-conformance-tier.generated.cjs` by exactly one
  entry, `tests/code-review-fix-pipeline-regression.test.cjs`, a test this
  PR adds;
- `tests/fixtures/compact-content-benchmark-baseline.json` by the
  `execute-phase` entry (offTokens 26344 -> 26280) and the aggregate that
  sums it, which this PR moves by editing `execute-phase.md`.

The other 18 install-tree goldens, `docs/FEATURES.md` and
`docs/INVENTORY-MANIFEST.json` were regenerated too and came back
byte-identical, so they carry no change here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* test(#3829): follow #4748's REVIEW.md-lookup gate to the step that now owns the lookup

Self-found while rebasing onto `029acd915`, not raised in review.

#4830 (`8a5166598c`) re-landed #4768's letter-suffix work on `next`, restoring the
`#4748 — execute-phase.md resolves the REVIEW.md path from init's padded_phase`
block in `tests/nsegment-phase-grammar.test.cjs`. That block had been removed by
#4781, which is why round 13 re-homed this PR's own four tests out of it. The
restored block anchors on a line this PR deletes:

    expected 1 line(s) containing "REVIEW_FILE=\"${PHASE_DIR}/${PADDED}-REVIEW.md\"", found 0

It fails at the describe level, so all four of its tests go with it. It was green
before this rebase only because the base did not carry the block yet.

Putting the line back is not available. #3829 moved the lookup into the lazily-read
step file because `execute-phase.md` did not fit under ADR-857's frozen pre-phase-6
ceiling (93600); the parent is at 93461, and the three lines this gate anchors on
cost 183 (measured, not computed: `git show 029acd915:... | sed -n '1168,1170p' | wc -c`). The block would also be dead code — the step performs the lookup.

So the gate follows the lookup. Two of its four assertions are properties of the
lookup and are re-pointed at the step file: that no fence hands PHASE_NUMBER to
`printf "%02d"`, and that both lookups are preceded by a PADDED binding that pads
the digit run through `10#` and carries the letter verbatim. The regression control
on the lookup line itself comes along, now over both fences. The init-parse-list
assertion stays on `execute-phase.md`, which still names `padded_phase`.

Two assertions do NOT come along, and they are the two that were properties of the
INLINE site rather than of the lookup: the `PADDED="{padded_phase}"` literal binding
(the step derives PADDED itself, validating PHASE_NUMBER for shape and traversal
first), and the composition run over the three live lines. The step's executable
coverage — a composition run plus a padding-agrees-with-the-canonical-normalizer
matrix over letter ids — already exists in `tests/code-review-pipeline-regression.test.cjs`
under "#3829 — the step's REVIEW.md lookup resolves a letter-suffixed phase without
a shell re-pad". Mirroring it here would be a second implementation of one grammar.

That coverage is NOT equivalent, and the difference is worth stating rather than glossing. At its
original site #4748's gate was a DATAFLOW pin: `execute-phase.md` bound init's own
`{padded_phase}`, so the lookup could not disagree with the canonical normalizer because it never
computed anything. The step reconstructs the value in shell, so that pin is not available at this
address and AGREEMENT with the normalizer is what replaces it. A follow-on commit adds a
fast-check property asserting that agreement over generated ids, because the existing 13-shape
matrix samples 2 of 26 letters and cannot see a divergence outside its own points.

One consequence is disclosed rather than absorbed: `padded_phase` is now parsed but unused in
`execute-phase.md` (`:95`). Removing it from that parse list is #4830's call on its own site, not
this PR's, so the assertion that it is still named stays.

#4748's property is unchanged: a letter-suffixed phase resolves its own REVIEW.md,
and an already-padded `08` does not read as octal.

Negative-controlled rather than asserted. Against the shipped step: 162 pass, 0 fail.
Dropping `$_let` from both PADDED bindings reds "every lookup is preceded by a PADDED
binding that carries the letter run"; dropping `10#` reds it too. The step file was
restored byte-identical after each control.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* test(#3829): assert the step's padding against the canonical normalizer by property, not by 13 points

From this round's own pre-push adversarial review, not from the maintainer's.

The prior commit re-points #4748's REVIEW.md-lookup gate at the step that now owns the lookup. The
review's finding was that this is not coverage-equivalent, and it is right: at the original site the
gate was a DATAFLOW pin — `execute-phase.md` bound init's own `{padded_phase}`, so the lookup could
not disagree with the canonical normalizer because it never computed anything. The step reconstructs
the value in shell, so agreement with the normalizer is what has to replace the pin.

That agreement was already asserted, but over a 13-shape matrix. Its words: "future canonical
grammar changes could therefore diverge without this gate detecting them." Correct — the matrix
samples 2 of 26 letters and a bounded set of segment shapes, and this PR has twice been told that a
generator which cannot reach the interesting input is a fixture with extra steps (round 3's
SOURCE_CELL, round 5's title generator). Same defect, third address.

So the agreement is now a property over generated ids: a digit run inside the step's own 8-digit
bound, an optional single A-Z, and up to two dot segments, asserted equal to
`normalizePhaseName(id)` through the shipped shell derivation of BOTH fences. Milestone `N-N` forms
are outside the step's accepted domain and are asserted nowhere here rather than silently passed.

Negative-controlled on two mutations, and the second is the one that justifies the property rather
than the matrix:

  %02d -> %03d                      matrix RED,   property RED
  drop the letter when outside {A,B} matrix GREEN, property RED

The second is the added coverage, demonstrated rather than argued: the matrix is structurally
unable to reach a letter it does not enumerate. The step file was restored byte-identical after
each control.

numRuns is 25 — each case spawns bash twice through the process seam, once per fence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* fix(#3829): pad the phase id as a string, so the step agrees with the canonical normalizer

Found by the property the previous commit added, on its second adversarial pass. This is a real
divergence in shipped behaviour, not a test-only correction.

`normalizePhaseName` (src/init.cts, via gsd-core/bin/lib/phase-id.cjs) left-pads a phase id's digit
run to a MINIMUM of two and otherwise PRESERVES it — `padStart(2, '0')`. The step re-derived the same
value arithmetically, `printf "%02d" "$((10#$_dig))"`, which does not preserve: it collapses every
leading-zero run longer than two.

  id          normalizePhaseName   the step (before)
  8           08                   08
  08          08                   08
  008         008                  08     <-- diverges
  0008        0008                 08     <-- diverges
  00000008    00000008             08     <-- diverges
  0008A       0008A                08A    <-- diverges

Consequence: for such an id the gate resolves `08-REVIEW.md` while init emits `008-REVIEW.md`, so it
finds no review and says so — advisory, and therefore silent. That is the failure class #4748 exists
to close, reached by a different road: not a letter this time, but a leading-zero run.

The fix is a string pad that implements `padStart(2, '0')` exactly, at both fences:

  case "${#_dig}" in 1) PADDED="0${_dig}${_let}${_sub}" ;; *) PADDED="${_dig}${_let}${_sub}" ;; esac

It is strictly less machinery than what it replaces. `10#` existed only to stop bash reading a
leading zero as octal inside `$(( ))`; with no arithmetic there is no octal hazard to guard, so the
remedy is retired rather than kept. `scripts/lint-phase-id-drift.cjs` — which exists to flag
unsanctioned `printf "%02d"` re-pads of phase-carrying variables — is green, and now has one less
re-pad to tolerate.

Two dependent assertions move with it, and both are now stated as the PROPERTY rather than as one
spelling of the remedy: the round-14 gate in `tests/nsegment-phase-grammar.test.cjs` asserts the
binding does no arithmetic and carries both `${_dig}` and `${_let}`, and the regression pin in
`tests/code-review-pipeline-regression.test.cjs` follows the new form.

The property's generator is widened in the same commit, because its first cut could not have found
this: it built the digit run with `String(fc.integer(...))`, which can never produce a leading zero,
so it had silently LOST the `08`/`09` coverage the 13-shape matrix beside it already had. The run is
now generated as a digit string, and segment depth goes to four (the repo exercises `1.2.3.4`). The
step's grammar is unbounded in depth; four is a stated bound, and it is this property's residual.

Negative-controlled, and the matrix is the control's control — it stays GREEN on both:

  revert to the arithmetic pad          matrix GREEN, property RED
  drop the letter when outside {A,B}    matrix GREEN, property RED

The step file was restored byte-identical after each. 407 pass / 0 fail across both test files;
lint:ci, gen-features --check, gen-platform-conformance-tier --check, benchmark-compact-content
--check and gen-install-tree-fixtures all clean, with no regenerated artifact moving.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* test(#3829): hold #4748's gate by execution, and retire the comments the string pad made false

Third adversarial pass on this round. Two findings, both fair, neither a correctness defect.

FIRST — the gate pinned a SPELLING, not the property. Its objection was concrete: an equivalent
multi-line string pad would have failed a regex that matches one `case ... esac` line. That is a
false-positive generator, and a guard that false-fires is a guard that gets deleted.

The split is now honest about what each layer can hold. A STATIC gate can hold the DEFECT SHAPE —
no arithmetic in the binding above each lookup — and that is all it asserts. Correctness is held by
EXECUTION: the base gate regains a composition test that runs the shipped derivation slice of both
fences against `normalizePhaseName`, over `3A 8 9 08 008 0008A 23A.1.2`.

That composition test is the one I removed two commits ago, and removing it was the weaker call. At
#4748's ORIGINAL site the gate could be static because the property was a literal binding of init's
own `{padded_phase}`; nothing could disagree, because nothing computed. At this site the step
derives the value, so the property is behavioural and only execution holds it. `008` is in the list
because it is the case the arithmetic pad got wrong and no prior fixture covered.

Controlled three ways, and the middle one is the finding being answered:

  arithmetic pad restored            4 fail   caught
  EQUIVALENT multi-line string pad   0 fail   no false positive
  drop the letter outside {A,B}      1 fail   semantic drift caught

SECOND — the step carried comments the fix had made false, in four places. Two explained `10#` as
part of the live phase derivation; two justified the eight-digit bound by bash integer overflow.
Neither described the code any more. They are rewritten to current truth rather than annotated,
because a fragment carries no supersession marker and overturned prose reads as canon:

  - the octal rationale now says the pad performs no arithmetic and needs no `10#`, and notes that
    `10#` survives in this step only on the severity COUNTS, which really are numbers being added;
  - the length bound now states that overflow is unreachable since the pad stopped converting, and
    that the bound stays for the reason it always also had — every component is interpolated into a
    filename, and filesystem components are finite.

408 pass / 0 fail across both files. lint:ci, lint-phase-id-drift, gen-features --check,
gen-platform-conformance-tier --check, benchmark-compact-content --check and the
prompt-injection-scan security suite are all clean, and no regenerated artifact moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* docs(#3829): correct four comments, including one this round's own rewrite got wrong

Fourth adversarial pass. Comment-only; no code, no test logic, no regenerated artifact moves.

ONE OF THESE IS MY OWN ERROR, introduced two commits ago. Rewriting the length-bound rationale, I
replaced the dead integer-overflow justification with "the bound stays because every component is
interpolated into a filename and filesystem components are finite". That is false, and it was driven
false: the bound is PER SEGMENT, every segment is joined into ONE filename component, and depth is
unbounded. Thirty 8-digit segments yield a 278-character PADDED and a 294-character name against a
NAME_MAX of 255. Replacing a dead rationale with a wrong one is worse than leaving the dead one, so
the comment now states what the bound actually does and names the composite-length gap as a residual
of this validator that predates the pad change. It is not fixed here; it is stated.

The other three are stale rather than wrong:

- `overflow guard` named the length check in two places. Nothing overflows any more -- the pad does
  no arithmetic -- so it is the digit bound, and is called that.
- The `#4748` block header in `tests/nsegment-phase-grammar.test.cjs` still said the step pads
  through `10#`, still said the composition run did not survive the move, and still said the
  executable coverage was "cited rather than copied" -- while the composition test sat twenty lines
  below it. All three were true when written and none survived this round. The header now records
  why a STATIC assertion cannot hold a BEHAVIOURAL property, and that the PR's fast-check property
  is a different instrument over the same contract rather than the same test twice.
- The severity-count comment said a base-inference failure "takes the whole advisory step down under
  `set -e`". It does not: the arithmetic sits inside an `if` condition, a TESTED context, where
  `set -e` is inert, and the consistency check is SKIPPED instead -- which the regression suite
  already records. `10#` stays; only the account of what it prevents is corrected. This one predates
  the round and is corrected because it is adjacent and factually wrong, not because it blocked
  anything.

408 pass / 0 fail across both files; lint:ci, lint-phase-id-drift, gen-features --check,
benchmark-compact-content --check and the prompt-injection-scan security suite all clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* chore(#3829): regenerate the compact-content benchmark baseline after the rebase onto fac0e9de8

The rebase onto current `next` conflicted on this generated fixture, as it has in every
recent round: the base regenerates it for its own token deltas and this branch regenerates
it for `execute-phase.md`'s, so both sides rewrite the same keys. Resolved arbitrarily
during the replay and regenerated with its own producer
(`scripts/benchmark-compact-content.cjs --write`), never hand-merged.

Key-level drift against the base's committed copy is exactly two entries and both are this
PR's own: `splits.execute-phase` (the workflow this PR edits) and `aggregate`, which is the
sum over the splits and therefore moves whenever any split does. Zero foreign keys moved.

The full generator sweep was re-run after the replay -- gen-features, gen-inventory-manifest,
gen-platform-conformance-tier (both targets), gen-install-tree-fixtures and the benchmark --
and this fixture is the only artifact that moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V5xnzzipRHPHq8Z3doQvn3

* chore(#3829): regenerate the compact-content benchmark baseline after the rebase onto b956bb7c6

`next` moved again while this round was running -- #4902 landed at 06:27Z and touches this same
generated fixture -- so the branch went back to CONFLICTING within minutes of the previous push.
This is the second rebase of the round, not a correction of the first.

Resolved arbitrarily during the replay and regenerated with its own producer
(`scripts/benchmark-compact-content.cjs --write`), never hand-merged. Key-level drift against the
new base's committed copy is again exactly `splits.execute-phase` and `aggregate` (the sum over
splits) -- zero foreign keys.

The full generator sweep was re-run after this replay as well; this fixture is the only artifact
that moved. Build inputs were untouched by the base range, so the lane's existing build stands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V5xnzzipRHPHq8Z3doQvn3

* fix(#3829): refresh the launcher preamble in place, without hoisting it above the guard

Found by this round's own post-rebase validity gate, not by review: a reader flip-test over the 85
tests that read the two at-risk workflow files was clean before the replay and failed after it, on
`runtime-launcher-parity (#373)` invariant (B).

The cause is a real interaction. #4902 landed on `next` mid-round and rewrote the canonical launcher
preamble; invariant (B) counts occurrences of that exact snippet, so this step's older copy matched
zero times even though it sat in the right place. The remedy the invariant names --
`node scripts/sync-runtime-launcher.cjs` -- fixes the count but also HOISTS the preamble to the top
of the block, and that is wrong here: it moved the shim ahead of the status guard, and six of this
PR's own tests exist to pin that ordering (`runDispositionGuard` asserts the block opens with its
guard, then the shim). Running the tool verbatim turned one red into seven.

So the preamble text is refreshed to the current canonical snippet IN PLACE, at the offset it
already occupied. Both constraints hold at once, verified by execution rather than by reading:
invariant (B) sees exactly one canonical occurrence and it precedes the first `gsd_run` call, while
the guard still opens the block (shim at offset 19897 of the second fence, and the test wants > 0).
`runtime-launcher-parity` + `code-review-pipeline-regression` together: 282 tests, 281 pass.

Not fixed here, and not ours: `(K2) end-to-end: the resolved local tool honors
git.allow_default_branch_commits (#4834)` -- the one remaining failure -- fails identically on a
detached worktree at pristine `b956bb7c6` carrying none of this PR's content (37 tests, 36 pass,
same single failure). Reported rather than chased.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V5xnzzipRHPHq8Z3doQvn3

* fix(#3829): record the shortfall when NO finding in a review parses

Found by this round's own pre-push adversarial review, not by the maintainer -- and it is the
same defect class round 2 raised as Blocker 4, surviving in the one corner that round's fix did
not reach.

The unparsed reconciliation exists so a finding the CR|BL|WR|IN heading parser cannot match is
SURFACED rather than dropped. It reported faithfully whenever SOME findings parsed. It reported
nothing at all when NONE did, on a phase with no prior ledger and no fix report: a REVIEW.md
declaring two Criticals, both written under a prefix the alternation does not carry, produced no
ledger, no console line and no diagnostic. That is precisely the silent drop this reconciliation
was added to close, reachable exactly where the evidence is weakest -- the run in which not one
finding was understood.

Two exits discarded it, and the first one is the one that actually fired. The shortfall was
derived beside the render, while the exit that stands down for "nothing to record" keys on
order.length and sits ~200 lines earlier; it returned before the value existed. The later
rows.length exit had the same hole but was unreachable for this input. So the derivation moves
above the earlier exit -- order is final from the heading walk and never grows again, so the value
is unchanged -- and both exits now decline to fire while a shortfall is outstanding. The result is
a zero-row ledger carrying an unparsed key: an honest record that the review declared findings and
none of them were understood, which is strictly better than the file not existing.

Scoped, not removed. A genuinely clean review is untouched: a declared total of 0 is not greater
than order.length, so unparsed is 0 and both returns still fire exactly as before. The new test's
companion pins that, and it is why the exit was relaxed conditionally rather than deleted.

Driven at every step rather than reasoned about. Before the fix, three cases through the shipped
script: two unmatched findings on a first run wrote NOTHING; two unmatched plus one matched
reported unparsed: 1; two unmatched against an existing ledger reported unparsed: 2. Only the
first was silent, which is why the mechanism read as covered. After the fix the first renders a
ledger with unparsed: 2 and names it on the console, and the other two are byte-unchanged.

The new regression test was negative-controlled against the pre-fix step file and goes red there
(1 pass / 1 fail over the pair); post-fix both pass. Its companion clean-review control is green
on both sides, so the pair is not passing by accident.

The PR's six test files: 554 tests, 554 pass, 0 fail, 0 skipped. lint:ci exits 0. The generator
sweep still produces no drift.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): pin the SECOND exit that discarded the shortfall, and prove the pair kills it

Self-found by the round review of the previous commit, not by the maintainer: that commit added
the !unparsed conjunct to BOTH exits but tested only one of them. Deleting the later one left
every new test green -- a surviving mutant, which is coverage in name only.

The two exits are reached by different inputs, which is why one fixture cannot pin both. The
earlier exit stands down the moment a fix report exists, so an input carrying one sails past it
and lands on the later rows.length return. The fixture therefore needs a fix report that
contributes NO row: an id the alternation CAN match becomes a carried row, rows.length is 1, and
the later guard never decides. The first draft of this test used CR-99 and was vacuous for exactly
that reason -- it passed with the guard deleted. It names SEC-03 now.

Mutation-controlled in all three directions, since a test that kills no mutant pins nothing:

  earlier exit loses !unparsed   -> test 1 RED,  test 2 green, test 3 green
  later exit loses !unparsed     -> test 1 RED,  test 2 RED,   test 3 green
  both exits neutered (over-fire)-> test 1 green,test 2 green, test 3 RED
  unmutated                      -> all three green

Every test kills at least one mutant and no mutant survives all three, so the pair covers both
roads to the drop and the clean-review control covers the over-fire the relaxation could have
introduced.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): close the fourth cell — the later exit's own over-fire

Self-found again by the round review, which drove the mutant set rather than trusting the matrix
the previous commit asserted: neutering ONLY the later exit survived all three tests. The previous
commit's matrix was accurate and incomplete, which is the more dangerous shape -- it reads as a
closed argument.

The guards form a 2x2 and only three cells were pinned. T1 pins the earlier exit's under-fire, T2
the later exit's under-fire, T3 the earlier exit's over-fire. Nothing pinned the LATER exit's own
over-fire, and it is reachable: a fix report -- even one naming no matchable id -- makes the
earlier exit stand down, so control arrives at the later exit carrying a genuinely clean review and
no shortfall. Neuter that exit and a phase with nothing to report grows a zero-row ledger reading
"0 of 0 finding(s) open", with all three earlier tests green through it. T4 is that case.

Five mutants driven over the four tests:

  earlier exit loses !unparsed        -> T1 RED
  later exit loses !unparsed          -> T1 RED, T2 RED
  both exits -> if(false)             -> T3 RED, T4 RED
  ONLY later exit -> if(false)        -> T4 RED          (the survivor this commit kills)
  ONLY earlier exit -> if(false)      -> all four green

The last one is reported as an EQUIVALENT mutant rather than an open cell, and the distinction is
the point of stating it: the earlier exit's over-fire condition is a strict subset of the later
exit's -- it adds only fixReports.length === 0 -- so removing it is subsumed and changes no
observable behaviour. A test cannot kill a mutant that does not alter output, and pretending
otherwise would mean writing one that asserts on internals.

The PR's six test files: 556 tests, 556 pass, 0 fail, 0 skipped. lint:ci exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* fix(#3829): keep the embedded record-builder inside a Windows command line

Block 2 runs the disposition record-builder as `node -e "<script>"`, so the entire
script is a single argv entry. Windows caps a command line at 32767 characters
(CreateProcess) and Node surfaces the overflow as ENAMETOOLONG from spawn -- the
process never starts. Linux's ~2 MB ARG_MAX cannot see that cliff at all.

The previous fix (bd83bc88e, hoisting the shortfall derivation above the earlier
exit) grew the extracted script from 31941 to 33353 characters. There were 826
characters of headroom; it spent 1412. On the next CI run ubuntu and macOS stayed
green and `conformance test (windows-latest, 24, shard 1/3)` went red with 83
failures -- 79 reporting `spawn_failed` out of runShippedDisposition, the other 4
asserting on a ledger that was never written. That same shard was SUCCESS at
596968aa0, the head before that commit.

Measured on native Windows (node v25.2.1), bisected: the largest `-e` argument that
still spawns is 32728 characters; 32729 fails. Both payloads driven directly:

    pre-fix   33353 chars  ->  SPAWN-FAIL ENAMETOOLONG
    post-fix  15916 chars  ->  SPAWN-OK

Note the shape of it: the fix that makes this gate report a silently-dropped finding
was itself silently dropped on Windows, because the whole script stopped launching.

What changes here is placement -- not content, not behaviour. 17 long rationale
comment blocks move out of the quoted payload into a new "Design notes for the
embedded record-builder" section in this file's prose, each anchored to the code line
it preceded so the pairing survives the move. Comment runs shorter than five lines
stay inline, where adjacency is cheap. Nothing is deleted.

    extracted script  33353 -> 15916 chars (16812 under the measured limit)
    executable code   byte-identical at 10853 bytes, verified by diffing the payload
                      with all comment lines stripped from both sides
    this file         78542 -> 79538 bytes -- the prose moved, it did not grow

Verified green: this PR's own test file at 250 tests across all 36 suites, 0 fail;
changeset-lint, docs-lint, default-flip-documentation, lint:ci, gen-emitted-baseline,
workflow-size-budget (133 tests), and lint-workflow-shellcheck (204 pre-existing
findings, 0 new).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): pin the embedded script under the Windows command-line budget

Nothing guarded the size of the `node -e` payload, so the regression the previous
commit fixes was invisible to every Linux gate and surfaced only as 83 Windows
failures that said `spawn_failed` and never mentioned length. Without a guard the
next addition to the script re-breaks Windows exactly the same way, and finds out
the same expensive way.

The test asserts the extracted script stays under a 24 KiB budget -- the 32767
CreateProcess cap less roughly 8 KiB of deliberate headroom, so the script has
somewhere to grow before this fires.

It is bounded from BELOW as well, and that half is the point: a pure length
assertion passes when the extractor returns '', which is exactly what a moved fence
or a renamed delimiter would produce. A guard that reports a comfortable 0 bytes is
the vacuous-oracle shape. The lower bound makes a broken extractor fail loudly here
instead of reporting success.

Negative-controlled rather than assumed. Against the PRE-fix step file the test goes
red on the real payload (33353 > 24576); against the fixed file it passes at 15916.
A new test that has only ever been run against fixed code can be green because it
hit a branch the bug never lived on.

This PR's own test file: 250 tests, 250 pass, 0 fail, 0 skipped, across all 36
suites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): say what the command-line guard does not prove

The round review's MISSED, adopted. The guard counts the extracted JavaScript, but the
32767 cap applies to the whole serialized command line -- executable path, quoting and
backslash escaping included -- so a quote-heavy payload expands on the way out, and the
32728 figure it cites is one host's measured threshold rather than the CI runner's.

The budget is unchanged and still correct; only the claim around it moves. 24576 leaves
roughly 8 KiB for both effects, which is a practical margin, not a proof that everything
the guard admits will spawn. Stating that in the test is cheaper than having a future
reader infer a guarantee the assertion cannot make.

Comment-only. No assertion, no budget and no behaviour changes.

This PR's own test file: 221 tests, 0 fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-22 21:04:09 -04:00
Tom Boucher
55fba5f7ce feat(#4917): add the PlanningDoc parse → mutate → serialize seam — Phase 1 of #4906 (#4918)
* feat(#4917): add the PlanningDoc parse -> mutate -> serialize seam

Phase 1 of epic #4906, implementing ADR-4910 and its 2026-09-21 amendment. Net-new
leaf module; NO call site is migrated, so nothing in the twelve absorbed issues
changes behavior yet.

src/planning-document.cts composes the seams that already exist rather than
reimplementing them: markdown-sectionizer for structure (fences and code spans
come from stripFencedCode / scanInlineCodeSpans, never a second scanner),
markdown-table for tables, frontmatter for frontmatter, write-set for Result<T>.

What is structural rather than conventional:

- A field node carries labelSpan, valueSpan and trailingSpan separately, and the
  only write entry point takes a node id and writes into valueSpan. trailingSpan
  is readable and has no exported writer, so the #2853/#3584/#4852 rule ("the verb
  owns the count token ONLY") stops depending on an author remembering a third
  capture group.
- Mutation is node-addressed. A handle is minted by the parser, so a caller cannot
  name a node the parser did not find. No path strings — a path is a grammar, and a
  grammar needs a parser.
- serialize splices staged spans into the ORIGINAL buffer. A no-edit serialize is
  byte-identical, which is what eliminates the #4499 defect without a targeted fix, and which also
  makes an already-escaped table cell impossible to double-escape on a round trip.
- A node that fails to parse carries its own error and span; siblings stay readable.
- Per the ADR amendment, serialize REFUSES whenever any node carries a parse error,
  even with zero staged edits, naming the offender.

One defect was found while building and fixed in place rather than deferred:
PLANNING_ARTIFACTS first derived from isCanonicalPlanningFile unfiltered, so
config.json, state.json, milestone.lock and skill-manifest.json were accepted and
returned {ok:true, nodes:[]} — "this document records nothing" when the truth was
"I have no grammar for this file". That is the exact empty-vs-error confusion this
epic exists to remove (#4899, #4900), reproduced inside the seam built to prevent
it. The registry now filters to .md while still DERIVING from
isCanonicalPlanningFile, because hand-writing a second list is the divergence this
epic is about, and parsePlanningDoc refuses a non-markdown kind at the document
level per ADR-4910 section 5.

New bin/lib module bookkeeping: .gitignore, eslint.config.mjs ignore (ADR-457 —
lint the .cts, not the emitted .cjs), docs/INVENTORY.md row, regenerated
docs/INVENTORY-MANIFEST.json, and the CONTEXT.md glossary entry.

Refs #4906

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4917): cover the PlanningDoc seam across 28 input classes

30 cases in 14 describe blocks, one per row of the phase test matrix.

The two load-bearing tests are the fast-check properties (seed 20260921,
numRuns 200):

- a single node mutation leaves every byte outside that node's valueSpan
  identical to the source
- serialize with zero staged edits is the identity function

Both are DOCUMENT-SHAPED per CONTRIBUTING.md fixture provenance (#2371): the
generator assembles arbitrary frontmatter, heading, label and value text with
join(), and never calls serialize or any other function from the module under
test to build a fixture. Seeding the generator from the module's own writer
would make the document shape a constant, and the property could then never
explore a document the writer would not itself emit.

Verified the byte-range assertion is not vacuous with a control run: it passes
against the real writer and FAILS against a simulated #4852 writer (one capture
group, replace-to-end-of-line), which visibly drops the trailing annotation.

Boundary coverage is zero / one / two staged edits. Negative space carries its
own rows — bold emphasis in prose, a field-shaped line inside a fenced block,
the same inside an inline code span, and a horizontal rule mid-body all
correctly mint no node. Row 24 is the Generative-Fix-Divergence parity
assertion: PLANNING_ARTIFACTS must not diverge from isCanonicalPlanningFile.
Row 28 covers the non-markdown canonical file found during the build.

Assertions are structural throughout — a typed Result / NodeRead /
SerializeOutcome shape, or a byte range computed from the node's own Span.
Full-string equality appears only where the contract IS byte equality.

Refs #4906

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4917): apply review findings — refuse unrepresentable values, compose the layers the seam claimed

Three review passes ran against this branch: /security-review, an isolated
adversarial pass, and a standards+spec pass. Five findings, all fixed here. Every
one is the epic's own failure class reproduced inside the seam built to end it,
which is the thing worth noticing.

1. setFieldValue accepted a value containing a newline. It survived serialization
   and reparsed as a REAL sibling field — one write to Plans forged a second
   Owner into a document that already had one. Content became structure, which
   defeats ADR-4910 Decision 2's "cannot reach past its own token by
   construction": the token boundary is a LINE boundary.

2. setFieldValue accepted a value containing the trailing separator and SILENTLY
   TRUNCATED it. Staged "sneaky - annotation", read back "sneaky", with the
   remainder reclassified as trailing prose. No error, nothing unreadable, both
   resulting nodes parsing perfectly. Worse than (1) because it loses the
   caller's own value rather than adding something visible.

   Both are fixed by ONE general check, deliberately not a blacklist:
   setFieldValue rebuilds the candidate line, re-parses it through the same field
   grammar, and refuses unless the value reads back identical. Blacklisting the
   separator would close this instance and leave the class open for whatever
   separator the grammar grows next. The round-trip check is ADR-4910 Decision 4
   stated executably.

3. The module reimplemented two layers it claims to compose. Checklist detection
   hand-rolled a checkbox regex that markdown-sectionizer's iterateBullets
   already owns. Frontmatter span detection re-derived fence handling because
   frontmatter.cts's frontmatterRegion was module-private — so ADR-4910 section 1's
   stated layering was UNREACHABLE as written, and the first implementation
   routed around it silently instead of surfacing the gap.

   frontmatterRegion is now exported (additive only; ADR-2143 section 2's
   extend-never-mutate lock is inherited) and both layers are consumed.

4. The CONTEXT.md glossary entry asserted "the Frontmatter Module supplies
   frontmatter" while zero frontmatter.cts code was invoked. That was a false
   claim in the repo's vocabulary of record, written by me, and it is now true
   rather than edited away.

5. Adopting iterateBullets narrowed GFM coverage: it classifies only
   dash-prefixed task items as checkboxes, so "* [ ] x" and "+ [x] y" stopped
   becoming checklist nodes. Widening the sectionizer is forbidden by the
   inherited lock, so the task-list MARKER is interpreted in this module while
   bullet STRUCTURE still comes from the sectionizer.

The sharpest finding was not a defect. The fast-check generator constrained
values to [A-Za-z0-9 .,!?], so it could not emit an em-dash, newline, backtick,
pipe or asterisk — precisely where (1) and (2) lived. The property was real,
seeded and non-vacuous, and structurally blind to the module's actual bug class.
The generator now spans the grammar's own metacharacters, and a new property
asserts that every value setFieldValue ACCEPTS round-trips identically. Proven
able to fail: against a scratch copy with the guard stripped it fails after 7
cases on newValue "\n".

Recorded as a measured boundary, not fixed: a bare CR inside a field line leaves
that field unrecognised. Measured — sibling fields still parse, no error node,
and serialize stays byte-identical, so the worst case is an unreadable field and
never a damaged document. Flagged for Phase 4's empty-vs-error census.

Refs #4906

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4917): backfill changeset pr number to 4918

Refs #4906

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:17:58 -04:00
Tom Boucher
c9a5cc3e12 fix(#4683): detect cross-plan threat-ID duplicates before execution (#4828)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; two orthogonal agent reviews ran (isolated adversarial REQUEST-CHANGES with all six findings dispositioned, plus a bypass/consumer-lens APPROVE) and the sha-pinned bench passed 46362/0 on the merged head.
2026-09-17 19:06:55 -04:00
Tom Boucher
2bfff17ff8 fix(#4682): route stale verification to the verifier regeneration path (#4818)
* test(#4682): add failing-first coverage for stale verification routing

* fix(#4682): route stale verification to the verifier regeneration path

The stale routing entry sent users to /gsd-verify-work — but verify-work
never rewrites VERIFICATION.md (its only write is the human_needed
canonicalization), so following the advice re-ran UAT, reached the same
stale check, and looped. init's projector and execute-phase's generic
next_command presentation both mirror this entry, so the dead end appeared
on three surfaces.

The stale entry now routes to execute-phase, and execute-phase's
all-plans-complete resume tree gains a stale arm (as a steps/ part, keeping
the spine under its frozen ADR-857 ceiling) mirroring the missing route:
skip cross_ai_delegation/execute_waves/checkpoint_handling, continue at
aggregate_results, and let verify_phase_goal re-dispatch the gsd-verifier —
regenerating VERIFICATION.md and its digest, marked phase or not. The
non-stale fall-through. Staleness detection, the digest format (#4623),
every other routing entry, and the #3684 resume arms are untouched.

Emitted-Drift-Ack-Growth: verify-work.md — stale stop rewritten to dispatch the verifier and re-check (#4682)
Emitted-Drift-Ack-Growth: execute-phase.md — VERIFY_STATUS == stale resume arm added to condition 3 (#4682)

* test(#4682): register the stale-reverification part and align projected commands

The new steps/ part must be registered in the inventory manifest and the
per-runtime golden install trees (regen:derived); the projected stale
next_command is /gsd-execute-phase <phase> (formatGsdSlash prefixes the
runtime surface), the human_needed bare-report probe keeps routing to
verify-work (unchanged semantics), and init-manager's recommended action
follows the new command.

* test(#4682): prefix the remaining stale routing assertions with the runtime surface

Nine stale next_command assertions and the human_needed bare-report probe
still carried the unprefixed or flipped forms from the earlier line-number
edit; all now assert the shipped /gsd-execute-phase <phase> projection,
with the human_needed probe reverted to its unchanged verify-work routing.

* test(#4682): align the last stale projection assertions with the execute-phase route

* docs(#4682): backfill changeset PR number

* test(#4682): refresh the compact-content baseline after the rebase

The rebase onto the #4670 squash brought verify-work.md's bounded
reconciliation text into this branch; the committed compact-content
baseline now reflects the post-rebase split sizes. Local --check is
clean; the previous bench drift (+243) was the baseline, not the diff.

* fix(#4682): carry the response_language directive in the stale-reverification part

The new steps/ part is its own coverage unit for lint-response-language-coverage;
it takes the shared canonical directive line like its sibling execute-phase
parts.

---------

Co-authored-by: sim <sim@local>
2026-09-17 02:23:03 -04:00
Tom Boucher
740ba0d8a3 fix(#4628): expose DAG-ready plans and restrict dispatch to them (#4781)
Emitted-Drift-Ack-Growth: execute-phase.md — #4628 consumer wiring: ready_plans parse pointer, not-ready named skip, and waiting condition 2b reference to the ready-wave-gate step file

Co-authored-by: sim <sim@local>
2026-09-16 02:38:43 -04:00
BeeHiggs
0967358b8b enhance(#3638): render bracket phase IDs on progress, stats, manager and statusline surfaces (epic #612 PR-5) (#4111)
* enhance(#3638): render bracket IDs on display surfaces

Gate progress, stats, manager, and statusline projections on the bracket convention; validate phase_id_convention and single-source the convention card.

Forward note: the uat.cts bracket co-change remains deliberately deferred to its owning slice.

* chore(#3638): point the changeset at PR #4111

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3638): close bracket display review gaps

* docs(#3638): register phase display modules

* chore(#3638): re-trigger CI after macOS shard SIGTERM

`full test (macos-latest, 24, shard 3/3)` failed on 20ce98cd1 in
`tests/lint-compiled-artifact-sync.test.cjs` — the spawned
`scripts/lint-compiled-artifact-sync.cjs` was killed at 60024ms
(`exited null (signal SIGTERM)`, stdout and stderr both empty), 24ms past
the test's own `TSC_COMPILE_TIMEOUT_MS`. That is the failure mode the
constant's comment already documents ("under CI shard load that compile
can exceed the budget, dying to a SIGTERM with empty piped stdout").

No content change; this empty commit exists only to re-run the matrix,
since re-running a job needs write access on the upstream repository.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 04:47:47 -04:00
Tom Boucher
37b965c0d1 enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code (#4553)
* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code

ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills
(src/init.cts) now selects between a canonical agents/<name>.md and a
token-minimized agents/<name>.compact.md sibling based on
workflow.compact_content, resolved in code (a real function call with a real
exit code) rather than a prose config-get gate — the same precedent stream 1's
spine/detail split established for a load-bearing seam, applied here because
this seam already runs through TypeScript instead of an eager @-include.

A missing compact sibling falls back to the canonical persona and discloses
the fallback in the served payload itself (a leading HTML-comment provenance
line), so the Done-when contract — compact when on, canonical when off, never
silent or empty — holds even for an agent nobody has compacted yet.

Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md),
each an independent, complete rewrite (not an extraction — nothing is "moved"
the way spine/detail moves text) that preserves frontmatter, every @-include,
every output-format contract, and every guardrail verbatim while cutting
restatement and verbose framing. Verified mechanically: every pair registers
(a canonical sibling exists), every compact file is strictly smaller, and the
full @-include set matches canonical's — including which references are
standalone eager-load lines versus inline prose mentions, since demoting one
to inline changes what the host actually substitutes.

Traced the install path before writing any code (.gsd/phase/.../40-design.md):
stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no
stem filtering under the default full profile, so the new .compact.md files
install for free with zero installer changes — matching issue #4407's stated
scope. A tiered agent profile that doesn't stage a compact sibling degrades
through the same fallback-with-provenance path already required for an
unauthored one, so no installer change is needed there either.

Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export
(deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are
reached by a generic code construction rather than a literal path in prose,
and checkReachability's markdown-search shape has nothing to find there).
Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new
"#4407 compact payload selection" describe block spawns gsd_run agent-skills
against real compact/canonical fixture pairs and asserts on the served
payload, which can only pass if the seam genuinely wires through.

Fixed a pre-existing test whose agents/*.md glob incidentally matched the new
.compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift
guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's
roster, both real, unrelated-to-content defects the new files' mere existence
surfaced.

Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files
under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token
benchmark baseline (npm run benchmark:compact-content-variants --write).

Closes #4407.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): apply orthogonal review findings from the compact-payload seam

Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath)
to collapse the duplicated read-and-empty-check shape between the compact
and canonical branches in cmdAgentSkills, and updated the adjacent comment
enumerating flat JSON extras to name agent_payload_variant alongside
source/degraded (added by the prior commit, comment left stale).

Security review and the Spec axis found no defects requiring a code change;
their non-blocking observations (a pre-existing, unmodified path-construction
pattern; the reasoned, documented substitution of a behavioral test for the
literal reachability check) are recorded in
.gsd/phase/enhance-4407-agent-skill-seam/60-review.json.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents

Root-caused via a real gsd-test run (93 failures) rather than guessing which
tests glob agents/ naively. Two classes of defect, both genuine:

1. Identity-roster confusion (11 files/areas): many tests and one production
   script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f
   => f.endsWith('.md'))`, which incidentally matched the new .compact.md
   variant siblings too — a compact file is a rendering of an EXISTING agent
   identity, not a new one. Fixed at the shared root
   (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests
   already consolidated on) and at each independent glob that didn't use it:
   agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix
   before checking XL/LARGE membership, so a compact file inherits its
   canonical sibling's tier instead of silently falling through to DEFAULT),
   agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual
   script, not just its test), codex-config.test.cjs (confirmed directly
   against generateCodexAgentToml that a compact role's derived sandbox_mode
   is byte-identical to its canonical sibling's before excluding it — not
   assumed), and copilot-install.test.cjs (two counts that legitimately DO
   need both files — an installed-file count and a full-conversion smoke test
   — fixed to expect 70, not stay pinned to 35).

   no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix:
   two compact files reproduce descriptive prose already allowlisted at their
   canonical file's line number; added matching entries at the compact files'
   own line numbers rather than excluding them from the scan (a genuine bare
   gsd-tools command-position bug in a compact file would be as real a defect
   as in canonical).

2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree
   run): six agents' compact renditions (gsd-debugger, gsd-executor,
   gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed
   the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction —
   confirmed structural, not a compaction-quality gap: each is dominated by
   content this phase's own rules require verbatim (the ~2.6 KB gsd_run
   bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in
   every agent that calls gsd_run, output-format contracts, guardrails).
   ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing
   spot in cmdAgentSkills's single-file synchronous read. Removed these 6
   compact files rather than ship an over-cap file or invent a multi-part
   read mechanism out of scope for this phase; recorded by name with the
   reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and
   50-test-matrix.md, per #4407's own "or explicitly recorded as not worth
   covering" allowance. Their canonical personas are served correctly today
   via the fallback-with-disclosed-provenance path this phase's own Done-when
   #2 already requires — 29 of 35 agents now have a compact variant.

Also fixes an unrelated, genuinely pre-existing defect this gsd-test run
surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own
DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in
this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md
| wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed:
two meaning-preserving trims in the <success_criteria> block (a repeated
parenthetical replaced with a same-exception reference; one redundant
qualifier dropped) bring it to 40,940 bytes.

Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant
benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6
now-orphaned roster rows removed alongside them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): make .compact.md-aware roster checks resilient to partial coverage

Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent
has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception),
breaking once 6 stems legitimately have none.

- tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was
  picking up the "### Compact Payload Variants" subsection's rows as
  phantom/uncounted entries in the primary/advanced/inventory-only
  classification this test validates — a compact row documents an existing
  agent's alternate rendition and never gets its own AGENTS.md heading, so it
  was never meant to participate in that classification. Excluded at the
  parser, not per-assertion.
- tests/copilot-install.test.cjs: the derived expected-file-list generator
  assumed every listAgentFiles() stem has a .compact.md source sibling;
  checks disk per stem now instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4407): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 12:38:59 -04:00
Tom Boucher
a27cb6b2fa enhance(#4139): Phase 6 — the lazily-read remainder and the artifact templates (#4540)
* enhance(#4406): the lazily-read remainder and the artifact templates

ADR-4139 Decision 3, Phase 6 of the #4139 Compact Content epic. Covers stream 1b
(gsd-core/workflows/<name>/{modes,steps,templates}/*.md) and stream 4
(gsd-core/templates/**) with a variant-swap mechanism, confirmed with the user:
two independent, complete files per covered path (canonical + .compact.md
sibling), with the gate picking which one gets Read at the call site. This is
a different shape from Phase 5's spine+detail partition, and is safe here
specifically because these files are already reached only by a runtime Read —
a missed Read already means zero overlay content today, with or without
workflow.compact_content, so selecting between two independently-complete
files introduces no new failure mode (documented in
gsd-core/references/compact-content-gate.md's new "Streams 1b and 4" section).

Disposition, after inspecting every candidate rather than trusting a byte-size
threshold (same rigor Phase 5 applied to review.md):

- Stream 1b: 1 of 78 files compacted (help/modes/full.md, a user-facing
  reference doc emitted verbatim, not orchestrator instruction). The other 9
  size-threshold candidates are dominated by fail-closed guards, exact CLI
  invocations, or output-format contracts (AskUserQuestion blocks) — recorded
  not-worth-compacting, same reasoning as Phase 5's review.md.
- Stream 4: a ground-truth reachability audit replaced the initial size-only
  candidate list. Two files (summary.md, user-setup.md) got compact variants;
  a third (spec.md) was drafted, then dropped after discovering its only two
  call sites are eager @-includes, not a runtime Read — stream-1 material
  hiding under gsd-core/templates/, not stream-4's actual mechanism. summary.md
  itself has 3 eager call sites and only 1 genuine runtime-Read call site
  (execute-plan.md); only that one was wired, so the compact variant's savings
  apply to the sequential single-plan execution path only.
- Discovered while auditing reachability: 12 gsd-core/templates/** files with
  zero references anywhere in workflow/agent/command prose, compiled source,
  or tests — dead scaffolding predating this phase. Deleted in this same PR
  per this repo's no-defer policy, after re-verifying against a computed
  path.join(...) pattern (not just a plain-string search) that nearly caused
  two genuinely load-bearing templates (user-profile.md, dev-preferences.md)
  to be misclassified as dead.

New checker (tests/helpers/compact-content-variant.cjs): registration,
reachability, protected-content-preserved, size-smaller — replacing Phase
3/5's disjointness/completeness checks, which assume a partition rather than
two deliberately-overlapping documents. The reachability check's own
"unprefixed match" guard had a real bug (rejected the repo's own
`~/.claude/gsd-core/...` convention), caught by running it against the
already-wired help/modes/full.compact.md pair rather than only synthetic
fixtures — fixed to anchor on the nearest `gsd-core` path segment instead.

Template consumer parity (tests/compact-content-template-variant-parity.test.cjs):
proves each compact variant's `## File Template` fenced block — the actual
output-format contract a generated SUMMARY.md/USER-SETUP.md is parsed
against — is byte-identical to the canonical file, then runs the one real
deterministic consumer (gsd-core/bin/lib/coverage.cjs's classifyContent,
backing `gsd-tools uat classify-coverage`) against content built from that
shared contract.

Added a sibling benchmark script (scripts/benchmark-compact-content-variants.cjs)
rather than extending the existing spine/detail one — different data shape,
and the existing script's own contract deliberately isolates it from a
test-only helper's shape changing.

Emitted-drift acknowledgement: not needed. Every changed/added path in this
diff is hand-authored and present in the diff itself, so diffEmitted's
attribution loop resolves `via` to the path's own source before reaching the
ack-lookup branch (same reasoning Phase 5 verified for its own diff).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* enhance(#4406): address code-review findings on the variant-swap gate

- docs/CONFIGURATION.md and gsd-core/references/planning-config.md's
  workflow.compact_content entries described only the spine+detail mechanism
  (Phase 5) and were missing this phase's variant-swap mechanism and its
  benchmark:compact-content-variants script entirely — required since this
  PR's changeset is type Added (CLAUDE.md's "Missing Docs for Changesets"
  rule). Both now describe both mechanisms and which call sites are wired.
- Added the missing RED^-1/no-op fixture for checkProtectedContentPreserved:
  a canonical file with zero <!-- gsd:protected --> blocks must be a
  no-op, not a violation — the only branch of that function the existing
  fixtures didn't exercise.
- Collapsed findCompactFiles/findMarkdownFiles in
  tests/helpers/compact-content-variant.cjs into one findFilesWithSuffix
  helper — the two were identical recursive walks differing only in the
  extension predicate (minor Duplicated-Code finding).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4406): restore copilot-instructions.md, a false-positive dead-template classification

gsd-test caught this, not static analysis: 10 real failures in
tests/copilot-install.test.cjs, tests/installer-migration-install.integration.test.cjs,
and tests/repo-layout.test.cjs — all downstream of bin/install.js's Copilot install
path, which does
fs.readFileSync(path.join(targetDir, 'gsd-core', 'templates', 'copilot-instructions.md'))
after copying gsd-core/templates/** into the target project, then merges it into both
.github/copilot-instructions.md and (local installs) AGENTS.md. The reachability audit
that flagged this file as dead checked src/*.cts and gsd-core/bin/*.cjs but never the
repo-root bin/install.js — a separately maintained installer bundle outside the
src/-to-gsd-core/bin/lib/ compiled-output convention. The fs.existsSync guard around
that read degrades to a silent skip rather than a crash when the template is missing,
which is why this surfaced only once the real E2E install test ran, not from any
static check.

Re-verified the remaining 11 deleted filenames against bin/install.js specifically
(plain substring and quoted-filename search) before trusting that list — all 11 have
zero hits there, confirmed dead by the same standard this one file failed.

Regenerated the installer emitted-tree goldens (tests/fixtures/install-tree/*.json) to
reflect the restored file, and corrected the "Removed" changeset (jolly-lynx-sprint.md)
and the phase design doc from 12 to 11 deleted files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: execute-plan.md — call-site wiring for the summary.md and user-setup.md .compact.md variants
Emitted-Drift-Ack-Growth: help.md — call-site wiring for full.compact.md, same variant-resolution rule

* docs(#4406): backfill changeset PR numbers

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4406): resolve removed-but-needed lint findings on the dead-template deletion

CI's own full-test matrix (not gsd-test's matrix, which does not run this
check) caught 4 more false-positive dead-template classifications via
tests/removed-but-needed-lint.test.cjs / scripts/lint-removed-but-needed.cjs
— a literal, word-boundary basename check across .github/workflows/,
gsd-core/, and docs/ (excluding docs/adr/** and docs/research/**) for every
file a PR deletes. It has no semantic awareness, so a deleted template's
basename colliding with something else entirely still fires:

- claude-md.md: gsd-core/templates/README.md had a stale table row claiming
  /gsd-profile reads this template to generate CLAUDE.md. Verified false (no
  code reads it anywhere, same search that already covered bin/install.js) —
  fixed the row to *(inline)*, matching every other command-generated
  artifact in that table. File stays deleted.
- codebase/testing.md: collided with docs/guides/testing.md, an illustrative
  example row in docs-update.md's sample output table (an unrelated real
  generated-docs path). Swapped the example topic to "contributing" — the
  row is illustrative, any topic works. File stays deleted.
- codebase/architecture.md, codebase/stack.md: collided with docs/reference/
  planning-artifacts.md's directory listing of a user's own generated
  .planning/codebase/architecture.md and stack.md output — the same
  semantic mismatch already investigated and dismissed as unrelated earlier
  in this phase's audit, now caught by a gate instead of judgment. That
  listing repeats across 5 locale copies of the doc.
- continue-here.md: collided with the real .continue-here.md pause-work
  artifact, referenced across 15+ locale and workflow files.

For the last two, the lint's own error message offers "restore the file or
update every consumer in the same commit." Rewording 15+ files across
languages I cannot verify translation quality for, to shave 2 already-tiny
templates that were merely presumed dead, is disproportionate to this PR's
actual scope — restored codebase/architecture.md, codebase/stack.md, and
continue-here.md instead, and corrected docs/ARCHITECTURE.md's Templates
section accordingly.

Final confirmed-dead set: claude-md.md, codebase/concerns.md,
codebase/conventions.md, codebase/integrations.md, codebase/structure.md,
codebase/testing.md, debug-subagent-prompt.md, discovery.md — 8 files, down
from the original 12. Verified locally: GSD_REMOVED_BUT_NEEDED_BASE=next
node scripts/lint-removed-but-needed.cjs now passes clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: docs-update.md — swapped an illustrative example-table topic (testing -> contributing) to avoid a removed-but-needed basename collision with the deleted codebase/testing.md template; net +10 bytes

* fix(#4406): split codex-config.test.cjs to fix a genuine Windows CI timeout

Root cause of the `full test (windows-latest, 24, shard 2/3)` failure the
user asked to be actually fixed, not just re-run past: PR #4497 (landed
2026-09-07, one day before this PR's CI run) isolated
tests/codex-config.test.cjs into its own dedicated chunk because its
measured weight (17.87, ~45% of the post-cut Windows budget) made it unsafe
to share a chunk with any other file. That isolation was necessary but not
sufficient — even alone, with zero companion-file contention, the file's
real Windows execution time sits right at the 600s per-chunk ceiling. Two
independent CI runs on two unrelated PRs (this one and #4154) were both
killed within ~1.4s of the identical 600000ms mark — not random contention,
a deterministic near-miss the isolation fix couldn't address because it
never reduced the file's own cost, only removed the risk of a companion
file's cost stacking on top of it (which the PR #4497 comment explicitly
anticipated: "if a future profiling pass genuinely speeds up
codex-config.test.cjs itself, this isolation can be revisited").

The file itself explains why it's this heavy: 11,262 lines / 433 tests / 79
describe blocks, accumulated over dozens of bug-fix PRs (#2695, #2760,
#3245, #3285, #3346, #3426, #3427, #3562, #3566, #3582, #3808, and more),
several of which are explicitly documented as "folded" in from separate
files that were never actually split back out ("Verified non-duplicate
against both the pre-existing target and the other three folded sources").

Split into 4 files by top-level AST statement boundaries (never a naive
column-0 regex — an early attempt at that overcounted 79 apparent
"describe(" matches when only 21 are genuinely top-level; the rest are
nested inside a handful of large folded-in blocks, which a regex can't tell
apart from real top-level statements). Verified lossless twice: the split
script asserts byte-for-byte reconstruction of every source character, and
independently, total test()/describe() call counts match exactly between
the original file and the sum across all 4 new files (433/79 both sides).
Each new file carries the complete original shared header (imports/helpers)
for safety; per-file unused-import warnings from that duplication are
resolved via ESLint-precise alias renames (`{ foo: _foo }`, the standard
form for an intentionally-unused destructured binding — never a bare `{
_foo }`, which would destructure a different, nonexistent property).

No change needed to scripts/run-tests.cjs's ISOLATED_HEAVY_FILES or its
pinned test in tests/run-tests-harness.test.cjs: the file that keeps the
original name (tests/codex-config.test.cjs) is now only ~28% of the
original's size and safely isolated in its own chunk as before; the other
three new files re-enter normal weight-balanced packing, none individually
close to disproportionate. Confirmed no other file hardcodes the hardcoded
filename anywhere that would silently stop these tests from running (the
CI test-selection scripts determine scope algorithmically, not by literal
filename).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 14:31:04 -04:00
Tom Boucher
e03921c7d8 enhance(#4405): split the rest of the eager-window workflows worth splitting (#4536) 2026-09-08 00:17:22 -04:00
Dennis Alexis Valin Dittrich
18c899def5 enhance(#4209): optional external source reviewer lanes for /gsd:code-review (#4323)
* test(01-01): define reviewer-support trait contract

Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02):
validator rejects non-boolean values with an exact field path, accepts
missing/true/false, and the real code-review capability.json steps
must declare supportsReviewerLanes: true. Add loop-resolver projection
coverage proving the trait reaches activeHooks verbatim for a
provider-neutral synthetic step (not code-review-specific), and that
omitted/false values stay inert (no key on the active hook).

All 8 new assertions fail today: the validator has no such field, and
loop-resolver has nothing to project. RED before GREEN.

* feat(01-01): declare reviewer-capable steps

Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional
boolean opt-in trait, step-scoped (not capability-wide). Only a
literal true validates and projects; false/omitted stay inert (no
key on the projected active hook), and every non-boolean type fails
capability-validator.cjs with an exact field-path error.

Opt both existing code-review steps (execute:post, execute:wave:post)
into the trait in capabilities/code-review/capability.json. Project
the validated field through src/loop-resolver.cts into activeHooks
so a provider-neutral generic interpreter can read it without any
code-review-specific knowledge. Document the field in
docs/reference/capability-manifest.md and regenerate
gsd-core/bin/lib/capability-registry.cjs via the generator (never
hand-edited).

Makes all 8 RED assertions from the prior commit pass.

* test(01-02): define shared reviewer dispatch

- Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes:
  inert when the supportsReviewerLanes trait is off or nothing is selected,
  exactly-once plan/invoke per selected lane, duplicate-alias dedup, the
  bounded metadata-only source-review prompt (repo root, paths+baseSha,
  depth, four fixed prohibitions), and capability-neutral reuse via a
  second synthetic step context.
- RED: module under test (src/reviewer-step-dispatch.cts) does not exist
  yet, so require() fails and every assertion is unreached.

* feat(01-02): dispatch reviewers for opted-in steps

- Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps),
  ONE interpreter for a step's supportsReviewerLanes trait. Reuses
  resolveReviewerSelection for selection and resolveLanePlan for planning
  (both already-existing, pure building blocks); invocation is the one
  required, caller-injected seam (deps.invoke) since runLane needs
  OS-aware spawn plumbing this module does not own.
- trait !== true, or a selection resolving to zero lanes, dispatches
  nothing (zero plan/invoke calls). Each selected lane is planned and
  invoked exactly once, in the selector's deduped/sorted order.
- buildSourceReviewPrompt assembles a metadata-only bounded prompt
  (repo root, canonical paths + base SHA, depth, four fixed
  prohibitions) — never file contents — written once per dispatch and
  shared across every invoked lane.
- GREEN: tests/reviewer-step-dispatch.test.cjs now passes.

* test(01-02): define reviewer dispatch failures

- Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed
  matrix: an explicitly requested lane the selector could not resolve
  still lets the OTHER resolved lane run, but the aggregate result must
  never read as a clean success (and 'every explicit lane unavailable'
  must be distinguishable from the plain no-flags-passed inert case);
  request-level validation (path traversal, absolute paths outside
  repoRoot, empty/non-string paths, missing depth/base SHA) halts the
  whole dispatch before any lane is planned or invoked; a per-lane
  prompt-budget overflow hard-fails only that lane before invoke while
  its sibling still runs.
- RED: src/reviewer-step-dispatch.cts does not yet implement any of
  these guards, so 9 of the new assertions fail against the current
  (Task 1) implementation.

* fix(01-02): fail closed in reviewer dispatch

- src/reviewer-step-dispatch.cts: add the fail-closed guards the prior
  commit deliberately left out. An explicitly requested lane the
  selector could not resolve no longer lets the aggregate read as a
  clean success — lanes that DID resolve still run and keep their
  results (never narrow the requested set), but selection.errors now
  flips the aggregate ok to false, and 'every explicit lane
  unavailable' is now distinguishable (SELECTION_FAILED) from the
  plain no-flags-passed inert case (NO_LANES_SELECTED).
- Add request-level validation (validatePaths, depth/baseSha presence)
  that halts the WHOLE dispatch before any lane is planned or invoked:
  path traversal, absolute paths outside repoRoot, empty/non-string
  paths, and missing provenance are all rejected up front.
- Add per-lane prompt-budget enforcement (resolveBudget, mirroring
  gsd-tools.cjs's budgetFor convention including budget 0 = unbounded):
  a lane whose resolved budget the prompt exceeds hard-fails before
  invoke runs for it, without cancelling a sibling lane already
  planned.
- Document the supportsReviewerLanes trait and its dispatch-step
  interpreter in gsd-core/references/loop-hook-dispatch.md.
- GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass;
  no regressions in the review-lane/reviewer-selection/prompt-budget
  suites (356 passing).

* test(01-03): define optional source reviewer flow

RED: assert code-review.md dispatches roster-derived reviewer-lane flags
through a single review-lane dispatch-step call (DISP-01..05), that the
no-flag path stays byte-for-behavior unchanged (COMP-01), and that
external evidence reaching the internal reviewer prompt is marked
unverified (CONS-02). Also covers the CLI contract directly: no-op with
no explicit selection, and fail-closed on an explicit unknown lane
(SAFE-07) via real gsd-tools.cjs subprocess calls.

* feat(01-03): route optional source reviewers

GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches
canonical reviewer-lane flags against the merged first-party + installed
roster (never a hand-maintained list) and, only when at least one is
present, calls the shared reviewer-step interpreter exactly once with the
already-resolved repo root, file scope, depth, and base SHA. Its evidence
paths are appended to the internal reviewer prompt via
${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane
flag leaves the internal-only dispatch byte-for-behavior unchanged
(COMP-01).

Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane
dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI
route `dispatchReviewerLanes` wires through, but never implemented the
gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add
it to the existing review-lane router, reusing the same effort-aware plan
building and runner deps `plan`/`invoke` already use (factored into
buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard
the CLI's own `detected` set on whether an explicit flag was passed:
resolveReviewerSelection's no-explicit-selection fallback is "select every
detected reviewer" (the correct default for /gsd:review), and passing it
an unconditionally non-empty detected set would silently invoke the whole
roster on every no-flag code review, violating COMP-01.

* test(01-03): define external finding consolidation

RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as
untrusted input — independently re-verifies every claim against the actual
current source, resists a prompt-injection attempt embedded in evidence
text, and folds a verified claim into the existing Narrative Findings
section with no second REVIEW.md schema (CONS-01..03). Also assert
code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* feat(01-03): consolidate external review evidence

GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence>
as untrusted data, independently re-verifies every cited claim against the
actual current source before it can appear in REVIEW.md, and explicitly
resists prompt injection embedded in evidence text (never a command, no
matter what it claims to be). A verified claim folds into the existing
Narrative Findings section with (external: {slug}) provenance — one
REVIEW.md schema only, no separate external-findings section.
code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* fix(01-02): gitignore the reviewer-step-dispatch build artifact

01-02 added src/reviewer-step-dispatch.cts but never added its
npm run build:lib output to .gitignore, unlike every sibling
gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked
noise in git status.

* docs(01-04): publish user and command contract for reviewer-lane source review

- Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md
  and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on
  failure, findings independently consolidated into the single REVIEW.md
- Add the same contract to the docs/features/code-review-pipeline.md
  fragment and regenerate docs/FEATURES.md from it
- Preserve /gsd-review as the plan-review command; cross-reference it
  rather than duplicating the reviewer roster
- Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md
  drift owned by source already shipped in Plans 01-01/01-03 but never
  regenerated (npm run regen:derived had not been run in this worktree)

* docs(01-04): align architecture and agent ownership docs for reviewer-lane trait

- ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes)
  through the shared dispatchReviewerLanes interpreter to the existing
  review-lane plan/invoke machinery, ending at gsd-code-reviewer as the
  sole REVIEW.md consolidator
- AGENTS.md: document gsd-code-reviewer's full-context verification scope
  and its treatment of external reviewer evidence as unverified input
- No new diagram, abstraction, or config key; docs/CONFIGURATION.md is
  unchanged since the feature adds no setting or default

* fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact

Same gap as the earlier .gitignore fix: 01-02 added
src/reviewer-step-dispatch.cts but never added its generated
gsd-core/bin/lib/reviewer-step-dispatch.cjs output to
eslint.config.mjs's ignore list like every sibling generated file,
so tsc's emitted __importDefault CommonJS-interop var tripped
no-var.

* fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md

01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists
cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster
row in docs/INVENTORY.md — required by design, since a role sentence
cannot be generated — was never added.

* fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes

refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook
strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added
supportsReviewerLanes: true to that step and this fixture was not updated.

* chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md

Both files grew as a direct, intended consequence of wiring optional
reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes
step and the untrusted-evidence consolidation contract) — not
incidental drift.

Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209)
Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209)

* test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes

From internal code review: dispatched must be false when zero lanes
actually reached plan(), and a throwing plan()/invoke() for one lane
must not discard results already collected for a sibling lane —
matching the fail-closed pattern gsd-tools.cjs already uses for the
same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794).

Refs: gsd-core-dks.16, gsd-core-dks.17

* fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review

- WR-01: dispatched now tracks whether any lane actually reached
  plan(), not results.length — an unresolvable selected slug no
  longer reports dispatched:true.
- WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a
  throw for one lane can never discard results already collected
  for a sibling lane, matching the same guard gsd-tools.cjs already
  has around the identical resolveLanePlan call.
- IN-01: documents the intentional budget===0-is-unbounded
  convention (#2797) the caller already relies on.
- IN-02: review-lane dispatch-step no longer blocks indefinitely on
  an un-piped interactive TTY; fails closed to empty paths instead.

Refs: gsd-core-dks.16, gsd-core-dks.17

* docs(01-05): add changeset fragment for PR #17

* fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract

agents/gsd-code-reviewer.md's untrusted-evidence section and its
pinning regression test both quote injection phrases as the exact
attack they defend against/detect — same
DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing
allowlist entries, not an actual injection vector.

* test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip

From CodeRabbit review: WR-02's earlier fix only wrapped plan() —
writePromptFile()/deps.invoke() still ran unguarded, so a throw
there still aborted every later selected lane. Also covers the
dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit
lanes silently not running when no prior review and no phase-start
commit exist).

* fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved

Previously an explicit reviewer-lane request with no prior review and
no resolvable phase-start commit reached dispatch-step with an empty
--base-sha, which fails closed via missing_provenance — correct, but
silent about why explicitly requested lanes didn't run. Now skip
dispatch entirely in that case with a stderr warning naming the
actual cause.

* fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan()

WR-02's original fix only guarded plan() — a throw from
writePromptFile() or deps.invoke() still aborted the whole dispatch,
discarding results already collected for lanes processed earlier in
the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it.

* fix(01-05): WR-02b mock must throw only on the first writePromptFile() call

The committed mock threw unconditionally, so codex's retry also threw and
failed for the same reason as claude's — the test could not distinguish
'sibling still runs' from 'sibling also breaks'. Gate the throw to the
first call, matching WR-02/WR-02c's single-failure intent.

* fix(#4209): close review findings from adversarial + critical-code-reviewer pass

Two independent reviews (agy adversarial review, Opus critical-code-reviewer +
ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch
wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and
fixed here:

- dispatch-step's reducer silently swallowed whole-dispatch rejections
  (invalid paths, missing provenance, etc); it now checks parsed.ok/reason.
- spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the
  LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review;
  now shares the single compute_file_scope derivation.
- the external reviewer prompt had no actual review request or citation
  requirement, only prohibitions; added both.
- removed the supportsReviewerLanes trait plumbing (capability registry,
  validator, loop-resolver, docs, tests) — it was never consulted by the
  real dispatch path, which gates on explicit CLI flags instead.
- flag-resolution require() was a fragile cwd-relative literal that failed
  silently on non-vendored installs; now resolves via GSD_TOOLS's own
  directory and warns instead of swallowing failure.
- reducer didn't unwrap the @file: overflow protocol for large payloads.
- deduplicated resolveBudget/budgetFor into one resolveLaneBudget.
- lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a
  second dispatch can't overwrite prior evidence.
- validatePaths rejects control characters, closing a markdown-injection
  vector into the external prompt via crafted filenames.
- reworded the one line that tripped prompt-injection-scan.sh instead of
  allowlisting the whole production prompt file.
- fixed a stale docstring range and a dispatched-field ordering bug.
- added 3 integration tests executing the actual reducer against synthetic
  dispatch-step JSON, replacing markdown-substring-only assertions.

771/771 tests pass across every touched suite; tsc --noEmit clean.

* fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait

The maintainer's approval on issue #4209 explicitly redirected implementation
shape: reviewer-lane dispatch must be a reusable capability/step-dispatch
trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call
itself. My previous commit (e2558326) deleted that trait entirely after
finding it declared-but-never-consulted, which was backwards — the fix was to
wire it, not remove it.

Restores the trait (capability.json, generated registry, validator,
loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes
now resolves its own active hook via `gsd_run loop render-hooks` for the
configured workflow.code_review_point and only proceeds to CLI-flag matching
when supportsReviewerLanes reads true. Explicit flags no longer bypass the
trait; a matching flag with the trait false resolves zero slugs (proven by a
new integration test executing the real fence with both trait states).

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the
dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer
redirect requires the capability layer, not the workflow, own the opt-in
decision).

* fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point

Both an agy adversarial review and an Opus critical-code-reviewer pass
independently found the same gap in my previous commit (9b2c3773d): the trait
check I wired into code-review.md only protected code-review's OWN
invocation — gsd-tools.cjs's dispatch-step handler still hardcoded
`trait: true` unconditionally, so a second capability declaring
supportsReviewerLanes would get zero enforcement from the shared CLI unless
it correctly re-implemented the ~15-line render-hooks scrape itself. That is
exactly the "each workflow.md hand-wiring the call" the maintainer's redirect
said to eliminate.

Moves the trait check into dispatch-step itself: given --cap-id/--point, it
self-invokes `loop render-hooks <point>` (relocating the one subprocess
code-review.md used to spawn for this, not adding a new one) and derives the
real trait from that capId's active hook, rather than trusting a
caller-passed boolean. code-review.md now only passes
--cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or
gates on the trait itself — the ~20-line scrape it previously carried is
gone. Any other capability opts into the identical enforcement by declaring
the trait and passing the same two flags.

Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input
variable (they proved a bash branch honors a variable, not that the variable
reflects the real capability manifest) with three integration tests that
invoke the real dispatch-step CLI against the real first-party capability
registry: the real code-review trait resolves true, an unknown --cap-id
resolves false (trait_not_enabled, fail-closed), and omitting
--cap-id/--point entirely resolves false (no context means no opt-in).

Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check
(agy-F1 was incomplete), and delete the promptWritten per-lane coupling
flag — the prompt write is idempotent, so writing it once per lane instead
of gating on "did any lane write it yet" removes a latent bug where a
deps.plan override that ever varies promptPath per lane would silently skip
writing for a later lane.

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count
drops (the trait scrape moved into dispatch-step), but the file still grew
this session across multiple commits; acknowledging per the growth-tracking
convention.

* fix(#4209): remove per-run token waste from the shipped prompts

Runtime prompt content, not session tokens: two real, per-invocation token
costs in the code that ships.

1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of
   load_context step 5's ~180-word untrusted-evidence contract in ~90 more
   words, breaking this section's own established terse one-liner style
   (every other rule here is 1-2 sentences). This prompt loads fresh on
   every /gsd:code-review invocation. Shrunk to a one-line cross-reference,
   matching how write_review's own reference to step 5 already does it.

2. buildSourceReviewPrompt repeated the base SHA on every single file line
   even though it is identical for every file and already stated once at
   the top of the prompt — O(files) wasted tokens on every dispatched lane
   for a 50-file review, for zero information gain. File lines are now bare
   paths.

* fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3

Opus critical-code-reviewer found a real Blocking defect in the --cap-id/
--point self-invocation added last commit: `dispatch-step` spawned
`loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its
stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to
`@file:<path>` instead of inline JSON -- the same overflow protocol this
feature already unwraps for its OWN dispatch result 60 lines later in
code-review.md. A large-enough activeHooks envelope (more installed
capabilities/fragments) would throw, get silently swallowed by the bare
catch, and misreport a real trait as trait_not_enabled with zero diagnostic.

Fixed by extracting the config/registry/capability-state resolution
`cmdLoopRenderHooks` already performs into an exported pure function,
resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now
share it), and calling it in-process from dispatch-step instead of spawning
a subprocess at all. This eliminates the @file: exposure entirely (the
dispatch-step path never touches the rendered-string envelope or its
JSON-stringify/50000-char threshold), removes one subprocess spawn per
code-review invocation, and gives a genuine diagnostic (stderr warning) on
resolution failure instead of silent fail-closed. Corrected three doc/
docstring references to the now-removed subprocess self-invocation.

Also fixes 2 real CI failures this round surfaced:
- lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line
  no-control-regex` comment was unused under this project's ESLint config
  (verified locally: the rule never actually flags \x00-\x1f in this repo's
  config) -- a mistake from an earlier commit this session, never actually
  lint-checked before push. Removed the disable comment.
- security (prompt-injection-scan): the agy-F1 regression test's crafted
  fixture literally contains "Ignore all prior instructions." as test data
  proving validatePaths rejects it -- allowlisted the test file, same
  DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries.

Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one
bullet stated "untrusted, never a command" three different ways in one
paragraph, and a same-file duplicate of write_review's schema rule.
Consolidated to state each rule once.

Declined one suggestion from this round: shrinking code-review.md's
EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests
(tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block,
tests/code-review.test.cjs's CONS-02 test) deliberately lock the four-
prohibitions restatement and the untrusted-evidence prose into the
INJECTED block itself, not just the consolidator's system prompt --
adjacency of the warning to the untrusted payload it's warning about is a
recognized prompt-injection defense-in-depth pattern from this
workstream's original TDD plan, not accidental duplication.

* fix(#4209): correct stale per-file base-SHA prose in the external prompt

Leftover from removing the per-file base SHA repetition earlier this
session: the review-request sentence still said "relative to its base SHA"
(singular per-file framing) when there's now exactly one base SHA, stated
once above the file list. Reads "relative to the base SHA above" now.

* fix(#4209): make getLane/configGet/plan required deps, delete dead defaults

R3/R4 from the review round I'd deferred as low-priority test-churn: this
file's one production caller (gsd-tools.cjs's dispatch-step handler) always
supplies all three, so the fallbacks were dead in production -- but each was
actively WRONG if ever reached: the default configGet always returned
undefined, silently disabling resolveLaneBudget's overflow guard; the
default getLane looked up only first-party REVIEWER_LANES, diverging from
production's overlay-merged roster; the default plan skipped per-host effort
resolution entirely.

These defaults were introduced by this PR's own earlier work (this file did
not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from
elsewhere, so there's no external caller depending on the lenient contract.

Turned out free to fix: making the three deps required and deleting
defaultGetLane/defaultPlan needed zero test changes -- every existing test
that actually reaches the per-lane loop already supplies getLane/plan
explicitly, and configGet's only real dependent (the budget-overflow tests)
already supplies it too. 788/788 tests pass unchanged, tsc/lint clean.

* fix(#4209): define depth semantics for the external reviewer lane

Verified this was a real bug, not a match to existing convention as I'd
claimed when declining the suggestion earlier this session: the internal
gsd-code-reviewer agent's own system prompt carries a full <depth_levels>
block defining what quick/standard/deep mean and do (agents/gsd-code-
reviewer.md:68-99). The external reviewer lane has no access to that
persona at all -- it only ever sees buildSourceReviewPrompt's bounded text,
which sent the bare depth label with zero definition to a third-party CLI
with no other source of truth for what "standard" means.

Added depthMeaning(), condensed from the internal reviewer's own
<depth_levels> definitions so the two stay consistent, and interpolated it
into the review-request sentence. 150/150 tests pass, tsc/lint clean.

* fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation

CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the
roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a
SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide
whether to dispatch at all. This file's own documented rule (its
depth-resolution guard, stated explicitly a few hundred lines earlier) is
that a guard and the extraction it protects must run as one shell
control-flow decision, because markdown-fenced blocks do not share shell
state -- this step violated its own file's rule for the entire feature's
gating condition.

Merged the roster-resolution fence and the dispatch-decision fence into one
continuous bash block, removing the intervening prose that split them.
Fixed the stderr-based failure detection in the same edit (RQ-01: checking
whether stderr is non-empty misfires on any benign Node warning; now checks
the actual exit status of the roster-resolution command).

Verified by extracting the merged fence and executing it standalone, driving
both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a
real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty,
SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests
pass, tsc/lint clean.

* fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write

Batch of Required/Suggestion fixes from the Opus critical-code-reviewer +
writing-for-agents pass:

- CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch
  blocks, commented-out code) and deep (error propagation, state mutation
  consistency, circular dependencies) relative to the real <depth_levels>
  block, and had zero test coverage. Restored full accuracy and added tests
  that read the real agents/gsd-code-reviewer.md file directly, so drift
  between the two can't recur silently. Unrecognised depth now normalizes to
  standard's definition, matching that agent's own documented rule, instead
  of rendering an undefined bare label.

- RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt
  `paths` does, but weren't checked for control characters like paths were
  (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and
  applied it to all four fields at the same provenance-check boundary.
  runDir previously had zero validation at all.

- S1: deleted the dead `identity` parameter on `invoke` -- the one production
  caller already ignores it, no test read it by name.

- S2: hoisted the shared prompt write above the per-lane loop -- promptPath
  is derived from runDir alone (constant across lanes by construction), so
  writing it once is both correct and cheaper than the per-lane write R1
  introduced earlier this session. Discovered and fixed a real regression
  from the naive version of this hoist: an unguarded throw would have
  escaped dispatchReviewerLanes as an uncaught exception instead of a clean
  per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason,
  matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with
  a dedicated regression test.

- S3: moved `planned = true` past the budget-overflow gate, so `dispatched`
  only reports true once a lane has cleared BOTH plan and budget checks.

- S5: relayed gsd-code-reviewer.md's own "performance issues are out of
  scope unless also correctness issues" policy into the external-lane
  prompt, which previously had no such guidance and could return findings
  the internal reviewer's own contract excludes.

- RQ-05 (partial): shrunk this file's own header docstring's restatement of
  the trait-reuse architecture to a pointer at
  gsd-core/references/loop-hook-dispatch.md, the canonical home.

234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion

RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the
SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke`
already share. code-review.md's ~18-line inline `node -e` reimplementing
`loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in
gsd-tools.cjs) is now a single call to this subcommand -- the exact
violation code-review-flags.cjs's own header warns against ("this is the
canonical flag-parsing surface -- do not replicate inline bash parsing").

RQ-03: an empty --cap-id XOR --point now warns distinctly from the
legitimate no-context opt-out (both absent) -- a caller that named a
capability without its point was silently indistinguishable from a correct
opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only
ever fires when the config-get COMMAND ITSELF fails (config-get already
resolves the manifest's own schema default in the normal case), but that
failure was previously silent.

RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait
resolved inside dispatch-step" explanation was restated in full in 5
places across this session's own review cycles. Consolidated to ONE
canonical statement in gsd-core/references/loop-hook-dispatch.md; the other
4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md,
code-review.md's step-opening comment) now point at it instead.

W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two
inert cases when capability-validator.cjs already rejects non-boolean at
load -- restated as the two cases that actually reach this code. Removed a
"do not hand-roll trait resolution" prohibition whose target no longer
exists once the positive description precedes it.

W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing
block means proceed as normal") -- an absent optional block already means
proceed as normal without being told.

W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the
made-up compound "byte-for-behavior [un]changed" with the token this
session's own docs already coined for this concept (inert) and the word
that means what byte-for-behavior was reaching for (unchanged).

W-10: dispatch_reviewer_lanes had no completion criterion -- added one
sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set,
either populated or empty). This exact sentence would have caught the
cross-fence bug fixed two commits ago at authoring time.

Declined from this round, with reasoning: W-02/W-03 (trim the
untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) --
two tests deliberately lock this as intentional adjacency-based
prompt-injection defense-in-depth, not accidental duplication (see this
branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site
`trap ... EXIT`) -- would fire at the end of the CREATING fence, before
spawn_reviewer's agent ever reads the evidence files, given this file's own
documented fenced-block execution model; the existing named cross-reference
between creation and cleanup already satisfies the co-location concern
without introducing that regression.

853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex

Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed
for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get
fallback lived in an earlier, separate fence from the fence that consumes it
via --point, split only by prose (not a guard, per this step's own documented
rule). Merged into the single continuous fence and added a structural test
asserting exactly one bash fence in the step.

The new end-to-end regression test for this used --codex, which drives the
fence's real `review-lane dispatch-step` call and, with the codex binary
present on PATH, spawns the real external CLI — which then blocks on
interactive auth with no stdin (BL-01). Stubbed gsd_run for
`review-lane dispatch-step` only (captures argv instead of executing),
keeping the real config-get/explicit-from-argv calls the test is actually
about.

* fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment

Round-5 review (Opus) warning-tier findings:

- WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but
  a control-character injection attempt" — a caller distinguishing a config
  problem from a security event couldn't tell them apart. Split into
  MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid).
- WR-05: validatePaths' containment check was lexical only (path.resolve),
  so a symlink whose own path sits inside repoRoot could still point outside
  it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can
  legitimately name a file already deleted in a stale worktree), realpathing
  repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't
  false-positive-reject its own real children.
- WR-08: a comment in the per-lane loop still said a throwing writePromptFile()
  was caught there — stale since the prompt write was hoisted above the loop
  in an earlier round.

WR-03 (validate depth against the quick/standard/deep enum) was considered
and declined: this dispatcher is deliberately capability-neutral (see the
existing "synthetic step context" test, which passes a non-code-review depth
label on purpose to prove no code-review-specific special-casing exists).
WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and
WR-07 (reason omitted on the aggregate return) were verified against source
and are not bugs — see review notes.

* docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap

Round-5 review (Opus, BL-03) flagged that an early exit between
dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A
trap-based cleanup was considered and rejected: if a step genuinely runs as
a separate process, a trap set at creation time would fire at the end of
that SAME fence, deleting the directory before spawn_reviewer/commit_review
ever read it — worse than the leak it would fix.

review.md's own gather_context/cleanup pair for the identical resource class
(a run-scoped reviewer temp dir) already makes and documents this exact
trade-off: cleanup runs only on a documented success path, and a leftover
$TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording
that precedent here so this isn't re-raised as a live gap in a future review.

* fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path

reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes
paths: ['docs/spec.md'] as a synthetic, never-read path proving the
dispatcher has no code-review-specific special-casing. lint-docs-guard-
registration correctly flagged this as an unregistered docs/ path reference —
add the docs-guard-exempt marker and its pinned baseline entry, the same
pattern every other synthetic docs/ literal in this test suite already uses.

* fix(#4209): backfill changeset pr: field with the real upstream PR number

changeset-lint's fail_pr_field_drift caught the fragment still pointing at
the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this
branch is now also open against.

* docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam

trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every
decision in ADR-2782 (D1-D9) and every prior dated amendment governs the
`role: "reviewer"` capability body and its one consumer, /gsd:review. This
PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary
feature capability's `steps[]` entry, projected through loop-resolver.cts
and resolved in-process via resolveActiveHooksForPoint - is a different
capability axis (steps/gates/contributions) that the ADR's own scope note
explicitly places out of reach. Per docs/contributor-standards.md's
"Amending an accepted ADR", an in-place dated section is the established,
lighter-weight path for an addition that stays within the ADR's existing
decisions - used twice already in this same file - so this appends a third
dated entry documenting the new seam, its consumer, and why it reuses the
existing D1-D9-governed plan/invoke machinery rather than adding a second
one. No decision is reversed; no new Amends/Amended-by pair is needed since
the steps/gates/contributions axis already carries reciprocal links to
ADR-857 and ADR-894.

* fix(#4209): close two test-quality gaps trek-e's review found

Minor 1: validatePaths (a path-shape parser guarding the prompt-
injection/path-traversal trust boundary) had only example-based coverage,
violating ADR-456's rule that parsers/budget limits carry at least one
fast-check property test. Adds three: safe-segment paths are never
rejected, a single leading "../" always escapes the one-segment repoRoot,
and a control character anywhere is always rejected - one property per
rejection reason validatePaths owns.

Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only
ever exercised far below budget or at budget:0 (unbounded), never at the
exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds
three exact-boundary tests using the real estimateTokens/
buildSourceReviewPrompt the module calls internally, so the resolved
token count is exact rather than approximated: budget == estimate (must
pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must
pass).

Also extracts okPlan()'s fixture timeoutMs into a named constant -
local/no-adhoc-timeout-literal (#4446) landed on next after this branch
was authored and flagged the pre-existing literal on rebase; it is fixture
data for a synthetic plan object dispatchReviewerLanes never waits on, a
distinct class from tests/helpers/timeouts.cjs's real subprocess norms.

* fix(#4209): update docs-guard-registration baseline for the new ADR citation

reviewer-step-dispatch.test.cjs's new fast-check property tests cite
docs/adr/456-test-rigor-architecture.md in a justifying comment (never a
real read). lint-docs-guard-registration fingerprints every docs/ path
string an exempted test file mentions and fails on drift so a human
re-confirms the exemption still holds - re-confirmed, and the baseline is
updated to match.

* fix(#4209): point changeset pr: field at the fork PR for CI validation

changeset-lint's fail_pr_field_drift check compares the fragment's pr:
field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH),
not a fixed target. Rehearsing this branch on fork PR
davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's
pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but
fails here. Backfill to 4323 happens again, as the last commit, immediately
before the approved push to open-gsd#4323 - never leaving pr: 17 on the
branch that ships upstream.

* fix(#4209): reject promptChannel:none lanes from source-review dispatch

CodeRabbit found a real scope mismatch: coderabbit's lane declares
promptChannel: 'none' and reviews the working tree on its own terms,
fed nothing (review.md:367). Silently dispatching it through
dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope
buildSourceReviewPrompt promises and let the lane review whatever it
independently sees fit, violating this interpreter's own scoped,
metadata-only contract. Reject before plan()/invoke(), same as an
unresolved slug.

* fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file

CodeRabbit found the whole-file match on workflowContent would still
pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts
of this 1000+-line workflow, proving nothing about the actual evidence
block's contract. Line-filtered via splitLines (not a bare-\n regex
spanning readFileSync content) so this stays CRLF-portable and passes
local/no-unbounded-quantifier and local/no-crlf-fragile-split.

* fix(#4209): guard DISPATCH_JSON substitution and capture its stderr

CodeRabbit found the dispatch-step command substitution unguarded: a
non-zero exit could leave DISPATCH_JSON empty (or halt the step under
errexit with no warning), and the downstream reducer would only ever
report the generic unparseable_dispatch_output reason, discarding the
command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/
EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface
it in a warning on failure, and fall back to a parseable dispatch_
command_failed JSON stub so the reducer's existing reason-reporting
path still fires.

* docs(#4209): fix byte-for-behavior wording and missing colon, regenerate

CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the
established repo term for output-identical unchanged behavior) and a
missing colon after the bold "Optional external reviewer lanes (#4209)"
lead-in in docs/features/code-review-pipeline.md. Fixed in the two
hand-authored sources (commands/gsd/code-review.md, docs/features/
code-review-pipeline.md) and regenerated the two derived projections
(skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/
FEATURES.md via gen-features.cjs) so they stay in sync.

* fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI)

The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds
double-quoted JSON keys inside a single-quoted shell literal. That
extra quote density, inside an already quote-heavy ~8KB driver string,
passed bash -n and the full local suite on Linux but broke Windows
Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end
to end (#4209 round 5)` failed on two Windows CI shards with `bash -c:
unexpected EOF while looking for matching '''` — a Windows argv-to-
command-line re-quoting edge case, reproducible on rerun, not a flake.
Root-caused via gh api job logs plus a byte-identical local
reconstruction of the test's own driver script.

Fix: drop the fabricated stub. The downstream node -e reducer already
falls back to reason `unparseable_dispatch_output` on any JSON.parse
failure, so an empty/partial DISPATCH_JSON on command failure is still
handled correctly, with zero new quoting risk.

* revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI)

Two materially different mechanisms for the same CodeRabbit Nitpick
("Trivial | Quick win") both broke Windows Git-Bash reproducibly:
a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ...
matching '''") and, after removing that, a plain `head -1 "$VAR"`
inside a nested command substitution ("unexpected EOF ... matching
'"'"). Both passed bash -n and the full local suite on Linux every
time; both failed the SAME test deterministically on Windows CI. Two
attempts at the same class of fix (nested-quote construction near
this exact step) is the retry limit - reverting to the original,
already-shipped, Windows-verified unguarded form rather than
continuing to guess at a third quoting mechanism for a Trivial-
severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for
anyone attempting this again: the fix belongs outside this specific
markdown-fence-driver test harness (e.g., a real .sh helper script)
if it's worth doing at all.

* fix(#4209): backfill changeset pr: field to the real upstream PR before push

Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy
changeset-lint's PR-number check while rehearsing there; this is the
last commit before the approved push to the real upstream PR
(open-gsd/gsd-core#4323), so the field points at that PR number again.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 22:52:33 -04:00
Tom Boucher
a0f8f956c4 enhance(#4139): Phase 3 — partition rules + the five checks (#4497)
* enhance(#4139): Phase 3 — partition rules + the five checks

ADR-4139 Decision 5, epic #4139 Phase 3. Issue #4403's own "Proposed behavior"
section lists four checks; the ADR's Decision 5 and its own phase table ("partition
rules + the five checks") list five — the same four plus "boundary moves are
declared, ongoing". Same issue-vs-ADR drift Phase 2 hit on the detail.md vs
detail/*.md layout: the ADR is the locked, reviewed document, so it wins. This PR
implements all five.

docs/PARTITION-RULES.md (new) is the partition-rules document: the partition rule
itself, the protected-content list and <!-- gsd:protected --> sentinel syntax
(relocated unchanged from gsd-core/references/compact-content-protected-content.md,
now deleted — it was never referenced by any runtime workflow Read, only by the
predecessor test as documentation, so nothing at runtime regresses, and removing it
from gsd-core/references/ also drops it from all 19 installed-project shipped-content
trees for a file nothing ever read), and the five checks explained for a human
reader. Referenced from a new CONTRIBUTING.md subsection under "Editing shipped
content".

tests/helpers/compact-content-split.cjs (new) is the shared mechanics: split
discovery (any gsd-core/workflows/<name>/detail/*.md paired with <name>.md — no
registry file, a pair is registered by existing on disk), line normalization
(carries forward Phase 2's bare-label-line isTrivial fix and the canonical
gsd_run-launcher-preamble exclusion), sentinel extraction, and a
Boundary-Move-Declared commit-trailer reader that is a direct structural port of
tests/helpers/emitted-runtime.cjs's Emitted-Drift-Ack-Hash/-Growth trailer reader
(ADR-3942) — same merge-base range, same fail-closed throw on an uncomputable range,
same dedupe/conflict rules.

tests/compact-content-partition-guard.test.cjs (new) is the actual guard, superseding
tests/plan-phase-compact-split.test.cjs (deleted — its per-pair checks are now the
general guard's job for plan-phase specifically). Checks 2 (disjointness) and 3
(registration + size cap) run unconditionally against every registered split. Checks
1 (completeness, fires once per split on the PR that introduces a new detail/ path),
4 (protected content — no trailer can ever excuse this one, unlike check 5) and 5
(boundary moves declared) are PR-diff-scoped against the resolved base ref and skip
cleanly when there's nothing to compare (a fresh clone, no PR in flight) — a
deliberate asymmetry from check 5's trailer reader, which must throw rather than
silently pass when ITS range is uncomputable, since that function is answering "did
this PR declare its moves" rather than "is there even a diff to look at". Each of the
five checks carries a RED (deliberately broken fixture) / GREEN (fixed) test pair,
built against synthetic temp files or real throwaway git repos, per this repo's rule
that a guard nobody has seen go red is not yet a guard. Building the real fixtures
caught and fixed one real bug before it shipped: check 4's line-presence test was
using the trivial-line-filtered normalizer, so a byte-identical spine falsely
reported its own protected code-fence line as "deleted" — fixed with a
non-filtering membership check.

Extending docs/INVENTORY.md's "Workflow Sub-Files" table for `detail` surfaced a
pre-existing, unrelated gap in the SAME area: gsd-core/workflows/<name>/templates/*.md
is a fourth workflow sub-file kind that already existed on disk and was already known
to lint-response-language-coverage.cjs's FRAGMENT_DIRS, but was invisible to
gen-inventory-manifest.cjs and undocumented in that table. Fixed alongside it, same
pattern, same PR, rather than deferred.

Also, mechanically required by the new fourth sub-file kind:
- scripts/lint-response-language-coverage.cjs: `detail` added to FRAGMENT_DIRS
  alongside modes/steps/templates — a detail/<part>.md inherits its parent's
  response_language coverage through the same per-file proof, not a parallel one.
- tests/workflow-size-budget.test.cjs: explicit regression test locking that
  detail/ files are governed solely by the hard, non-waivable NEW_FILE_CAP
  (tests/helpers/emitted-diff.cjs) and never by the XL/LARGE/DEFAULT spine tiers —
  true by construction (measureWorkflows/listWorkflowStems don't recurse), made
  explicit per the issue's own Done-when item rather than left true-by-omission.
- scripts/gen-inventory-manifest.cjs: `workflow_detail` and `workflow_templates`
  NESTED_FAMILIES entries; docs/INVENTORY-MANIFEST.json regenerated
  (plan-phase/detail/elaboration.md, discuss-phase/templates/*.md now tracked);
  docs/INVENTORY.md's table updated to four kinds.

Verified: `npm run lint:ci` clean with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4403): review findings + a real gsd-test failure in the new guard

Two orthogonal review passes (Standards + Spec, isolated sub-agents) plus a
separate security review ran against the prior commit. Fixed everything each
surfaced:

- Security (Low, path-traversal existence oracle): checkRegistration's
  dangling-reference check extracted detail-path-shaped substrings from spine
  PROSE via a regex that permits `.`/`/` freely, then joined them onto repoRoot
  and probed fs.existsSync with no containment check — a spine file containing
  `../../../etc/detail/passwd.md`-shaped text could make the guard test file
  existence outside the repo. Added a path.relative-based containment check
  before the fs.existsSync call; anything that resolves outside repoRoot is now
  reported as a dangling reference directly, never probed on disk.
- Standards (Boundary Coverage): the size-cap fixtures covered NEW_FILE_CAP and
  NEW_FILE_CAP-1 but not NEW_FILE_CAP+1 — added the third boundary-point case
  CLAUDE.md's TEST RULES require (limit-1/limit/limit+1).
- Standards (Property-Based Testing): extractProtectedBlocks (a sentinel
  parser) and the new parseBoundaryMoveTrailerValues (a declare/dedupe/conflict
  parser, bijective-shaped) had no fast-check property test. Added three: a
  render/parse bijectivity property for the trailer parser (mirroring the exact
  ADR-3942 sibling test's alphabet/idiom), a dedupe-is-idempotent property for
  the same parser, and a well-formed-sentinel-round-trips property for
  extractProtectedBlocks.

Then dispatched gsd-test on the resulting commit. It found a real bug the
reviews couldn't have caught (none of them can run inside gsd-test's sandbox):
checks 4/5's real-repo assertion failed against plan-phase's own split,
reporting DISK_PLANS/#3218-comment lines as "undeclared boundary moves" —
content Phase 2 (#4402) legitimately moved into detail/elaboration.md months
before this PR's Boundary-Move-Declared mechanism existed to require a
trailer for it. Root cause: `resolveBase()`'s own doc comment already documents
that no `origin/*` remote-tracking ref exists inside the gsd-test sandbox
container, and its fallback candidate (a bare `next` branch) can resolve to a
point in history that predates an already-merged, already-reviewed split —
making that split look "newly introduced" from the sandbox's vantage point.
Check 1 (completeness) already scopes itself correctly to only genuinely-new
detail paths (git diff status 'A'); checks 4 and 5 did not share that scoping,
so a stale base made them re-litigate a settled split retroactively. Fixed by
having checks 4/5 skip any split name check 1 already counted as newly-split —
their own premise ("did an EXISTING split shed/undeclare something") does not
apply to a split that is, from the resolved base's vantage point, brand new;
that is check 1's domain alone. Verified locally (25/25 tests pass via a
direct `node -e` require, since `node --test` is blocked in this repo) and via
re-reasoning through the exact real-repo scenario the gsd-test failure showed.

Also regenerated all 19 tests/fixtures/install-tree/*.json goldens — the
prior commit's deletion of gsd-core/references/compact-content-protected-content.md
was never reflected there, which is what golden-install-tree.test.cjs's other
19 failures in the same gsd-test run were.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4403): backfill changeset pr number to 4497

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4403): isolate codex-config.test.cjs into its own chunk, root-causing the Windows CI failure

PR #4497's "full test (windows-latest, 24, shard 2/3)" job failed: run-tests
killed chunk 3/8 at the 600s per-chunk backstop, with codex-config.test.cjs
(weight 17.87, by far the chunk's dominant cost) packed alongside 39 other
files. Traced, not assumed:

- scripts/run-tests.cjs's own timeout-headroom comment for the OUTER
  per-shard timeout documents that "adding one test file reshuffled 115 of
  268 unit files between shards" — shard/chunk composition is architecturally
  known to be unstable to single-file additions, which is exactly what this
  PR's own new tests/compact-content-partition-guard.test.cjs is.
- A second comment, dated 2026-09-06 (one day before this PR, PR #4428's own
  CI), already documents the SAME chunk hitting the SAME 600s backstop with
  the SAME file (codex-config.test.cjs, "a genuinely MEASURED weight of
  17.87 — not a stale-table miss") dominating it — the fix then was cutting
  the Windows per-chunk budget from 60 to 40. That cut clearly was not
  enough: two documented incidents in two days, at two different budget
  settings, both centered on one file that alone consumes ~45% of even the
  reduced Windows budget.
- tests/test-timings.json's own header confirms its source data
  (test-events-linux-node22/24.jsonl) is Linux-only, and run-tests.cjs's own
  chunk-timeout diagnostic already prints "real Windows cost runs ~2.2x the
  recorded figure" — the packer's weight-balancing is working off data that
  is both stale (table last regenerated 2026-08-07) and known to
  underestimate the platform where the failure occurs.

Given codex-config.test.cjs is disproportionately heavy AND every companion
sharing its chunk is decided by a packing algorithm already documented as
reshuffling unpredictably on any new file, tuning the shared budget a third
time only moves the marginal line to wherever the next new file happens to
land — it does not remove the gamble. Isolating codex-config.test.cjs into
its own dedicated single-file chunk, unconditionally and on every platform,
removes it at the source: the file never enters the pool packChunks balances,
so no other file's packing changes, and no future single-file addition
(mine or anyone else's) can silently reintroduce this exact failure by
landing in its chunk.

Extracted as a small pure function, partitionIsolatedFiles (mirroring this
file's existing pattern of pulling packing/analysis logic out of main() for
in-process unit coverage — see computeSweepProtectSet, analyzeChunkEvents),
with 6 new tests in tests/run-tests-harness.test.cjs covering basename
matching across path separators, near-miss non-matches, the empty-list case,
and the isolated-set contents.

Root cause is now closed rather than papered over with a retry: this failure
is a property of one specific heavy file's chunk placement, not something
that recurs randomly. If codex-config.test.cjs itself is ever genuinely sped
up, this isolation can be revisited — this is a packing-side mitigation for
a known file's cost, not a claim the cost is irreducible.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 16:44:53 -04:00
Tom Boucher
8b7a0b696b enhance(#4139): Phase 2 — one shared gate, one pilot split, one accuracy spot-check (#4471)
* enhance(#4402): split plan-phase into a spine + detail, add the shared compact-content gate

ADR-4139 Decisions 3-5, Phase 2 of the #4139 Compact Content epic. Pilot
split for plan-phase.md, the largest of the 58 eagerly-@-included workflow
files (98,290 bytes): the spine keeps every happy-path step, every
protected-content block (planner/checker prompt templates, quality gates,
the failing-direction few-shot example, the two ScheduleWakeup guardrail
paragraphs — each marked with a <!-- gsd:protected --> sentinel), and
condensed one-paragraph summaries of five rare/opt-in fallback paths
(planner and checker filesystem-hang recovery, phase-split recommendation,
source-audit gaps, the thinking-partner conditional, and plan bounce). The
full text of those five moves verbatim to gsd-core/workflows/plan-phase/detail.md
(9.9KB, well under the 32,768-byte NEW_FILE_CAP), read by the spine only
when workflow.compact_content is false (the default) — the exact same
resolution rule now stated once in the new shared
gsd-core/references/compact-content-gate.md, which every future split
references instead of restating.

Verified mechanically (tests/plan-phase-compact-split.test.cjs, scoped to
this one split — Phase 3/#4403 owns the generalized guard): the union of
spine + detail contains every non-trivial line the parent commit carried
(0 missing), no non-trivial line is duplicated between them (0 duplicated),
and every declared protected block is well-formed and non-empty. The spine
shrinks from 98,290 to 93,206 bytes (-5.2% of the eager-window cost this
epic exists to reduce); detail.md's 9,853 bytes are only ever paid by a
project that has NOT opted in.

Verified live, end to end, twice, against this actual repo (not a
synthetic fixture) — real gsd-planner and gsd-plan-checker subagent
spawns, real PLAN.md output:
- workflow.compact_content=false: planned a real disposable phase
  (a docs/how-to page for enabling the key itself); planner returned
  PLANNING COMPLETE, checker returned VERIFICATION PASSED, all fact-checks
  against real repo state confirmed.
- workflow.compact_content=true (detail.md never read): planned a second
  real disposable phase; planner returned PLANNING COMPLETE with
  frontmatter.validate and verify.plan-structure both clean, again fully
  grounded against real repo state. The five condensed fallback sections
  were independently re-read spine-only and confirmed sufficient to act on
  correctly without detail.md's elaboration.

Also drafts gsd-core/references/compact-content-protected-content.md — the
protected-content category list and <!-- gsd:protected --> sentinel syntax
ADR-4139 Decision 5 calls for, written to move to Phase 3 (#4403) unchanged
once it lands there.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4402): move detail.md into the ADR-4139-mandated detail/ subdirectory

Two independent review sub-agents (Standards and Spec axes of /code-review)
caught the same structural defect: ADR-4139 Decision 6 mandates
gsd-core/workflows/<name>/detail/*.md ("one or more parts... individually
skippable"), and this PR had shipped a flat plan-phase/detail.md instead,
copying issue #4402's own (inconsistent) restatement rather than the
locked ADR text. Fixed by git-mv to plan-phase/detail/elaboration.md and
updating every cross-reference (the spine's step 0.5 gate pointer, the
shared compact-content-gate.md's own resolution-rule wording, and the
completeness test's path constants).

Also, from the same review pass:
- docs/CONFIGURATION.md and gsd-core/references/planning-config.md's
  workflow.compact_content rows said "nothing branches on it yet" — no
  longer true now that plan-phase.md's spine does. Updated both to name
  plan-phase as the pilot and note the rest of the corpus is still pending.
- Regenerated all 19 tests/fixtures/install-tree/*.json golden fixtures
  (npm run gen:install-tree) — the three new shipped files were missing
  from the installer emitted-tree goldens.
- Found via a cache-busted `eslint . --max-warnings 0` (this repo's
  eslint --cache has produced false-greens before): the split test's
  `git show` call had a bare `timeout: 10000` literal, tripping
  local/no-adhoc-timeout-literal. Extracted to the existing GIT_TIMEOUT_MS
  constant from tests/helpers/timeouts.cjs instead of a second guessed
  copy of the same class of timeout.

Verified NOT needed, by tracing the actual mechanism rather than asserting
(tests/helpers/emitted-provenance.cjs's gsd-core-verbatim rule attributes
every gsd-core/{workflows,references}/** path to itself as an identity
source): an Emitted-Drift-Ack-Hash/-Growth trailer. Every changed/added
path in this diff is hand-authored and present in the diff itself, so
diffEmitted's attribution loop resolves `via` to the path's own source
before ever reaching the ack-lookup branch — there is no unattributed
delta to acknowledge. The spine also shrank (98,290 to 93,206 bytes), so
the growth ratchet has nothing to ack either.

Re-verified after these changes: the completeness/disjointness self-check
(0 missing, 0 duplicated) still holds against the relocated detail file,
and a full `npm run lint:ci` passes clean with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4402): restore literal content the pre-existing drift guards pin on

The first gsd-test run against this split (19 failures) surfaced real
regressions: several pre-existing structural guards pin the EXACT text
of the sections this split condensed, and paraphrasing broke them.

- tests/plan-phase-drift-guard.test.cjs expects the literal
  `DISK_PLANS=$(gsd_run query find-phase ...)` bash assignment inside
  plan-phase.md itself, not a prose description of the same check.
  Restored the exact line into both §9a and §11a's spine summaries.
- tests/thinking-partner.test.cjs expects plan-phase.md to literally
  offer "No, I'll decide" as the skip option. Restored that exact
  phrase into the condensed thinking-partner paragraph.
- Both restores would have duplicated the same text into
  plan-phase/detail/elaboration.md (which still carries the full
  elaboration). Removed the now-redundant restatements from the
  detail file instead of leaving them duplicated — the spine already
  computes DISK_PLANS before the detail elaboration is ever read, so
  the detail file references it rather than recomputing it.
- Re-running scripts/sync-runtime-launcher.cjs after that edit found
  the canonical gsd_run preamble had also become an unintentional
  spine/detail duplicate (both files call gsd_run and each is
  required, by runtime-launcher-parity's own contract, to carry its
  own copy). That's sanctioned duplication under a DIFFERENT
  contract, not lost/copy-pasted content, so
  tests/plan-phase-compact-split.test.cjs now excludes it from the
  disjointness check the same way it already excludes trivial
  fences/headings.
- Applied the adversarial-review finding on tests/plan-phase-compact-split.test.cjs's
  own isTrivial(): a blanket `line.length <= 15` cutoff silently
  swallowed real content (e.g. the 14-char `<quality_gate>`
  sentinel). Replaced it with a specific bare-label-line pattern
  (`Options:`, `Display banner:` etc.) — verified 0 missing / 0
  duplicated against the actual split, an improvement over both the
  original cutoff and a naive full removal (which produces
  false-positive "duplicates" on generic recurring labels).
- gsd-core/references/planning-config.md's own workflow.compact_content
  row used `/gsd-plan-phase` (hyphen). That file is Claude-facing
  source text (gsd-core/references/), which tests/slash-command-namespace.test.cjs
  requires in colon form; docs/CONFIGURATION.md's use of the hyphen
  form is correct as-is since docs/ is human-facing and outside that
  test's scanned directories. Fixed to `/gsd:plan-phase`.
- tests/plan-phase-compact-split.test.cjs's own `git show` of the
  parent commit failed inside the gsd-test sandbox ("detected dubious
  ownership") because the checkout is mounted under a UID the
  invoking user doesn't own. Scoped `-c safe.directory=<repo-root>`
  to that one git invocation rather than touching global git config.
- docs/INVENTORY.md still had one outstanding "detail.md part" wording
  fix from the earlier adversarial-review pass, staged now.

Re-verified locally against the exact assertions in all four affected
test files (all pass) before dispatching a fresh gsd-test run — no
change here should have broken any of the other 18 gates; `npm run
lint` is clean with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4402): restore the full marker enumeration to §9a's spine trigger line

The isolated Spec-axis review flagged that §9a's "Triggered when" line was
condensed to "Agent() returns but the return contains no recognized
marker" — dropping the literal `## PLANNING COMPLETE` / `## PHASE SPLIT
RECOMMENDED` / `## ⚠ Source Audit` / `## CHECKPOINT REACHED` /
`## PLANNING INCONCLUSIVE` enumeration, which is exactly the "machine-
parsed structural headings" category compact-content-protected-content.md
lists as protected. The load-bearing use of that same list (the
gsd_stall_watch call and the Handle Planner Return bullets a few lines
above) was never touched — only this one descriptive restatement was
genericized — but leaving any instance of a protected category
unsentineled is the silent erosion ADR-4139 Decision 4(c) warns
sufficiency isn't machine-checkable enough to catch on its own. Restored
the full enumeration into the spine.

That reintroduced an exact duplicate into plan-phase/detail/elaboration.md,
which still stated the same trigger sentence verbatim. Reworded the
detail file's version to reference the spine's trigger condition instead
of restating it, since the spine is now the single place that sentence
lives in full — mirroring the DISK_PLANS/"already computed above" pattern
from the previous commit.

Re-verified locally: completeness/disjointness (0 missing, 0 duplicated)
and all previously-fixed literal-content assertions still hold.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4402): backfill changeset pr number to 4471

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 12:18:38 -04:00
Tom Boucher
476394689a fix(#4254): pin sequential executor to the orchestrator's validated root (#4476)
* test(#4254): sequential executor root pin — failing-first regression + matrix

The new suite executes the shipped supplied-root-pin guard against real git
fixtures (drifted primary-checkout cwd halts before the write and the FATAL
names both roots; matching cwd permits it; unexpanded/empty pins halt;
normalization forms; submodule and sibling boundaries; metacharacter quoting;
drive-letter form gate) and locks the dispatch contract across execute-phase.md,
its sequential-root-pin step fragment, and worktree-path-safety.md. The #2772
per-plan serialization assertion retargets to the fragment that now carries
those rules (ADR-857 Phase 6 ceiling), plus the host-step wiring.

* fix(#4254): pin sequential executor to the orchestrator's validated root

Sequential-mode dispatch told the executor to self-derive PROJECT_ROOT from its
own cwd; every existing guard is worktree-mode-only or self-referential, so an
executor spawned with a drifted cwd committed onto the wrong checkout silently.

- worktree-path-safety.md step 0p: mode-agnostic supplied-root pin guard,
  composed by the orchestrator at build time with the literal $ORCHESTRATOR_WT
  (git-vs-git comparison on both sides — representation-safe on Windows, the
  #4296 lesson), fail-closed on empty/unexpanded pins, registered-submodule
  allowance, warn-and-proceed only when the dispatch carries no pin block.
- execute-phase.md sequential branch: build-time embed of the bound
  <project_root_pin> via the new execute-phase/steps/sequential-root-pin.md
  fragment (ADR-857 Phase 6 frozen ceiling — the host step cannot grow; the
  wave serialization rules move with the fragment, verbatim in substance) plus
  the per-write/commit pin instruction in <sequential_execution>. Worktree-mode
  dispatch untouched (its self-derived toplevel IS correct there).
- INVENTORY rows (5 locales) + INVENTORY-MANIFEST + install-tree goldens
  regenerated for the new fragment; changeset added.

* chore(#4254): backfill changeset PR number

* fix(#4254): accept backslash-separated Windows drive pins

CI on windows-latest showed every permit-path test failing with
"Actual root: <none>": pins composed from Node's path.join arrive in the
backslash drive form (C:\Users\RUNNER~1\...), which the guard's absolute-form
gate rejected before the cwd-side root was ever computed — a legitimate
matching pin could never pass. The gate now accepts either separator
([A-Za-z]:[\\/]); git -C resolves both forms (and 8.3 short names) to the
same canonical toplevel, so the git-vs-git comparison is unaffected. Form-gate
tests cover the emitted (C:/…) and produced (C:\…) spellings plus short names.

* fix(#4254): portable drive-form gate for MSYS bash

The bracket class [\\/] that accepted backslash drive pins parses
inconsistently on MSYS bash (the Windows CI leg still rejected C:\ pins —
every permit-path test red with "Actual root: <none>"). Replace it with
standard pattern escaping outside brackets: [A-Za-z]:/*|[A-Za-z]:\\* —
the escape form is version- and build-portable. Verified across all forms:
both drive spellings accepted; bare "C:", relative, empty, and unexpanded
rejected.

* fix(#4254): runtime-generated backslash comparator + self-describing FATAL

The Windows CI legs failed every #4254 permit-path row with
'Actual root: <none>' across two prior pattern spellings ([\\/] and \\*).
Stage misattribution: <none> appears whenever the FATAL fires BEFORE the
cwd-side capture assigns ACTUAL_ROOT — the absolute-form gate was what fired.

Mechanism: the test harness spawns bash -c <script> through the Windows
command-line boundary; that round-trip applies one extra shell-quoting pass
with double-quote semantics — a backslash written twice in the script text
arrives halved, while a lone backslash survives (the pin displays intact;
row 9's pure-bash gate independently showed the halved pattern rejecting
C:\ while C:/ still passed its surviving arm). On windows-latest every pin
carries backslashes (os.tmpdir() is the 8.3 short form C:\Users\RUNNER~1\...),
so the gate ate every pin before the actual root was ever computed.

Fix, robust by construction:
- the drive-form gate generates its backslash comparator at RUNTIME
  (BS=$(printf '\134'); match [A-Za-z]:"$BS"*) — the shipped guard now
  contains no doubled backslash anywhere, enforced by a regression
  assertion on the extracted guard text;
- the FATAL self-describes: Guard stage (pin-unbound / form-gate /
  actual-capture / pinned-capture / root-mismatch) plus a Diagnostic line
  carrying git's own stderr for capture failures and both compared values
  for mismatches — future platform failures name their stage in the log;
- row 9's hand-rolled duplicate case gate (transit-fragile copy, #4296
  Minor 1 duplication smell) is replaced by driving the SHIPPED guard and
  asserting the stage; rows 2/4 pin the new stage machinery.

Validated on darwin across drift/match/relative/unbound/empty/bare-drive/
forward-and-backslash drive forms, each also re-run under a simulated
Windows transit (every doubled backslash halved) with identical outcomes.

* fix(#4254): close the empty-comparator fail-open seam in the drive-form gate

Self-review of the runtime-generated backslash comparator: if printf's
octal escape ever returned empty, the drive arm [A-Za-z]:"$BS"* would
widen to drive-RELATIVE pins (C:foo) — the construction's one theoretical
fail-open path. Fail closed with a self-describing diagnostic instead of
trusting the shell's printf.

---------

Co-authored-by: sim <sim@local>
2026-09-07 10:54:30 -04:00
Tom Boucher
2cf119f57e fix(#4217): reconcile artifacts before classifying an abnormally-ended executor (#4442)
* fix(#4217): reconcile artifacts before classifying abnormal ends

* test(#4217): pin the completion-reconciliation contract

* chore(#4217): regen derived inventory and install-tree fixtures

* test(#4217): follow the #4003 anchoring pins into the reconciliation fragment

Emitted-Drift-Ack-Growth: execute-phase.md — the runtime-neutral completion-reconciliation pointer, the two Codex wait-rule bindings, and the step-7 reconcile-first gate net +33 bytes over the extracted fallback block (#4217)

* chore(#4217): add changeset fragment

* chore(#4217): backfill PR number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-09-06 18:57:38 -04:00
Michel Moreira
19b66c3ec8 fix(#4218): stop the orchestrator steering an executor that is still working (#4391)
* fix(#4218): stop the orchestrator steering an executor that is still working

An executor with recent RED/GREEN/REFACTOR commits and passing verification had
not yet written its SUMMARY because it was finishing closeout. The parent saw no
local OS test/build process, inferred an "idle tail", and sent "Finalize
immediately" into a working child; in CLI runs the same inference interrupted an
executor before GREEN, leaving a RED commit and an uncommitted edit.

The stall block said only "if no completion signal, no SUMMARY.md, and no
expected-branch commits appear for N minutes" — it never said what to do when
commits DO exist and only the SUMMARY is outstanding, never defined the
threshold as a period without progress rather than a total runtime, and never
ruled out a process listing as an idleness signal. Four rules close that:

- the threshold measures a period WITHOUT MEANINGFUL PROGRESS, from the last
  sign of progress, not from dispatch — a long verification tail is not a stall;
- commits + missing SUMMARY + recent activity resolves to KEEP WAITING, with
  steering, interrupting and re-dispatching each named and forbidden;
- urgency/finalization messages ("Finalize immediately" and family) are
  forbidden outright — they arrive mid-verification and truncate a correct run.
  The existing user-facing pause is the only sanctioned stop, and `kill and
  retry` is a clean restart, not a nudge;
- the absence of a local OS test/build process is NOT idleness: a native
  subagent runs in the runtime's own session, and an executor between two tool
  calls shows no process at all. Progress is judged only by the signals this
  workflow names.

Five prose-contract assertions in tests/execute-phase-wave.test.cjs, all red on
next.

* fix(#4218): extract the progress policy to a step fragment

CI's #1168 gate caught it: execute-phase.md sits 77 bytes under a frozen 93600
ceiling and the four rules added ~2.3 KB. "Extract, not bump" is the repo's
stated remedy, and this workflow already carries policy detail that way.

execute-phase/steps/executor-progress-policy.md owns the policy. The
worktree-recovery arm moved with it — `kill and switch to inline execution`
qualifies the stop this policy governs, so it belongs beside the rule about when
stopping is sanctioned at all, not stranded in the host. The #3212 recovery
OPTIONS stay in the host, where tests/config.test.cjs pins them.

The host keeps what must be read before the orchestrator acts: the verdict, the
threshold definition, and a pointer that fires before any message is sent to the
child. execute-phase.md is now 93475 bytes — 48 SMALLER than next.

* chore: add changeset for #4218

* chore(#4218): regenerate the inventory manifest for the new step fragment

docs/INVENTORY-MANIFEST.json is the authoritative per-file list behind
INVENTORY.md's `<workflow>/steps/*.md` row, so a new fragment has to appear
there or gen-inventory-manifest --check reds the lint-tests lane.

* chore(#4218): restore the issue ref on the allow-test-rule marker

ADR-456 requires a #NNN on a new exemption; the block rewrite that moved the
policy into the fragment dropped it.

* chore(#4218): regenerate the install-tree fixtures for the new step fragment

The fragment ships with the workflow, so every runtime's golden install tree
gains one path — gen:install-tree is the generator that owns those fixtures.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-06 17:51:34 -04:00
Tom Boucher
6adf3098ac fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans (#4364)
* test(#4145): regression rows for hash-matching prefix-less pristine baselines

RED skeleton: src/pristine-baseline.cts exports findPristineByHash as a
null-returning stub so the new rows fail behaviorally, not at require time.
Failing-first rows: verifier resolution (no_baseline must drop to 0 when an
exact-hash orphan exists), findPristineByHash unit row, and the two
saveLocalPatches relocation rows. Negative-space rows pin today's behavior:
missing baselines still report ok_no_baseline, mismatching orphans are never
adopted or deleted, canonical precedence and the #3657 drift posture are
untouched.

* fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans

Both pristine readers joined the manifest-keyed path strictly, so a snapshot
stored without the gsd-core/ prefix (an earlier release's writer) was reported
as ok_no_baseline by the verifier and pushed into regeneration by
saveLocalPatches — where incoming-release candidates can never satisfy the
recorded outgoing hash, leaving the correct baseline permanently unconsumed.

- src/pristine-baseline.cts (new, ADR-457): shared findPristineByHash —
  deterministic sorted scan of gsd-pristine/, exact sha-256 equality with the
  recorded pristine_hashes entry (the same authority the #3657 drift guard
  trusts), symlink-skipping, canonical path excluded via skipRel.
- verify-reapply-patches.cjs verifyFile(): on canonical miss with a recorded
  hash, adopt byte-identical content found anywhere under gsd-pristine/ before
  reporting OK_NO_BASELINE. Drift posture (#3657), canonical precedence, and
  the frozen REASON/report shapes are untouched; the verifier stays read-only.
- install.js saveLocalPatches(): preserve-check rescue — relocate a
  hash-matching orphan to the canonical path (copy, hash-verify, then remove
  the orphan) so the state self-heals on the next update instead of repeating
  forever. Honest accounting: new non-overlapping rescued counter.
- Workflow doc: one-sentence note on hash-based snapshot resolution.
- Derived ripples: INVENTORY-MANIFEST.json regen, eslint ignore + .gitignore
  entries for the compiled artifact, seedFixture mkdir fix in the new rows.

Emitted-Drift-Ack-Growth: reapply-patches.md — one-sentence note on hash-based pristine snapshot resolution (#4145)

* fix(#4145): review follow-up — orphan scan never consumes a canonical path

Adversarial review finding: with two modified files sharing byte-identical
outgoing content, recoverOrphanedPristine could adopt the OTHER file's
canonical pristine as its rescue source — relocating it (copy + delete at
its home path) and ping-ponging the single baseline between the two files
across updates. findPristineByHash's skip parameter now accepts a Set, and
saveLocalPatches passes the normalized manifest keys so every canonical
path is excluded; only genuine non-canonical orphans are eligible for
removal (no strict-join reader ever consults those). Adds the
canonical-theft regression row, a Set-skip unit assertion, and tightens the
workflow doc sentence the same pass flagged as overstated.

* fix(#4145): INVENTORY roster row + symlink-fixture correction

Two leftovers from the ab17b7a1e5 bench run, both root-caused:
- docs/INVENTORY.md roster row for cli_modules/pristine-baseline.cjs
  (#3762 gate: every manifest entry carries a row).
- The findPristineByHash symlink unit fixture placed its symlink target
  INSIDE the scanned root, so the walk legitimately matched the real target
  file. The implementation skips the symlink itself; the fixture now keeps
  the target outside the scanned tree so the assertion tests what it claims.

* changeset(#4145): fixed fragment for pristine baseline hash resolution

---------

Co-authored-by: gsd-agent <agent@gsd.local>
2026-09-06 02:04:18 -04:00
Cody Anderson
77e2472ca0 enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration

Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on
Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets
(the .env.example/.sample/.template/.dist templates stay readable).
Read checks file_path; Grep checks an explicit path and judges the glob
per brace alternative; Bash runs a two-pass token scan (quotes, comments,
redirects with fd digits, separators, $( )/backtick/<( ) recursion,
heredoc bodies never scanned as commands, nested bash -c/eval rescans,
git <ref>:<path> shapes) with a closed non-reading exemption set for
existence checks. Fail-open crash policy; 1 MiB commands are denied as
command-too-large; more than 64 glob alternatives as glob-too-complex.

Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt
for approval whenever any Read() deny rule exists, even in auto mode. A
hook denial is not a permission rule and never arms that check. The
installer-written deny rules are retired in the follow-up commit.

Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks
HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking
guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell),
shell-command-projection managed sets, installer-migration-report,
OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch),
docs tables in five locales, ADR-766 always-on list, regen:derived
fixtures, and a new table-driven unit suite.

* test(#4221): pin the secret-read guard in existing hook gates

Register gsd-secret-read-guard.js in every existing hook gate: the
hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest
REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity
EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal-
hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS,
kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and
typed-payload floors, the OpenCode adapter (grep mapping, include ->
glob, three dispatch tests) and a Kimi TOML matcher assertion.

* fix(#4221): retire installer Read() deny rules (legacy filter)

Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS
and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets)
strings. mergeClaudePermissions now only filters them out of an existing
permissions.deny: an absent deny key stays absent, a malformed one is
still repaired to [], and an array emptied by the filter is deleted so
no `"deny": []` residue is left. Uninstall filters the same legacy list
and, symmetric with the Antigravity branch, drops an emptied allow or
deny key and an emptied permissions object.

Unlike the #2278 allow-side migration there is no surviving current
deny list, so the constant is renamed rather than mirrored. Removal is
byte-exact: a hand-written identical rule is indistinguishable from the
installer's and is removed too (the manifest never recorded permission
strings). USER-GUIDE and CONTEXT.md updated.

* test(#4221): flip install-regressions deny-rule assertions to the retired shape

The fresh-merge, non-destructive merge, idempotency, end-to-end install,
reinstall and uninstall assertions now expect no Read(.env*) deny rules
and no permissions.deny key on a fresh install; the deny:null repair case
is kept. A new describe block covers the legacy filter: retired strings
removed with a user entry kept, partial sets, near-miss strings
untouched, idempotency, GSD-only deny array deleted, a pre-existing
empty deny preserved, and uninstall symmetry for allow/deny/permissions.

* chore(#4221): add changeset fragment for PR #4236

* fix(#4221): case-fold names; scan shell stdin and xargs pipes

Review round 1 (trek-e):

- Blocker: secret-name matching is now case-insensitive in the Read,
  Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a
  case-insensitive filesystem are recognized as the same secret file.
- Major: a shell interpreter's script is now scanned wherever it comes
  from. The tokenizer keeps heredoc bodies as per-segment tokens and
  records separator operators; pass 2 groups by segment id and resolves
  bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined
  `-lc`) scans the script operand, a file operand is checked as a file
  (a `<( )` operand's echo/printf output is reconstructed), otherwise
  stdin is the script and heredocs, here-strings and a piped echo/printf
  source are scanned. `eval` joins all its operands; `source`/`.` handle
  process substitution. Data heredocs (`cat <<EOF`, the commit-message
  shape) stay unscanned.
- Major: `… | xargs <cmd>` checks the upstream segment's operands as
  file names when the sub-command reads (`echo .env | xargs cat`,
  `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the
  inference; a shell sub-command's `-c` script is scanned.

Header, USER-GUIDE bullet and changeset updated; documented gaps now
include piped scripts from non-echo sources and `exec`/`timeout`
wrappers. 60 new suite cases pin the block and allow shapes.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 04:00:08 -04:00
JusticeWay
0fca71eaae enhance(#2529): cover every workflow with response-language directives + CI lint (#2558)
* enhance(#2529): cover every workflow with response-language directives + CI lint

Every workflow now carries response-language coverage in one of three forms,
and a CI lint keeps it that way.

- 43 workflows load the new shared reference,
  `gsd-core/references/response-language-directive.md`, by eager `@`-import.
- Lazy-loaded modes/steps/templates, which cannot rely on an eager import,
  carry an exact inline directive; 35 such paths are pinned by exact path.
- Fragments dispatched by a covered parent inherit coverage, proven per file
  rather than granted per directory.

The 45 workflows whose directive covered only "questions, prompts, and
explanations" now name inter-tool narration, which is the defect #2529
reports: the running commentary between tool calls stayed English while the
answers around it were translated.

`scripts/lint-response-language-coverage.cjs` enforces it and fails closed on
three independent discovery failures (unreadable catalog, empty catalog,
unfollowed symlink). It resolves which reference a workflow imports and applies
the same four-predicate test to that file, so a weakened shared reference
uncovers its importers instead of passing silently, reported once as a systemic
failure rather than 43 times. The walk follows symlinked subtrees with a
realpath cycle bound. `lint:ci` invokes it by name.

REQ-LANG-03 and REQ-LANG-04 state the contract in docs/FEATURES.md;
REQ-LANG-04 names the two forms that satisfy it ("narration", "between tool
calls") rather than enumerating class members an author cannot use verbatim,
and a test pins that text to what the matcher accepts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): register the coverage test in the docs-guard lane

`107eb8c1` (#3787) landed the docs-guard lane on `next` while this PR was
open: a test that reads a `docs/` path must be named in
`scripts/docs-guard-registry.cjs` or carry a `docs-guard-exempt` marker,
so the guards that read a doc run on the PR that changes it.

`tests/response-language-coverage.test.cjs` reads `docs/FEATURES.md` -- it
extracts every form REQ-LANG-04 offers an author and runs each through the
matcher that enforces it. Registration, not exemption, is the correct side
of that gate: a reword of the requirement with no code change is precisely
the diff this test exists to catch, and it is the diff the lane would
otherwise skip.

Registered narrowly (`['docs/FEATURES.md']`) rather than with the `'*'`
sentinel, so an unrelated docs change does not pull this test into the lane.

Verified: lint-docs-guard-registration 0 violations, tests/ci-docs-guard-registry.test.cjs
51/51, lint:ci exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): consolidate this PR's emitted-growth acks into its own fragment

This PR ripples emitted bytes across 85 workflow paths. Until now each ripple
was acknowledged by appending to whichever live fragment owned that path,
because two ack sources may never name the same path.

`a84f7563` (#3078) swept all 45 fully-spent fragments off `next`. Forty-two of
the paths this PR grows were owned by swept fragments, so those keys are now
unowned and this PR's own fragment declares them directly -- one path, one
source, and no dependence on a fragment that no longer exists. Each adopted
entry keeps its measurement and records where it came from.

Two paths are handled differently, because the sweep did not free them:

- `review.md` is now owned by `3034-parallel-reviewer-lanes.json`, which
  landed on `next` after the sweep. Its entry is live, so the old route still
  applies: this PR's note is appended to that entry rather than declared a
  second time.
- `plan-review-convergence.md` keeps the arrangement made in round 24.

Result: 3 fragments in the directory, 85 keys in this PR's own,
0 cross-source duplicates. `lint-emitted-drift-ack` exit 0,
`tests/emitted-attribution.test.cjs` green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): move REQ-LANG-03/04 into the feature fragment that now generates them

`36375513` (#3845) made docs/FEATURES.md a generated projection of
docs/features/*.md, marked "do not edit by hand". This PR wrote REQ-LANG-03
and REQ-LANG-04 straight into the generated file, so the rebase left the
requirement present in the projection and absent from its source -- the next
regeneration would have deleted both, and `tests/features-index-gate.test.cjs`
was already red on the mismatch.

Both requirements now live in docs/features/response-language-config.md
alongside REQ-LANG-01 and -02. Regenerating produces a docs/FEATURES.md that is
byte-identical to the committed one, so the text this PR shipped is unchanged --
only its source of truth moved to where #3840 put it.

The docs-guard registration is widened to name the fragment as well as the
projection. The requirement's source is the fragment now, and an edit there
that skips regeneration would otherwise reach this guard through neither path.

Verified: features-index-gate 68/68, lint-docs-guard-registration 0 violations,
ci-docs-guard-registry + response-language-coverage 142/142.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): hand the plan-phase ack back to its new live owner

`c933184b` (#3825) landed `3172-stated-failing-direction.json` on `next` after
fragment had adopted that path when the sweep left it unowned, so the merged
tree named it from two sources -- a hard failure in
`scripts/lint-emitted-drift-ack.cjs`.

The path has a live owner again, so the append route applies: this PR's note
joins that entry, carrying its own measurement, and the key is dropped from
this PR's fragment (84 keys left, the others untouched). The provenance
sentence written for the swept-fragment case is removed rather than reused --
this path was never orphaned, so that account of it would be false.

Same shape as `review.md` and `plan-review-convergence.md`: ownership is a
property of the merged tree, and a fragment landing upstream after a push can
reclaim a key no local check would have flagged.

Verified: lint-emitted-drift-ack exit 0, lint:ci exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): state byte figures that are true against the tree

The reference claimed `execute-phase.md` has "2 bytes of headroom under the
ceiling named below". That was true when the sentence was written -- the file
sat at 93398 against the 93400 comfort assert -- and upstream has since shrunk
it to 91493 against a 93600 hard ceiling, so the figure now understates the
headroom by three orders of magnitude. The rationale the sentence supports does
not depend on the number, so the number is gone rather than refreshed: a
restated figure would go stale again on the next upstream edit, and nothing
parses it.

Audited every other numeric claim this PR ships the same way, mechanically
against the merge base: all 82 FILE-delta claims in the ack fragment match the
real per-file delta exactly, and the 1,629-byte reference and 63-byte import
line check out. One class was imprecise: the 41 notes for workflows whose
inline directive was rewritten in place quoted the conversion counterfactual as
"+1,692 bytes more loaded context", which is the reference form's whole weight,
not the increase over the inline directive those files already carry. Each now
names both quantities and the net (+1,605 / +1,609 / +1,584).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): one rule for pinned vs inherited coverage, and the docs to pick it

Review measured that 14 of the 35 pinned fragments would pass by inheritance
anyway, and that the PR asserted both readings at once: inheritance is real
coverage (so those 14 pins are noise) or it is not (so 30 inheriting fragments
are green-but-uncovered). Only one can be true.

Inheritance is real: the predicate proves it per file -- the parent must
dispatch this exact path from a read/execute context AND be covered itself --
so the parent's directive is in the loaded context by the time the fragment is
read. The 14 pins are therefore removed along with the directive lines they
pinned, and those files inherit like the 30 structurally identical ones. The
rule is now stated where the set is declared, and enforced from the other side
by a test: no member of the pinned set may be one that would have inherited.
That is what decides the form for the next fragment.

- pinned set 35 -> 21; 14 workflow files revert to their base content
- `findViolations` no longer returns early on a pinned path: a file that becomes
  eagerly loaded and takes the shared reference is strictly better off, and the
  gate must not red that. The reference form is admitted because its own wording
  is validated in turn; an arbitrary reworded inline line still fails.
- the reference-directive cache is keyed by size and mtime, not by path alone,
  so a rewritten reference re-asked in one process no longer returns the stale
  verdict
- `carriesInlineDirective` names its negation blindness: four independent hits
  read vocabulary, not polarity
- the real-tree scan asserts each source produced files instead of `> 152`, a
  constant that read as the workflow count and would have passed a scan that
  lost one of its two directories
- the pinned-set size assertion goes the same way: the size follows from the
  rule, so the rule is what the suite asserts

Docs, for the gate that now governs every future workflow:
- `docs/contributing/response-language-coverage.md` -- why the narration class
  is the discriminator, the four coverage forms, the decision order that picks
  one, the pinned line, and what each failure message means
- a row in CONTRIBUTING.md's CI checks table, matching the docs-guard row
- `docs/CONFIGURATION.md` points at it from the `response_language` entry

Also: the changeset said 45 reworded workflows; it is 44 (42 @-reference + 21
pinned + 44 rewritten = 107 touched). That text ships to CHANGELOG.md.

`3707-parse-gap-reporting.json` landed on `next` reclaiming `audit-uat.md` and
`progress.md`; both handed back by the append route, leaving 82 keys here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): correct the reference-taker count, 43 -> 42

The ack notes said the import line is byte-identical "in each of the 43
workflows that take the reference" and that the alternative would be "43 inline
copies". The shared reference has 42 importers; the 43rd file in review's table
is `execute-phase.md`, which imports the OTHER reference. Corrected in all 41
notes that carry the sentence, across this PR's fragment and the two it appends
to.

Found by re-running the numeric audit from the previous round after the rebase,
which also re-verified all 84 FILE-delta claims against the new base -- all
exact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): migrate the emitted-drift ack from a fragment to commit trailers

ADR-3942 (#3954) landed while this PR was open: the acknowledgment is now a commit
trailer and tests/emitted-drift-acks/ no longer exists. The fragment is deleted and
each key it declared becomes one trailer, reasons unchanged.

The four keys this PR had handed to 3034-*, 3172-* and 3707-* under the one-source
rule come home here. That rule was the whole reason for the hand-backs, and the
trailer model has no shared namespace to collide in -- five of this PR's rounds were
spent on exactly those collisions.

Emitted-Drift-Ack-Growth: add-backlog.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: add-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: add-tests.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: add-todo.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: ai-integration-phase.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by ddf85287 (fix(#3357), #3513), and no fragment on `next` declares this path now. The ack therefore returns to this PR's own fragment, which is the only live source for it — the change to the path is this PR's.
Emitted-Drift-Ack-Growth: analyze-dependencies.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: audit-fix.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 19: this PR declared the path in its own fragment, and `3602-workflow-subagent-model-resolution.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3602-workflow-subagent-model-resolution.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: audit-milestone.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed from `2962-zsh-nomatch-for-glob-portability.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: audit-uat.md — A live/archived split was added and then reverted on this branch (see `$comment`): the split's extra rule in `initialize`, the narrowed Unparsed-table filter, and the separate 'Unparsed UAT Files in Archived Milestones' informational section are all removed, so the file settles at origin/next 5582 -> 7124 bytes (+1542, final). #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: autonomous.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 16: `3210-autonomous-precondition-gate.json` landed on `next` in 8fc88f66 (fix(#3210), #3528) and declares this path today. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3210-autonomous-precondition-gate.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: check-todos.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: cleanup.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 18: this PR declared the path in its own fragment, and `2142-quick-task-archival.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `2142-quick-task-archival.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: code-review-fix.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 13: `3190-code-review-fix-auto-rewrite-review.json` landed on `next` in 1d5d7795 (fix(#3190), #3434) and declares this path too. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3190-code-review-fix-auto-rewrite-review.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: code-review.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 19: the fragment that carried this sentence (`3503-diff-base-scope-anchor.json`) was retired on `next` by 2fca0e17 (enhance(#2554), #3695), and `2554-code-review-depth-overrides.json` declares this path today. One path takes exactly one ack source, so the sentence follows the path to its live owner. Re-homed from `2554-code-review-depth-overrides.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: complete-milestone.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 15: the fragment carrying it (`3458-audit-open-acknowledge-wiring.json`) was retired on `next` and the path is declared by `3409-unreachable-guard-arms.json` today. One path takes exactly one ack source, so the sentence follows the path to its live owner. Re-homed from `3409-unreachable-guard-arms.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: debug.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 14: the fragment that carried this sentence (`3149-init-debug-entry-point.json`) was retired on `next` by 26f8015c (fix(#3448), #3476), and `3448-debug-autoresume-next-action.json` declares the path today. One path takes exactly one ack source, so the sentence follows the path to its live owner rather than being dropped or re-armed under a retired number. Re-homed from `3448-debug-autoresume-next-action.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: diagnose-issues.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: an inline copy in every workflow would be that many places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 19: this PR declared the path in its own fragment, and `3602-workflow-subagent-model-resolution.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3602-workflow-subagent-model-resolution.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: discuss-phase.md — #2529 round 36: this workflow's inline directive was rewritten in round 10 to name inter-tool narration, but in the compressed form, and that rewrite came to −1 byte against `next` — so it declared no growth and this key was absent from this PR's ack set until now. Round 36 replaces the compressed clause with the same enumeration the other rewordings carry — narration between tool calls, status updates, progress notes, findings, questions, prompts, and explanations — because `discuss-phase.md` started from the identical upstream sentence as `verify-work.md` and `new-milestone.md` and those two took the full list, so the shorthand was an inconsistency rather than a decision. +87 bytes against `next`, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. The directive stays INLINE rather than becoming an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have cost 1,692 bytes of loaded context against the 87 this sentence costs. `commands/gsd/discuss-phase.md` dispatches this workflow lazily (`Read and execute ...`) rather than `@`-importing it, so the 87 bytes land in the installed file and are read once the workflow is dispatched, not on every command invocation.
Emitted-Drift-Ack-Growth: discuss-phase-assumptions.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 15: `3409-unreachable-guard-arms.json` landed on `next` in #3558 and declares this path too. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3409-unreachable-guard-arms.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: discuss-phase-power.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: do.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: docs-update.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +83 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +83 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 83 bytes this inline directive costs, a net +1,609. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,609 bytes more loaded context per invocation. Re-homed in round 19: this PR declared the path in its own fragment, and `3602-workflow-subagent-model-resolution.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3602-workflow-subagent-model-resolution.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: edit-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 13: `3262-editphase-milestone-scope-guard.json` landed on `next` in fd4715f8 (fix(#3262), #3446) and declares this path too. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3262-editphase-milestone-scope-guard.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: eval-review.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by ddf85287 (fix(#3357), #3513), and no fragment on `next` declares this path now. The ack therefore returns to this PR's own fragment, which is the only live source for it — the change to the path is this PR's.
Emitted-Drift-Ack-Growth: execute-plan.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 14: the fragment that carried this sentence (`2652-quick-diagnose-dispatch-isolation.json`) was retired on `next` by 362d0434 (fix(#3370), #3478), and `3370-execute-phase-gate-conflation.json` declares the path today. One path takes exactly one ack source, so the sentence follows the path to its live owner rather than being dropped or re-armed under a retired number. Re-homed from `3370-execute-phase-gate-conflation.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: explore.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed from `2229-explore-claim-disposition.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: extract-learnings.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: fast.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 18: this PR declared the path in its own fragment, and `3585-planning-commit-guard.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3585-planning-commit-guard.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: forensics.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: graduation.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: health.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 13: the fragment that carried this sentence (`2573-state-head-freshness.json`) was retired on `next` by 7ddcc198 (fix(#3309)), and `3309-health-docs-generated.json` declares the path today. One path takes exactly one ack source, so the sentence follows the path to its live owner rather than being dropped or re-armed under a retired number. Re-homed from `3309-health-docs-generated.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: help.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: import.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 18: this PR declared the path in its own fragment, and `3576-references-canonical-cites.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3576-references-canonical-cites.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: inbox.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: ingest-docs.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed from `2658-trae-instruction-file-path.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: insert-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: list-phase-assumptions.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: list-seeds.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: list-workspaces.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 …

* fix(#2529): read the catalog-relative dispatch spelling, and the plural of "output"

Two false positives in the coverage lint, both surfaced by this round's work
rather than by a red gate finding them for us.

#3552 landed `execute-phase/steps/protected-branch.md` on next while this PR was
open, dispatched from execute-phase.md's `"none"` arm with the path written
RELATIVE to the catalog. `namesFragmentAsEntryPoint` only ever looked for the
`gsd-core/workflows/`-rooted spelling, so it read a live dispatch as no dispatch
and the new fragment as uncovered. It now accepts both spellings and matches the
relative one on a path boundary, so `vendor/<path>` cannot vouch for `<path>`.

Recognizing that spelling makes one pin redundant: execute-phase.md dispatches
executor-isolation-dispatch.md the same way, so the fragment inherits and its
own copy of the sentence comes back out. That is the rule round 29 encoded,
enforced by the test that measures it rather than by hand.

`output` was the one term in USER_OUTPUT_RE without an `s?`, so "translate all
outputs, including narration between tool calls" read as uncovered. The new
property tests caught it on their first run.

Those properties pin the rule the hand-written cases are instances of: four
signals on ONE line accept, dropping any one rejects, spreading them across
lines rejects. The vocabulary is written out in the test rather than read back
from the script's regexes, per CONTRIBUTING.md "Fixture provenance (#2371)" -- a
generator seeded from the matcher can only re-derive what the matcher already
believes, and that independence is what caught the plural. Both new properties
are mutation-verified: dropping the narration predicate reds the necessity
property, and collapsing the document to a single line reds the cross-line one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#2529): spell the narration enumeration out in the last two shorthand directives

discuss-phase.md and plan-phase.md were the only two of this PR's 44
rewordings that abbreviated the inserted clause to "narration between tool
calls included" instead of naming the output classes the way the rest of them
do. Both forms satisfy the lint's four predicates, so nothing was broken --
but the point of #2529 is that an author reading one workflow should not have
to infer what the neighbouring one means by "included".

Both abbreviations were size decisions rather than wording ones, and both
reasons have since expired because next shrank the files. discuss-phase.md
sat 25 bytes under the 32,000-byte #717 dispatcher budget and now has 1,825;
plan-phase.md sat 87 bytes under the 94,519-byte ADR-857 capstone ratchet
against a +108 clause and now has 3,180. Neither budget is raised here and no
unrelated prose is trimmed; workflow-size-budget and
phase6-capstone-conformance both pass.

discuss-phase.md started from the identical upstream sentence as verify-work.md
and new-milestone.md ("All user-facing questions, prompts, and explanations in
this workflow"), and those two received the full enumeration; it now matches
them exactly. plan-phase.md keeps its own scope word ("orchestrator output") and
its subagent pass-through instruction, both upstream's, and only trades the
shorthand for the enumeration.

The shorthand now appears nowhere in the catalog. The two remaining variants
(plan-review-convergence.md, spec-phase.md) keep upstream's own verb and scope
and end on "report prose", which is what those workflows actually emit --
rewriting those would change a directive's strength, not its wording.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): match the workflow extension case-insensitively in coverage discovery

`findMarkdownFilesRecursive` filtered on `entry.name.endsWith('.md')`, so
`SETTINGS.MD` — the same file to Windows and macOS, a different one to Linux —
was skipped on the only platform whose verdict gates the merge. The direction of
that failure is the problem: a workflow the walk declines to see is a workflow
this lint certifies by omission, which is the same vacuous pass `main()` already
refuses when discovery returns nothing at all.

The filter is now an allowlist keyed on the lowercased `path.extname`.
`.mdx` stays out on purpose: admitting an extension states what a workflow IS,
and that claim has a second half — `inheritsParentCoverage` resolves a
fragment's parent as `<workflow>.md`. An `.mdx` entry belongs here next to the
parent resolution it would have to move with, not ahead of it.

Two tests: an uppercase-extension file is discovered AND lands as a violation
rather than an exemption, and every admitted extension is spelled so the
lowercasing match can reach it (an uppercase or dotless entry would be dead
configuration that reads like coverage).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#2529): cover the quick-batch workflow that #3676 landed uncovered

`next` gained `quick-batch.md` and nine step fragments in 2f64e6230 (#3676,
PR #4212) with no response-language directive, so the merge result reds this
PR's own lint with 10 violations. The lint is doing exactly what it exists to
do; the coverage is what has to move.

`quick-batch.md` takes the shared @-reference on line 1, the same as the other
42 top-level workflows, and eight of the nine fragments then inherit through
its `read and execute` stubs. The ninth does not:
`quick-batch/steps/plan-checker-loop.md` is dispatched by a SIBLING fragment
(`planner-wave.md:134`) and named in the parent only inside a parenthetical
with no dispatch verb, which is the shape round 29's rule already covers for
`execute-phase/steps/regression-gate-run.md` and
`plan-phase/steps/prd-express-path.md`. It carries the pinned inline directive
and joins `EXACT_INLINE_DIRECTIVE_WORKFLOWS`; the comment above that set now
names four such fragments instead of three. Coverage: 163 workflows.

`FULL_BUDGET` in tests/skill-frontmatter-contract.test.cjs moves 844 -> 846.
The same commit grew `help/modes/full.md` from 834 to 844 lines, landing it
exactly on the ceiling with zero slack, and the two lines this PR adds there
are its pinned directive and the blank separating it. That is a coverage
contract every workflow carries, not the content creep the budget guards.
The #597 ratchet rule holds: actualMax 846, slack 0, well inside LARGE_GRACE.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: quick-batch.md — #2529: the workflow arrived on `next` in 2f64e6230 (#3676, PR #4212) with no response-language directive, so this PR's lint reds on the merge result; covering it is the PR's whole contract, not an optional extra. It gains the shared directive as a single eager `@`-reference line, the identical form the other 42 top-level workflows take. FILE delta: +62 bytes, as the gate measures it. LOADED-CONTEXT delta: +1,691 bytes — the import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates read the FILE and not the transitive inline, so they see 62 of those 1,691 bytes; the remaining 1,629 are declared here because no gate reads them. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Nine `quick-batch/steps/*` fragments are covered without a byte of their own — eight inherit through the parent's dispatch stubs, and the ninth takes the pinned inline sentence, which the emitted surface does not measure.

* fix(#2529): scope row 48 by what a diff says, not by which paths it names

`tests/gsd-quick-batch-quick-regression.test.cjs` treats any branch touching a
`quick-batch` path as #3676 phase work, then forbids it from editing ordinary
`quick.md`. This PR covers EVERY workflow with the shared response-language
directive — quick-batch.md and its fragments included — so the scope check
turned true, and the row read this PR's one-line directive on `quick.md` as a
phase violation.

That is the false positive the row's own #3730 note already scoped away from,
arriving by the other door: not an unrelated branch that misses the surface,
but a catalog-wide sweep that touches all of it. A path now counts as phase
work only when its diff says something other than the coverage contract, and
the two accepted directive forms are read from
`scripts/lint-response-language-coverage.cjs` rather than restated, so a
reworded contract cannot leave the carve-out matching prose the lint no longer
recognizes. A file the branch ADDED still counts — every line is new, which is
what a real #3676-phase branch looks like.

The invariant is unweakened in the direction that matters: a phase branch that
edits `commands/gsd/quick.md`, `gsd-core/workflows/quick.md` or anything under
`quick/steps/` for any reason other than the directive still fails the row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-04 21:12:12 -04:00
Tom Boucher
8249ebcf6e fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate

RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do
not exist yet; every row fails on require. Per #3770 only an intentional
target-test failure may authorize GREEN; zero-test discovery, fixture crashes,
unrelated failures, and unexpected green are INVALID_RED.

* fix(3770): require intentional RED evidence before GREEN

Only an intentional failure of the TARGET test (distinctly named, TAP-reported
assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test
discovery, fixture/load crashes (file-named failures), nonzero exits without a
failing test, unrelated failures, unexpected greens, and malformed/missing
records are INVALID_RED and block GREEN.

- src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses
  the prohibition-enforcement TAP primitives; fail-closed, never throws)
- check tdd-red-evidence <record.json>: validates the persisted record
  (command, exit code, failing test, expected, actual)
- gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now
  requires the evidence record + gate verdict, not a nonzero exit or a RED: tag

* chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs

* fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib

- gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B <
  49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid)
- tests: the row-6 fixture used String.replace (first-occurrence), so the
  `not ok` line still named the target test and the classifier was right to
  accept it; replaceAll makes the failure genuinely unrelated
- eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs
  (lint the src/*.cts source, per ADR-457 migration rule)

Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line

* chore(3770): add changeset

* chore(3770): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-04 16:44:38 -04:00
Tom Boucher
2f4f7538e9 fix(#4264): wire both TDD dispatch backends to phase.tdd-applicable (#4284) 2026-09-04 16:08:42 -04:00
Tom Boucher
2f64e6230a feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core

Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch
binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the
new pure decision-logic module (arg validation, effective concurrency,
deterministic merge order, spawn backpressure, verification/merge
routing, cleanup-entry construction — design doc rows 3-15,24,26-28,
30-36,39; property rows 51-53). quick-batch-update-items.test.cjs
covers the new updateBatchItems export on src/quick-batch.cts (rows
15,22-23, including the negative cycle-rejection case).
quick-batch-command-router.test.cjs covers the new
gsd-tools quick-batch CLI family (rows 46-47). These reference modules/
exports that do not exist yet.

* feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router

Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision
layer — CLI verbs and pure orchestration logic only; no workflow
markdown, no Agent()/git-worktree I/O.

- src/quick-batch-dispatch.cts (new): pure decision functions consumed
  by the (separate, follow-up) /gsd:quick-batch workflow markdown —
  parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder,
  computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome,
  buildCleanupManifestEntry (the last parses caller-supplied plan text
  via the existing parsePlanDocument; no filesystem access).

- src/quick-batch.cts: adds updateBatchItems, resolving the design
  doc's Open Question 1 as ONE additive export on this module instead
  of the second, independent BATCH.json writer the design doc
  originally proposed. Reuses the same withPlanningLock transaction
  shape, computeWaves, and platformWriteSync call resumeBatch/
  completeQuickItem already use; fails closed without persisting on
  an unknown item, an unknown/self dependency, or an introduced cycle.

- src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI
  family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs)
  as a first-party always-on command (like /gsd:quick), not the opt-in
  capability-registry path graphify uses. Verbs: create/update/resume/
  complete (wrap quick-batch.cts) and effective-concurrency/
  merge-eligible/spawn-plan/verification-routing/merge-routing/
  cleanup-entry/parse-args (wrap quick-batch-dispatch.cts).

Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property
rows 51-53. Rows covering workflow markdown / Agent() dispatch /
`git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the
follow-up markdown-authoring pass, per the phase brief's explicit
scope boundary.

* docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces

New-.cts-module ripple for the two Phase 4 modules (epic #3344,
ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts,
ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source,
not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json
(via `node scripts/gen-inventory-manifest.cjs --write`, after
`npm run build:lib`), and CONTEXT.md glossary entries for
"Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router
Module", plus an update to the existing "Quick-Batch Core Primitives
Module" entry documenting the new updateBatchItems export.

* test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count)

scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs
file under the quick-batch production module by longest-prefix match,
and that module is already at its 2-file cap (quick-batch.test.cjs +
quick-batch.property.test.cjs). The standalone
tests/quick-batch-update-items.test.cjs added in the prior commit
pushed it to 3 and failed `npm run lint:ci`. Fold its content into
quick-batch.test.cjs (append-only — no existing test in that file is
modified) and update the CONTEXT.md glossary reference to match.

Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after
`npm ci` (this worktree previously had no local node_modules, which
also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-*
unable to resolve typescript — resolved by npm ci, no code change
needed there). `npm run lint:ci` and
`npx tsc -p tsconfig.build.json --noEmit` are both green after this
fix.

* test(#3676): add failing tests for the quick-batch command/workflow markdown

Failing-first tests for Phase 4's markdown-authoring pass (epic #3344,
ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs
covers commands/gsd/quick-batch.md's frontmatter/objective/process,
gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR
1610 NEW_FILE_CAP) and step-fragment count, the isolation model
(rows 20-22), the executor single-writer invariant (row 18), merge
validation reusing the existing bounded primitive (row 25), the
optional research/plan-checker/verification leaves (rows 16,17,19,
30,31), planning-failure blocking execution (row 29), the submodule
guard (rows 36,44), and the new agents/gsd-planner.md quick-batch
mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers
row 48 (ordinary /gsd:quick stays byte-identical). Named
`gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's
longest-prefix bucketing doesn't fold these markdown-only tests into
the already-capped quick-batch/quick-batch-dispatch/
quick-batch-command-router production-module buckets from the CORE
pass. These reference files that do not exist yet.

* feat(#3676): author the quick-batch command, workflow, and planner mode

Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch
binding") — the orchestration layer that calls into Pass 1's CLI
verbs (src/quick-batch-command-router.cts).

- commands/gsd/quick-batch.md (new): frontmatter/objective/process,
  delegates argument validation to `quick-batch parse-args`
  (parseQuickBatchArgs) rather than re-deriving the grammar.

- gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR
  1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded
  step fragments under gsd-core/workflows/quick-batch/steps/:
  resume-mode, batch-init, research-phase (flag:--research),
  planner-wave (+ nested plan-checker-loop when --validate),
  worktree-dispatch, merge-wave, verification-wave (flag:--validate),
  completion. Covers design doc rows 3-45: capacity/isolation
  resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG-
  layer planning with full-task-catalog prompts and always-required
  depends_on/files_modified frontmatter, serialized worktree create/
  merge/cleanup via the existing worktree.cleanup-wave primitive,
  deterministic wave-order merging, verification routing
  (human_needed/gaps_found), the executor single-writer invariant,
  submodule fail-loud guard, and #1941 fork-base auto-degrade.

- agents/gsd-planner.md: additive new `load_mode_context` bullet for
  `**Mode:** quick-batch`, pointing at the new
  gsd-core/references/planner-quick-batch.md reference (documents the
  always-required depends_on/files_modified contract, reusing the
  existing frontmatter grammar — no new keys). Existing modes
  byte-identical, only a new bullet added.

- src/init.cts (+init-command-router.cts, +command-aliases.cts):
  cmdInitQuickBatch / `init.quick-batch` — model profiles,
  commit_docs, roadmap/planning existence checks, and the
  section_manifest field gating research-phase/verification-wave
  (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY
  atoms — no new atom needed).

Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the
prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43,
45 covered by construction (verb wiring, single-writer prompt
constraints, crash-window resume via unmodified Phase 3 primitives).

* docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern

npm run regen:derived output for the new command/workflow/reference
(epic #3344, ADR-1239 "Quick-batch binding"):
- skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md)
- docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md,
  planner-quick-batch.md, and the quick-batch-dispatch.cjs/
  quick-batch-command-router.cjs CLI-module rows' now-live
  `/gsd-quick-batch` cross-reference (was "(separate, follow-up)")
  + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`)
- gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`)
  — research-phase/verification-wave gsd:section entries for the new
  quick-batch workflow
- tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) —
  the new command/workflow/skill/reference files now ship to every
  runtime

scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for
gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS
word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional-
flag pattern gsd-core/workflows/quick.md already carries baselined
(e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2);
quoting would break the intended "omit this arg when the flag is
false" splitting.

* fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch

Security review pass findings, both confirmed real:

1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment
   interpolated the raw, attacker-influenced task ${description} (and
   the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw
   description into every planner's prompt in the layer) straight into
   Agent() prompt bodies with no boundary. Fixed by wrapping every such
   interpolation in a <security_context> + DATA_START/DATA_END
   boundary, matching the CONCRETE convention already implemented in
   this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md,
   gsd-core/workflows/debug.md) — commands/gsd/quick.md's own
   <security_notes> only asserts this convention in prose, so the
   debug-agent files are the real precedent followed here. Added a new
   <security_notes> block to commands/gsd/quick-batch.md (it had none)
   documenting both this fix and the one below.

2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection.
   gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both
   ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED,
   causing shell word-splitting and pathname expansion on raw task-list
   text before the parser ever saw it. Fixed at the source: added a
   `--text <string>` form to the `parse-args` verb
   (src/quick-batch-command-router.cts) that accepts the ENTIRE
   $ARGUMENTS as ONE quoted argv element and does the whitespace split
   itself, in Node — which is never glob-aware, unlike the shell.
   Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form
   is kept for direct/test callers that already have a real argv array.

The SC2086 baseline entry added for the original unquoted line is now
stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports
it) and has been removed; the two SC2046 entries for the UNRELATED,
still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style
conditional-flag splitting remain — that line only ever expands to one
of a few known-safe literal strings (never raw user text), matching
quick.md's own already-baselined convention exactly.

Tests: quick-batch-command-router.test.cjs covers the new --text form
(token splitting, glob-shaped text passing through literally
unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs
asserts the DATA_START/DATA_END boundary on every leaf prompt
(research-phase/planner-wave/plan-checker-loop/verification-wave,
including the shared task catalog) and the quoted --text call sites.

* fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35

Spec review pass findings — the test matrix claimed "yes" coverage
these assertions did not actually support:

- Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection
  only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs,
  committed alongside the security fix that touches the same file) that
  .planning/quick-batches/ is never created for any rejected value —
  createBatch is genuinely never reached.
- Row 18 (--resume <unknown-batch-id>): previously only exercised a
  hand-corrupted BATCH.json, never a genuinely nonexistent batch
  directory. Added the real nonexistent-id case (also in
  quick-batch-command-router.test.cjs).
- Row 24 (post-planning updateBatchItems racing a concurrent
  completeQuickItem for a different item, both through
  withPlanningLock): zero test existed. Added a property test
  (tests/quick-batch.property.test.cjs, appended — Phase 3's own file,
  no existing test touched) exercising both call orders and asserting
  no lost update in the final on-disk manifest — the same technique
  Phase 3's own row-15 lock-contention property test uses (sequential
  calls through the real lock; a working mutex makes any interleaving
  equivalent to some serial order, so this is the same claim a literal
  concurrent-thread test would make without OS-level threading).
- Row 34 (worktree preserved on merge_failed) and row 35 (undeclared-
  deletion detection): both were previously asserted only at the pure
  routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs
  using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs
  already establishes for executeWorktreeWaveCleanupPlan (real repo,
  real worktree, a REAL merge conflict / a REAL file deletion diffed
  against declared_deletions) — asserting the actual worktree directory
  survives on disk, not just that a pure function returns a
  preserveWorktree:true field. Named gsd-quick-batch-* so lint-test-
  file-count's bucketing doesn't fold it into any capped module bucket.

Row 48 (/gsd:quick regression) intentionally left as-is per the
reviewer's own framing: the byte-identity claim is already
mechanically proven by the changed-path diff (git diff --name-only
empty on those two paths IS byte-identity), and a genuine execution-
level regression test would require actually running the workflow —
out of scope for this repo's unit-test model (no other quick.md
regression test in this repo does that either).

* docs(#3676): add the changeset and user-facing docs the command needed

Standards review pass findings — both HARD:

- Missing changeset. None of the 6 prior #3676 commits touched
  .changeset/*. /gsd-quick-batch is a new user-facing command;
  CLAUDE.md/CONTRIBUTING.md require one. Added
  .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder —
  backfilled after the PR opens, matching CLAUDE.md's own documented
  convention and Phase 3's own precedent, #4190's
  .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen
  form `/gsd-quick-batch` throughout, never the source-artifact colon
  form (`scripts/lint-docs-command-form.cjs` confirms 0 violations;
  that check scans docs/**, not .changeset/, so it was never actually
  in scope for the fragment itself, but the wording still follows the
  doc convention for consistency, matching how Phase 3's own fragment
  named the not-yet-shipped command).
- Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis
  how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's
  existing convention for /gsd-quick /gsd-fast) covering --jobs,
  --validate, --research, --resume, --file, the capacity/isolation
  interaction, and resume/failure recovery. Cross-linked from
  docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's
  own "Related" section. Added a /gsd-quick-batch section to
  docs/COMMANDS.md (same table format as the existing /gsd-quick
  entry) and docs/features/quick-batch.md (REQ-QB-01..12, same
  frontmatter shape as docs/features/quick-mode.md) — regenerated
  docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md
  via the standard generators.

* fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught

gsd-test's real run against 155e8975b3 found 43 failures, all rooted in
this phase's own new command/workflow never being registered across
~10 independent generated/hand-maintained registries this repo keeps
in parity by convention. Root-caused each, no test weakened or
special-cased.

- help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live-
  registry.test.cjs): added a /gsd:quick-batch entry to
  gsd-core/workflows/help/modes/full.md (the real help.md content;
  gsd-core/workflows/help.md is a thin dispatcher) documenting every
  flag (--file/--jobs/--validate/--research/--resume), matching the
  existing /gsd:quick entry's format.

- gen-section-manifest.test.cjs: quick-batch.md's
  `gsd_run query init.quick-batch` invocation used inline
  `$([ ... ] && echo --flag)` substitutions, which never satisfy the
  test's exact-whitespace-token / assigned-variable detection (the
  trailing `))` glued onto `--research` in the compound substitution
  broke the "exact token" match). Rewrote to the same
  VALIDATE_PARAM/RESEARCH_PARAM two-line pattern
  gsd-core/workflows/quick.md's own Step 2 already uses.

- runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md
  fragments that call gsd_run each needed their OWN embedded copy of
  the canonical shim preamble (every workflow .md that calls gsd_run
  carries its own copy — reading one file does not persist shell state
  into another). Ran `node scripts/sync-runtime-launcher.cjs`, which
  inserted it before each file's first gsd_run call.
  plan-checker-loop.md correctly has none — it never calls gsd_run
  directly.

- Namespace routing (skill-manifest.test.cjs, install-nested-
  layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added
  `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and
  routing table (same namespace `quick` already routes through), and
  to src/clusters.cts's `utility` cluster (same cluster `quick`
  already belongs to). Verified by hand-running installRuntimeArtifacts
  + applySurface for augment/cline against a real temp install: exactly
  6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested
  under gsd-ns-workflow/skills/, never re-flattened.

- mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a
  brand-new command is a real count change, not a bug this test should
  hide).

- model-omit-when-inherit-guard.test.cjs: added the canonical
  `<!-- #2517 model-omit-on-inherit -->` marker block to
  gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/
  researcher/checker/executor/verifier — lives in a steps/ fragment,
  read combined with the host by this test's own readWorkflowCombined,
  same as quick.md's own research-phase.md carries it for its gated
  section). Also fixed a genuine pre-existing inconsistency in the
  test's own "#2711: the guarded set is derived from dispatch sites"
  check: its `nonDispatching` computation read the BARE host file while
  `derived` (the set it's checked against) reads the combined
  host+steps content — inconsistent with that same test file's own
  #2994 doc comment explaining why the combined read is necessary.
  quick-batch.md is the first workflow whose EVERY model="{...}"
  dispatch site lives in a mandatory (never gated) steps/ fragment —
  extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a
  brand-new file — which is what exposed the mismatch. Fixed by using
  the same readWorkflowCombined read in both places.

- skill-frontmatter-contract.test.cjs: shortened
  commands/gsd/quick-batch.md's frontmatter `description` from 107 to
  91 chars (<=100 budget), and added `quick-batch.md` to the hand-
  maintained KNOWN_SKILLS consolidation allowlist with a #3676
  justification comment (a genuinely new first-party command, not a
  consolidation of an existing skill).

- workflow-fragments-emission.install.test.cjs: added `quick-batch.md`
  to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is
  deliberately NOT a no-op for it — its research-phase/verification-
  wave sections are gated).

- Regenerated all downstream artifacts (npm run build:lib && npm run
  regen:derived && npm run gen:plugin-skills -- --write && npm run
  gen:features -- --write): skills/gsd-quick-batch/SKILL.md,
  skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for
  augment/cline/hermes/qwen/trae/zcode.

- emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition
  (one new `load_mode_context` bullet pointing at the new
  gsd-core/references/planner-quick-batch.md reference) grew the file
  124 bytes without an acknowledgment trailer. Acknowledged below —
  the growth is the deliberate, additive, single-bullet change from
  the earlier feat(#3676) commit, not drift.

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green
(includes lint-workflow-shellcheck, lint-test-file-count,
lint-docs-command-form). The deep install/spawn/registry tests gsd-test
actually runs (docs-parity-live-registry, gen-section-manifest,
runtime-launcher-parity, install-nested-layout,
runtime-artifact-layout-surface, skill-manifest, skill-frontmatter-
contract, mcp-server-catalog, model-omit-when-inherit-guard,
workflow-fragments-emission) are not part of lint:ci — each fix above
was independently verified by hand-invoking the exact production
function the failing test calls (installRuntimeArtifacts, applySurface,
composeWorkflow, the CLUSTERS union, the section-manifest forwarding
regex) against the real repo tree and confirming the expected shape.

Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift.

* fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget

skill-frontmatter-contract.test.cjs's "feature #3039: tiered help —
size budgets" enforces a SEPARATE line-count ceiling for
gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines,
tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's
assertTightCeiling) — independent of the skill-frontmatter description-
length budget and consolidation allowlist I touched in the prior round;
those are unrelated checks in the same test FILE, not the same check.

Root cause: the /gsd:quick-batch entry I added to full.md in the
docs-parity fix round was 17 lines, pushing the file from 834 to 851
lines — 7 over the 844 ceiling. Condensed the entry (merged the
per-flag bullet list into one dense "Flags:" line, dropped from 3
Usage examples to 1) to 844 lines exactly — at the ceiling with zero
slack, which assertTightCeiling accepts (it only fails on
actualMax > ceiling, or on slack > grace when the ceiling is too
LOOSE — zero slack triggers neither).

Verified after trimming: full.md still contains a live /gsd:quick-batch
reference (bidirectional parity) and all 5 argument-hint flags
(--jobs/--validate/--research/--resume/--file) still appear as literal
tokens (docs-parity-live-registry.test.cjs's own flag-coverage check,
re-run by hand against the trimmed content).

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green.

* docs(#3676): backfill changeset pr number to 4212

Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's
pr:0 placeholder backfilled with the real PR number now that
gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR
Number Handling convention and Phase 3's own #4190 precedent
(708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt
from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE.

* fix(#3676): resolve prompt-injection-scan false positive on test fixture

tests/quick-batch.test.cjs:232's row 11b regression proves the task-list
parser carries a prompt-injection-shaped task description through
createBatch as inert data, never interpreted. The fixture has to be a
real "ignore all previous instructions..." phrase or the test asserts
nothing, but the full-file --diff scan flagged it once unrelated edits
in the same file pulled it into the changed-file set.

Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the
sanctioned, precedented exemption already used for other legitimate
security-regression fixtures (tests/windsurf-conversion.test.cjs,
tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs)
per DEFECT.PROMPT-INJECTION-SCAN-COLLISION.

---------

Co-authored-by: sim <sim@local>
2026-09-02 22:38:31 -04:00
Tom Boucher
91ed46882a feat(#3675): quick-batch core primitives and resumable manifest (#4190)
* test(#3675): add failing tests for quick-batch core primitives

Adds the full behavioral (tests/quick-batch.test.cjs) and property-based
(tests/quick-batch.property.test.cjs) coverage for #3675's quick-batch core
primitives per the phase's 35-row test matrix — task-list parsing (inline +
--file, with path-confinement/symlink-escape/non-regular-file rejection),
collision-safe quick-id preallocation under withPlanningLock, BATCH.json
schema/validation/resume, dependency-DAG + partitionByFileOverlap wave
construction, and exactly-once STATE.md completion (including the
STATE-row-written-but-manifest-not-yet-updated crash window).

The import target (gsd-core/bin/lib/quick-batch.cjs, compiled from a
not-yet-written src/quick-batch.cts) does not exist yet — every test in both
files fails at the top-level require() before any assertion runs. Five
fast-check properties cover collision-freedom under lock contention, resume
idempotency, exactly-once STATE completion, wave totality, and DAG-respecting
wave order, per the design doc's property-based-coverage requirement.

* feat(#3675): implement quick-batch core primitives

Adds src/quick-batch.cts (ADR-457 build-at-publish, compiled to
gsd-core/bin/lib/quick-batch.cjs) implementing #3675's quick-batch core
primitives per the phase design lock — pure/state primitives and
CLI-testable core operations only, no agent dispatch, no worktree creation,
no user-facing command (Phase 4/#3676's job):

- parseTaskList / parseTaskListFromFile: inline bulleted/numbered task-list
  parsing (>=2 items required) and a --file variant strictly confined to the
  planning workspace root via requireSafePath, rejecting non-regular-file
  targets.
- allocateQuickIds / createBatch: collision-safe YYMMDD-xxx quick-id
  preallocation under withPlanningLock, checked against both on-disk
  .planning/quick/ entries and sibling .planning/quick-batches/*/BATCH.json
  manifests (never on-disk-only, which would miss another in-flight batch
  that hasn't dispatched any real quick directory yet) — replicates
  cmdInitQuick's own grammar rather than delegating to it (that function's
  2-second granularity is not batch-safe).
- computeWaves: deterministic wave construction combining dependency-DAG
  layering with partitionByFileOverlap (#3674), called per DAG layer over
  path-separator-normalized planned_files — normalization happens at this
  module's boundary, never inside the Phase 2 helper.
- loadBatch: fail-closed BATCH.json schema validation (corrupt/truncated
  JSON, wrong types, missing fields, out-of-batch dependency references,
  dependency cycles, a worktree path absent from disk).
- resumeBatch: skips complete items, never auto-retries failed items,
  propagates/reverses blocked status along the DAG to a fixed point, and
  detects a STATE.md row that already exists for a non-complete item (the
  "STATE written, BATCH.json not yet updated" crash window) — completing it
  without re-appending. Idempotent across repeated calls.
- completeQuickItem / hasQuickTaskRow: exactly-once STATE.md completion —
  appendQuickTaskRow (unmodified) is called at most once per quick id, gated
  by hasQuickTaskRow's own idempotency check re-parsing the real "Quick Tasks
  Completed" table, since appendQuickTaskRow itself carries no idempotency.

BATCH.json lives at .planning/quick-batches/<batch-id>/BATCH.json, a sibling
of .planning/quick/ — never inside it, so scanQuickTasks never misreads a
batch manifest as a broken quick task.

* docs(#3675): register the new quick-batch module

New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained
registrations beyond the code itself: .gitignore (compiled artifact),
eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs),
docs/INVENTORY.md's CLI Modules roster row plus the regenerated
docs/INVENTORY-MANIFEST.json cli_modules entry, and a CONTEXT.md glossary
entry matching the convention set by the sibling File Overlap Partitioner
Module (#3674) entry it sits beside.

NOTE: docs/INVENTORY-MANIFEST.json was updated BY HAND (alphabetically
sorted single-entry insertion into families.cli_modules, matching the
existing file's structure) rather than via
`node scripts/gen-inventory-manifest.cjs --write` — this session's
MEMTRACE-FIRST guard hard-blocks direct execution of that indexed script
path from Bash, with no available Memtrace tool to route through instead.
The orchestrator should re-run `node scripts/gen-inventory-manifest.cjs
--check` to confirm this hand-edit is byte-identical to the generator's own
output before merging.

* fix(#3675): resolve lint findings in quick-batch primitives and tests

Unsafe `any[]` assignment from `new Array(n)` in the DAG cycle-check color
array, two unnecessary `as string[]` casts TS 5.5's inferred type
predicates already narrowed, raw `fs.rmSync` in test cleanup (needs the
Windows-EBUSY retry budget `helpers.cleanup` carries), an unused `loadBatch`
import, an unbounded `mkfifo` subprocess spawn missing a timeout, and a
CONTEXT.md glossary illustration that looked like a real file reference.

* feat(#3675): close acceptance-criteria gaps found in review

Standards- and spec-axis review (plus a self-caught race) surfaced real
gaps against issue #3675's own acceptance criteria and this repo's test
conventions:

- BATCH.json was missing options, base_revision, per-item wave, and
  per-item commit — the issue's AC explicitly lists all four as things
  the manifest must track. Added them: createBatch persists caller-supplied
  batchOptions/baseRevision verbatim and assigns each item its computed
  wave index; completeQuickItem now persists the commit onto the item,
  not just the STATE.md row. All four are backward-tolerant on load (an
  older/hand-built manifest without them still validates).
- resumeBatch had no "incompatible base divergence" check at all, despite
  the AC and the ADR's own "Base divergence" section requiring one. Added
  an opt-in currentBaseRevision comparison that fails closed with a
  recoverable diagnostic on mismatch, and touches nothing on refusal.
- resumeBatch read-modify-wrote BATCH.json OUTSIDE withPlanningLock — the
  only durable write path in this module that wasn't lock-protected,
  a real lost-update race against a concurrent completeQuickItem or
  another resume. Now runs inside the same lock createBatch/
  completeQuickItem use.
- loadBatch and collectExistingBatchQuickIds used raw JSON.parse with no
  size cap (security review, Low/informational); switched to the
  existing safeJsonParse (1MB cap) for defense-in-depth.
- Parser (parseTaskList) had only example-based tests; CLAUDE.md requires
  a fast-check property test for parsers. Added one plus a companion
  reject-property for <2 items.
- The id-exhaustion fail-closed ceiling (MAX_TIME_BLOCK) was untested at
  any boundary. Exported the pure allocateIdsGivenUsed/MAX_TIME_BLOCK for
  direct limit-1/limit/limit+1 testing without needing 46k fixture dirs.
- Issue AC explicitly asks for prompt-injection-payload test coverage,
  distinct from the existing shell-metacharacter test; added one.
- Test row 9 (FIFO skip) silently returned instead of calling t.skip(),
  so an unsupported platform would report a pass rather than a documented
  skip; fixed to bind the test-context param and skip properly.
- Extracted toWaveInput to remove a 2-site production duplication of the
  QuickBatchItem -> computeWaves reshape (Standards-axis smell).
- Added the required .changeset/ fragment (CONTRIBUTING.md: editing src/
  is user-facing even though the compiled .cjs is gitignored).

* fix(#3675): restore "not valid JSON" wording in loadBatch's parse-failure reason

gsd-test caught this: switching loadBatch to safeJsonParse changed the parse-
failure message shape ("... parse error — ...") without preserving the
"not valid JSON" substring row 27's own test asserts on. Re-wrap
safeJsonParse's error into the original diagnostic phrasing regardless of
which of its three failure modes fired.

* docs(#3675): backfill changeset pr number to 4190

* fix(#3675): detect a silently-no-op mkfifo on Windows, not just a throwing one

CI caught this on windows-latest: row 9's platform-skip only caught mkfifo
throwing (command not found). On this runner mkfifo resolves to something
that exits 0 without creating a file (NTFS has no FIFO concept), so
execution fell through to parseTaskListFromFile against a path that
doesn't exist, producing an ENOENT stat error instead of the expected
"not a regular file" rejection. Check the artifact actually exists before
trusting a zero exit code, and skip with a documented reason either way.

---------

Co-authored-by: sim <sim@local>
2026-09-02 15:38:27 -04:00
Tom Boucher
b0572c0108 feat(#3674): extract shared file-overlap wave partitioner (#4166)
* test(#3674): characterize existing wave-dispatch output and add tests for the extracted partitioner

Pins resolveWaveDispatch's and emitWorkflowScript's current, unextracted
output (chain-overlap, disjoint-empty-set, and a multi-wave/multi-stage
golden script) as a regression safety net ahead of extracting
partitionStages into a standalone module. Also adds the new module's
unit and property tests (test matrix rows 1-11) against its expected
public API, which does not exist yet and is added in the next commit.

* feat(#3674): extract file-overlap partitioner into a shared, generic module

Moves partitionStages' greedy first-fit file-overlap algorithm into a new,
dependency-free src/file-overlap-partitioner.cts module (partitionByFileOverlap),
generalized over a plain {id, files}[] shape rather than claude-orchestration.cts's
Plan/Wave interfaces. partitionStages becomes a thin adapter mapping its own
Plan[] shape onto the generic input and back — behavior-preserving, no dependency
ordering, no path normalization, no filesystem access moved or added. Enables a
future consumer (quick-batch, #3675 / ADR-1239) to reuse the same primitive
without pulling in orchestration internals.

* docs(#3674): register the file-overlap-partitioner module bookkeeping

New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained
registrations beyond the code itself: .gitignore (compiled artifact),
eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs),
docs/INVENTORY.md's CLI Modules roster row (regenerated via
gen-inventory-manifest.cjs --write), and a CONTEXT.md glossary entry
matching the convention set by similarly-scoped leaf modules
(text-lines.cts, plan-dependency-graph.cts, spec-section.cts).

* fix(#3674): alphabetize INVENTORY.md row, manifest regen no-op, fast-check import already correct

- docs/INVENTORY.md: move file-overlap-partitioner.cjs row to alphabetical position
- docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest.cjs --write,
  produced no diff (manifest is keyed by content, not row order)
- tests/claude-orchestration.test.cjs's direct require('fast-check') is correct as-is:
  tests/helpers/fast-check-setup.cjs's own docstring scopes the shared-seed wrapper to
  "every *.property.test.cjs file"; claude-orchestration.test.cjs is not a .property.test.cjs
  file, and every .property.test.cjs file sampled uses the wrapper consistently. No outlier.

* fix(#3674): constrain the no-overlap property test to unique ids, fixing an ambiguous duplicate-id reconstruction

The `no two plans in the same stage share a modified file` property
reconstructs which physical item produced each output id via
`remaining.findIndex(r => r.id === id)`. Under duplicate ids (an
explicitly-supported input shape for `partitionByFileOverlap`) that
reconstruction can pick the wrong physical occurrence, producing a
false-positive overlap failure (observed counterexample: p0(f1),
p208(f1), p208([]) — correctly staged as [[p0,p208#2],[p208#1]], but
misread by id-order as [[p0,p208#1],...], which do overlap).
Properties (a) determinism and (b) totality already exercise
duplicate ids correctly and are left unchanged; only this property's
generated items are now constrained to unique ids via
`fc.uniqueArray`, where the reconstruction is unambiguous.

---------

Co-authored-by: sim <sim@local>
2026-09-02 07:33:42 -04:00
Cody Anderson
8c9265d4e5 fix(#3724): warning-only Dimension 3b findings no longer force the revision loop (#3758)
* fix(#3724): stop advisory Dimension 3b findings from forcing the revision loop

Dimension 3b (undeclared/temporal coupling, #1954) is spec'd "never a
blocker" but tagged severity: warning — the tier plan-phase's revision
loop counts as must-fix — and the planner is never taught the rule, so
every multi-wave phase touching shared mutable state replans at least
once, and intentionally coupled plans re-flag identically every
iteration to the stall prompt.

Three coordinated changes:
- gsd-plan-checker: retag 3b to severity: info, the tier
  references/revision-loop.md already exempts by design; recognize a
  coupling_justified frontmatter declaration in the Do-NOT-flag list so
  deliberate pairs converge. Additions are offset by trimming 3b
  motivation prose — the checker sits 45 bytes under its LARGE hard cap.
- plan-phase step 12: INFO-only accept — an issues block with zero
  BLOCKER/WARNING entries accepts the plan and surfaces the advisories
  instead of re-entering the revision loop. Real blockers and warnings
  still gate unconditionally.
- gsd-planner: slim pointer in assign_waves to the new
  progressive-disclosure reference gsd-core/references/planner-coupling.md
  (the planner sits 19 chars under its own cap), which carries the
  shared-mutable-state rule and the coupling_justified escape hatch so
  first-pass plans avoid the finding when the coupling is unintentional.

Documented the coupling_justified field in docs/reference/plan-md.md.
Growth acks per #2914; inventory manifest and install-tree fixtures
regenerated for the new reference file.

Closes #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin Dimension 3b at severity: info

The severity retag makes the old assertion (severity: warning) stale;
lock the advisory tier from both directions — info must be present,
warning must not — so a future edit cannot silently re-arm the
revision-loop trigger.

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* chore(#3724): changeset fragment for PR #3758

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* docs(#3724): roster planner-coupling.md in docs/INVENTORY.md

The new reference was enumerated in the manifest and all 19 install-tree
fixtures but missing its row in the Modular Planner Decomposition table —
the roster half the manifest-sync test cannot check. (Review Blocker.)

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): cover all four acceptance criteria (review round 1)

- plan-checker-coupling: the 3b severity assertion is now a PARITY check
  deriving the exempt tier from revision-loop.md's flow instead of
  hardcoding info — editing either side alone reds the suite. New
  describe pins the other three criteria: plan-phase's INFO-only accept
  clause (proven failing-first), the BLOCKER + WARNING count staying
  intact, the coupling_justified Do-NOT-flag exemption + fix_hint, and
  the planner pointer + planner-coupling.md content.
- ack fragment: $comment's plan-phase figure corrected to +79B; the 2775
  pin note carried forward into the gsd-planner.md entry, updated for
  upstream's #3761/#3764 Rule-paragraph anchor (which this diff leaves
  verbatim).

The parallel-dependent-plans re-anchor this commit originally carried was
superseded by upstream #3764 during review; this branch no longer touches
that file.

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 2 — align the stance enumeration, complete the template contract

MAJOR: <adversarial_stance>'s severity enumeration gains the INFO bullet so it
agrees with Dimension 3b's 'ALWAYS INFO' mandate instead of contradicting it.
Funded by extracting the inline <examples> block to the new progressive-
disclosure reference gsd-core/references/plan-checker-examples.md (@-inlined
from the same spot; #1949 precedent), which also restores the 3b motivation
clause round 1 traded away (Nit 4) and nets the agent file SMALLER than base
(49107 -> 48486) — the extraction the byte pressure was owed.

MINOR: gsd-core/templates/phase-prompt.md now carries coupling_justified, and
the field's shape becomes one 'plan-id: reason' string per coupled peer so a
plan justified against two peers can express it; docs/reference/plan-md.md's
Type column names the shape.

NIT: the 3409 ack's plan-phase entry no longer calls the #1168 workflow
ratchet an 'XL tier'.

Acks and derived artifacts updated accordingly (checker entry removed — a
shrink needs no ack; INVENTORY roster row + regen:derived for the new file).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): derive the 3b negative severity assertion (review round 2)

Every severity token in the 3b span must BE the tier revision-loop.md exempts,
replacing the hardcoded severity:warning negative — if the loop's exemption
ever moves, the failure names the real conflict instead of blaming the agent
file with a mutually-unsatisfiable pair.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): refit the planner coupling pointer under the char cap

Upstream #3299 (PR #3390) grew agents/gsd-planner.md to 49146 chars at the
base, leaving 5 chars of headroom where the +16-char pointer was measured
against 13 more. The pointer prose shortens to 'Non-file coupling:' —
49150 chars, back under the strict 49152-char cap — and the ack figures
follow. The @-path the tests pin is unchanged.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): re-home the plan-phase ack after the #3823 spent-fragment sweep

Upstream #3078/#3823 deleted all fully-spent ack fragments, including
3409-unreachable-guard-arms.json, which carried this PR's plan-phase.md
+79B append. Per the collision remedy that sweep added: take the deletion
and home the still-live entry in this PR's own fragment. Figures
re-measured at this merge base (90871 -> 90950 LF bytes).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): absorb the spent #3172 plan-phase fragment into this PR's ack

Upstream #3825 shipped 3172-stated-failing-direction.json naming only
plan-phase.md, now spent at the base — colliding with this PR's live
plan-phase entry. Per the #3003 pattern the fully-spent single-path
fragment is deleted and this fragment stays the path's one source;
figures re-measured at this base (93073 -> 93152 LF bytes).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 3 — true up the ack figures, restore the wave comment

The fragment's absolute sizes are re-measured and anchored to base
e40e9670 (planner 47259 -> 47330 chars, checker 45537 -> 44916 B,
plan-phase 91186 -> 91265 LF bytes), with a note that absolutes rot as
next moves — the deltas are the durable claims. The round-1 removal of
the '# Implicit dependency: files_modified overlap forces a later wave.'
pseudocode comment offset headroom base drift had already returned, so
it is restored (findings 2-3). Changeset gains the (#3724) backlink
(finding 4).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 4 — close the verify-work surface, harden the boundaries

BLOCKER: verify-work.md's verify_gap_plans is the second multi-plan
consumer of the checker's sentinels, and its ISSUES FOUND handler entered
revision_loop with zero severity parsing — the guaranteed replan #3724
fixed in plan-phase, alive on the gap-closure surface. The handler now
counts BLOCKER + WARNING and accepts INFO-only returns with advisories
displayed. The checker's INFO stance bullet is reworded to the claim that
is true everywhere ('revision gates count only BLOCKER + WARNING').

Minor 1: plan-phase's iteration_count >= 3 arm recounts severities, so an
INFO-only third check accepts instead of halting on a '0 issues remain'
user gate. Minor 2: the coupling_justified exemption now requires the
entry to NAME the other plan, closing the blanket-suppression reading.
Nit 1: INVENTORY row states the extraction buys cap headroom, not context.
Nit 2: the advisory display gains a concrete format on both surfaces.

Ack fragment re-anchored at base ddde001a: verify-work.md +264B (new
entry), plan-phase.md +395B, checker still net negative (-512B).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin the verify-work accept and the iteration-cap boundary (review round 4)

Two wiring assertions: verify_gap_plans' ISSUES FOUND handler gates on
BLOCKER + WARNING and accepts INFO-only blocks, and plan-phase's
iteration_count >= 3 arm recounts severities instead of gating advisories
— the limit+1 boundary of the gate this PR fixes.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 5 — fail closed at the gates, surface the advisory

Blocker 1: the checker's step-10 status rule routes an INFO-only result to
## ISSUES FOUND (with a new ### Advisories (info) template section and a
severity-aware recommendation) so the orchestrator receives the block and
displays the advisory instead of silently accepting a bare PASSED.

Blockers 2+3: all three gate surfaces (plan-phase step 12 both arms,
verify-work verify_gap_plans) carry one canonical clause verbatim — an entry
whose severity is missing or unrecognized counts as a BLOCKER (fail closed) —
making the accept condition an explicit-INFO whitelist while keeping
issue_count coherent for stall math.

Major 1: the INFO stance bullet scopes its claim to the plan-phase and
verify-work gates (quick mode's loop still revises on any ISSUES FOUND).
Major 2: INVENTORY row and ack $comment state the extraction's real trade
(readability, +0.6 KB eager runtime context), not a cap remedy.
Minor 1: plan-md.md marks coupling_justified as prompt convention, unvalidated.
Nit 1: ack absolutes re-anchored at base 1e67ec97; checker now +120B and acked.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin the round-5 contract — fail-closed parity, INFO-only return shape

New: three-surface verbatim parity test for the fail-closed clause (Blockers
2+3); checker return-contract test for the INFO-only ## ISSUES FOUND route and
advisories section (Blocker 1). All seven newly pinned tokens are absent at
f3a5682d, so each new assertion fails pre-fix.

Updated: accept-clause regexes track the explicit-INFO whitelist wording;
the severity sweep scopes to the span's fenced yaml examples via
yamlSeverityTiers (round-5 Minor 3, applied to the blocker negative too);
the iteration-cap comment states it is a prose pin, not an executed boundary
check (Minor 4); splitLines call sites document the line-pin coupling (Nit 2).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): adopt next's line wrap in the 3b motivation clause — drops a wrap-only hunk from the diff

Byte-identical content; the wrap difference was an artifact of the round-1
base adaptation predating upstream's #3003 landing.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 7 — gate every checker consumer, not just the two audited ones

Blocker: quick/steps/plan-checker-loop.md (issue-named in #3724) gets the
same canonical fail-closed clause and explicit-INFO whitelist accept as
plan-phase/verify-work — an INFO-only result proceeds instead of entering
quick mode's revision loop.
Major: import.md plan_validate handles the checker return by severity
(INFO-only never blocks an import) and is added to agent-contracts.md's
consumer enumeration, which had omitted it.
The checker's INFO stance bullet drops the quick-mode carve-out — the claim
is universally true again now that every consuming gate is severity-aware.
Minor: an applied coupling_justified exemption is surfaced as its own info
advisory so a stale one-sided declaration stays observable.
Nit: plan-phase's revision-iteration Display line is explicitly conditioned
on not having already proceeded to step 13.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM
Emitted-Drift-Ack-Growth: import.md — #3724 round 7: the plan_validate step's checker-return handler becomes severity-aware — counts BLOCKER + WARNING failing closed and accepts an explicitly-INFO-only return with advisories displayed instead of blocking the import

* test(#3724): pin the round-7 surfaces — five-gate parity, quick/import accepts, exemption visibility

The verbatim fail-closed parity test extends to quick/steps/plan-checker-loop.md
and import.md plan_validate; new assertions pin quick mode's INFO-only proceed,
import's never-blocks accept, import.md's presence in agent-contracts.md's
consumer row, and the surfaced coupling_justified exemption advisory. All four
newly pinned token families are absent at the pre-fix head, so each new
assertion fails first.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-31 19:08:11 -04:00
Tom Boucher
6beaa66b25 enhance(#3304): gate re-verification blockers on deterministic evidence (#4085)
* test(#3304): add failing-first suite for the convergence evidence gate

Content-assertion suite for the Step 7 re-verification evidence gate
(agents/gsd-verifier.md / gsd-core/references/verifier-evidence-gate.md).
Committed before the implementation to prove RED via gsd-test.

* enhance(#3304): gate re-verification blockers on deterministic evidence

Step 7's anti-pattern scan re-runs at full, unbounded scope on every
re-verification pass, independent of the must-haves established in Step 2.
A blocker it finds — other than the self-evidencing debt-marker check —
previously reverted a completed gap-closure round and started another
--gaps cycle on nothing more than the verifier's own new judgment call,
with no bound on how many times that could repeat.

A Step 7 blocker now blocks unconditionally in re-verification mode only
if it is a carried-forward gap (present in the prior VERIFICATION.md's
gaps: list) or the flagged file was git-modified since the prior pass
(a regression; fails closed toward blocking when history is unresolvable).
Otherwise it predates the gap-closure round unflagged and needs
deterministic evidence — a named test run red, or another concrete
reproducible artifact — to stay blocking. Unevidenced, it downgrades to
a new advisory: frontmatter list and report section instead of setting
status: gaps_found, and never reverts a completed must-have.

Maintainer approval was narrowed to this evidence condition only,
explicitly rejecting the broader "advisory whenever untraceable to a
requirement/decision/prior-gap" proposal — implemented and pinned by
tests/verifier-evidence-gate.test.cjs and documented as rejected in
gsd-core/references/verifier-evidence-gate.md so it can't silently
re-expand.

Closes #3304

* fix(#3304): correct window-truncation and indentation bugs in evidence-gate tests

gsd-test's GREEN checkpoint caught 3 real bugs in the test file itself
(not the production prose): a {0,600} match window was shorter than the
724-char paragraph it was scanning (the "exclude from Step 9 Rule 1"
phrase starts at offset 662), and two regexes assumed no indentation
after a markdown list-continuation line break. All three phrases are
confirmed unique across agents/gsd-verifier.md, so the windowed
submatches are replaced with direct whole-string assertions instead of
just widening the window.

Also acknowledges the deliberate byte growth in agents/gsd-verifier.md
that the differential-attribution check (ADR-2719) correctly flagged.

Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the #3304 re-verification evidence gate (Step 7 rule, Advisory bucket, advisory: frontmatter, report section); 1488 bytes, still within the LARGE-tier 48 KiB cap (48751/49152).

* docs(#3304): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-30 16:51:52 -04:00
Dennis Kim
8487f0ed42 enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage

- pin configured, absent, and malformed branch-list behavior
- require opposite CLI and execute warning outcomes

* feat(01-01): warn on configured protected branches

- resolve the base branch union configured protected branch names
- expose exact boolean CLI comparison output for workflow callers
- keep execute-phase warning advisory and within its byte budget

* test(01-01): add failing protected branch config coverage

- cover valid list persistence and null unset
- reject hostile shapes while preserving the prior value

* feat(01-01): validate protected branch configuration

- register git.protected_branches as a canonical config key
- require a non-empty array of non-blank branch names

* test(01-02): add failing ship protected-branch controls

- Execute both workflow warning blocks with exact predicate arguments
- Require true and false results to produce opposite warning outcomes
- Preserve the none-strategy feature-branch offer contract

* feat(01-02): warn at ship on protected branches

- Reuse the typed protected-branch predicate in ship preflight
- Keep raw base resolution for PR targeting and advisory branch creation
- Prove execute and ship warning blocks with opposite-result controls

* test(01-02): add failing protected-branch docs parity

- Require the canonical schema key in both English config references
- Pin the non-empty string-array type and absent default
- Require synchronized multi-branch examples and advisory semantics

* feat(01-02): publish protected branch configuration contract

- Document the optional non-empty string-array field in both references
- Explain resolved-base union and absent-field compatibility
- Keep execute and ship warnings advisory under branching_strategy none

* fix(01): CR-01 honor active workstream branch policy

* fix(01): WR-01 assert protected config path selection

* docs: add changeset fragment for #3648

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx

* fix(#3648): resolve base_branch precedence inversion and round-1 findings

Blocker 1/2: production config resolution was flat-first, so a project
that migrated to git.base_branch but still carried a stale flat
base_branch got the old value back. Add base_branch to
normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos
pattern: canonical nested wins) and route readEffectiveGitConfig's
test seam through the same normalization so it can't silently diverge
from production again. Adds a regression test with both keys set that
fails without the fix.

Blocker 3/4/5: restore the handle_branching case-selector prose and
"none" contract sentence that #3389's tests anchor on, and revert the
unrelated prose/comment compaction in the same step — both were
drive-by edits outside #3552's scope.

Also addresses review majors/minors: delete readConfigBaseBranch and
readConfigProtectedBranches (dead in production, only self-tested);
--is-protected now fails closed (reports protected) instead of
silently answering false when the base branch can't be verified;
trim configured protected-branch names; fix HOME-without-USERPROFILE
vacuous isolation on Windows; correct the drift-ack's byte accounting.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N

* test(#3648): add failing legacy-key hoist safety coverage

Round-2 review found normalizeLegacyKeys block 5 records a normalization
carrying the DISCARDED flat value on the canonical-wins branch. Probing
that turned up a second, unreported defect in the same helper shape:
blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no
object guard, so a config whose section key holds a string is spread into
index keys —

  {"git":"main","base_branch":"release"}
    -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}}

The resolved value is accidentally still correct, so nothing fails and no
diagnostic fires. But normalizations.length > 0 sets configDirty, and
config-loader then serializes that shape back into the user's
config.json — a read that silently corrupts config.

The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"}
explicitly; this is the input it would have caught.

Covers both defects across blocks 1 and 5, with object/array/null
negative controls that must stay green in both phases, and a fast-check
property over arbitrary `git` values.

* test(#3648): pin fail-closed handling of malformed protected_branches

Replaces the test that pinned the fail-OPEN behaviour. The old
assertion — ['develop', 42] yields isProtected === false for 'develop' —
locked in the exact failure #3552 exists to close: config-set validation
is bypassable by a direct edit of .planning/config.json, so a user who
believes 'develop' is protected got a silent false and no warning.

It was also inconsistent with the fail-CLOSED direction twelve lines
away, where an unverified base reports protected and writes a
diagnostic. A protection predicate must not have two opposite failure
directions depending on which input is bad (#3648 review Blocker 3).

New coverage: a bad element drops only itself, a non-array contributes
no names, an empty list is well-formed rather than malformed, and
--is-protected surfaces the rejection. Both negative controls — a clean
list reports nothing rejected and writes no diagnostic — must stay green
in either phase, so the reject channel cannot fire unconditionally.

* fix(#3648): drop only invalid protected_branches and report them

Partition git.protected_branches instead of discarding the whole list on
one bad element, and carry the rejections out through
ProtectedBranchStatus so --is-protected can name them on stderr. Valid
names keep protecting; the user finds out the rest were ignored.

A non-array value still contributes no names — a bare string is not a
list of branch names — but is now reported rather than swallowed. An
empty array stays silent: declaring no extra protected branches is a
valid choice, not a misconfiguration.

writeDiagnostic is hoisted out of the unverified-base branch since both
arms now use it.

* test(#3648): prove the predicate diagnostic survives both call sites

The workflow bash stub now emits a stderr diagnostic the way the real
command does, which is what makes a swallowed `2>/dev/null` visible to a
test — previously the stub was silent on stderr, so discarding it changed
no observable behaviour and the call sites could drop the explanation
undetected.

Adds the Minor 2 binding check as well: ship must expose the predicate
result as IS_PROTECTED rather than only echoing a warning, asserted by
running the extracted bash and reading the bound value, not by grepping
the workflow source.

Both tests carry opposite-outcome controls — an empty diagnostic must
leave the text absent, and a false predicate must bind false.

* fix(#3648): surface the predicate diagnostic and bind ship's result

Drop `2>/dev/null` from the --is-protected call at both call sites. The
fail-closed explanation and the new rejected-entry warning both go to
stderr, so discarding it left the user with a bare "protected branch"
warning on a branch that is not protected and no way to tell a real
match from a degraded-git guess. `git branch --show-current` keeps its
own redirect — that one is genuine noise.

ship.md binds IS_PROTECTED and its prose now branches on the variable,
so the following steps have evaluable state instead of having to infer
it from warning text in tool output.

execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth
319 bytes (was 331 before the redirect came out). Baseline re-verified
against the current rebase base by blob id; the ceiling check passes
with 755 bytes of margin.

* test(#3648): restore negative space for the readFile config seam

The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm
it pinned survives verbatim in readEffectiveGitConfig's readFile branch —
the JSON.parse catch, the non-object guard, the git-section object guard,
.trim() and blank-string rejection — and the four surviving readFile
injections were positive-path only. protected_branches was never driven
through this seam at all.

Restores nine cases against the seam, including protected_branches
partitioning, plus a control proving loadConfig still wins when both
seams are supplied.

Records honestly what the suite pins. Mutating the built lib shows
.trim() is KILLED, while the non-object guard and the blank-string
rejection SURVIVE — both are unreachable through this entry point for
the same reasons the deleted suite documented against its own
equivalents: a JSON-parsed non-object carries no relevant own-property
either way, and a blank value is rejected a second time downstream by
the resolver's truthiness check. They stay as defence-in-depth and are
labelled known-unkillable rather than left looking like coverage this
suite does not provide.

* test(#3648): distinguish detached HEAD from a missing branch argument

`args[1] ?? ''` collapsed two different situations into one: a detached
HEAD, where `git branch --show-current` legitimately prints nothing, and
the flag being called with no argument at all. Both answered false, so
the right outcome arrived by an unintentional path and a caller bug was
indistinguishable from normal operation.

Asserts the detached case stays silent and the missing-argument case
reports, with a control that the two diagnostics differ.

* fix(#3648): report a missing --is-protected branch argument

Answer false either way, but say so when the flag arrives with no
argument. A detached HEAD passes an explicit empty string and stays
silent, since that is a normal state rather than a misconfiguration.

* docs(#3648): state exact-name matching and per-entry rejection

isProtected is exact string equality, so a git-flow project must
enumerate every release/* and hotfix/* by name. #3552 only asked for an
integration-branch field, so the implementation satisfies the letter of
the issue while leaving its git-flow motivation partly unserved — say so
where users will meet it rather than leaving them to discover it.

Also documents the Blocker 3 behaviour change: an invalid entry is
ignored with a warning naming it and the remaining names still apply.

Both statements land in docs/CONFIGURATION.md and
gsd-core/references/planning-config.md, and the config-field-docs parity
test asserts each in both so the two cannot drift.

* refactor(#3648): extract isValidProtectedBranches for cross-surface pinning

The `git.protected_branches` check inside `cmdConfigSet` and the resolver's
per-entry filter in `git-base-branch.cts` are deliberately different shapes —
all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot
fail the guard open. Nothing structural keeps their two definitions of "usable
branch name" in step.

Lifting the write-side check into a named, exported predicate lets a property
test ask both surfaces about the same value and assert they agree, which is the
fast-check gap the round-2 review flagged. No behaviour change: the predicate is
the same expression, called from the same place.

* fix(#3648): stop --is-protected rewriting the config it is asking about

`gsd_run query git.base-branch --is-protected` runs on every execute-phase and
every ship. It resolved config through `loadConfig`, whose normalize-then-write
path rewrites `.planning/config.json` whenever any legacy key normalizes — so a
boolean question was silently editing the user's checked-in config. This PR had
widened the trigger by adding a fifth normalization block (top-level
`base_branch` -> `git.base_branch`), making it fire for exactly the projects the
feature targets.

`loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution
is unchanged, only the two write-back side effects are suppressed. The predicate
passes `persist: false`; the ~30 other callers are untouched, so a legacy config
is still migrated by ordinary use.

Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and
reflows whitespace even when the values are equivalent. Three tests, each with
its own control: the end-to-end CLI leaves the file byte-identical while still
answering `true` from the legacy key (proving the config WAS read); an ordinary
persisting load of the same fixture DOES change the bytes (proving the fixture
is live rather than inert); and `persist:false` vs default over one directory
returns deep-equal config while differing on the write. Reverting the one-line
`persist: false` fails the first of those and only that one.

Also from the review:

- `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through
  the same precedence authority production uses". It does not, and cannot — it
  reproduces two of production's steps over a single file. The comment now names
  what the seam covers and what it does NOT (root/workstream deep merge, builtin
  and global defaults, federated merge), and the seam now applies production's
  flat-then-nested lookup so it stops disagreeing about a surviving flat key.

- The missing-argument diagnostic promised "answering false", which the
  fail-closed guard on the same call can contradict by printing `true`. It now
  states what it did with the argument and leaves the answer to stdout.

* test(#3648): re-pin block 5 on #3760's refusal contract

#3767 landed on next while this PR was in review and fixed the non-object
config-section defect properly: a present-but-non-object section now BLOCKS its
own migration — value preserved, no Normalization pushed, refusal reported via
`skipped[]` — rather than being rebuilt from a plain-object view. That supersedes
this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread
but still dropped the section value silently, and which the round-3 review
correctly called out as destruction in place of corruption. The rebase drops that
commit and routes block 5 through the upstream helper.

This file's tests asserted the superseded design, so they are rewritten to pin
block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's
suite was written — against the contract that now governs it: ordinary hoist into
an absent/null/object section, canonical-nested-wins, and refusal for each of
string/number/boolean/array sections with the exact `skipped` entry.

Two controls keep it from passing vacuously: the refusal must be scoped to block
5 (an unrelated block still normalizes in the same call), and a property over
arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually
exclusive per key, that a refusal leaves both the section and the legacy key
untouched, and that a hoist manufactures no index key the input did not carry.

* docs(#3648): correct the Git Query and Config Loader module contracts

CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct
`.planning/config.json` read. Since this PR it is the EFFECTIVE configuration
resolved by the Config Loader — a materially different authority, carrying the
root/workstream deep merge, flat-then-nested lookup and builtin/federated
defaults. The `--is-protected` predicate, `git.protected_branches`, and the two
invariants that distinguish the predicate from the plain query (fails closed on
an unverified base; must not write) were undocumented entirely.

The Config Loader entry now states that loading is not side-effect-free by
default and documents `options.persist`.

docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and
no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write`
was run and produced no diff: the manifest indexes roster NAMES, not row prose,
so a description edit cannot move it.

Also closes the global-defaults minor: `git.protected_branches` is inert in
`~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key
appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is
section-wide and predates this PR, so the fix is to state the scope where users
meet it rather than to quietly extend the resolution set for two new keys.

* fix(#3648): close four defects found by the round-4 external review

Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially
against this branch. Four findings reproduced against source; each is fixed with a
failing-first test and a control, and each fix was verified by reverting it and
watching exactly the intended test fail.

1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was
   only half closed. `loadConfigResolved` re-enters itself with a bare
   `{ workstream: null }` when a workstream has no config.json of its own, and
   that literal discarded every other option — so the recursive pass ran at the
   DEFAULT persistence and rewrote the ROOT config. Reproduced: with
   GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected`
   rewrote `.planning/config.json` despite `persist:false`. Both recursions now
   forward `options` and override only `workstream`; the explicit override still
   wins the hasOwnProperty check, so spreading cannot let `workstreamContext`
   reintroduce a workstream.

2. Both workflow call sites failed OPEN, and aborted under `set -e` (both
   reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty
   string when the query fails, so `[ "$X" = true ]` was simply false: no
   warning, no trace — a silent hole in the guard whose only job is to warn. The
   bare assignment also aborted the step under `set -e`. Both sites now degrade
   VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the
   check did not run. Deliberately not fail-closed — claiming "protected" on no
   evidence would warn on every branch whenever gsd-tools is unavailable.

3. `isValidProtectedBranches` and the resolver disagreed on a sparse array
   (antigravity). `.every()` skips holes; the resolver's `for...of` yields
   `undefined` for them, so `["main", , "develop"]` was accepted by config-set
   and rejected by the resolver. The cross-surface property passed only because
   `fc.array` cannot generate a hole. The predicate now indexes, and the
   generator punches holes so that axis is actually falsifiable. JSON cannot
   express a hole, so this is unreachable in production — but two definitions of
   one predicate must not contradict each other.

4. A top-level `protected_branches` silently outranked `git.protected_branches`
   (antigravity). Routing the key through `get(key, {section, field})` gave it
   flat-then-nested precedence, which is back-compat for keys
   `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has
   no legacy form, so that invented an undocumented alias. It now resolves
   nested-only through a new `getNested`, in production and in the test seam.
   `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's
   refusal path can leave behind — and a control pins that distinction.

Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed
only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`);
a git command that runs and exits non-zero counts as a clean negative, so a cwd
that is not a repository answers `false`, not `true`. Verified pre-existing on
next @ 738f42f4, so the documentation was over-claiming rather than the code
regressing — but an over-broad contract is exactly what the module docs must not
carry.

Both workflow byte figures re-derived after the call-site change:
execute-phase.md 92356 -> 92865 (+509), ship.md 36784 -> 37227 (+443).

* test(#3648): pin git config read parity

* docs(#3648): document git query contracts

* fix(#3648): expose protected branch default

* test(#3648): snapshot planning tree for read-only query

* test(#3648): pin planning snapshot stray-write detection

* fix(#3648): resolve merge conflict from #3078's ack-fragment sweep

next swept the fully-spent 2818/3003 ack fragments this branch had
appended to (#3078, a84f7563). Rebased onto upstream/next and took
the deletions on both, then moved the #3552 append into a new
fragment of its own.

Rebasing onto the current base also left execute-phase.md only 34
bytes under the frozen ADR-857 Phase 6 margin ceiling (93400 bytes) —
intervening next PRs consumed the rest while this PR was in review.
Extracted the "none" arm's protected-branch-warning bash block into
gsd-core/workflows/execute-phase/steps/protected-branch.md (content
unchanged, matching the existing steps/ extraction pattern used
elsewhere in this file) so the inline growth is a one-line pointer
instead of the full block. 93366 -> 93385 bytes (+19), 15 bytes
inside the ceiling.

* fix(#3648): drop stale ack entry for the new step file

The extracted execute-phase/steps/protected-branch.md needed no
acknowledgment of its own — the differential-attribution check flagged
the entry as stale once the build ran, so removed it and kept the two
growth entries (execute-phase.md, ship.md) that actually needed one.

* fix(#3648): follow the step-file reference in the bash-extraction test helper

extractProtectedBranchWarningBash() read the "none" arm's bash block
directly out of execute-phase.md. That block now lives in
execute-phase/steps/protected-branch.md (byte-ceiling extraction);
the helper follows the step-file reference and extracts from there
when no inline block is found, so the three execute-phase tests that
execute this bash for real keep exercising the actual behavior.

* fix(#3648): regenerate INVENTORY-MANIFEST.json and satisfy the CRLF-fragile lint rule

- gen-inventory-manifest.cjs --write to pick up the new
  execute-phase/steps/protected-branch.md entry (already covered by
  docs/INVENTORY.md's generic workflow_steps wildcard row, so no
  INVENTORY.md edit is needed).
- Reworked the step-file-reference lookup in
  extractProtectedBranchWarningBash() to avoid a bare-\n regex split
  on file content (local/no-crlf-fragile-split), using the same
  line-array scan the function already uses elsewhere.

* fix(#3648): regenerate golden install-tree fixtures for the new step file

npm run gen:install-tree, adding gsd-core/workflows/execute-phase/
steps/protected-branch.md to all 19 runtime install-tree fixtures.
CI's tests/golden-install-tree.test.cjs caught this on push — I'd
verified the differential-attribution and INVENTORY-MANIFEST checks
but missed this separate golden-fixture check for the new file.

* fix(#3648): add the canonical gsd_run preamble to the new step file

CI's runtime-launcher-parity suite requires exactly one canonical
resolver preamble in every workflow .md that calls gsd_run. The
inline "none"-arm block never needed one (execute-phase.md already
carried a preamble elsewhere in the same file), but the extracted
execute-phase/steps/protected-branch.md is now its own file with no
preamble of its own. Ran node scripts/sync-runtime-launcher.cjs to
insert it (execute-phase.md itself is untouched — still 93385 bytes,
inside the ADR-857 ceiling).

That preamble defines its own gsd_run(), which shadows the mock
tests/git-base-branch.test.cjs injects for the three #3648 tests that
execute this bash for real — without stripping it, those tests reached
the real gsd-tools.cjs on the machine running them instead of the
test's fixture. Preamble correctness is already covered by
tests/runtime-launcher-parity.test.cjs, so extractProtectedBranchWarningBash()
now strips the preamble line before handing the block to the harness;
it only needs to exercise the #3552 warning logic.

* fix(#3552): address PR 3648 review feedback on protected branch warnings

- Fix execute-phase handle_branching branching_strategy=none instruction
  to "Read and execute execute-phase/steps/protected-branch.md"
- Use io.error(..., ERROR_REASON.USAGE) for cmdGitBaseBranch usage errors
- Align git.protected_branches schema default to (none) without fallback []
- Relocate CONTEXT.md forward-referencing sentence into module body
- Sanitize control and ANSI characters in renderRejected diagnostics
- Clean up out-of-scope whitespace hunks in gsd-tools.cjs

Emitted-Drift-Ack-Growth: execute-phase.md — #3552: execute-phase handle_branching adds a pointer to execute-phase/steps/protected-branch.md for branching_strategy=none so the protected-branch check executes while keeping execute-phase.md within the ADR-857 Phase 6 margin ceiling (93400 bytes). 93392 bytes, 8 bytes inside the ceiling.
Emitted-Drift-Ack-Growth: ship.md — #3552: ship preflight step 3 now asks the same typed git.base-branch --is-protected predicate as execute-phase, binding IS_PROTECTED and warning without refusing execution or blocking the branching_strategy=none feature-branch offer; it degrades visibly (rather than silently reading an empty result as "not protected") when the query itself fails to run. 36841 bytes, well inside the XL cap (98304, tests/workflow-size-budget.test.cjs).

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-29 17:00:52 -04:00
Tom Boucher
dd4f179672 feat(#3970): per-task external-tracker content-resolution seam (#4000)
* feat(#3970): per-task external-tracker content-resolution seam

Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute
plus a new optional `taskContentResolver` capability-manifest field let a
capability resolve a task's action/verify/acceptance-criteria/read_first/done
content from an external issue tracker instead of PLAN.md's inline body.

- src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId`
- src/task-content-resolution.cts: new leaf module — split/find/build/resolve,
  with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed
  resolution, never a silent fallback to possibly-stale inline text
- src/task-command-router.cts: new `task resolve-content --plan --task-id --raw`
  CLI verb wiring the module into a real process exit code
- gsd-core/bin/lib/capability-validator.cjs: validates the new
  `taskContentResolver` manifest field (feature-role only, cross-capability
  trackerPrefix uniqueness)
- gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md,
  docs/reference/capability-manifest.md: wire the seam into the per-task loop
  and document it as a new `execute:task` point outside the existing
  contribution/step/gate vocabulary (unconditional in autonomous mode)

Closes #3970

* fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard

Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646
Phase 1) found three defects:

1. execute-plan.md's task-content-resolution bullet fired on any
   tracker-id-bearing task with no check that it wasn't type="checkpoint:*",
   contradicting ADR-3646 Decision 1 (a checkpoint task must never enter
   resolve-content). plan-document.cts already parses trackerId: null
   unconditionally for checkpoint tasks; only the workflow prose needed the
   fix, so the bullet now explicitly excludes checkpoint tasks.

2. task-content-resolution.cts's parseResolverDeclaration accepted any
   non-empty trackerPrefix with no grammar check, while capability-
   validator.cjs's KEBAB_RE enforces kebab-case at install time — a
   Generative Fix Divergence gap. Added the same grammar (as a literal
   regex, documented as intentionally not shared across the .cts/.cjs build
   boundary) plus a parity test asserting the two surfaces agree across a
   valid/invalid trackerPrefix table.

3. task-command-router.cts's routeResolveContent path-traversal guard on
   --plan had zero test coverage. Added a test exercising a
   ../../../etc/passwit-shaped path and asserting the USAGE rejection names
   the offending path.

* fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs

Two findings caught by an isolated security-review pass on the task
content resolution seam:

- ResolverFailedError/ResolverMalformedOutputError embedded raw,
  unsanitized subprocess stderr/stdout (attacker/model-influenced via
  the tracker-id argv token) into .message. A hostile or buggy resolver
  could smuggle a newline plus a forged "Error: " line, or terminal
  escape sequences, into a diagnostic io.cjs's error() writes verbatim
  to stderr. Fixed at the constructor (task-content-resolution.cts) via
  io.cjs's existing formatDiagnosticToken(), so every caller of
  resolveTaskContent gets a safe .message by construction.

- capability-validator.cjs's validateTaskContentResolverFields had no
  upper bound on taskContentResolver.invoke.timeoutMs, letting a
  manifest declare an effectively unbounded value and defeat the
  "bounded subprocess" design intent. Added a 120000ms ceiling specific
  to this field, without touching the shared isPositiveIntegerMs()
  helper (still used unbounded by the reviewer lane's timeoutFloorMs
  and probe timeoutMs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion

gsd-test (remote dockerized matrix) came back red with 5 failures on this
PR; all five are real defects, fixed here.

- tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's
  execute-plan.md entry pointed at line 415, which ffc190df4's
  checkpoint-exclusion caveat (added near line 221) shifted down by one
  line. The actual "validated downstream by gsd-tools uat
  classify-coverage" descriptive mention now sits at line 416. Updated
  the allowlist entry's line number to match.

- tests/task-command-router-resolve-content.test.cjs: the path-traversal
  test asserted the outside-project-scope diagnostic against the thrown
  ExitError's own .message. io.cts's error() (ADR-3889) writes its
  human-readable message to fd 2 via writeAllSync and then throws a bare
  `new ExitError(1)` with no message argument — by design, so the
  exception carries no duplicate text and the thrown ExitError's message
  defaults to "process exit 1" (cli-exit.cts's ExitError constructor).
  Root cause was the test, not the source: task-command-router.cjs's
  outside-project-scope rejection already calls error() correctly and the
  diagnostic text is genuinely emitted, just on fd 2, not on the
  exception. Fixed the test to capture fd-2 writes (mirroring
  tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this
  same file's own captureStdout for fd 1) and assert against the captured
  stderr text instead of err.message. This was masked locally because a
  manual `node -e` sanity check that only inspects the caught exception's
  .message cannot see what the real node:test run actually failed on.

Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3970): backfill changeset PR number (pr:0 -> pr:4000)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 13:17:04 -04:00
Tom Boucher
39673ae9ff fix(#3738): antigravity global skills/agents install to ~/.gemini/config (#3921)
* test(#3738): antigravity global skills/agents must resolve under ~/.gemini/config

Regression tests (RED first): --skills-root and gsd-tools query surfaces,
install-plan dest dirs, and converter skills-path rewrite.

* fix(#3738): antigravity global skills/agents install to ~/.gemini/config

Antigravity's machine-local discovery scans ~/.gemini/config/{skills,agents};
the configHome (~/.gemini/antigravity) is deprecated for artifacts. Declare the
ADR-1239 skills/agents 'home' override on the antigravity global layout — the
same mechanism codex uses (.agents) — and divert ~/.claude/skills/ references
in converted global content to ~/.gemini/config/skills/. configHome, settings,
probe/migration semantics, and the local .agents layout are unchanged.

* fix(#3738): retire deprecated configHome artifacts via installer migration 010

Next install converges an existing antigravity install: manifest-managed
skills/gsd-*/ and agents/gsd-*.md under the configHome (a location AGY does
not scan) are removed — modified files backed up first, unmanifested and
non-gsd entries preserved — and now-empty containers retired. Global scope
only; the local .agents surface is live. Docs + inventory updated.

* fix(#3738): converter sync in bin/install.js, harness emit-root coverage, migration baseline

- bin/install.js converter gains the same ~/.claude/skills → ~/.gemini/config/
  rewrite as src (ADR-1508 dual copy must stay in sync).
- Parity-manifest walk covers home-override emit roots (extraEmitRootsFor) so
  antigravity's emitted skills/agents stay differential-visible at their new
  install root; install-tree fixture regen confirms an unchanged key set.
- skills-from-commands rule declares the antigravity converter as a
  runtime-scoped transform; one ack fragment covers the identity-classed
  workflow whose antigravity copy embeds the old skills path.
- Migration 010 checksum baseline + home-override set doc updated; existing
  tests updated to the #3738 contract (global dest, golden parity via layout
  dest, integration expectations).

* fix(#3738): tolerate an absent extra emit root on baseline-side measurement

The base tree's installer predates the home override, so <HOME>/.gemini/config
does not exist there; walk() threw ENOENT and the in-job baseline build failed.
An absent extra root is the legitimate pre-override shape — skip it.

* fix(#3738): review findings — manifest agents root, bare skills-path rewrite, guard comment

- writeManifest resolves the agents-kind home override (_kindDestDirSafe), so
  the manifest records agents at their actual install root and drift detection
  keeps working (isolated review finding 1, major).
- Converter bare forms ~/.claude/skills and $HOME/.claude/skills (no trailing
  slash) divert to ~/.gemini/config/skills instead of falling through to the
  retired configHome path (finding 2).
- real-home-guard comment updated: antigravity's global agents kind is the
  first agents-kind home override (finding 3, doc-only).
- Regression tests for both behavioral findings.

* chore(#3738): changeset fragment (pr number backfilled after PR creation)

* chore(#3738): backfill changeset PR number (3921)

* fix(#3738): sandbox HOME in tests that install antigravity global artifacts

antigravity is the first home-override runtime in the golden-parity and
skills-wrapper suites (codex is not in their runtime lists), so those tests
never needed HOME sandboxing — the real-home guard now (correctly) refuses
their un-sandboxed global installs on CI, where HOME is the passwd home.

* fix(#3738): stop the K3 sequential-sandbox env leak; sandbox L2's home-override plans

K3's two back-to-back sandboxHome calls leave HOME pointing at the first
sandbox once the after-hooks restore (each call saves the env as it found
it, so the second saves the first's sandbox as 'original'). On the windows
matrix that leaked gsd-k3-qwen-* home into the L2 property, whose
antigravity/global run then (correctly) refused via the #3712 real-home
guard — antigravity is the runtime that made L2's plan escape into
os.homedir(). K3 now manages the env with a single restore; L2 sandboxes
HOME per run, mirroring L1.

* fix(#3738): L2 property's HOME sandbox must exist on disk

The #3712 guard's sandbox exemption fails closed when identify(effectiveHome)
is 'absent' — L2 never created its configDir, so on the windows matrix (tmpdir
under the real home) the antigravity/global run refused even with HOME
sandboxed. Create the per-run sandbox dir and clean it up.

---------

Co-authored-by: sim <sim@local>
2026-08-27 02:24:03 -04:00
Tom Boucher
878f25025c enhance(#3905): the exit-code registry — one number, one meaning, enforced at build (#3920)
* feat(#3905): the exit-code registry — one number, one meaning, enforced at build

A generated registry replaces locally-invented exit codes. Every entry records code, name, meaning, owning module and the decision that authorized it. The generator refuses to build a table where two entries claim one code, two claim one name, a code falls in a range Node or the shell reserves, 2 is claimed by anything but the hook adapter, or an allocation carries no justification. exitCodeFor is pure and total: it throws rather than returning undefined, including for prototype-chain names.

Inert by design — nothing emits a registered code until #3906. Every registered code is non-zero, asserted over the whole table, so a caller testing for failure behaves identically for pass and trips for everything else.

* feat(#3905): make the registry generator's failures machine-readable

Adds a --json mode carrying {ok, reason, context, detail}, where context is a typed payload naming the specifics the prose embedded - which code collided and under which names, which band rejected a code, which field was missing. The tests now assert on that structure instead of regex-matching the generator's stderr, which CONTRIBUTING prohibits, and the CONTEXT.md glossary gains the entry the issue's scope requires.

* test(#3905): refresh the install-tree fixtures for the new declaration

The registry declaration ships in the install tree, so all 19 golden fixtures needed regenerating. Caught by the remote matrix, not by lint:ci - the install-tree goldens are verified by a test rather than a lint, so a newly shipped file clears every local gate and fails only under the suite.

* chore(#3905): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-27 00:10:11 -04:00
Tom Boucher
6b7df61938 enhance(#3881): one YAML parser — vendored js-yaml replaces the hand-rolled dialect (#3888)
* docs(#3881): answer §8.1's open question and correct three wrong premises

ADR-3473 §8.1 carries a blocking open question with a forcing function: it must
be answered before any implementation PR for the rule opens. Answered here as (a),
a string-coercing adapter, with the measurement that settles it.

The sequencing note bet that §8.8's schema would make (b) tractable. Measured
against merged reality it does not: only 33 of extractFrontmatter's 78 non-test
call sites read STATE.md, and two of the five compensating mechanisms §8.1 lists
survive real types, leaving ~31 lines across 3 call sites as the actual prize.

Also corrects three claims verified false while answering it. §8.1's justifying
sentence names #3349 and #3360 as defects a real parser would fix; both are
already fixed on next, confirmed by executing the compiled parser rather than
reading it. The guard roster calls lint-frontmatter-scalar-broad-grep.cjs an
expected casualty of this rule, but it guards shell grep idioms in workflow bash
fences and never touches our parser. The same roster calls lint-vendored-deps.cjs
reusable as-is; it is hardcoded to re2js throughout.

The last two were caught by applying the rule this amendment records -- a factual
claim in this ADR is a hypothesis until the implementing phase executes it -- on
its first use.

Refs #3881

* docs(#3881): record that §8.1's fork is ill-posed and (a) is not implementable

An adversarial pass on the Phase 4 design established by execution that
extractFrontmatter is not a YAML parser but a line-oriented scanner whose output
is a function of raw source text. Four spellings of the same value collapse to
one js-yaml tree but produce four distinct legacy strings, one of them mangled.
No adapter over a tree can choose among outputs the tree does not distinguish,
so fork (a) -- keep a string-coercing adapter so the existing contract holds --
cannot be built. For any document with a non-scalar value, (a) collapses into
(b); about 26 percent of frontmatter-carrying documents have one.

Also records three design defects and one new attack surface, all confirmed by
execution: catching a parse failure and returning {} would delete the frontmatter
block on the next write at eight call sites that conflate empty with unparseable;
an empty value yields null where legacy yields {}, and reconstructFrontmatter
omits null-valued keys, so the shipped state template's empty progress key would
vanish; the #1882 truncation probe is parseYamlRegion itself rather than a
pre-parse heuristic, so it cannot both stay unchanged and survive that deletion;
and FAILSAFE_SCHEMA still resolves aliases, expanding seven lines to 22.8 MB.

The rule is not deferred. The measurement is the deliverable and the re-scoping
is recorded as an open question with a forcing function, per section 8's own rule.

Refs #3881

* test(#3881): failing-first rows for block scalars, unicode keys and the missing #3594 matrix

Creates tests/feat-3594-parser-adversarial-frontmatter.test.cjs, the file the fixture README instructs contributors to register fixtures in but which never existed.

Section C: table-driven ownership check over tests/fixtures/adversarial/frontmatter/ so a fixture with no matrix entry fails loudly; six existing fixtures (duplicate-keys, crlf-mixed, unclosed-block, unicode-keys-and-values, null-byte-value, huge-bounded) each get the invariant its README states.

B1 blockScalarValueIsNotTheBlockIndicator: parsing commands/gsd/add-tests.md must give argument-instructions the instruction text, not the literal '|'. RED today.

B2 blockScalarDoesNotInventATopLevelKey: same parse must not produce a top-level Example key scraped from inside the block body. RED today.

B3 unicodeKeyRoundTripsAsIs: the 相 key in unicode-keys-and-values.md must survive parsing; today it is silently dropped. RED today.

Refs #3881

* chore(#3881): vendor js-yaml and generalize the vendored-deps guard to a manifest

Packaging step for ADR-3473 §8.1: makes js-yaml available to gsd-core/bin/** without promoting it out of devDependencies (promoting broke every installed tree, #3496).

gsd-core/bin/lib/vendor/js-yaml.cjs is a verbatim copy of node_modules/js-yaml/dist/js-yaml.js (the self-contained UMD dist bundle, not index.js), exposing load/dump/FAILSAFE_SCHEMA/YAMLException with zero require() calls of its own.

src/vendor/js-yaml.d.cts is hand-authored, not copied, because js-yaml ships no upstream .d.ts and @types/js-yaml is not installed. It is deliberately narrow, declaring only the four symbols in use, so anchors/aliases/custom types/loadAll are unreachable from typed code -- a compile-time enforcement of ADR-3473 §8.1's refusal to expand alias resolution for security reasons. Because it has no upstream counterpart it is excluded from the byte-compare.

scripts/lint-vendored-deps.cjs is refactored from a script hardcoded to re2js into a table-driven VENDORED manifest (one row per package: upstream/vendored .cjs paths, optional .d.cts paths, twin kind upstream-verbatim vs hand-authored) so a second vendored package does not require a second hardcoded check block, per ADR-3473 §8.3 'one implementation per rule'. The four existing re2js checks (vendored .cjs vs node_modules, vendored .d.cts vs node_modules, src/vendor twin vs bin-side twin, devDependency version pin vs installed version) are preserved unchanged; verified pass/fail identical before and after the refactor, and the guard's ability to fail was re-proven with a deliberate one-byte append to both re2js.cjs and js-yaml.cjs, then restored.

docs/INVENTORY.md and docs/INVENTORY-MANIFEST.json (via gen-inventory-manifest.cjs --write, run after build:lib) register vendor/js-yaml.cjs. gsd-core/bin/lib/vendor/README.md documents both vendored packages and the two twin kinds.

Refs #3881

* feat(#3881): parse .planning frontmatter with the vendored js-yaml

ADR-3473 §8.1: extractFrontmatter's read path is no longer a hand-rolled
line scanner. parseYamlRegion, escapeDoubleQuoted, unescapeDoubleQuoted and
parseQuotedScalar are deleted (not patched); parsing now goes through the
vendored js-yaml (./vendor/js-yaml.cjs) under { schema: FAILSAFE_SCHEMA,
json: true }. Everything js-yaml does not do is layered on top, in one
place, carrying the seven design-doc consequences:

1. Empty value: a null js-yaml value is coerced to {} (matching legacy's
   own empty-value contract) so reconstructFrontmatter — which omits
   null-valued keys — still round-trips a bare `key:` line instead of
   deleting it. Verified live: progress: with no value survives
   parse -> reconstruct -> re-parse.

2. Unparseable no longer collapses to a bare {}: a new FRONTMATTER_UNPARSEABLE
   Symbol (exported), keyed exactly like the existing #3257 FULL_LINE_COMMENTS
   channel, is carried on the {} returned for malformed/refused YAML. Invisible
   to Object.keys/entries/JSON.stringify/for-in, so the 70 call sites that
   never inspect it are unaffected; wiring the 8 hasFrontmatter sites to
   consult it is a separate change, not done here.

3. Non-scalar object-list items (the four spellings of `- test: a b` that
   js-yaml collapses into one tree shape) are rendered as a canonical
   `key: value[, key2: value2]` string per item, keeping the existing
   array-of-strings value SHAPE. A full corpus differential over all 1702
   tracked markdown files found 11 residual divergences from the legacy
   parser (enumerated in the PR/report), most of them the parser now being
   MORE correct (a dropped quoted top-level key, the block-scalar/phantom-key
   defect, a dropped Unicode key).

4. The #1882 truncation probe still runs the one real parser, but derives
   its key count from js-yaml's own thrown error and mark.line when the
   whole region doesn't parse cleanly (the dominant real truncation shape:
   fence opened, well-formed keys, no closing fence). Verified against both
   the clean-parse and the exception-fallback path.

5. The #3257 comment channel now attributes each pending column-0 comment
   against js-yaml's own parsed top-level key list (matched by literal key
   text, in document order) instead of the legacy ASCII-only key regex, so
   a comment above a Unicode key attaches correctly.

6. Anchors, aliases and merge keys are refused outright (a raw-text
   pre-scan, since FAILSAFE_SCHEMA still resolves them) — corpus occurrences
   today: zero. A 7-line billion-laughs fixture is verified refused rather
   than expanded.

7. A literal U+0000 is swapped for a private-use sentinel before the parse
   and restored in every resulting string afterward, since js-yaml rejects
   NUL unconditionally under every schema.

escapeDoubleQuoted is deleted and reimplemented via js-yaml's dump()
(forced double-quoted style), with control-char hex escapes lowercased to
keep serialized output byte-stable (#1779 emitted lowercase); it keeps its
exported name and signature for its two other call sites (commands.cts,
runtime-artifact-conversion.cts), which need no change.

frontmatterDeepEqual, the comment channel, sliceTopLevelFrontmatterSegments,
regenerateFrontmatterKey's guard, noOpObjectListSetError and
parseMustHavesBlock are all unchanged — retiring them is fork (b) and is
not this phase.

Refs #3881

* fix(#3881): quote template placeholders and preserve unparseable frontmatter

SECURITY.md/UI-SPEC.md/VALIDATION.md wrote frontmatter placeholders as
bare {N}/{phase-slug}/{date}, which is valid YAML flow-mapping syntax
under the vendored js-yaml parser, not the literal placeholder text
intended. Quote them so they parse as strings.

Wire the FRONTMATTER_UNPARSEABLE Symbol (exported but unused) at the
8 call sites in state.cts/state-transition.cts that compute
hasFrontmatter via Object.keys(extractFrontmatter(...)).length > 0 and
reassemble the document without a frontmatter block when false. That
check conflated 'no frontmatter' with 'unparseable frontmatter' (both
parse to {}), so a document with a merge-conflict marker or refused
alias in its frontmatter had that block silently dropped on write.
Each site now preserves the exact raw bytes stripFrontmatter removed
when the marker is set, leaving the genuinely-empty case unchanged.

Refs #3881

* test(#3881): consequence and boundary coverage for the js-yaml migration

Rows: A1 emptyValuedKeySurvivesAWrite, A2 unparseableDocumentKeepsItsFrontmatterBlock, A3 unparseableIsDistinguishableFromEmpty, A4 nonScalarValuesCanonicalize, A5 truncationProbeStillFiresOnAnOpenFence, A6 commentsStayOnTheirOwnKey, A7 anchorsAndAliasesAreRefused, A8 aliasExpansionCannotExhaustMemory, F1 UNTERMINATED_KEY_THRESHOLD boundary, F2 alias/nesting refusal bound, F3 frontmatter size boundary (huge-bounded.md + larger). Adds tests/fixtures/adversarial/frontmatter/anchor-alias-bomb.md and its entry in the feat-3594 fixture matrix.

Refs #3881

* docs(#3881): document the vendored parser, correct a stale rationale, add a vendoring how-to

Refs #3881

* docs(#3881): correct the frontmatter glossary entry

Two errors in the entry as first written: it named parseYamlRegion as part of
the read path when that function is deleted, and it recorded the eight
hasFrontmatter call sites as unwired follow-on work when they were wired in
e35ac2a2c. Also records the scope caveat that the CLI write path rebuilds the
frontmatter block independently, so the marker binds at the transform layer.

Refs #3881

* docs(#3881): record the semantic-migration decision and the counted guard ledger

The maintainer chose the full semantic migration over splitting the rule into
its own epic or patching the scanner, so section 8.1 is answered as "the fork
was ill-posed and the migration is semantic" rather than as (a) or (b).

Also replaces the pre-implementation guess that this phase would shrink the
guard surface with the counted result: excluding vendored third-party lines the
hand-maintained surface is net +307, and frontmatter.cts grew by 68 lines
despite four functions being deleted, because the compatibility layer over
js-yaml is larger than the scanner it replaced. Section 8.1's stated benefit is
therefore not delivered as written; what improved is the kind of code
maintained, not the amount. Decision 6 requires recording that rather than
netting it away.

Refs #3881

* chore(#3881): changeset for the vendored YAML parser migration

Refs #3881

* test(#3881): golden parity, round-trip property and packaging coverage

Refs #3881

* fix(#3881): refuse anchors structurally and fold in review findings

ADR-3473 §8.1 review findings, addressed inline:

Finding 1 (BLOCKER): refuseAnchorsAndAliases was a raw-line regex that matched
only the bare-key spelling (key: &x). A quoted key ("a": &x), a flow mapping
({b: &x}) and a flow sequence ([&x, *x]) all define/use the SAME anchor
mechanics while never matching that line shape, so the exact expansion the
guard exists to stop went straight through unrefused (a 303-byte quoted-key
bomb expanded to ~35.8MB). Replaced with js-yaml's own `load` `listener`
callback, which reports `state.anchor` for every event belonging to an
anchored node in every spelling, and throws from inside the callback to abort
before any expansion (~1-2ms vs full expand-then-discard). A merge key with
an alias is still refused (merge always requires a previously anchored node,
so the alias itself trips the listener); a bare merge key with NO alias is no
longer separately refused, documented as intentional: FAILSAFE_SCHEMA never
resolves `!!merge`, so it carries no expansion risk. Table-driven tests added
for all four bypass spellings + merge key, plus a quoted-key-spelled
billion-laughs fixture registered in the adversarial matrix and README.

Finding 2: src/vendor/js-yaml.d.cts's docblock falsely claimed anchors/
aliases were "simply UNREACHABLE from typed code" through the twin. Corrected
to state the truth: anchor/alias resolution is document-level `load`
mechanics reachable through exactly the declared surface, and refusal is
enforced at RUNTIME (Finding 1's listener), not by the type surface.

Finding 3 (MAJOR): the null-byte sentinel (U+E000) round-trip was
non-injective — restoreNullBytesDeep rewrote every U+E000 in the parsed tree
back to NUL, including one the document author legitimately wrote, silently
corrupting it. Now refuses outright whenever the raw region already contains
U+E000 (consistent with the existing anchor/merge-key refusal path), making
the substitution provably injective. Tests added for a real NUL alone
(preserved), a pre-existing U+E000 alone (refused, not corrupted), and both
together (refused, not merged into one byte).

Finding 4 (MAJOR): scripts/lint-vendored-deps.cjs's `srcTwin` field was dead
for a hand-authored row (only read inside the upstream-verbatim branch) —
exactly how Finding 2's stale docblock drifted unnoticed. Added
checkHandAuthoredTwin: every value-level export the twin DECLARES must be an
actual own property of the vendored runtime module at require-time. Tests
added, including a sensor that a declared-but-nonexistent export IS caught.

Finding 5: the existingFm/hasFrontmatter/stripFrontmatter/fmPrefix/
unparseableFm/reassemble preamble, copy-pasted at 7 sites in
state-transition.cts plus a sixth hand-inlined copy in state.cts's
cmdStateCompletePhase, is now one exported helper
(beginFrontmatterReassembly) every site routes through, including the
hand-inlined one. Three call sites (beginPhaseCore, patchCore, updateCore)
keep a literal `body = stripFrontmatter(content)` assignment alongside the
helper call so scripts/lint-state-write-path-drift.cjs's single-hop backward
scan (which does not chase aliases) still sees the strip; stripFrontmatter is
pure/idempotent so the extra call changes nothing observable.

Finding 6: corrected the frontmatter.cts docblock's stale "wiring is a
separate change" claim (the 8 call sites are wired on this branch) and the
changeset's backlink from (#3473) to (#3881).

Finding 7: fixed the lint:ci failures blocking the gate — an
@typescript-eslint/only-throw-error violation from throwing a bare Symbol as
the anchor-detected signal (now a real Error subclass), unused-var warnings
left over from the Finding 5 refactor, a lint-test-file-count cap exceeded by
two migration-specific test files (allowlisted with justification), and the
lint-state-write-path-drift false positive from Finding 5's helper (fixed
above). tests/frontmatter-golden-parity.test.cjs:117's execFileSync already
carried an explicit timeout; no change was needed there.

Golden fixture: added a golden entry for the new
anchor-alias-bomb-quoted.md fixture ({} — matches what the legacy line
scanner would also produce, since it independently dropped every quoted
top-level key). No other corpus document diverges: real .planning/ documents
carry zero anchors/aliases/merge keys/U+E000 today.

Refs #3881

* fix(#3881): fold in second-round review findings

Finding 1 (BLOCKER): tests/frontmatter.test.cjs pinned the pre-migration
ASCII-only key regex for the Unicode fixture; updated to require the 相
key's value now that js-yaml has no such restriction. Audited the rest of
the file for other pre-migration pins (block scalars, quoted keys,
flattened values, empty values, duplicate keys, unclosed blocks, null
bytes) by execution against real fixtures; found none regressed.

Finding 2: parseYamlRegion and escapeDoubleQuoted renamed to
parseGuardedYamlRegion and escapeDoubleQuotedScalar in src/frontmatter.cts
so no function still answers to the deleted hand-rolled scanner's name
(ADR-3473 §8.1 "deleted, not patched"). escapeDoubleQuotedScalar's three
external call sites (src/commands.cts, src/runtime-artifact-conversion.cts)
updated in the same change — a mechanical rename, not an ADR-amendment
matter.

Finding 3 (BLOCKER): fixed a real crash and a silent data-loss bug found
by execution. A top-level key named constructor/__proto__/toString/
valueOf/hasOwnProperty crashed reconstructFrontmatter (bracket read
resolving an inherited Object.prototype member); a key literally named
__proto__ was silently DROPPED entirely (bracket assignment on an
ordinary {} invoked the inherited __proto__ setter instead of creating a
data property). Fixed by building every parsed Frontmatter object with
Object.create(null), and replacing an `in` check with hasOwnProperty.call
in propagateCommentChannel. Added round-trip tests for all five hostile
keys, each with its own leading comment.

Finding 4 (MAJOR): escapeDoubleQuotedScalar's docstring falsely claimed
full byte-stability across the migration. Verified by execution: BEL/NUL/
NEL/NBSP/LS/PS/BOM now emit YAML-named escapes instead of the old hex/raw-
literal forms. Proved round-trip equivalence (each escape re-parses to the
exact source codepoint) and corrected the docstring. Found and fixed a
related real defect while verifying: a lone UTF-16 surrogate was emitted
BARE (scalarNeedsDoubleQuoting didn't trigger), producing genuinely
unparseable YAML that silently collapsed to {} on re-read — extended
scalarNeedsDoubleQuoting to route surrogates through the quoted+escaped
path.

Finding 5 (MAJOR): countKeysBeforeTruncation went silent on 4 real
truncation shapes (unquoted colon, open flow collection, mis-indented
sibling key, refused anchor). Root cause: the mark-based prefix recovery
excluded the very line whose key needed counting, and a mark-less refusal
never entered the recovery branch at all. Fixed by taking the max of two
lower bounds: the longest parser-verified line-prefix, and a raw-text
count of key-shaped lines (reusing the same key-shape pattern this file
already uses for isFrontmatterShaped). Extended test-matrix row A5
table-driven over all 4 regressed shapes.

Finding 6: the design doc's claim that no test owned the #3594 adversarial
fixture corpus was false — consolidation epic #1969 had already folded it
into tests/frontmatter.test.cjs. An earlier commit on this branch
re-created a standalone duplicate under that false premise; folded its
genuinely-new coverage (fixture-ownership check, anchor-bomb fixtures,
block-scalar B1/B2 rows) into frontmatter.test.cjs and deleted the
duplicate file. Corrected the false claims in 40-design.md §3.3.1 and the
ADR's §8.1 note, including the roadmap-sibling claim (no such file exists).

Finding 7: the golden serializer sorted object keys, making it structurally
blind to the key-order-parity invariant ADR-3473 §8.1 actually claims.
Made it order-preserving and regenerated the golden fixture from a
standalone compile of the legacy (pre-#3881) parser at ddde001af; the
current parser matches it with zero undocumented divergences, confirming
key-order parity genuinely holds. Extended row A2 table-driven across 6 of
the remaining 7 transitionCore kinds (all pass) plus documented, by
execution, a newly-discovered 8th-site regression: state.cts's
cmdStateCompletePhase calls the same preservation helper but its result is
clobbered by a later unconditional resync — filed as a distinct finding
rather than fixed here (touches syncAndPreserveStateMd, outside this
change's verified scope).

Refs #3881

* fix(#3881): preserve unparseable frontmatter through the CLI write path

Characterization (executed, before/after shown): case (b), not (a). The
frontmatter FENCE survives — `state complete-phase` on a conflict-marked
STATE.md returns success and a well-formed, freshly-derived frontmatter
block, not a document with no frontmatter at all. But the block's actual
content (the merge-conflict markers, and with them any signal to a human
that the document was in conflict) is silently discarded and replaced.

Root cause was two clobber sites, not one:

1. syncStateFrontmatter (src/state.cts) re-parses the already-preserved
   `transformedContent` from readModifyWriteStateMd, finds {} + the
   FRONTMATTER_UNPARSEABLE marker, and unconditionally rebuilt a fresh
   frontmatter block from the body anyway.
2. Even after (1) is fixed, applyPostSyncPreservation's own
   postFm/applyStatePreservation/authoritativeFm-reassertion machinery
   re-extracts frontmatter from syncedContent, restores curated fields
   from the pre-write snapshot, and reconstructs a NEW block again —
   confirmed live via `state begin-phase`, which still lost the markers
   after fixing (1) alone.

Both are now guarded by the same predicate (isUnparseableFrontmatter,
checking FRONTMATTER_UNPARSEABLE): when the ORIGINAL frontmatter did not
parse and the caller is not on ADR-3408 §8.3's closed "body wins" list,
both functions return their input content unchanged rather than
re-deriving over it. The closed list (cmdStateSync #905,
/gsd-health --repair's REGENERATE_STATE, both routed only through
writeStateMd, which never reaches applyPostSyncPreservation and passes
sanctionedPermanentEmptyFallback=true to syncStateFrontmatter) is
untouched — neither widened nor narrowed; verified by execution that
`state sync` still overwrites the conflict-marked block exactly as before.

Other verbs sharing the same readModifyWriteStateMd path were checked and
were equally affected before this fix: state update, query state.patch,
and state begin-phase all lost the conflict markers (RED, shown by
execution), and all three now preserve them (GREEN). Covered table-driven
in tests/feat-3881-yaml-parser-consequences.test.cjs's new A2b describe
block, which drives the real CLI verbs via runGsdTools — not just the pure
transitionCore layer the earlier A2 rows exercised — plus a control
asserting state sync's body-wins contract is unchanged.

Refs #3881

* fix(#3881): restore the parse surface's prototype and fix remote-runner failures

Root cause of the bulk of the 88 remote-runner failures: extractFrontmatter/parseGuardedYamlRegion handed back Object.create(null) trees for prototype-pollution safety, but assert.deepStrictEqual compares prototypes, so every assertion against a plain object literal failed (57 frontmatter.unit.test.cjs + 5 frontmatter.test.cjs + others). Fixed by keeping the internal construction null-prototype (unchanged) and converting to a plain-prototype tree via Object.defineProperty (never bracket assignment, so __proto__/constructor/toString keys stay safe) at the parseGuardedYamlRegion/unparseableResult return boundary only; the internal FULL_LINE_COMMENTS Symbol channel is copied by reference, not recursed, so its own __proto__-safety is untouched.

Per-class fixes: (1) bomAcrossArtifactTypes was the same prototype bug, no separate code change needed. (2) frontmatter-cli #1660: added objectListFieldWouldLoseData, a broader lossy-field detector alongside the existing byte-identical noOpObjectListSetError -- js-yaml's flattenObjectListItem now correctly includes every sub-key of an object-list item (a real bug fix over the legacy scanner, which silently dropped every field but the first), so a set that drops that now-included data is no longer byte-identical to the original and needs its own guard. (3) uat.test.cjs: updated the pinned expectation for the human_verification quote-stripping artifact -- js-yaml resolves quoting correctly where the legacy regex left an unbalanced quote; documented as an intentional, non-lossy behavior change. (4) smart-entry: added a fallback-only loadWithAmbiguousColonRepair so a column-0 key: value line whose value itself contains an unquoted colon (the #2571 hand-edited-STATE.md shape) round-trips instead of failing the whole frontmatter block closed. (5) frontmatter.unit.test.cjs bracket-array leniency: added a second fallback, repairMalformedInlineArrays, restoring the legacy scanner's tolerant inline-array handling (consecutive/blank commas, unclosed bracket) -- both repairs run ONLY after the primary parse already threw, so well-formed documents are unaffected. (6) prompt-injection-scan: src/frontmatter.cts had a literal U+FEFF BOM embedded in a comment illustrating the #2977 fix; replaced with the U+FEFF text escape. (7) eslint-glob-coverage: allowlisted the new src/vendor/js-yaml.d.cts vendored type declaration, same precedent as the existing re2js.d.cts entry. (8) frontmatter-golden-parity: git ls-files *.md now runs with -c safe.directory=* (process-scoped) so it survives the remote runner's dubious-ownership check without a persistent git config write.

Refs #3881

* chore(#3881): backfill changeset PR number

Refs #3881

* test(#3881): make golden parity resistant to unrelated tree churn

A corpus-wide snapshot keyed to every tracked *.md file was coupled to mutable-by-design files: .changeset/*.md's pr:0 -> real-PR-number backfill is a required workflow step, not a parser change, yet it turned this suite red. Training people to 'just regenerate the golden' on that kind of failure defeats the point of the snapshot. Exclude .changeset/** from the golden corpus entirely, tolerate tracked *.md files with no golden entry (they postdate the capture) instead of failing on them, keep hard failures for a golden entry whose file has vanished from the tree and for any real parity divergence, and add a coverage floor so the enumeration cannot quietly degrade to comparing a handful of files. Golden regenerated by recompiling the legacy pre-migration parser (git show ddde001af:src/frontmatter.cts) standalone, independent of the current parser, over the same non-changeset corpus.

Refs #3881

* test(#3881): make the parser golden hermetic instead of tree-keyed

This repo merges ~21 commits/day; a 14-day sample measured 937 touches of the
exact files (commands/gsd/*.md, gsd-core/workflows/*.md, agents/*.md,
docs/*.md) the prior golden pinned by tracked path. Any PR editing one of
those files' frontmatter for reasons unrelated to the parser (an
argument-hint addition, an allowed-tools tweak) turned the suite red, and the
reflex fix -- "regenerate the golden" -- overwrote the very snapshot meant to
catch a real regression. Excluding .changeset/** was not enough; the design
itself was wrong: a regression fixture must not be keyed to mutable repo
paths, and a single 376-entry JSON every such PR touches is also a
guaranteed merge-conflict surface.

Rebuilt the fixture to carry its own documents: each of 51 entries stores a
stable id, literal documentText (shrunk from a real ddde001af-era corpus
document), and an expectedParse captured independently from the
pre-migration legacy parser (git show ddde001af:src/frontmatter.cts,
compiled standalone against its byte-identical sibling modules). The test
reads no tracked path, shells out to no git command, and enumerates no tree
-- a PR editing commands/gsd/help.md cannot affect it. Every entry's
reconstruction was verified at capture time to reproduce both the current
and legacy parser's output on the original document; 0 of 51 candidates
were dropped by that check (1, the deliberately-unterminated
unclosed-block.md adversarial fixture, has no closing fence to truncate at
and is stored unshrunk). Kept the 5 documented DIVERGENCES rows (now
diverges:true entries) and the D2 order-preserving structural serializer
that keeps the comparison from passing vacuously; dropped the
tree-enumeration helpers, the coverage floor, the post-capture-skip logic,
and the vanished-file check -- all artifacts of the path-keyed design.

Refs #3881

* fix(#3881): resolve vendored-deps paths independently of cwd shape

Five rows in tests/lint-vendored-deps-manifest.test.cjs failed on
windows-latest CI: the test passed absolute scratch-file paths into
compareFiles()/checkRow(), whose helpers joined every input onto ROOT
via path.join(ROOT, rel), producing garbage when the input was already
absolute. It surfaced on windows-latest specifically because GitHub's
Windows runners checkout the repo on a different drive than TEMP, so
path.relative(REPO_ROOT, tmpFile) returned the absolute path unchanged
(no relative traversal is representable across drives) rather than the
relative form the test assumed. The remote gsd-test runner this repo
gates pushes on is Linux-only and could never have caught this;
GitHub CI's windows-latest job is the only signal that does, and it did.

Fixed the helper itself (scripts/lint-vendored-deps.cjs's new
resolvePath()) to treat an already-absolute input as absolute-in,
absolute-out instead of silently mis-joining it, and updated the test
to pass the scratch file's absolute path directly rather than relying
on a relative conversion that is not always representable. Kept every
mutation-sensor assertion intact and added coverage proving
resolvePath is a no-op for relative inputs and correctly passes
absolute ones through unchanged.

Refs #3881

* fix(#3881): warn when state sync regenerates over unparseable frontmatter

state sync (ADR-3408 §8.3's sanctioned regenerate path) correctly
overwrites an unparseable frontmatter block per its 'body wins'
contract — that overwrite behavior is unchanged here. The defect was
the silence: synced:true/exit 0 gave no signal that the existing
block (including git merge-conflict markers) could not be parsed and
was destroyed, per ADR-3473 §8.5 ('a derived conclusion may not be
reported as authoritative when the derivation dropped input it could
not resolve') and §8.4 ('failure is a value').

Adds a gsd: warning — ... (#3881) line on stderr, matching the
existing #3573 precedent, and surfaces the same disclosure in the
JSON result's existing changes[] array so a machine consumer sees it
too. Exit code and synced:true are left unchanged — sync did what its
contract says.

REGENERATE_STATE (/gsd-health --repair's sibling on the same
sanctioned-regenerate list) is DESTRUCTIVE-risk and unconditionally
refused by applyRepairs's dispatcher before runRepairAction ever runs
(src/health-diagnostic.cts), so it is not a live path today and is not
in scope for this fix.

Refs #3881

* fix(#3881): exit non-zero when a state command returns an error

Refs #3881

* chore(#3881): changeset for the state exit-code fix

Refs #3881

* fix(#3881): honor the documented --project-dir flag

Refs #3881

* revert(#3881): restore exit-0 result envelopes for state errors

Reverts 9638f2936 and its changeset. The change was wrong and the revert is
the correction.

This repo distinguishes two error mechanisms deliberately. error() in
src/io.cts writes to stderr and calls process.exit(1) -- the hard-failure
path. output({error: ...}) writes a JSON result envelope to stdout and returns
normally with exit 0. The reverted commit converted 23 result-envelope sites
into hard failures, which is a different contract, not a bug fix.

tests/state-contract.test.cjs's errorPathDoesNotPublish asserts the envelope
contract directly -- a failing command exits 0 with a JSON error envelope and
must not publish state.json -- and the remote matrix run caught it along with
four cases in the QA scenario walk. Thirteen tests in tests/state.test.cjs that
the original commit rewrote were encoding that real contract, not the bug it
claimed; they are restored.

Whether an error envelope on stdout with exit 0 is the right CLI design is a
genuine question, and it is section 8.4's rule ('failure is a value') with its
own phase. It is not something to flip inside this PR.

Refs #3881

* chore(#3881): backfill changeset PR number for the project-dir fix

Refs #3881

* test(#3881): keep the frontmatter mutation shard inside its time budget

The Stryker (frontmatter) shard hit the documented 15-minute (900s) shard
cap. Root cause is NOT row-level spawn overhead (contrast the #2790/
core-utils precedent): the three shard test files' own logic runs in
~413ms total (356+30+27ms) with all 392 assertions passing. Instead,
src/frontmatter.cts grew from ~825 to 1496 lines (+671/-187) migrating to
the vendored YAML parser, proportionally growing the mutant count Stryker
generates for gsd-core/bin/lib/frontmatter.cjs. Stryker's command runner
bills the full 'node --test <3 files>' invocation once per mutant, and
node:test's default per-file process isolation forks a child process for
each of the three files on every one of those invocations — pure fork
overhead multiplied by a much larger mutant population.

Fix: scripts/mutation-matrix.cjs COVERED.frontmatter now declares
isolation: 'none', and .github/workflows/mutation.yml passes
--test-isolation=${{ matrix.isolation }} (defaulting to 'process' — i.e.
unchanged behavior — for the other 8 shards, which were not individually
audited for cross-file state leakage under shared-process execution).
Measured locally via node:test's run() API on the exact 3-file set:
isolation:'process' took ~593ms vs isolation:'none' ~478ms for the same
392 passing assertions. The true CI-shard number can only be confirmed
on the GitHub Actions run (Stryker cannot run locally, and 'node --test'
is hard-blocked in this environment).

Refs #3881

* test(#3881): register the vendored-parser tests in the frontmatter mutation shard

stryker.config.mjs's own rule ("Keep this list in sync with the tests
arrays in scripts/mutation-matrix.cjs COVERED") was violated: #3881 grew
src/frontmatter.cts from ~825 to 1496 lines but its new tests
(tests/feat-3881-yaml-parser-consequences.test.cjs,
tests/frontmatter-golden-parity.test.cjs,
tests/frontmatter-roundtrip.property.test.cjs, and +167 lines in
tests/frontmatter.test.cjs) were never added to the frontmatter shard's
tests array, so Stryker's mutants in the new vendored-js-yaml adapter had
nothing constraining them. PR #3888 measured 55.8% against the 65 floor
(748 killed / 593 survived / 17 timeout) and the shard was separately
cancelled at 15m04s against the 15-minute per-shard cap.

Registers all four files (each earns its slot on evidence of a unique
constraining assertion, documented inline), gives the shard a
measured/projected 180-minute budget via a new per-module
timeoutMinutes field threaded through mutation.yml's job-level
timeout-minutes the same way isolation is threaded, and removes the
prior isolation:'none' override (re-measured at this file-set size, its
savings are within run-to-run noise, not worth the unaudited
cross-file-state-leakage risk).

Refs #3881

* feat(#3881): derive the mutation test list and ratchet the score floor

Refs #3881

* test(#3881): ratchet five stale mutation floors and close the frontmatter gap

Raised five module minScore floors per CI run 33012034388 (floor(achieved)-1):
config-schema 75.51%->74, prompt-budget 88.95%->87, context-composer 79.92%->78,
context-utilization 92.31%->91, active-workstream-store 87.42%->86. Updated both
scripts/mutation-matrix.cjs COVERED entries and tests/mutation-matrix-ratchet.test.cjs
RATCHET_BASELINE in the same diff per the ratchet's own contract.

Closed the frontmatter shard's 63.03%-vs-65 gap with new behavioral tests in
tests/feat-3881-yaml-parser-consequences.test.cjs, each paired with a documented
near-miss: frontmatterDeepEqual's array-order/length/type-mismatch/key-order
semantics (via spliceFrontmatter's no-op guard), scalarNeedsDoubleQuoting's
leading/trailing-whitespace and dash/surrogate triggers (via reconstructFrontmatter),
repairAmbiguousColonValues' already-quoted vs ambiguous-colon repair paths (via
extractFrontmatter), and the null-byte sentinel round-trip surviving at region
offset 1. Did not lower minScore.

Refs #3881

* test(#3881): decouple the ratchet test from real module floors

The CLI end-to-end rows in tests/mutation-score-ratchet.test.cjs hardcoded config-schema's real floor (52), which commit 973321541 legitimately ratcheted to 74 -- breaking a test pinned to the exact value the mechanism under test exists to change. Add an injectable --matrix seam to scripts/check-mutation-score-ratchet.cjs and point the CLI rows at a synthetic module + synthetic floor built via a temp fixture, so the rows are indifferent to any real module's floor moving while still exercising the same fail/pass behaviour.

Refs #3881

* refactor(#3881): parse must_haves with the vendored parser and drop re-implemented leniency

Refs #3881

* fix(#3881): restore the ambiguous-colon repair its hand-edited-STATE.md contract needs

A tracked-document sweep of 910 *.md files cannot see this dependent: repairAmbiguousColonValues's one real caller is user hand-edited STATE.md content that never lives in this repo's tree, only on end users' machines, and is pinned by tests/smart-entry.unit.test.cjs. Restores the function plus its post-throw fallback path (loadWithAmbiguousColonRepair) only; repairMalformedInlineArrays and splitLegacyInlineArrayItems stay deleted, reverified against the full frontmatter test shard. Adds a frontmatter-level regression row in tests/feat-3881-yaml-parser-consequences.test.cjs so the dependency is visible where the function lives.

Closes #2571
Refs #3881

---------

Co-authored-by: sim <sim@local>
2026-08-26 19:29:32 -04:00
Tom Boucher
ddde001af6 enhance(#3873): the STATE.md schema — one owner, generated artifacts (#3880)
* test(#3873): failing-first locale parity, plus tripwires for what must not move

Pins ADR-3473 §8.8 at the artifact a reader actually sees. The English STATE.md
reference carries a Status lifecycle section that is missing from all four
translations — the section documenting the status enum whose clobbering is
#3853. The test derives the heading set rather than hard-coding the missing
one, and names the locale and the heading when it fails.

Two tripwires that must pass today and after. The field-drift guard still
catches a re-derived fallback ladder: §8.8 instructs deleting that script, and
that instruction rests on a wrong premise about what it guards, so the test
stops a future reader from deleting it on the ADR's word. And last_activity's
label resolution is pinned to what ships today, because it is declared in one
of the two tables this phase consolidates and not the other — the
consolidation must not silently pick a side.

The locale test buckets under docs rather than state, which is what it tests;
that bucket is allowlisted with justification rather than folded into an
unrelated docs suite. It reads only markdown, so it carries no allow-test-rule
marker — a marker there would suppress nothing and would grow the unverified
pool against its ceiling.

Refs #3873

* feat(#3873): one schema owns the STATE.md key set, three tables become projections

ADR-3473 §8.8. The key set was declared in four places that had to agree by
hand and already did not: FIELD_CLASSIFICATION, FRONTMATTER_BODY_SOURCE,
FRONTMATTER_KEY_TO_BODY_LABEL and buildStateFrontmatter's emit behavior. One
frozen null-prototype schema now declares each key's type, enum, cardinality,
source, preservation, body source, body label, accepted parse shapes and
whether it is emitted unconditionally; the three tables are derived from it at
module load.

The projections are byte-identical to the literals they replace, key order
included, and the parity tests compare against verbatim copies of today's
tables rather than re-deriving both sides from the schema — a parity test fed
from one source proves nothing, which is how a consolidation ships a changed
policy under a green test.

last_activity was the live disagreement: present in one table, absent from the
other. The schema declares what ships today rather than the tidier answer, and
a test pins it.

The schema is a leaf module and owns the four field-policy types, re-exported
from state-transition so existing importers are untouched — the same split
health-diagnostic-types made to break a CJS require cycle.

Refs #3873

* feat(#3873): generate the schema-derived regions, parity-check the prose tables

ADR-3473 §8.8's generator half. gen-state-md-docs.cjs owns marked regions in
the shipped template and all five reference docs, follows gen-features.cjs's
fail-closed contract, and is wired into regen:derived and lint:generated-sync.

The Status lifecycle section was missing from all four translations — the
section documenting the status enum behind #3853 — and is now generated into
every locale. Field cardinality is a new generated table: pure schema data,
no prose, so nothing to lose.

The Field-reference and Status-values tables are parity-CHECKED rather than
generated. Their Purpose, When-populated and Matched-text columns are
genuinely hand-translated per locale, and §8.8 itself says prose stays
hand-translated; generating them from an English registry would overwrite four
locales' translations on every write. The row set is checked against the schema
instead, so a key added to one and not the other fails, which is what field
drift actually means. Building that check found last_activity_desc
undocumented in all five tables.

Three keys the docs describe are absent from the schema — active_phase,
next_action, next_phases. They are grandfathered by name, not by wildcard, so a
fourth fails: a declared gap with a forcing function rather than a silent one.

Refs #3873

* fix(#3873): declare what the parsers do, and close the shape-parity gap

Two declarations in the new schema described intended behavior rather than
actual — the defect class this epic exists to end, committed inside the epic.
Both were caught by executing the parsers instead of reading their docstrings.

current_plan.acceptedShapes claimed ['N', 'N of M']. Standalone, the hybrid
shape errors; the path that looks like support is parseInt truncating '2 of 5'
to 2 and discarding the rest. Narrowed to ['N']. The parser is deliberately NOT
fixed here: that is #3784 and PR #3791 is already doing it. When #3791 lands
this row must widen, and the shape test will go red until it does — the schema
and the parser cannot drift apart quietly, which is what §8.8's checked-not-
generated rule is for.

STATUS_LIFECYCLE_ENUM claimed to be the closed set status can hold.
normalizeStateStatus passes unrecognized prose through unchanged, so it is not
closed at runtime. The seven members are the canonical values it maps onto; the
docstring now says that and the test asserts the real lenient contract.

Closes the acceptance item that a test asserts the parsers accept exactly the
declared shapes: the check is table-driven over every row carrying
acceptedShapes, guarded against passing vacuously on an empty set, and fails
loudly if a future row has no registered driver. Adds the unwired-label throw
and the fast-check property that every projection agrees with its schema row.

Refs #3873

* fix(#3873): keep the shipped template's frontmatter first, and make row 27 able to fail

The remote matrix caught 12 failures with one cause. Making the template's
frontmatter a generated region wrapped it in its own yaml fence ahead of the
markdown fence, so extractFileTemplate and readShippedStateTemplateBody — which
both match the single markdown block — found the heading first, not the
frontmatter. That breaks the contract every new project's STATE.md is created
from: bug #21 and epic #1969 B8 pin that the File Template block starts with
frontmatter and carries gsd_state_version.

The markers now sit inside the single markdown fence, so the fence opens before
the frontmatter and the region still ends ahead of the heading. Same layout as
before this phase, with markers embedded rather than a second fence.

Row 27 existed to catch exactly this and did not, because it was writer-seeded:
it asserted against the generator's own output shape, so it passed on the broken
template. It now parses the fence the way production does and was verified to
fail against the broken shape before being trusted against the fixed one. A test
that would not have caught the bug it exists to prevent is worse than no test.

The emitted-attribution failure was separate and the fragment was the wrong
remedy: gsd-core/templates/state.md self-attributes under a verbatim-copy
identity rule, so a diff touching it needs no acknowledgment. Fragment deleted
rather than left explaining nothing.

Refs #3873

* docs(#3873): how to change the STATE.md schema

The phase gate was right and my docs artifact was wrong. I listed
lint:generated-sync as the second enablement step, which is a verification
command dressed as one, and then claimed a one-step sequence owed no how-to.

The real sequence is build:lib then regen:derived, and the ordering is a trap:
the generator reads the COMPILED schema, so regenerating before building
regenerates against the previous schema and commits artifacts that look
plausible while disagreeing with the code just written. A reference table
cannot carry an ordering dependency; that is what the how-to test is for.

The page covers adding, changing and removing a key, every reason code the
check emits and what to do about each, what is generated versus hand-translated
and why the two prose-bearing tables are parity-checked instead of generated,
adding a language, and the three grandfathered keys. Indexed from docs/README.md.

Refs #3873

* chore(#3873): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-26 01:57:47 -04:00
Tom Boucher
63abcface9 feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools (#3831)
* feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools

The predecessor package get-shit-done-cc publishes a colliding gsd-tools bin whose phases.clear DELETES where this package's ARCHIVES, and both print success-shaped output against a gitignored .planning/ -- which is how #3129 cost a user 43 phase directories with no error and nothing recoverable from git.

The launcher's PATH branch now resolves gsd_run, published only by this package and self-locating via its own symlink chain to the sibling shim, instead of the colliding gsd-tools. A foreign handler becomes unreachable from PATH, and when no gsd_run is reachable the resolver fails closed rather than falling back -- that fallback was the vulnerability. This is smaller than the branch it replaces, which matters: the preamble is inlined into 113 shipped files and agents/gsd-verifier.md sits 2 bytes under a red-line size cap.

unset -f gsd_run leads the preamble so a re-source is idempotent. Without it, command -v finds the shell function, returns a bare name, and the resolver falls through to an exit 1 that kills a sourced caller's shell.

Adds gsd-tools runtime-identity, a manual diagnostic reporting this runtime's package coordinates over the baked package-identity (#498) and readHostVersion, with a strict total classifier: only a JSON object with an exact packageName verifies, since JSON.parse admits 0/"str"/[]/null/true.

An inlined identity assertion was built and reviewed first, then withdrawn -- it breaks five frozen size ceilings and no assertion fits in 2 bytes.

Closes #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3146): stop sync:launcher relocating a deliberate preamble placement

Pre-existing defect, surfaced by this PR because sync is a no-op unless the snippet content actually changes. transformFile inserts the preamble into the first block that CALLS gsd_run, but gsd-core/workflows/explore.md deliberately places it in a bootstrap-only block that DEFINES gsd_run without calling it -- its own comment explains why: declining the research offer must not leave Step 5's commit call unbootstrapped. Stripping empties that block of calls, so the preamble migrated forward and broke the define-before-use invariant tests/explore-command.test.cjs pins.

Reproduced on a pristine origin/next checkout with the base snippet and base file, so this was not introduced here. The insertion target now honours a block that already carried the preamble, falling back to the first calling block for files that have none yet. Adds a behavioral regression test over a two-block fixture.

Also updates three runtime-launcher-parity tests that pinned the removed PATH fallback to gsd-tools. Their intent is preserved -- the PATH stub is renamed gsd_run so it is reachable by the new resolver, and the RUNTIME_DIR-wins test still asserts the stub is never invoked. Fixture shebangs move to an absolute /bin/sh, because the fixture PATH is deliberately restricted and #!/usr/bin/env sh could not resolve.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3146): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3146): document the FEATURES.md section-numbering practice

The monotonically increasing section number in docs/FEATURES.md is the most frequent merge-conflict source in this repo, and it has TWO conflict cells, not one: the ### N. heading and the hand-maintained table of contents. Two PRs adding differently numbered features still collide on the TOC, so renumbering alone does not make a branch safe. This branch alone was renumbered 165 -> 166 -> 167 -> 168 across successive rebases.

Adds a CONTRIBUTING section stating the practice: allocate the number last, never pre-emptively renumber, take max+1 after a rebase and update the TOC in the same commit, and never renumber someone else's section. Fork contributors are told explicitly they may leave the number to a maintainer at merge rather than chasing the counter. Agents are told to lease the allocation and to include the file in their published touched set.

Records the durable fix as planned rather than pretending it exists: FEATURES.md should be generated from per-feature fragments the way CHANGELOG.md is generated from .changeset/, and the way tests/emitted-drift-acks/ works (#2914).

Also renumbers this branch's own section to 168, leaving 167 to the PR already in flight.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 20:57:16 -04:00
Tom Boucher
c933184b97 enhance(#3172): require a stated failing direction for every automated acceptance command (#3825)
* test(#3172): failing-first suite for the stated failing-direction probe

Pins the <fails_when> pairing walk, placeholder denylist, MISSING sentinel
exemption, degraded-read contract, CLI arm and the plan-authoring contract text.
RED by construction: the module exports it requires do not exist yet.
Executed on the remote runner.

* feat(#3172): require a stated failing direction for every automated acceptance command

Every runnable <automated> command now carries a <fails_when> sibling naming
what output constitutes failure. A command with no expressible failure mode is
not an acceptance test: it reads as rigour and is not falsifiable.

- verify-command-grounding gains a failing-direction probe sharing the existing
  <automated> grammar, MISSING sentinel and walk guard rather than copying them
- gsd-tools check verify-failure-directions <N> backs it; plan-phase dispatches
  it and hands the JSON to gsd-plan-checker check 8f
- Dimension 8 detail extracted to references to stay under the agent size cap

Verified on the remote runner.

* fix(#3172): close four review findings in the failing-direction probe

- MISSING_SENTINEL_RE matched an env-var assignment prefix (MISSING=1 cmd), so
  a real command was exempted from the new blocking gate. Tightened the SHARED
  constant rather than adding a second copy.
- Both token regexes scanned to EOF on unclosed openers (O(n^2), 1562ms at 40k).
  Bodies are now non-crossing; 1ms, byte-identical on well-formed input. The
  pre-existing AUTOMATED_BLOCK_RE carried the same defect and is fixed here too.
- probePhaseFailingDirections reported status 'ok' when one plan was unreadable,
  conflating 'could not look' with 'nothing to report'.
- Extracted the phase-resolution block both check arms had copied verbatim.

Also corrects a docs/AGENTS.md dimension list stale since #2401.
Verified on the remote runner.

* fix(#3172): project the planner rule onto the spawn contract, settle emitted bookkeeping

The remote runner refuted the planner-side edit. agents/gsd-planner.md is frozen
under a 49152-LF-char cap asserted by four suites and sat at 49,146 — six chars
of headroom — so the +537 of authoring rule blew it. #3297/#3645 already settled
where such a rule goes: the planner spawn contract in plan-phase.md, beside
<tracked_source_paths>. The agent file is reverted to origin/next verbatim.

- plan-phase.md gains <failing_direction_contract>; tests row 30 now asserts the
  contract there and row 30b guards the freeze in both directions
- plan-phase.md growth acknowledged by APPENDING to the 3409 fragment, per the
  precedent that two ack sources may never name the same path
- install-tree fixtures regenerated for the three new reference files

Verified on the remote runner.

* chore(#3172): backfill PR number into the changeset fragment

pr:0 -> pr:3825 now that the PR exists.

---------

Co-authored-by: sim <sim@local>
2026-08-24 19:05:11 -04:00
Tom Boucher
596540f864 feat(#3227): publish machine-readable state contract at step boundaries (#3824)
* feat(#3227): publish machine-readable state contract at step boundaries

Adds src/state-contract.cts, a best-effort publisher that writes
.planning/state.json (contract 1.0.0) at 11 step-boundary commands, so
external tools read a versioned contract instead of parsing STATE.md and
ROADMAP.md heuristically.

Composes existing owners rather than re-deriving: phase rows come from a
new locateProgressTable extracted from deriveProgressFromRoadmap (so the
snapshot can never disagree with GSD's own progress counters), milestone
identity from getMilestoneInfo, and next from classifyProject. Owners are
required lazily to avoid the state -> state-contract -> smart-entry ->
state require cycle.

Also fixes a pre-existing defect in scripts/lint-test-file-count.cjs
(maintainer-approved as a second concern): testEffectivePrefix never
stripped the suite qualifier, so 65 dotted test files counted against no
module and 9 mis-bucketed into a shorter one. Allowlist re-baselined for
the 74 files the gate can now see.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3227): backfill PR number into the changeset fragment

pr:0 -> pr:3824 now that the PR exists. Doc-only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3227): shape hostile-name fixtures away from the scan corpus

The two hostile-input fixtures used a literal phrase from
scripts/prompt-injection-scan.sh's corpus, so CI's Security Scan redded on
this file. These tests assert that an arbitrary phase name round-trips into
state.json as inert data -- the property holds for any string, so the
injection flavor is illustrative, not load-bearing.

Reshaped to a hyphenated fake instruction tag, which stays hostile-looking
while matching none of the scanner's patterns. Allowlisting the file was
rejected: that mechanism is for suites whose subject IS injection defense,
and it would blind the scanner to this whole file permanently.
See DEFECT.PROMPT-INJECTION-SCAN-COLLISION.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3227): ratchet the state-contract mutation floor to its measured score

The module was registered at minScore 50, the ratchet's minimum permitted
floor for a newly-registered module whose score had not been measured. This
PR's own Stryker shard measured 66.25% (run 32769289750, job 97565813640),
so the floor moves to floor(measured) - 1 = 65, per the rule the registry
documents.

66.25 is below TARGET_MUTATION_SCORE (80), so this stays a ratchet
candidate: raise as the tests improve, never lower.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 17:56:02 -04:00
Tom Boucher
4af59f8dd3 fix(#3662): resolve managed hook node runners at hook-fire time (#3790)
* test(#3662): failing-first suite for runtime-resolving hook runners

* fix(#3662): resolve managed hook node runners at hook-fire time

* fix(#3662): close review findings and document the resolver

* fix(#3662): close adversarial and security review findings

* chore(#3662): backfill changeset pr number

* test(#3662): honor win32 skip return and platform-aware sh runner pin

* test(#3662): pin the bare win32-claude sh-hook shape omitting the bash runner

---------

Co-authored-by: sim <sim@local>
2026-08-24 00:07:06 -04:00
Tom Boucher
107eb8c1d9 feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on dacae9273 while the PR that caused it (#3746) was green on every
check.

The docs-lint job in .github/workflows/docs-required.yml -- an ALREADY-REQUIRED
context -- now selects and runs the docs guards that read the specific docs files
the PR changed.

  scripts/docs-guard-registry.cjs    test file -> the docs paths it reads (63)
  scripts/select-docs-guards.cjs     pure (changedPaths, registry) -> test files
  scripts/lint-docs-guard-registration.cjs   drift guard, wired into lint:ci

scripts/ci-test-scope.cjs is NOT touched -- `git diff origin/next --` on it is
empty -- so #764's saving stands and its 21 pinning tests are untouched.

Selection: exact path; trailing-slash directory prefix (boundary-checked --
docs/adrenaline.md does NOT match docs/adr/, which a naive startsWith gets
wrong); and '*' for the 6 entries that walk docs/ generally or read a computed
path. Unknown maps to '*' -- guessing narrow is how a guard silently stops
running. Measured: a typo fix selects 6 of 63; docs/AGENTS.md selects 12;
docs/COMMANDS.md selects 18.

Four things this got wrong first, each found by an independent reviewer or by
probe, and each having been asserted safe in a comment:

1. The registry started as a RULE in ci-test-scope.cjs's RULES, on the theory
   that classify()'s !codeChanged normalization made it inert. True for
   docs-ONLY diffs; false for MIXED docs+code diffs, where codeChanged is true
   and the normalization never runs:

     node scripts/ci-test-scope.cjs --files "docs/a.md src/semver.cts"
       with the RULE:  25 targeted_tests
       origin/next:     3 targeted_tests

   Category error: RULES is the scoped lane's input; a docs-guard registry is a
   lane manifest for a consumer that never calls classify(). Extracted; pinned
   by value.

2. The second attempt was a dedicated workflow with paths: [docs/**]. Such a
   workflow never reports on a non-docs PR, so it can never be a required
   context without hanging every non-docs PR -- and a non-required check does not
   block a merge, so the guard would have been advisory and #3753 unfixed.
   docs-required.yml already has no paths: filter, already supplies the required
   docs-lint context, already computes docs_changed, and already ran one docs
   guard gated on it. Generalizing that step needs no ruleset edit at all.

3. The registry and the drift lint were built from ONE path-segment heuristic, so
   both were blind identically -- and blind at the guard that motivated the issue.
   The reader-call regex required a character BEFORE its keyword, so a callee
   named exactly read( / load( / parse( / doc( / file( / content( could never
   match; and only an INLINE path.join(ROOT,'docs','X.md') argument was caught,
   missing the two-step-via-variable form -- the MAJORITY spelling -- plus
   template literals and concatenation. Detector 1 fired on 14 of ~450 files, so
   35 genuine guards sat unregistered while the lint reported 0 violations,
   including cursor-reviewer (reads docs/COMMANDS.md, asserts
   .includes('--cursor')) and inventory-headings-countfree. The "accepted blind
   spot" this shipped with was the common case, not a fringe.

4. With detection fixed the true population is 115 files: 63 genuine guards, 52
   incidental. Running all 63 in a REQUIRED check on a one-line typo fix is the
   cost #764 exists to avoid -- install.test.cjs is 7840 lines and reads exactly
   one docs file, docs/AGENTS.md, for its frontmatter. Dropping it reproduces the
   bug; running it for a typo elsewhere is waste. Hence the map.

Then a second review round found six more, all fixed here:

- fragment-single-edit-propagation.install.test.cjs was EXEMPTED as
  "overlay fixture only". False: it reads the real docs/registries/eos.json and
  asserts on a registry entry name, and reads the real ADR-0001 and asserts its
  H1. A docs-only PR touching either would have gone green and red next -- #3753
  shipping again, from inside the fix for it. Now registered against both paths,
  and all 52 remaining exemptions were re-audited one by one.
- The SUITES-collision guard compared RAW registry keys, but run-tests.cjs strips
  a leading `tests/` BEFORE its suite check. So it caught 'all' and missed
  'tests/all' -- the only spelling that can actually occur, since every key
  carries the prefix. One typo would have run all 824 test files inside the
  required job. Now normalized the same way run-tests.cjs normalizes.
- The lint failed OPEN on an unreadable tests dir or candidate file: 0 violations,
  ok:true. A guard that cannot read its input must never report success.
- The exemption ratchet gated identity only, so a baselined file that later
  STARTED asserting on shipped docs stayed exempt silently -- 52 permanently blind
  files. The baseline now fingerprints the docs paths each exempted file
  references and fails when that set changes, naming what changed.
- The exemption marker was still honored inside a multi-line template literal in
  the header window. The scanner now tracks template-literal and block-comment
  state.
- `git diff --name-only | grep '^docs/'` silently dropped C-quoted non-ASCII docs
  paths, making docs_changed=false a green zero-guard check. Both call sites now
  pass -c core.quotepath=false.
- The run step was gated on hashFiles(), which a force-committed
  .docs-guard-tests.txt would satisfy. The step now rm -f's both scratch files
  first and gates on an output it sets itself.

Three empty states, deliberately distinct, because conflating them rebuilds
#3753: an empty or malformed registry HARD-FAILS; docs changed with no guard
covering them logs and skips; no docs change is already gated. The middle state
must never be expressed as an empty --files-from, which prints `no tests in suite
"all"` and exits 0 -- a green check that guarded nothing. With the current
registry that state is unreachable, because the six '*' entries always match;
the branch is kept as defensive handling for a future registry and says so.

timeout-minutes: 15 bounds the required job against a hanging fork-supplied test;
it had none. npm ci was added because the job never installed dependencies -- the
previous single-file step got away without it, the registry does not.

docs/contributing/docs-guard-registration.md documents the rule, following its
sibling cross-platform-portability-rules.md, and CONTRIBUTING.md's CI Test
Quality Checks table links to it. It is also load-bearing: without a docs/ file
in the diff this PR would not have triggered its own lane, shipping an
unexercised change to a required check.

One unrelated fix, included because this PR surfaced it and CLAUDE.md forbids
deferring a defect found while working. On this branch's first CI run,
`full test (windows-latest, 24, shard 3/3)` was CANCELLED at exactly 30 minutes;
tests were still passing 0.8s before the cancel, so it is a wall-clock timeout,
not a hang, and a cancelled job reddens `Required tests`.

The cause is not this PR's test file, which costs ~60ms. Shard composition is
unstable: adding ONE file to the unit suite reshuffled 115 of 268 files between
shards, and shard 3 drew a heavier mix. Underneath that is a real pre-existing
defect. tests/ci-test-job-timeout-budget.test.cjs requires every lane's budget to
be >= 1.5x its MEASURED cost -- "a lane that got slower must be re-budgeted, not
excused" -- and its test-full entry recorded 19m from a windows-22 shard. That is
stale. Measured on `next` with none of this PR's changes present: 26m18s (run
32614439702, windows-latest/24 shard 3/3), 23m36s and 23m17s on shard 2/3. So the
lane costs ~26m and the 30-minute cap carried 1.14x headroom, not 1.5x. The gate
had been out of compliance with its own rule; this PR was merely the file
addition that reshuffled shard 3 past the cliff.

Fixed as that file prescribes: measuredMinutes 19 -> 27 with fresh evidence, and
test-full timeout-minutes 30 -> 45. The rule's minimum for 27m is 41; 45 is
deliberately above it because the reshuffle means per-shard worst case moves run
to run, and a budget pinned to the exact minimum would be re-breached by the next
test file anyone adds. Only that one job's timeout changed; test.yml's scope,
matrix and steps are untouched, so #764's saving is unaffected.

Raising that cap let the Windows shard finish (28m45s, inside 45) and uncovered
a real failure the 30-minute cancel had been masking:
`new quick-task branch branches off origin/main (#2916)` died with
`outcome=timed_out exitCode=null`, SIGTERM, at the 15000ms bound.

tests/quick-branching.test.cjs:149 `runStep` runs a `#!/usr/bin/env bash` script
executing MULTIPLE git commands, but was bound to GIT_TIMEOUT_MS (15000) -- the
norm for a SINGLE git plumbing call. tests/helpers/timeouts.cjs already documents
this exact failure and exists to fix it: HOOK_FANOUT_TIMEOUT_MS was created after
PR #3285 recorded "outcome=timed_out exitCode=null at exactly the 15000ms probe
bound while every other lane passed the same commit", and calls that "a bound
sized for the wrong class, not a slow machine". Our failure is that case
verbatim, so both sites move to the class norm rather than to a bigger number.

The same class also failed on `next` itself 21 hours earlier -- run 32608945654,
windows-latest/24 shard 1/3, `plan touching only src/ in a submodule project
keeps worktree isolation ENABLED` -- where tests/worktree-safety.test.cjs:5845
`runGate` fans out to `git config --file .gitmodules` under a hardcoded 30000.
Fixed too, since it is a defect in the tree regardless of which branch surfaced
it.

A survey of the whole tests/ tree found the same class-mismatch at further
bash fan-out sites bound under 60000ms, and the maintainer approved sweeping
them rather than leaving them latent to surface the same way one at a time. 16
fan-out sites across 16 files now use the class norm.

The sweep is class-correctness, not raising numbers until things pass. Sites
were moved ONLY where the bash body demonstrably spawns something (git, node,
npm, a CLI); self-contained shell snippets were left where they are, and are
listed as deliberately unchanged: pure if/printf bodies (copilot-install), pure
array/case builtins (code-review-pipeline-regression:638), a documented
pure-shell gsd_run stub (host-integration), single-process hook calls
(workflow-guard:222/271/302), and a deliberately tight 5000ms fast-check hook
(gsd-write-guard.property). Nothing was lowered. process-seam.test.cjs:513
(literal 300) is untouched on purpose -- it tests timeout BEHAVIOR, so raising
it would destroy what it asserts.

Shared file-level constants were the trap here, and were handled per file rather
than by redefinition: GIT_TIMEOUT_MS has ~15 users in git-base-branch and only 1
is a fan-out; WORKTREE_TIMEOUT_MS has 16 users in worktree.test.cjs and 3 are;
PROBE_TIMEOUT_MS has several in three more files. In each the CALL SITE was
changed and the constant left alone, so no single-plumbing-call site silently
inherited a 60s bound. The one exception is hooks-opt-in.test.cjs, where
HOOK_TIMEOUT_MS has exactly one consumer -- spawnHook, the fan-out itself -- so
redefining it is identical in effect and reads better.

Only two of these sites have actually been observed failing. The rest cite that
shared class and those two run ids rather than inventing evidence of their own.

Co-authored-by: sim <sim@local>
2026-08-23 21:21:21 -04:00
Behruz Nassre Esfahani
a44d513566 fix(#3712): confine in-process installs to a sandboxed HOME (#3725)
* fix(#3712): confine in-process installs to a sandboxed HOME

A runtime kind may declare a global `home` override resolved from os.homedir()
rather than from the caller's configDir — codex's skills kind (`home: ".agents"`,
ADR-1239 / #2088) is the only live case. Sandboxing configDir/targetDir does not
contain it, and assertDestWithinConfigHome cannot see the class: that gate
confines a destSubpath to whatever root it is handed, and here the root IS the
escaped home. So an in-process caller that forgot to sandbox HOME wrote to, and
pruned gsd-* entries from, the developer's REAL ~/.agents/skills.

tests/agent-descriptor-parity.install.test.cjs's K1 loop did exactly that: it
iterates every agents-kind runtime (codex included) with a sandboxed targetDir
and an un-sandboxed HOME. Reproduced against a canary home on next @ adb46cdd8 —
71 gsd-* skill dirs deleted, a foreign `cloudflare` skill surviving, suite still
exit 0. It is silent because the runtime's own config home is untouched, so the
manifest keeps reporting a healthy install.

FIVE writers resolve a kind `home` and then destroy under it. Three are reachable
today — installRuntimeArtifacts, uninstallRuntimeArtifacts (install-engine.cts)
and applySurface (surface.cts). Two are descriptor-dependent and guarded against a
future descriptor change rather than a present escape: installOpencodeFamilySkills
(behind the combined-family early return) and installAgentsKindStandalone. Those
two are scoped to the single kind each destroys — passing the whole layout made
codex's unrelated skills override trip a writer that never touches it.

- src/test-home-guard.cts: refuse when a run under a test runner cannot be shown
  to have sandboxed HOME. NODE_TEST_CONTEXT (set by `node --test`) gates it, so
  installs outside a Node test context are untouched; GSD_TEST_MODE is unusable,
  as several candidate files including the offender never set it. Homes are
  compared by FILESYSTEM IDENTITY (st_dev + st_ino), not by pathname:
  path.resolve() resolves neither symlinks nor case, and realpath returns a
  canonical pathname that two routes to one directory can still disagree on (bind
  mounts). Verified on macOS/APFS — HOME=/users/<name> made the strings differ
  while naming the same directory, and the lexical form ALLOWED a write into the
  physical real home. FAILS CLOSED: a pair is "different" only when both identify,
  or one is definitively absent (ENOENT/ENOTDIR) while the other identifies; every
  other errno is "cannot tell" and refuses. Only when neither home identifies is a
  marker consulted, and it carries the sandbox PATH and must equal the home in
  effect — a boolean checked first let an ambient or stale value disarm the guard.
- helpers: promote sandboxHome() out of its two byte-identical private copies,
  which is also what makes them record the sandbox; the three withFakeHome()
  helpers record it too. The marker NAME is duplicated as a bare string rather
  than required from the compiled guard, keeping helpers.cjs's documented
  no-built-lib-at-import-time contract; a test pins the two together.
- agent-descriptor-parity: sandbox HOME across the K1 loop.
- helpers-process-isolation: #3156's canary asserts on <home>/.gsd only, and its
  `--cursor --local` spawn cannot reach `.agents` at all, so an assertion added
  there would pass with all confinement removed. Add a discriminating row — a
  `--codex --global` spawn against a seeded ambient home — which also asserts the
  runtime still declares the override. Its check is a sampled inventory (dir names
  + each SKILL.md), not a tree compare.
- install-write-confinement: predicate rows through the deps seam, covering the
  symlinked HOME, ambient and stale markers, and each sameDirectory branch
  (both-identify, one-absent, neither-identifiable), plus wiring rows that drive
  the REAL entrypoints so deleting a guard call site is red.

Verified: guard fires end-to-end against a real un-sandboxed HOME (exit 1, zero
deletions); the case-variant fail-open reproduced on APFS before the fix and
refuses after; K1 file 29/29 green with skills intact; mutation-tested — each of
the three reachable call sites, lexical-only comparison, and treating an unknown
errno as "absent" each take exactly one row red, with every mutation echoed back;
the process-isolation row negative-controlled by reverting installerEnv to its
pre-#3156 leak (16/0 -> 13/3); a full npm test leaves ~/.agents/skills at 71.

Stated residual: the two descriptor-dependent writers have no wiring test, because
no runtime declares a `home` override on those kinds and neither can be exercised
without inventing a descriptor.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3712): add changeset

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3712): let a sandbox nested inside the real home through the guard

All six Windows shards of #3725 failed on legitimately sandboxed
destinations. On Windows os.tmpdir() is %LOCALAPPDATA%\Temp — inside the
user's home — so every sandbox a test creates is a descendant of the real
home, and "does this land inside the real home?" answers yes for the safe
case and the dangerous one alike. POSIX conceals this: /tmp and
/var/folders both sit outside $HOME.

Add the missing conjunct: a destination inside the real home is allowed
only when it also sits beneath a HOME that was sandboxed away from the
passwd home. Both halves are required — dropping the first re-admits a
plain un-sandboxed install, and dropping the second decays into the
"is HOME sandboxed?" check the module rejects, which a layout resolved
before the sandbox walks straight through. Each is mutation-proven by a
row that goes red without it.

Also covers the two fail-closed branches of the new exemption, which
survived mutation to `true` with the suite green, and avoids `<user>` in
a docblock — the prompt-injection scanner reads it as a delimiter tag.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): close three writer/rollback gaps found reviewing the whole PR

Cross-AI review of the full PR (not just the round's delta) surfaced
three ways the guard could still be defeated:

- The nested-sandbox exemption trusted the SPELLING of a destination.
  With HOME sandboxed to a directory inside the real home — legitimate on
  Windows — an aliased `.agents` (symlink, junction, subordinate bind
  mount) beneath it redirected an allowed path into the real home. Decide
  containment on the path the write RESOLVES to: walk up to the nearest
  existing ancestor, canonicalize, re-append the tail.

- `migrateLegacyDevPreferencesToSkill` is a SIXTH writer that resolves a
  skills-kind `home` override. It creates rather than prunes, which is
  why it was missed, and `_runLegacyInstallMigrations` runs it before
  `installRuntimeArtifacts`' own assertion. Guarded, scoped to that kind.

- Worst of the three: `bin/install.js` snapshots the resolved skills root
  before installing, and its outer catch rolls back by deleting and
  recreating every snapshotted `gsd-*` directory there. The guard's own
  throw landed in that catch, so refusing an un-sandboxed codex install
  provoked exactly the mutation the guard exists to prevent. Refusals are
  now marked and rethrown without rollback — nothing was written, so
  there is no partial install to undo. Every other error still rolls back.

Also carries the sandbox marker into `installSpawnEnv`, so spawned
installers are not refused on passwd-less CI images, and corrects three
claims that no longer hold: "every writer" (six, and named), the
unconditional "fails CLOSED" (the passwd-less marker branch is a
deliberate weakening, and TOCTOU is out of scope), and the assertion that
Windows os.tmpdir() is always %LOCALAPPDATA%\Temp (Node honors TEMP/TMP).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#3712): name the guard's two limits instead of overclaiming

Round-2 review found the prose had drifted ahead of the code. Corrected,
with no behavior change:

- The module still said FIVE writers; there are six, and the sixth is
  now named along with why it was missed (it creates rather than prunes)
  and why it carries its own assertion (it runs before the main one).
- The canonicalization docblock listed subordinate bind mounts among the
  aliases it closes. It does not close them: a bind mount is not a link,
  so realpath keeps the mount-point spelling. `sameDirectory` already
  recorded that limit; the new helper now inherits it explicitly rather
  than contradicting it. Closing it needs mount-table introspection.
- "FAILS CLOSED" was unqualified while the passwd-less marker branch is
  a deliberate weakening — with no passwd entry, nothing can contradict a
  marker naming the real home.
- "Refuses BEFORE any write" was too broad: legacy install migrations run
  ahead of the layout-driven ones, which is exactly why the two
  rollbackInstallerMigrations() calls still execute before the rethrow.
  Only the codex skills-root rollback is skipped, and that is the only
  _codexPreConfigRollback() call site — applySurface is never called from
  bin/install.js and uninstall cannot reach it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): sandbox HOME in the opencode-family home-override parity rows

The last two Windows failures, and the same platform asymmetry in a
different disguise. This row drives a skills-kind `home` override on
purpose — precisely what the guard polices — but relied on the override
temp dir happening to sit outside the real home. It does on POSIX
(/tmp, /var/folders); on Windows os.tmpdir() is under %USERPROFILE%, so
the guard correctly refused and only Windows went red.

Declare the sandbox instead of depending on the platform: HOME becomes
the override itself, which is the home the call writes under. This is the
fix the guard's own message prescribes, applied to the test rather than
to the guard.

Both failing Windows shards fail on exactly these two rows and nothing
else; every other shard is green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): make sameDirectory answer NO when it cannot tell

Review Major 1. sameDirectory()'s only caller is the passwd-less marker
branch, which reads a `true` as permission to PROCEED:

    if (marker && sameDirectory(marker, osMod.homedir())) return;

The fallthrough returned `true` whenever neither side identified — two
absent paths, or two stats failing EACCES/EPERM/EIO on a locked-down
host — on the reasoning that "cannot tell" should make the caller refuse.
That reasoning was inverted with respect to this caller: it turned the
passwd-less escape hatch into an unconditional bypass for any marker
value at all, on precisely the hosts the fallback exists to serve. Only
two things now answer yes: one resolved pathname, or two readable
identities that match. Restoring the old fallthrough takes the new row
red.

Also from review:

- Major 2 asked whether st_dev/st_ino discriminate directories on
  Windows, where Node derives them from BY_HANDLE_FILE_INFORMATION. The
  whole guard rests on that primitive, so assert it rather than argue it:
  a row comparing two distinct temp directories, and one directory
  reached by two spellings. It runs on every platform in the matrix, so
  Windows answers the question itself.

- Minor 1: the refusal now names the real home it compared against, not
  just the destination it refused. That is the one fact needed to tell a
  true positive from a false one, and its absence is what made the
  Windows case a CI-log dig rather than a glance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): refuse a HOME that merely spells the real home more widely

Review round: one Blocker, four Minors, a Nit.

N2 (the one with teeth) — a destination's ancestor chain is linear, so
"inside the real home AND inside the effective HOME" admits two
arrangements, not one. The intended `effectiveHome ⊂ realHome` is the
Windows temp shape; `realHome ⊂ effectiveHome` — HOME at /Users, /home,
C:\Users — is not a sandbox at all, it is the real home reached by a
wider spelling, and it was exempting a stale destination pointing
straight at ~/.agents. Third conjunct added; the docblock no longer
claims two conditions suffice. Removing the conjunct reds the new row
and nothing else.

N4 — the migration guard resolved its OWN layout, and without
capabilityRegistry, so a registry-dependent descriptor could make it
vouch for a path the migration does not write: a guard reporting safe
while the unsafe write proceeds. It now guards the destination already
resolved by _resolveDevPreferencesSkillTarget, keyed on
`installRoot !== targetDir` — which is exactly the condition under which
a `home` override was declared, read off that same result.

N1 — CONTEXT.md gains the Test Home Guard Module glossary entry that
contributor-standards.md requires of a new Module. Not CI-enforced, so
green CI was never evidence it was met.

N3 — the docs/INVENTORY.md row was misfiled between install-fs-adapter
and install-model-override-resolver; the table is alphabetical and the
manifest already had it right. Moved, and its text now names six writers
and the third conjunct.

N5 — applySurface's signature docblock was two parameters stale; this PR
added the second of them.

N6 — the duplicated rollbackInstallerMigrations() adjacent to the new
rethrow: two identical consecutive calls, not two phases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): guard the sixth writer, and close two false-ALLOW paths

Review round 3 (NEW-1, NEW-2) plus three defects Codex found in the
whole-PR pass, each reproduced before it was fixed.

NEW-1 — migrateLegacyDevPreferencesToSkill called the guard with no
`deps`, so it bound real os/process.env and could not be wiring-tested
the way the other three reachable writers were. It now takes the same
optional `deps: { os?, env? }` tail parameter. The wiring block gains
the missing fourth row, and a fifth pinning the ALLOW half; the test
file's header docblock said "FIVE writers ... the three reachable
today", contradicting the six/four statement this PR already put in
src/test-home-guard.cts, CONTEXT.md, docs/INVENTORY.md and the
changeset. Both directions of the guard's condition now fail a row
when broken — previously neither did.

NEW-2 — derivesFromSandboxedHome's docblock claimed "THREE conditions
are required, and no two of them suffice". False for {2,3}: isInside is
reflexive, so whenever conjunct 1 fires conjunct 2 already returns
false on its own. Reworded as a fast path, which is what it is.

Codex 1 (false ALLOW) — on a host with no readable passwd entry the
marker branch returned as soon as the marker matched the effective
HOME. That attests a caller sandboxed HOME and says nothing about where
an already-resolved destination points, so a layout captured before
sandboxHome() — still naming the real ~/.agents — was waved straight
through: the same stale-layout shape the primary branch refuses by
design. The marker must now identify AND contain every destination.

Codex 2 (false ALLOW) — `installRoot !== targetDir` was the stand-in
for "the skills kind declared a home override". The two are not
equivalent: the inequality is false when the override resolves onto
targetDir itself, which is exactly a configDir of $HOME/.agents. The
guard was skipped and SKILL.md written into the real home under a test
runner. _resolveDevPreferencesSkillTarget now reports hasHomeOverride
off the same resolution instead of inferring it from two paths.

Codex 3 (prose) — the shared refusal message claimed every guarded
writer prunes; the migrate writer only creates. The changeset headline
claimed in-process installer calls can no longer reach the real home,
which is wider than the guard: writeNonClaudeDefaults still writes
~/.gsd/defaults.json through os.homedir(). INVENTORY's and CONTEXT's
fail-closed sentences omitted sameDirectory's pathname-equality
shortcut. All four narrowed to what the code does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): let the sandbox marker follow an overridden HOME

Found by Codex in the whole-PR pass. installSpawnEnv spreads
`overrides` last so an explicit HOME wins — deliberate, and its
docblock tells callers needing per-spawn isolation to pass their own
{ HOME, USERPROFILE }. But the #3712 marker was set before that spread,
so such a caller got HOME=<theirs> and marker=<helper default>. On a
host with no readable passwd entry the guard compares the two and
refuses a legitimately sandboxed spawn — tests/install.test.cjs:7143
and install-shared.cjs's own runInstaller both take that path.

The marker is now derived from the final HOME unless the caller
supplied one explicitly. The contract test asserted HOME after an
override but not the marker, which is why it stayed green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#3712): name the shipped guard condition, not the deleted one

Review round 4 of #3725. Two artifacts this PR adds still described
`target.installRoot !== targetDir` in the PRESENT tense as the live guard
condition on `migrateLegacyDevPreferencesToSkill`. The shipped condition is
`runtime && target.hasHomeOverride` (src/install-engine.cts:510).

This is not ordinary doc drift. The named condition is the exact false-ALLOW
the previous round closed: a `home` override resolving onto `targetDir` — a
configDir of `$HOME/.agents`, which is where codex's override points — makes
the inequality FALSE while the override is declared, so the guard was skipped.
A maintainer reading CONTEXT.md:290 as authoritative would believe the guard
still skips that case.

  - tests/install-write-confinement.test.cjs — the ALLOW-half row's comment.
    Its "teeth" rationale is unchanged and still correct as written.
  - CONTEXT.md:290 — the Test Home Guard Module glossary entry, a documented
    PR gate. Now states the condition and names the inequality only as what it
    is NOT, with the reason.

The three surviving mentions of the inequality are all past-tense or negated
(src/install-engine.cts:448, :508 and the sibling test comment at :3698) and
are correct as they stand.

Verified: `npm run lint:ci` exit 0; full `npm test` 31327 tests / 31312 pass /
0 fail / 14 skipped, run with TMPDIR unset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): canonicalization fails closed, matching identify's errno split

Codex full-PR review of #3725, run against the round-4 head.

`resolveThroughLinks` caught EVERY realpathSync error and fell back to
`path.resolve(dest)` — the lexical spelling. That inverts the function's own
purpose. An aliased `<sandbox>/.agents` that cannot be canonicalized keeps its
sandbox spelling, satisfies the nested-sandbox exemption at :227, and the write
is ALLOWED into the real home — the exact escape this walk exists to close. The
module documents that it fails CLOSED with ONE named exception (the marker
branch); this was a second, unnamed one.

Split by errno, and deliberately by the SAME split `identify` already draws
rather than a second policy in one module — both answer "does this path exist
as named?", so they must not disagree:

  ENOENT / ENOTDIR -> walk up. The ordinary case: a fresh install resolves a
    destination nothing has created yet, so realpath fails on the leaf and on
    every not-yet-created ancestor. Refusing here rejects every install.
  anything else (EACCES, EPERM, ELOOP, EIO) -> refuse. The component exists but
    cannot be resolved, so the guard cannot tell where the write lands.

Three rows in the predicate block, beside the other aliasing rows:
  - a symlink CYCLE in the destination path (ELOOP)   -> REFUSE
  - a destination that does not exist yet (ENOENT)    -> ALLOW
  - a component behind a regular file (ENOTDIR)       -> ALLOW

Teeth checked against the artifact the test loads, not the source: reverting
the condition to the swallow-everything shape in the compiled
test-home-guard.cjs turns row 1 — and only row 1 — red. The ENOTDIR row caught
a stale build during development, which is the point of asserting on the
compiled file.

CONTEXT.md and the resolveThroughLinks docblock both record the new behaviour,
so this does not repeat the prose-vs-code drift the round-4 finding was about.

The changeset's existing scope sentence now bounds "six writers" to the
`installRuntimeArtifacts` call tree and names `cmdGenerateDevPreferences` —
which resolves the same codex `home` override through `getGlobalSkillsBase` and
writes SKILL.md beneath it unguarded. It has no in-process caller today (its
only direct require-and-call is a spawnSync with HOME sandboxed), so it is
latent rather than live, and whether it belongs in this PR is raised with the
maintainer rather than decided here.

Verified: `npm run lint:ci` exit 0; full `npm test` 31330 tests / 31315 pass /
0 fail / 14 skipped, TMPDIR unset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 18:27:02 -04:00
Tom Boucher
14679b866b enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability

Binds the approved triage shape before any of it exists:

- containment — the execute:wave:post hook must not render unless
  workflow.live_dom_uat is true AND the capability resolves active
  (fail-closed on a missing state entry, and on a non-boolean value)
- criterion 4 — agents/gsd-executor.md carries no browser MCP family;
  asserted as an absence, which is the only way it is observable
- Hyrum guard — the pre-existing mcp__playwright__* branch must stay
  outside the key-gated block, or upgrading silently removes working
  automated UI verification for every current Playwright-MCP user
- parity — the browser glob list now lives in two surfaces (agent
  frontmatter + workflow detection block); the assertion fails if
  either gains or loses a family without the other

Red by construction: the capability, agent and workflow block do not
exist yet. Verified on the remote runner.

Refs #2856

* enhance(#2856): add default-off live-DOM UAT capability

A phase whose acceptance criteria needed a live DOM could not be
finished by the agent that executed it: gsd-executor carries no browser
tools, so it correctly returned checkpoint:human-action even though the
work was not human-only, just tool-less. Every such phase degraded to
"executed, then finished by hand in the orchestrator", and autonomous:
false could not distinguish "a human must judge this" from "the executor
lacks the tool".

Implements the shape approved at triage, not the one reported. The
executor's tools: line is NOT widened, in any configuration: for a
first-party agent the static list is the only control that exists
(ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one
default-off capability owns the key, the agent, and the step:

- capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat
  (boolean, default false), one additive step at execute:wave:post
  (onError: skip, gates: []), so it can never halt a wave
- agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP
  globs, in its own tools: line, with no Bash
- verify-work automated_ui_verification — a gsd:live-dom-families block
  naming both new families AND the key; presence alone never activates

Two independent fail-closed gates: isCapabilityActive renders a hook
only on state.active === true, plus the step's own `when`.

The pre-existing mcp__playwright__* branch keeps the gating it already
had and stays outside the new block. Pulling it behind a default-off key
would have silently removed working automated UI verification from every
current Playwright-MCP user on upgrade.

Also closes a host gap this surfaced: execute:wave:post dispatched only
contribution + gate, so ANY registered step was declared and silently
never run — exactly the single-kind hand-roll loop-hook-dispatch.md
names. Step 5.75 now dispatches every kind == "step".

The browser-profile lock is tolerated, not coordinated: --isolated is a
flag on the operator's own MCP-server registration that GSD neither
launches nor parameterizes, so the verifier reports could_not_look /
profile_locked, names the flag, and stops. DOM-VERIFY.md keeps
could_not_look and nothing_to_report distinct behind a closed reason
enum — collapsing them is the ambiguous-run-notes defect reported.

Verified on the remote runner.

Closes #2856

* fix(#2856): apply review findings from the orthogonal passes

Correctness pass (blocker):
- delete detectionBlockIsCrlfSafe. It was pass-always: it read the file,
  replaced LF with CRLF, then indexOf'd marker strings that contain no
  newline, so the replacement could not change the result and the
  assertion could never fail for the reason it stated. There is no real
  CRLF risk on this surface either — the gsd:live-dom-families block has
  no parser, only human and agent readers. Deleted rather than replaced,
  per the repo's pass-always-test rule.

Isolated security pass (two minors, both real):
- execute-phase.md step 5.75: this change is what first activates
  kind == "step" dispatch at execute:wave:post, which newly opens the
  ref.command shell path at that loop point. Our own step uses ref.agent
  and never touches it, but the door is now open, so the step-dispatch
  line carries the same in-context validate-before-shell warning the
  sibling gate-dispatch line directly below it already carries.
- gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker
  influenced. Require it wrapped in inline code or a fence, kept short,
  and never left reading as a directive to the next reader.

Verified on the remote runner.

Refs #2856

* fix(#2856): settle the new-agent roster ripple

Checkpoint 2 returned 28 failures, none in the new suite — all of them
the guards that exist to make adding an agent a deliberate act. Each is
a real boundary that had to move:

- docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526),
  so the browser globs lose their backticks; primary-agent counts 21->22,
  roster 33/34->34/35, Verifiers category 1->2
- docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md
  to be classified exactly once
- gsd-dom-verifier: add the anti-heredoc instruction and the commented
  hooks: frontmatter pattern both agent gates require
- gsd-core/bin/shared/model-catalog.json: every shipped agent needs a
  profile entry (#3229)
- copilot-install / kilo-upgrades / qwen-upgrades: expected agent list
  and the 34->35 roster boundary
- execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately
  carries one step now. Asserted as an exact shape — one step, capId
  live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a
  real guard against accidental change rather than being relaxed

Two findings worth naming:

mcp-tool-inheritance (#2526) rejected the agent for documenting
mcp__playwright__* while its tools: line withholds it — a dead
instruction that invites the agent to claim a path it cannot take. The
prose now names the Playwright MCP family without the dispatchable
token, in both the agent and the capability fragment.

runtime-launcher-parity rejected the new gsd_run call: each fenced block
is its own shell, so a workflow step file invoking gsd_run needs its own
canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs.
That script also normalizes explore.md, which is unrelated pre-existing
drift the parity check tolerates, so it is reverted to keep this diff
scoped.

The emitted-drift ack supersedes the spent #3370 entry for
execute-phase.md — it is merged into next, so its ripple is absorbed at
the base and it can no longer clear anything. That is the same supersede
the #3370 entry itself performed on the spent #3324 fragment. Its
unrelated execute-plan.md entry is untouched.

Verified on the remote runner.

Refs #2856

* fix(#2856): drop the stale emitted-drift ack entry

The automated-ui-verification.md entry was written speculatively rather
than from a reported growth, and the check names that precisely: an ack
"written or reworded in THIS diff, but nothing here needed it, so it
explains nothing".

The growth tier keys on the bare filename as it appears under
gsd-core/workflows/ or agents/. automated-ui-verification.md is nested
under verify-work/steps/, so it was never in the tracked set — only
execute-phase.md was ever reported, both before and after the launcher
preamble landed.

Only ack what the check actually reports.

Verified on the remote runner.

Refs #2856

* chore(#2856): backfill changeset pr number

pr:0 -> 3716. The placeholder fails both changeset-lint
(fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by
design and can only be resolved once the PR number exists. Both now
report ok against GITHUB_BASE_REF=next.

Refs #2856

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:21 -04:00
Tom Boucher
8da2dd3ad2 feat(#2790): add read-only planning.inspect schema-v1 snapshot query (#3708)
* feat(#2790): add read-only planning.inspect schema-v1 snapshot query

Adds a read-only query emitting a schema-versioned JSON projection of .planning/
so downstream harness UIs can consume planning state without parsing GSD's
Markdown a second time.

Composed strictly from the ADR-3180 section 7 owners plus parsePlanDocument,
parseRequirements and parseUatItems; markdown structure is read through the
Markdown Sectionizer and Markdown Table Model seams. It declares its own flat
external schema rather than serializing PlanningSnapshot, which is the
diagnostic-rule subject and still growing.

Extracts plan-document parsing out of cmdPhasePlanIndex into a shared leaf
module so phase.plan-index and planning.inspect cannot drift, including the
plan-id derivation both surfaces report.

Also fixes parseRequirements dropping the separator delimiter used by the
shipped requirements template, surfaced while wiring the requirement rows.

* fix(#2790): close spec gaps and a raw-text test assertion found in review

Review findings from the standards, spec and security passes:

- phases[] rows carry goal and dependencies, the two per-phase elements the
  issue Summary names that had no corresponding field. Goal is bounded to the
  section's leading prose so the Depends-on line, the Plans checklist and the
  wave annotations are not duplicated into it.
- requirement rows carry their own diagnostic codes, so a consumer no longer
  has to string-parse the global diagnostics subject to correlate.
- roadmap_acceptance.checkbox is looked up through the phase-id key owners.
  It was compared raw against the on-disk directory name, so it read null for
  every real-world slugged phase directory and the evidence channel was inert.
- the hostile-input test asserts the structured payload instead of matching the
  raw stdout string. The absence proof over raw stdout is kept deliberately.

* fix(#2790): register planning in the runtime usage list and repair fixtures

Remote runner reported 9 failures on 9b3f9aa. Two root causes, both fixed:

- gsd-tools.cjs registered the planning family in HOST_COMMAND_ROUTERS but
  never added it to TOP_LEVEL_USAGE's Commands list. Those are two surfaces a
  parity test guards, and the top-of-file block comment is not the runtime
  help string. A real wiring gap that every local gate and three review passes
  missed.

- the new suite's fixtures could not produce a resolvable phase set. STATE.md
  frontmatter omitted the milestone field, which ADR-3180 7.2 rule 1 makes the
  primary milestone selector, so the phase set scoped unscoped and every
  percentage was correctly withheld. Separately declarePhase returned a path
  without creating the directory, so a phase declared but never written to left
  phases empty. Both reproduced against the built module before fixing.

No assertion was weakened. The withholding path is still exercised and still
returns null when the roadmap is absent.

* chore(#2790): backfill changeset pr number

* test(#2790): cover every enumerated matrix row and contain a symlink escape

Reverses a silent deferral. An earlier revision left 23 of the 78 enumerated
matrix rows unimplemented and 7 more as one-off manual checks, with a paragraph
in the artifact and the PR body describing the gap. CLAUDE.md is explicit that
such a note is not a fix and is not surfacing. The rows are implemented instead
and the manual-evidence bucket is gone: 49 test cases become 88, covering all 78.

Writing the symlink row proved a real leak: a *-PLAN.md symlinked outside
.planning/ had its content emitted into the payload, confirmed via a direct call
and the spawned CLI. readDocument now resolves target and planning root with
realpathSync and rejects an escape, returning the ordinary unreadable-document
shape. Tested both ways, because a containment check that over-rejects is its own
defect: an escaping symlink leaks nothing and degrades that plan alone, while a
legitimately relocated .planning/ symlink stays fully readable.

The three new modules are registered in the mutation COVERED registry, which had
been reporting has_work false and skipping the Stryker gate entirely. Provisional
non-binding floors so the shards run and report; raised to the measured value
before merge, since the registry forbids calibrating from a local run.

* fix(#2790): satisfy the mutation ratchet contract and scope the 1MB test

Remote runner reported 16 failures on 8c451ed. Two causes.

The COVERED registry has a paired contract the earlier commit violated: every
module needs a matching RATCHET_BASELINE entry, and minScore must be between 50
and 100 with minScore === baseline. The provisional floor of 1 was illegal on
both counts. All three modules now sit at 50 — the registry's own enforced
minimum — with matching baselines. The score cannot be measured locally: the
shard runs node --test, which this repo hard-blocks, so CI is the only source.
Floors are raised to the measured value once this PR's shards report; a shard
below 50 means the tests need strengthening, since the floor cannot go lower.

The 1MB test was measuring the test harness rather than the product. The command
handles the oversized payload correctly by spilling to a tmpfile and resolving it
back, but the resolved stdout then exceeds runGsdTools' maxBuffer and the helper
reports ENOBUFS. It now uses --pick so stdout stays one byte while the full 1MB
document is still read and parsed end to end.

* fix(#2790): wire containment across every document read this command drives

An isolated security review of the containment control found the boundary logic
sound but not comprehensively wired: two content reads reached the filesystem
without it.

An escaped phase DIRECTORY could enumerate external filenames into the file
fields and diagnostic subjects. Both enumeration sites now containment-check the
directory before reading. Worth recording that the leak was already prevented one
layer earlier than the review claimed: Dirent#isDirectory() reports false for a
directory symlink, so such a directory never becomes a phase row at all. The
guard is defense-in-depth for a direct caller and for platforms where a reparse
point reports as a directory.

A *-VERIFICATION.md symlinked outside the root leaked one frontmatter value
verbatim, because readVerificationStatus does its own read and copies an
unrecognized status into the payload's next_action. Closed from the consumer
side through that function's existing fs injection seam, so src/verification.cts
keeps its signature and its other callers are untouched.

The reviewer additionally rated a forged status: passed as an integrity bypass.
It is not: anyone able to plant the symlink can plant a real VERIFICATION.md
saying the same thing. The incremental risk is confidentiality, which is what
these fixes close.

src/plan-scan.cts is deliberately unchanged: isPlanSuperseded reads
symlink-followed content but yields only a derived boolean, no document text.

* test(#2790): give the mutation shards an in-process surface

Two Stryker shards were CANCELLED at the 15-minute cap, not failed on score.
CI log: 640 mutants instrumented, and the dry run reported 'Ran 1 tests in 20
seconds' because the shards pointed at the integration suite, where nearly every
case spawns a gsd-tools subprocess and Stryker's command runner treats the whole
test-runner invocation as a single test. 640 x 20s cannot finish in 15 minutes;
at the kill it was 27/640 with an ETA over an hour.

Every other COVERED module points at a property or unit file, and the workflow's
own paths filter lists exactly those two patterns. In-process is the intended
mutation surface; the shards were pointed at the wrong shape of test.

Adds tests/planning-inspect.unit.test.cjs — 39 cases in 10 describes that spawn
nothing and call the built modules directly. plan-document and the router need no
filesystem at all, one being a pure content-to-object parser and the other taking
an injected mock. The three shards now point here. The 91-case integration suite
is untouched and still runs in the normal test job.

* chore(#2790): ratchet mutation floors to the measured CI scores

CI run 32392791843 measured all three shards, which is the only source the
registry accepts — local runs count timeouts as kills and inflate badly.

  planning-command-router  95.65 -> floor 94
  plan-document            76.58 -> floor 75
  planning-inspect         57.03 -> floor 56

Applied the registry's own rule, floor(score) - 1, and updated RATCHET_BASELINE
to match, since the ratchet test enforces equality.

planning-inspect sits well below the file's target of 80 and is the obvious
ratchet candidate as its tests improve. planning-command-router already exceeds
the target. The placeholder comment about floors pending measurement is removed
rather than left standing as a false statement.

---------

Co-authored-by: sim <sim@local>
2026-08-20 13:42:43 -04:00
Tom Boucher
2fca0e17e4 enhance(#2554): resolve code review depth from path-scoped override rules (#3695)
* test(#2554): failing-first suite for path-scoped code review depth overrides

Binds the not-yet-built code-review-depth module: segment-aware path-prefix
matching of a changed-file set against ordered {paths,depth} rules, resolution
order flag > strongest matching rule > global > standard, typed validation
errors, and the large-scope downgrade boundary. Also proves behaviorally that
workflow.code_review_depth_overrides is not yet a registered config key.

Refs #2554

* feat(#2554): resolve code review depth from path-scoped override rules

Adds workflow.code_review_depth_overrides — an ordered array of {paths, depth}
rules matched against a review's changed-file set by segment-aware path-prefix
comparison. Resolution order is --depth= flag, then the strongest matching rule,
then workflow.code_review_depth, then standard; a matching rule replaces the
global rather than being max'd with it, so quick and standard rules stay
meaningful. Glob metacharacters are a hard configuration error rather than sugar
for a prefix, and malformed rules halt the review instead of degrading to
standard. The resolver is pure and reports its own provenance, so the workflow
can print the resolved depth and the rule that matched. The pre-existing
>50-file deep-to-standard downgrade moves into the module and now names the rule
it overrode.

The key is registered centrally rather than as a capability config slice: the
federated slice channel admits only boolean/string/number/enum, so an array
slice would be dropped as malformed.

Closes #2554

* test(#2554): correct depth-provenance assertions and pin out-of-repo paths

Two corrections to the failing-first suite. The source assertion for a
non-matching rule with no global configured expected 'config'; with no global
set the depth comes from the default, and a companion assertion tolerated
either value, so both passed against an implementation that derived provenance
from whether any rules existed rather than from where the depth came from.

The out-of-repo absolute-path case used a home-directory path that matched
neither implementation, so it never exercised the defect it named. It now pins
the discriminating cases: an absolute path outside the repo root must not match
a repo-relative rule, and one under the root must.

* docs(#2554): document path-scoped code review depth overrides

Reference rows for workflow.code_review_depth_overrides in the configuration,
features and commands references plus the locale copies that carry those tables,
and in the planning-config reference. Explanation of why escalation is
whole-review rather than per-file and why v1 is prefix-only. New how-to for
scoping review depth by path, carrying the configuration-error reason table and
the distinction between nothing to report and could not look. CONTEXT.md
glossary entry and the INVENTORY row for the new CLI module.

ja-JP and ko-KR CONFIGURATION.md carry no code_review keys at all, and ko-KR and
pt-BR FEATURES.md carry no code-review config table, so those files are
deliberately untouched.

* fix(#2554): make the depth-misconfiguration halt executable and reject control chars

Three review findings, all in this change.

The misconfiguration halt was prose rather than shell: the error-printing fence
was followed by an unconditional extraction fence, so an ok:false result threw
and left the depth empty instead of stopping the review. Prose is not a guard —
the two fences are now one block with a real conditional, and anything that is
not the literal string true fails closed.

An interior control character in a rule path survived validation and reached the
provenance string and the summary box; rule paths now reject control characters
via a new PATH_CONTROL_CHAR reason, after the glob check so precedence is
unchanged. That in turn makes the field record safe to delimit, so the seven
node invocations that each re-parsed the same result to read one field collapse
to one.

Also corrects the glossary entry's illustrative paths, which the glossary-ref
check read as real repository references.

* fix(#2554): use the fast-check v4 string API and acknowledge workflow growth

Two failures from the remote matrix on d3111f45, both this branch's.

The property block built its segment arbitrary with fc.stringOf, removed in
fast-check v4. Because the arbitrary is constructed in the describe body, the
throw took out all four property tests rather than one — they had never
executed. Rewritten to fc.string({unit, ...}), the form this repo already uses
in emitted-attribution.test.cjs. Every other fast-check helper in the file was
audited against the installed module.

The emitted-attribution growth arm needed an acknowledgment for code-review.md,
which grew 5376 bytes. The pre-existing 3503 fragment keying the same file is
spent — its ripple was absorbed when #3503 merged, and the base file is exactly
the 34435-byte baseline this growth is measured against — so it cannot clear
anything, while the ack lint hard-fails on a duplicate key across two sources.
Removed it in favor of the new fragment, which is exactly how #3503 itself
replaced the spent 3191 fragment.

* docs(#2554): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-19 22:45:35 -04:00
Tom Boucher
ea594300d9 fix(#3606): validate hook-kind coverage at call sites and dispatch generically (#3687)
* test(#3606): pin hook-kind coverage in the wired guard

* fix(#3606): validate hook-kind coverage at call sites and dispatch generically

* fix(#3606): address review - segment-granular narrowing, zero-coverage diagnosis, quick.md, fragment extraction

* fix(#3606): drop stale shrink-ack, export HOOK_GROUP_KINDS, dedupe scanner regex

* chore(#3606): regenerate install-tree fixtures for new wave-post fragment

* chore(#3606): sync canonical launcher preamble into new fragment

* fix(#3606): keep fragment preamble ahead of first gsd_run mention

* fix(#3606): revert sync script's preamble move in explore.md

* chore(#3606): regenerate derived manifests post-rebase

* chore(#3606): allowlist peer test files - base was red on the count lane

* chore(#3606): regenerate inventory for peer's verify-command-grounding doc

* chore(#3606): grounding test maps to its own module by longest prefix

* chore(#3606): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 16:41:27 -04:00
Tom Boucher
79781e68eb enhance(#2401): ground verify-command paths and inherit prior-phase commands (#3678)
* feat(#2401): ground <automated> verify-command paths and inherit prior-phase commands

Adds a deterministic resolvability probe over each PLAN.md <automated> verify
command and surfaces the nearest prior phase's proven commands to the planner
at every context window.

- src/verify-command-grounding.cts: recognizer (not a shell interpreter) that
  grounds a leading cd <literal> chain and npm --prefix <literal>, and reports
  unresolvable rather than guessing. Never executes command text.
- gsd-tools check verify-command-paths <N>: per-phase probe, wired into
  plan-phase.md before the plan-check pass.
- init.plan-phase gains prior_verify_commands, ungated by context_window.
- gsd-plan-checker: new Verify Command Path Resolvability dimension that
  reports the failing target and never prescribes a replacement.

Also fixes first-match-wins prefix bucketing in scripts/lint-test-file-count.cjs
(readdir order is not stable across platforms, so a module whose name extends
another's with a hyphen bucketed differently on Linux than on macOS).

Closes #2401

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2401): ground the canonical --prefix form, quoted paths, and absolute cd resets

Independent review found three defects in the recognizer:

- npm --prefix DIR run SCRIPT never reached the script-existence check,
  because the pattern required npm and run to be adjacent. That is the
  form the docs tell planners to prefer, so script_missing never fired
  for it. The prefix flag and its value are now stripped before matching.
- --prefix captured with \S+, so a quoted path containing a space was
  truncated to a stray opening quote and reported as a missing directory
  - a false blocker, worse than the bug this feature fixes. The capture
  is now quote-aware.
- A chained cd whose later segment was absolute concatenated instead of
  resetting, producing a nonsense path and another false blocker. The
  fold now resets on an absolute segment.

Also replaces the bespoke phase-directory regex with the canonical
phase-id helpers. Real phase directories are NN-slug, not phase-N-slug,
so the prior-command harvest matched nothing outside its own fixtures
and the planner-inheritance half of this feature was dead code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(#2401): source task blocks from the canonical sectionizer

The module carried its own copy of the <task>-block grammar - a fourth
hand-rolled mirror of the one markdown-sectionizer owns. verify.cts keeps
its copy only because it needs the type= attribute the canonical helper
discards; this module never reads that attribute, so it can share the
owner outright instead of adding a test around a copy.

extractAutomatedCommands now takes task bodies from extractTaggedBlocks
and the out-of-task remainder from stripTaggedBlocks. A task-grammar
parity test pins the attributed task-name set against the canonical
helper across six awkward task shapes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2401): extract agent-file overflow to references and repair the property arbitrary

The remote matrix run came back red with 19 failures, four root causes:

- agents/gsd-plan-checker.md and agents/gsd-planner.md both blew the
  49152 agent cap. Their bodies move to gsd-core/references/, leaving
  @-reference stubs, per the documented overflow pattern.
- The new checker dimension invoked gsd_run before the canonical
  preamble that defines it. The call is deleted outright: plan-phase.md
  already runs the probe and hands the result in as {VERIFY_PATHS}, so
  the dimension consumes that rather than re-running anything.
- fc.fullUnicodeString does not exist in fast-check 4.8.0. Replaced with
  fc.string({ unit: 'binary' }), which covers the same 0000-10FFFF range.
- Three runtime-loaded files grew; acknowledged in the existing ack
  fragments that already own those bare filenames, since two ack sources
  may never name the same path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2401): regenerate golden install-tree fixtures for the new references

Adding two files under gsd-core/references/ changes what the installer
emits into every runtime's tree, so all 19 golden install-parity
fixtures went stale. Regenerated with npm run gen:install-tree; the
delta is exactly the two new reference paths per runtime, no removals.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2401): backfill changeset pr number to 3678

* fix(#2401): treat ~ as a home expansion only at the start of a path

Windows CI caught this on both shards; the Linux-only remote matrix
cannot see it. The dynamic-path refusal rejected ~ anywhere, and a
GitHub Windows runner's tmpdir is an 8.3 short name -
C:\Users\RUNNER~1\AppData\Local\Temp - so a valid absolute Windows
path came back unresolvable/dynamic_path.

This was a production bug, not a test artifact: any Windows user whose
project path carries an 8.3 short name, or any literal ~, silently lost
the probe entirely - every command degrading to unresolvable with no
explanation.

~ is a home expansion only at the start of a path; elsewhere it is an
ordinary literal. The check is now split: $, backtick, *, ? and newline
stay refused anywhere (substitution and globs, and the glob characters
are illegal in Windows path components regardless), while ~ is refused
only leading, tolerating one leading quote since the check runs before
quote stripping.

The prior tests only caught this on Windows because only Windows puts a
~ in tmpdir. Four new tests pin it on every platform via a fixture
directory literally named RUNNER~1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 15:21:15 -04:00
Tom Boucher
3ab0007164 enh(#2875): materialization primitives — durable user-artifact staging and descriptor-authoritative agents (#3600)
* fix(#2875): stage user artifacts durably across install wipes (#1874-F19)

preserveUserArtifacts held user files only in an in-memory Map across the
wipe, so any process death between preserve and restore lost them outright.

Seven call sites, not the four the issue records. Three of them never called
the helper at all - they open-coded the same read/wipe/write - so searching
for callers under-counted by construction; the extra sites were found by
sweeping for the pattern instead.

The worst is the mainline install path, where the crash window spans the
entire gsd-core tree copy rather than a single rmSync.

Adds src/user-artifact-staging.cts: durable on-disk staging with a record
written after the copies land as the commit point, plus recovery of orphaned
batches on the next run - without recovery the staged bytes survive but the
user's file is still gone, which would pass its own test while delivering
nothing.

Routes copyPreservingSymlink through installFs() so staging cannot bypass the
install fs seam, and reunites its symlink-safety docblock with the function it
documents.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): amend ADR-3574 with four claims disproved by implementation

Implementing Phase 6 disproved four statements the ADR rests on. The central
decision - no single materializer - is unaffected and stands.

Corrected: decision 3 was already satisfied, so nothing was extracted; the
agents-bypass runtime set omitted claude, kilo and opencode, and closing it
needed three new pieces of descriptor contract rather than proceeding on its
own terms; three of the four blockers the layout comment names were already
stale; and F19 is seven call sites, not four.

Records the generalizable lesson: the defect is the pattern of holding user
data in memory across a wipe, not the helper, so searching for callers of the
helper under-counts by construction.

Also resolves the ADR's open question on USER_OWNED_ARTIFACTS membership, and
notes that copyPreservingSymlink needed routing through the install fs seam
before it could be reused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close dangling-symlink blind spot and harden staging recovery

An adversarial review found the F19 staging work shipped red and unsafe.

Root cause, shared by two arbitrary-write findings: hasExistingSymlinkBetween
missed dangling symlinks in both its root check and its per-segment walk,
because it probed with existsSync, which is false for a link whose target does
not exist. Fixing only the new module would have reused a guard that was
itself blind. This guard protects the whole install tree.

Recovery no longer throws: it degrades per entry and per file, so one bad
batch cannot block the others. Previously an unrecoverable entry propagated
out of the first statement of install and uninstall, before the cleanup that
would have removed it - wedging the installer permanently.

Partial fs adapters now throw on any omitted method instead of silently
reaching the real filesystem, closing the trap that let a test poison list
pass while real IO happened.

Staged names must be flat, recovery refuses a dangling destination symlink,
and a batch whose recovery genuinely failed is no longer swept - it was
discarding the only durable copy of the file it had just failed to restore.

Replaces three tests that could not fail, including the one labelled negative
proof.

Known limitation, documented not closed: concurrent installs sharing a staging
key can still lose a batch. A real fix needs a cross-process lock.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enh(#2875): make the descriptor authoritative for the agents kind

Deletes the inline agent-staging loop in bin/install.js and the
_DESCRIPTOR_AGENTS_RUNTIMES set, so every runtime materializes agents from
its capability descriptor instead of an inline hostBehaviors dispatch.

Closing it needed three pieces of contract the descriptor pipeline never had,
all reducible to one missing input - per-agent resolution context: a
frontmatter-extensions step for claude's effort and disallowedTools, per-agent
model-override resolution for kilo and opencode, and a named branding
converter for hermes, whose rewrite data was already declared.

Seven runtimes were on the loop, not the six the design recorded - kimi-code
was found by a golden fixture, not by analysis. claude-local and kimi-code
both silently lost their agents mid-change; the fixtures caught both and the
cause was fixed rather than the fixtures regenerated.

A parity harness gates the migration: both pipelines over identical inputs,
byte-identical output including filenames, per runtime. It is demonstrated
red before being trusted. Surface and install paths converge for all seven,
which also fixes surface previously writing no agents for these runtimes.

Codex's config.toml strip stays put - it mutates host config, which no
descriptor kind models.

Also routes install-model-override-resolver and install-effort-resolver
through the install fs seam. Both leaked real filesystem IO from the install
call tree; the stricter adapter is what exposed them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): record the agents-descriptor migration and correct the ADR count

The _DESCRIPTOR_AGENTS_RUNTIMES allow-list no longer exists, so the host
integration guide told readers to join a set that is gone. Replaces that with
what is now true - declare an agents entry and it installs, on the surface
path as well as install - and points anyone needing a per-agent transform at
the three extension points rather than at a new inline branch.

Corrects the ADR amendment: seven runtimes were on the inline loop, not six.
kimi-code was found by a golden fixture going red, not by reading. That is the
third short count this phase, all from enumerating by symbol or set membership
when the thing that matters is a behavior.

Adds the Changed changeset for the surface-path convergence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): amend ADR-2866 - claude global always wrote agents on disk

The claude row's global=[skills] described what capability.json declared, not
what the installer wrote. bin/install.js's inline agent-staging loop was never
scope-gated and never consulted the descriptor, so a claude --global install
has always written agents/gsd-*.md.

Phase 6 closes the gap by deleting that loop and declaring agents on claude's
descriptor at global scope. On-disk bytes are unchanged - the golden fixtures
did not move, which is the evidence that the descriptor, not the installer,
was incomplete.

#2218 is unaffected: agents are not trigger-bearing, so the wider row does not
introduce a new shadowing case.

Records the warning that an incomplete descriptor is invisible while a second
code path silently does its work, and only surfaces when the two are forced
into agreement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close review findings across staging, agents and the parity harness

Two independent reviews of this branch found defects the local gates missed.

Security: a dangling symlink at a migration destination allowed writing
outside configDir - the same class this change claimed to close, missed at the
terminal write of the flow being added. The staging-root resolver threw as the
first statement of install and uninstall, so a hostile symlink bricked both,
and symlinked-configDir users lost uninstall as well as install; it now
degrades instead of aborting. Recovery gained a source-side symlink check and
now refuses a relative destDir, which resolved against cwd. Converter dispatch
gained a runtime allowlist - lint-time validation stopped mattering once this
branch promoted that dispatch from the surface path to real installs.

Correctness: claude --local --minimal exited 1 because the minimal profile
legitimately yields zero agents and the new path treated that as a failure.
cline --local silently lost its agents - its descriptor declared none while
the deleted loop wrote them unconditionally. The agents prune was widened to
any gsd-* entry and destroyed user files it never owned.

The parity harness, on which the migration's safety argument rested, drove a
synthetic registry and never byte-compared the shipped descriptors; two of its
trap rows could not fail. It now drives the real registry across 13
runtime-scope rows including kimi-code and cline-local, and its red-proof is
demonstrated by corrupting a live capability.json. Three goldens that had
encoded the cline regression as expected behavior were corrected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close findings from both mandated review engines

/security-review found the staging source-side walk honouring
GSD_ALLOW_SYMLINKED_DEST, an opt-in documented as relaxing only the write
destination. A symlinked files/ component dereferenced because
copyPreservingSymlink lstats the leaf only, so an intermediate link is
followed. The source walk no longer honours the opt-in; the destination check
still does.

/code-review spec axis found this branch had reintroduced its own bug:
migrateLegacyDevPreferencesToSkill's new symlink refusal threw unguarded after
the legacy dir was wiped and before the staged batch was restored, so a
planted symlink bricked uninstall permanently and orphaned the batch. Refusal
kept, abort removed.

kimi-code local silently lost its agents, the same class as the cline bug, and
the parity harness recorded that exclusion as intentional - the third test in
this branch to pin a regression as correct.

--minimal now creates an empty agents/ dir that never existed. Behaviour
restored rather than softening the changeset, so its byte-identical claim
stays true.

Standards axis: try/finally removed from twelve test bodies, fast-check
properties added for parseOwnerPid, boundary coverage at the grace window and
the ancestor-probe depth, a parity assertion for the staging-root helper
duplicated across two files, and the 8-deep config walk deduplicated.

Records 60-review.json with every finding and disposition from five passes,
including the smells left unfixed and why.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): prune stale agents unconditionally in minimal mode

The previous round stopped an empty agents/ directory being created when the
resolved profile yields no agents. That was implemented by skipping the agents
kind entirely, which also skipped its stale-agent prune - so a full to minimal
downgrade left stale gsd-* agents behind.

The deleted inline loop pruned unconditionally and only skipped writing. Those
are three separate conditions, not one: prune always, write only when there is
something to write, create the directory only when writing.

Both call sites now run _removeGsdEntries before the empty-staged early exit.
The symlink-escape guard moved with it, since the prune also touches dest.
Codex .toml agents and the config.toml stanzas are cleaned again, and
user-owned agents are still preserved.

The agents/ directory is left in place after a prune empties it, matching
every sibling kind - none of them remove the destination directory itself.

Golden fixtures confirmed byte-identical: the prune is a no-op on a fresh
install, so fixture generation is unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): document interrupted-install recovery for user-owned files

The durable-staging fix is invisible to the user it protects. Someone whose
install died mid-flight has no way to know USER-PROFILE.md was staged before
the delete, that the next run restores it, or that recovery happens at the
start of that run rather than in the background.

Written as the task the user has - finish the interrupted command - rather
than as a description of the mechanism, and states what it will not do:
overwrite a file already present, or touch staging belonging to another
install still running.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2875): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2875): assert the J8 model override without building a regex

CodeQL flagged incomplete string escaping: the assertion interpolated the
override value into a RegExp while escaping only forward slashes, which is
meaningless in a constructor, leaving real metacharacters unescaped.

The failure direction was the dangerous one - a metacharacter would have made
the match more permissive, so the row would pass when it should fail. That
matters here because J8 exists precisely because an earlier revision was a
tautology; the rewrite reintroduced a different way for the same assertion to
stop discriminating.

Replaced with a line-wise exact match, so no regex is constructed at all.
Swept the other test files this branch adds; no sibling instances.

lint:ci passed on the original - lint-no-adhoc-regex-escape matches a full
metachar-escape copy, so a single slash replace slipped under it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 17:25:53 -04:00
Tom Boucher
a4a02a7a01 enhance(#2874): return the executed plan and route install IO through a seam (#3568)
* test(#2874): add failing-first gate for the executed-plan return

Four rows from the matrix's red-first order. E3 pins the one early return,
for the opencode family, where a void-shaped hole would otherwise survive
unnoticed. E13 sweeps every runtime in the registry - enumerated from the
registry rather than hardcoded, so a runtime added later cannot slip past.
F2 proves absence of real filesystem contact rather than merely that the
happy path ran, which is the difference between a complete seam and a
partial one.

G1 and G3 are the additive guard and must be green before and after. G3
deliberately leaves the two existing adapter test doubles untouched: if
this change required editing them it would not be additive, and the
acceptance criterion would be unmet.

No production code. All 19 runtimes install without throwing today, so
E3 and E13 fail on the undefined comparison alone.

Refs #2874

* feat(#2874): return the executed plan and route install IO through a seam

installRuntimeArtifacts returned void, so its correctness was observable
only by re-reading disk. It now returns what it executed - per kind, per
scope - including on the combinedFamilyInstall path, which was the one
early return where a void-shaped hole would have survived unnoticed.

Failure still throws rather than becoming an ok:false return, so control
flow is unchanged for both existing callers. A best-effort cleanup that
fails is still swallowed, but is now visible in the returned value rather
than silently absent.

The fs seam is ambient rather than threaded. Explicit deps through
install-profiles and the 3000-line conversion module was impractical; the
tradeoff, the synchronous-only re-entrancy assumption, the restore
guarantee and the partial-adapter fallback trap are all documented at the
seam. findInstallSourceRoot and its sibling stay unrouted by design -
they locate the package's own source, not the install destination.

readCmdNames keeps a second implementation because the standalone CLI
that owns the original cannot require the compiled adapter without a
build-order dependency on its own output. A parity test fails if the two
ever disagree.

Refs #2874

* chore(#2874): gitignore the new build artifact

install-fs-adapter.cjs is tsc output from src/install-fs-adapter.cts, not
a tracked source file. It was added to eslint's ignore list but not to
.gitignore, so it landed as a tracked file - the third time this step of
the new-.cts ripple has been missed on this epic.

Refs #2874

* fix(#2874): close two seam leaks and correct a false comment

A correctness review found the seam still leaked in two places, both
subtler than the three already closed.

readGsdCommandNames was routed when it should not have been: it reads the
package's own commands directory, which a destination-fake is never
seeded with, so under a fake adapter it returned an empty or wrong roster
instead of failing loudly. It now reads real fs, matching the precedent
already documented for findInstallSourceRoot.

cleanupStagedSkills ran raw rmSync from a process exit handler, which is
real filesystem work deferred past the point where withInstallFs has
restored - the one thing the synchronous-only contract exists to
exclude. Staging now captures the adapter that created each directory and
cleanup replays it, so a real install cleans up exactly as before and a
fake-staged path never reaches the real filesystem.

Also corrected a comment claiming the migration reads were an unrouted,
untested residual gap. They are routed and exercised; a comment
understating the seam is as corrosive as one overstating it in a module
whose trust rests on being honestly documented.

Refs #2874

* test(#2874): migrate the exemplar group and cover the matrix

AC3's exemplar migration lands in place: the qwen install group now
asserts skills and agents destinations from the returned plan in one
deepStrictEqual instead of probing the filesystem for each.

Nine facts the old probes established were enumerated first. Two moved to
the value assertion; seven were retained deliberately - per-file SKILL.md
existence, the VERSION file written outside this function, the manifest
content, and the post-uninstall absence checks all sit outside the plan's
per-kind contract. A migration that quietly asserts less looks like a win
and is a regression, so the enumeration is the guard rather than the
line count.

Also implements the rest of the matrix: the executed-plan shape, adapter
failure modes, the security-boundary rows including a fake that cannot
certify an install the real filesystem would refuse, cleanup visibility,
and two seeded property tests. Only the two external CI gates are left
unticked, because self-certifying them would be a claim rather than a
check.

Refs #2874

* fix(#2874): restore streaming hashes and derive F2 from the boundary rule

The checkpoint found three things reasoning had missed.

sha256File had been converted from raw-fd streaming to a single
readFileSync on the assumption that GSD artifacts are never large. A test
named for exactly that contract already existed and went red. Streaming is
restored, now routed through the adapter, which gains openSync, readSync
and closeSync. The contract was the specification; the assumption was not.

Three existing tests inject faults by monkeypatching real fs. They broke
because mkInstallTempDir stopped calling real mkdtempSync, not because of
any binding subtlety - the real adapter was already late-bound. It now
calls the real function when no fake is injected, so a monkeypatch applied
after import is still seen and the additive contract holds.

F2 poisoned real fs by method, so a deliberately unrouted package-source
read failed a correct design. It now poisons by path: destination IO is
forbidden, package-source IO is allowed and positively asserted. The claim
was always zero real destination IO, and the test now derives from that
rule instead of coincidentally matching it.

Refs #2874

* docs(#2874): add the contributor how-to for plan-based test migration

The phase gate caught a real gap. The docs plan was Reference plus
Explanation only, and every CI check would have passed, because the
docs-required lint only verifies that some file under docs/ moved.

But this phase exists to demonstrate a pattern for follow-on work, and
that work is other contributors migrating probing test groups. The
sequence has two live traps - a partial fake silently falls back to real
fs, and the seam is ambient and synchronous-only - plus one discipline
nobody infers: enumerate the facts before converting, or you assert less
and call it a win.

The page carries the qwen migration's arithmetic, nine facts enumerated
and only two converted, because a reader seeing only the diff would
reasonably conclude the pattern is to replace probes wholesale.

No locale mirrors: none of the four carries any contributor-only how-to,
so a single translated file would manufacture parity rather than provide
it.

Refs #2874

* chore(#2874): backfill changeset pr number

* test(#2874): normalize both sides of the G1 tree comparison

G1 failed on Windows only, deterministically on both shards. The defect
was in the test helper, not production.

_computePathPrefix posix-normalizes the resolved config dir
unconditionally, so on Windows the path embedded in every emitted
SKILL.md body is forward-slash form. hashDirTree stripped against the raw
backslash path from mkdtempSync, so the substring never matched and each
install's unique temp suffix stayed baked into every file - all fifteen
skill bodies hashed differently for two runs that had written identical
bytes.

Both sides are now normalized unconditionally rather than gated on
path.sep, matching the rule this repo already records: backslash paths
arrive on Linux too.

Production code is untouched and was verified correct. Normalizing this
away on the production side would have hidden a real portability bug if
one had existed.

Refs #2874

---------

Co-authored-by: sim <sim@local>
2026-08-16 02:48:24 -04:00
Tom Boucher
c5b83cb050 chore(#3560): delete two unreachable workflows, gate workflow reachability in lint (#3564)
* chore(#3560): delete two unreachable workflows, gate reachability in lint

discovery-phase.md and plan-milestone-gaps.md shipped to all 19 runtime
install trees with no command, agent, or skill referencing them.
plan-milestone-gaps' command was deleted by #2790 and the workflow was
left behind; discovery-phase's own header claimed a caller in
plan-phase.md's mandatory_discovery step, and that step does not exist —
plan-phase.md contains zero occurrences of "discovery".

docs/INVENTORY.md asserted discovery-phase.md was an alternate entry for
/gsd-new-project. new-project.md never referenced it. The row and the
matching note sentence are removed across all five locales rather than
corrected.

Adds rule 6 to lint-command-contract: every shipped workflow must be
reachable from a loader, walking the transitive closure over the three
reference shapes this repo uses. The closure seeds ONLY from
commands/agents/skills, so a workflow that references only itself and a
pair that reference only each other are both correctly reported rather
than satisfying themselves; a visited set makes reference cycles
terminate. The measure is a mention in a LOADER — docs/ and install-tree
fixtures deliberately do not count, because scan.md proved a file can be
documented and shipped while entirely unreached.

Ships blocking, not report-only: #3561 is in this branch's base, so the
tree reports 0 unreachable from the start.

Closes #3560

* test(#3560): drive rule 6 end-to-end, sweep a stale allowlist, update ADR-0002

Review findings.

Rule 6 had no end-to-end coverage: the tests exercised the pure closure
with in-memory data, so the wiring — file collection, exit code,
diagnostic — was unproven, and #3560's acceptance list explicitly wants
a fixture showing the rule FAILS on a planted orphan. Adds an optional
--root to lint-command-contract (default behavior unchanged) and four
tests driving the real CLI through the process seam against a temp
fixture: clean=0, planted orphan=1, orphan referenced only from docs/=1,
orphan reachable transitively=0. The docs/ case is what pins the
Goodhart defense — a mention outside a loader must not confer
reachability.

Deletes two tests that were byte-identical to a third and could not
assert anything loader-specific, since the closure is source-agnostic by
design; that distinction lives in the lint script's file collection and
is now covered above.

Removes a stale ALLOWLIST entry for discovery-phase.md in
planner-language-regression — the exact sweep-miss class rule 6 exists
to catch, found in the PR that adds the rule.

ADR-0002 described five per-file frontmatter checks; rule 6 is a
repo-level reachability graph, so the Decision section now says so.

Refs #3560

* test(#3560): cut the bug-3298 test pin on the deleted plan-milestone-gaps workflow

The remote runner went red with four failures: tests/phase.test.cjs
asserted the plan-milestone-gaps workflow exists and checked its mkdir
patterns, so deleting the file broke the test that pinned it. This is the
fence the epic describes — the content-sync test IS what keeps an
unreachable file alive — and cutting the coupling is what makes the
deletion safe.

Removes only that arm. The bug-3298 block guards three workflows against
phase-dir prefix drift; the import and add-backlog arms and both shared
mkdir-pattern helpers are untouched.

Worth recording where the sweep failed: my reachability walk covered
commands, agents, skills, gsd-core and docs, and lint-removed-but-needed
covers .github/workflows, gsd-core, docs and package.json. Neither looks
at tests/, so a test-pinned deletion is invisible to both and surfaces
only on the remote runner. The how-to added by this PR names that gap
explicitly so the next deletion searches tests/ by hand.

Refs #3560

* docs(#3560): add a how-to for resolving unreachable-workflow findings

* chore(#3560): backfill changeset pr number to 3564

---------

Co-authored-by: sim <sim@local>
2026-08-15 23:29:30 -04:00
sim
147856040b fix(#2873): close review findings across fences, sanitizer and docs
Isolated security review found resolveSpecRootReference's fence tracker
toggled on any delimiter, so a backtick fence could be closed by a tilde
one and an include in the gap was rewritten inside a code block. Fixed by
reusing scanFencedBlocks - the canonical engine already behind
stripFencedCode and extractFencedBlock - rather than carrying a fourth
copy of fence detection, which also closes the duplication the standards
review flagged.

sanitizeForRender now strips combining marks and zero-width characters
alongside the ANSI, control and bidi classes it already handled.

Adds the C, E and F matrix rows the spec review found missing, including
installer-level coverage that spawns the real install rather than calling
the report builder. Ships the how-to, the reference and command docs in
five locales, the changeset, the inventory and glossary entries, and
regenerates health.md for the new W028 rule.

Refs #2873
2026-08-14 23:48:39 -04:00
Tom Boucher
895d9df96d fix(#3477): run untrusted key_links patterns on a linear-time engine (#3496)
`cmdVerifyKeyLinks` compiled `must_haves.key_links[].pattern` from plan frontmatter with `new RegExp()` and tested it against whole file contents, so a nested-quantifier pattern such as `(a+)+$` hung `verify-phase` indefinitely (CWE-1333). JavaScript has no regex-execution timeout.

Untrusted patterns now run on RE2 (re2js), whose match time is linear in input length — the class is closed by the engine, not by a heuristic screen. The screen lost in the ADR-0174 consolidation was deliberately NOT restored: it never worked, since `(a|a)*$`, `((a+))+$`, `(a+){2,}$` and `(a{1,3})+$` all evade it. A refused pattern's matcher returns false for every input, so it cannot report a match no matter what the caller does.

The engine is vendored at gsd-core/bin/lib/vendor/re2js.cjs because gsd-core/bin/** is copied into installed trees with no node_modules; runtime dependencies are unchanged. New ESLint rule local/no-external-require-in-bin enforces that invariant, which had been documented in a comment since the #3024/#2071 bug class and enforced nowhere.

Backreferences and look-around are unsupported by RE2 by construction — disclosed in a Changed changeset.

Closes #3477

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 14:34:36 -04:00
Tom Boucher
29c27e948b fix(#3312): require static frontend evidence before ui-plan-gate blocks (#3451)
* fix(#3312): require static frontend evidence before ui-plan-gate blocks

* fix(#3312): fill changeset pr reference

* fix(#3312): align regression assertions with section heading and legacy artifact shape

---------

Co-authored-by: sim <sim@local>
2026-08-14 09:27:48 -04:00