Files
msd-core/docs/FEATURES.md
0xdhx b90eef28e8 enhance(#3829): report code review severity counts and record a per-finding disposition (#3861)
* enhance(#3829): report code review severity counts and record a per-finding disposition

`code_review_gate` extracted `status:` from REVIEW.md's frontmatter and discarded
the `critical`/`warning`/`info`/`total` values sitting in the same range, so its
output was byte-identical for a review with one `info` finding and a review with a
Critical. Nothing anywhere recorded what happened to a finding: no file under
`gsd-core/workflows/` branches on `issues_found`, and `gsd-verifier.md` has zero
references to REVIEW.md. A phase therefore reached `phase.complete` with Criticals
standing and no trace they had been seen.

Both halves were approved on the issue; the gate stays advisory.

A — severity surfacing, in `execute-phase.md`. The gate states the breakdown it
already parsed, accepting `blocker:` as the documented tier-equivalent of
`critical:`. The breakdown is shown only when all four counts are numeric
(`REVIEW_COUNTS_OK`); otherwise the countless message stands, because gating on
the total alone still emits `6 findings —  critical` for a review carrying a total
and nothing else.

Frontmatter is extracted by an `awk` that emits only when it saw the CLOSING
delimiter, after stripping CR. A `sed` range re-opens on a body `---` and runs to
EOF: first-match protects a key the frontmatter always carries, but not an
optional one, so a review with no `findings:` block and a body `total:` line would
have reported the body's number. An unterminated block would leak the whole body
the same way.

Every read is guarded and `|| true`-terminated. This step is advisory, and under
`set -e`/`pipefail` a non-matching `grep` exits 1 — an assignment whose command
substitution fails would take the step down with it. A REVIEW.md that is missing,
a directory, or unreadable now leaves the counts empty and execution continues.

B — per-finding disposition, in a new lazily-read step file,
`gsd-core/workflows/execute-phase/steps/code-review-disposition.md`, referenced
from the gate in the established plain read-and-execute form. One row per finding
ID, defaulting to `open`, and:

- `fixed`/`skipped` are reconciled from REVIEW-FIX.md, whose section headings are
  matched WHOLE — a prefix match let `## Fixed Issues Verification` classify every
  finding beneath it as fixed — and only when the fix report names the SAME
  finding. Finding ids are reused across re-reviews, so matching on the id alone
  let a stale fix report declare a brand-new CR-01 already fixed.
- headings inside fenced blocks are ignored; a quoted example is not a finding.
- an id listed under both sections resolves by first occurrence, not row order.
- a recorded disposition is preserved together with the reason in its Source cell,
  escaped pipes included, and a hand-mangled row missing its trailing pipe still
  keeps its decision.
- a decided finding the current review no longer reports is CARRIED and marked;
  `--auto` rewrites REVIEW.md each iteration, so this is routine, and dropping the
  row would erase the record that it was seen. An untriaged `open` row for a
  vanished finding is not carried. A review reporting nothing still reconciles an
  existing ledger rather than freezing it.
- a run that changes no disposition rewrites nothing, so a re-executed phase does
  not produce a docs commit whose only delta is a timestamp.

The record is a sibling artifact, not a section inside REVIEW.md: `--auto`'s
re-review loop rewrites REVIEW.md every iteration, so a ledger kept inside it
would not survive the next pass, and REVIEW.md has a single writer that this step
is not.

B lives in an extracted step file because `execute-phase.md` was 91,493 bytes
against a 98,304 hard cap the size-budget test calls a red line, and because
`scanWiredKinds` caps a call site's dispatch-coverage region at 6000 characters —
an inline version pushed the `kind == "gate"` paragraph out of that window, which
silently drops `gate` from the covered set and fails
`gen-capability-registry --check` while pointing at the capability rather than at
the prose that displaced it. Extraction is what that size test's own message
prescribes, and it leaves the file at 93,854 bytes.

The tests execute the shipped script rather than modelling it. Three adversarial
review rounds each refuted "the mirror is faithful", and mutation testing agreed:
with a hand-written model, deleting the carried-row logic from the shipped file
turned nothing red. The suite now extracts the embedded script — undoing exactly
the four shell double-quote escapes — and runs it, so all ten mutations of its
behaviour are caught.

* chore(#3829): set changeset fragment pr to 3861

* fix(#3829): keep execute-phase.md under both size ceilings and propagate the launcher probe

The first push failed `full test (macos-latest, 24, shard 3/3)`. Two things it
caught that the CI-selected scope for this diff does not run, and that I
therefore did not run either:

1. `execute-phase.md` is governed by TWO ceilings, not one. The XL hard cap in
   `tests/workflow-size-budget.test.cjs` (98304) was satisfied at 95179, but the
   frozen ADR-857 pre-phase-6 ceiling in `tests/claude-orchestration.test.cjs`
   (93600) was not. The whole budget from base is 2107 bytes, which the inline
   reporting half alone did not fit. That half now lives in the extracted step
   file alongside the disposition half, and the parent carries only the paragraph
   that reads and executes it — 91529 bytes, 36 over base.

2. The step file calls `gsd_run`, so it owes the hermes runtime-home probe that
   `tests/runtime-launcher-parity.test.cjs` (E) requires of every workflow file
   that does. Propagated with `node scripts/sync-runtime-launcher.cjs`, the
   remedy that test names.

Verified with the FULL unit suite this time rather than the scoped selection —
14 shards, 0 failures — plus `npm run lint:ci`, and a re-run of the ten mutations
of the shipped disposition script, all still caught.

* fix(#3829): stop the disposition step instructing the agent to execute itself

Blocker 1 and Minor 7 of the round-1 review are one defect. The step file
carried a copy of execute-phase.md's pointer paragraph, so it named its own
path as something to "read and execute" — unbounded self-recursion at runtime
— and that copy is also the duplicated paragraph, sitting immediately above
the full instruction it duplicates.

Removing the copy resolves both. execute-phase.md remains the only surface
that points here, which is what it always intended.

Two structural tests guard it. Both are red against the pre-fix file: no
behavioural test could see either defect, because they execute the node
script through the process seam and so never read the prose that tells the
agent what to load.

* fix(#3829): re-derive the ledger paths in the block that uses them

Blocker 2. The disposition block reads REVIEW_FILE, DISPOSITION_FILE and
PADDED, all derived in the step's FIRST shell block. Each fenced block is
dispatched as its own shell, so all three are empty by the time the second
block runs: the ledger write lands on a bare `-REVIEW-DISPOSITION.md` path
and the review read finds nothing. The step then reports success having
produced no artifact — the feature's central acceptance criterion, silently
unmet, with no error to notice.

The tell was already in the file: the gsd_run shim preamble is re-emitted in
the second block for exactly this reason. These three paths belong beside it,
and now are.

The guard test asserts the general property rather than the instance — every
block derives what it reads, inheriting only the step's declared inputs
(PHASE_DIR, PHASE_NUMBER) — so a third block added later cannot reintroduce
it. Red against the pre-fix file.

* test(#3829): assert the counts mirror against the shipped shell, and execute its guards

Major 4, with Minor 6 and part of Minor 9.

The disposition builder stopped being a mirror three rounds ago, and the
reason given then was that a hand model of a shell-embedded script drifts
while the tests stay green. parseGateCounts kept its mirror anyway. That
argument does not stop applying at the boundary between the step's two shell
blocks, so the mirror now loses its authority: it is asserted against the
shipped awk and greps, run under `set -euo pipefail` in a real shell, across
every fixture it is exercised on.

Negative-controlled in both directions. Dropping `blocker:` from the mirror
alone fails the parity test; replacing the shipped awk with the leaky
`sed -n '/^---$/,/^---$/p'` range fails it on the unterminated-frontmatter
fixture. Divergence in either half is now red, which is what the finding asks
for. Skipped on win32, where there is no bash to compare against.

Minor 6: the zero-count edge is covered — `0` is numeric, so a zero-finding
review reports `0 findings — 0 critical, …` rather than falling back to the
countless form. A guard written against truthiness would have failed here
silently, and now cannot.

Minor 9, partially: running the block makes its advisory guards behavioural,
so the four `src.includes()` assertions that stood in for them are retired —
a missing and an unreadable REVIEW.md are now proven not to abort under
`set -e`, rather than asserted to contain a string. The remaining docs-parity
assertions are kept deliberately; see the PR discussion.

* test(#3829): add the render/re-parse fixed-point property for the ledger

Major 3. RULESET.TESTS.property-based-testing asks for at least one fc
property on a parsing/transformation contract, and the ledger is one with a
fixed point stated in its own prose: re-running the gate preserves every
disposition except `open`, and rewrites nothing when nothing changed.

Two properties, both driving the SHIPPED script rather than a model of it:

  idempotency — a second run reports `unchanged` and leaves the file
                byte-identical. Without it, the timestamp alone dirties the
                tree on every phase re-run.
  round-trip  — a hand-recorded decision AND the reason beside it survive
                render -> re-parse -> render, escaped pipes included. The
                Source cell is where a human writes why something was
                deferred, so losing it loses the only thing that instruction
                asks for.

Negative-controlled per property: disabling the unchanged-check fails the
first and only the first; discarding the carried source cell fails the second
and only the second.

numRuns is 40 rather than the shared 200 because each case spawns the shipped
script twice through the process seam. The seed stays pinned, so a failure
still reproduces; the deviation is stated in the file header rather than made
silently.

* fix(#3829): state a stale fix-report match instead of dropping it silently

Minor 5, plus the finding-id census this round owes.

Exact-title coupling stays — ids are reused across re-reviews, so a stale
REVIEW-FIX.md must not mark a brand-new CR-01 as already fixed. What changes
is the silence. A row that stays `open` because the report named a different
finding under the same id is indistinguishable, to any reader, from a row
that stays open because no report mentioned it. The gate now names the ids it
could not reconcile, on both report paths, and stays advisory throughout.

The census (RV4, self-found — the review did not ask for this). The script
enumerates finding-id prefixes in three places: the heading matcher, the
ledger re-parser, and the severity map's keys. The DOMAIN those enumerate is
owned elsewhere — gsd-code-reviewer.md's body template and its
Label-equivalence paragraph — so it can acquire a member without this script
changing.

Reached: CR, BL, WR, IN — 4 of 4, all present. Not reached: none today. What
follows if that changes is the payload: an unlisted prefix is not mis-tiered,
it is INVISIBLE — the finding never enters the order list and gets no row at
all, so the artifact silently under-reports the review it is meant to record.
Adding a prefix to two of the three copies fails the same way, and additionally
drops carried rows on the next run.

Two guards rather than a rewrite: hoisting the alternation into one constant
means rebuilding three regexes inside a double-quoted shell string, which is
the exact class of edit that produced both of this round's blockers. The
guards make the drift loud instead, and are negative-controlled against each
of the two ways it can happen.

* docs(#3829): keep the feature reference descriptive, not instructional

Minor 8 — a Diataxis mode mix. "Set `deferred` by hand and put the reason in
the Source cell" is a how-to instruction sitting in a reference doc. The
information belongs there (a reader needs to know the field exists and what
preserves it); the imperative does not.

Rewritten to describe the field instead: `deferred` is the one disposition the
gate never writes, and the reason recorded beside it survives re-runs. The
same pass records Minor 5's new behaviour, since the reference described the
title coupling but not what happens when it misses.

The imperative form is kept where it belongs — inside the ledger the gate
renders, which is where a reader meets the field and the only place an
instruction has an audience.

docs/FEATURES.md regenerated from it; `gen-features.cjs --check` is green.

* fix(#3829): close six defects found by reviewing this round's own fixes

None of these came from the maintainer's review. They came from adversarially
reviewing the five commits above before pushing them, and two are worse than
anything the round was opened to fix.

1. A foreign fence marker swapped an example for a finding. The heading scanner
   toggled fenced/not-fenced on ANY fence marker, so a ~~~ line inside a ```
   example closed the fence and the example's real close reopened one. Driven:
   a review quoting ~~~ inside a fenced example produced a ledger recording
   CR-77, the illustration, and omitting CR-01, the actual finding. A
   confidently-written artifact wrong in both directions at once. The open
   marker's character and length are now remembered, and a fence closes only on
   the same character at least as long, per CommonMark.

2. The disposition block had no status gate at all. The prose above it says it
   runs only when the review reports issues — but block 1 computes
   REVIEW_STATUS, emits nothing, and its shell is discarded, so no later block
   could act on that condition even in principle. A prose gate on a value
   nothing downstream can see is not a gate, and a clean re-review would rewrite
   a ledger it was never meant to touch. Re-derived in block 2's own shell.

3. A numeric breakdown could still be internally false. `total: 0` beside
   `critical: 1` is four valid numbers rendering `0 findings — 1 critical, …`.
   Numeric was necessary and not sufficient; an inconsistent breakdown is now
   withheld for the same reason a partial one is.

4. The carried-marker strip ate hand-written prose. It removed a trailing
   `(not in the current review)` unboundedly and unconditionally, so a deferral
   reason that merely ENDED in that phrase lost it — the one field a human
   writes into this artifact. Now bounded to one occurrence, and only on rows
   the marker can legitimately be on. The no-growth property it exists for is
   re-pinned.

5. parseGateCounts diverged from the shipped pipeline in two ways no fixture
   reached. The shipped reads are `cut -d: -f2 | tr -d ' '`: `tr` removes
   INTERNAL spaces (`1 0` -> `10`) where `.trim()` keeps them, and `cut` takes
   only the second colon-field where a tail capture keeps the rest. The mirror
   models the pipeline now, and both counterexamples are fixtures — a parity
   assertion that agrees only on well-formed input asserts very little.

6. The prefix census guards were both partly vacuous. The drift guard read the
   two regex alternations and not the severity map, so a set could agree in both
   regexes while mis-tiering in the map. The domain guard scanned only `### XX-01:`
   headings — and BL appears in no heading at all, only in the Label-equivalence
   prose, so the guard passed purely because BL happened to be hard-coded and
   would have missed the next prose-defined prefix exactly as it missed BL. Both
   widened; the domain the guard now sees is BL, CR, IN, WR.

Each fix fails a named test on reversion and none fires on the ordinary path.
The property generator now deliberately produces the reserved suffix from (4),
which a generator drawn only from innocuous characters could never reach.

Also corrected: the previous commit's account of the empty-path failure. The
script did not write a bare `-REVIEW-DISPOSITION.md`; it threw on reading the
empty review path and the trailing `|| echo` swallowed it as a non-blocking
skip. Same silent outcome, different mechanism, and the comment said the wrong
one.

* fix(#3829): the tests now run what bash runs — and six fixes to the fixes

A second adversarial pass over the previous commit. It found a regression that
commit introduced, and the reason it slipped through is the finding worth
keeping.

THE FIDELITY GAP. Every test here extracts the embedded script as TEXT and
runs it. Bash does not: it expands the double-quoted `node -e "..."` argument
first, so a backtick inside it is COMMAND SUBSTITUTION. The previous commit put
one in a code comment. Bash duly ran it, failed with `+: command not found`,
and handed Node a script two bytes shorter than the one 122 green tests were
exercising. No behavioural test could see this, because none of them ever
asked bash what it would actually pass. One now does, and it is the general
guard: it catches an unescaped backtick, an unescaped $, and any other
expansion the extractor cannot model.

Then, in the shipped step:

- A padded count silently disabled the sum check. `$((08 + …))` fails on base
  inference; it does not abort — the expansion sits in an `if` condition, where
  set -e does not fire — so the check simply never ran and an inconsistent
  breakdown passed with a stray diagnostic as its only trace. `10#` on every
  operand.
- The status guard made the script's own reconciliation unreachable. A clean
  review with an EXISTING ledger must still be reconciled — decided rows
  carried, stale `open` rows dropped — or the ledger freezes showing findings
  as open that the review no longer reports. The guard now skips only when
  there is nothing to reconcile.
- The carried marker is no longer stripped at parse time at all. Bounding the
  strip still ate a carried row's human-written reason. No-growth is a property
  of the RENDER, so it is enforced there: a marker already present is not
  appended again. Nothing is stripped, nothing doubles.
- Fence openers are bounded to three leading spaces, per CommonMark.
- parseGateCounts matched `[ \t]` where the shipped grep uses `[[:space:]]`,
  which covers form feed and vertical tab. Third counterexample of the same
  class, and a fixture.
- The census drift guard checked only one direction, so a tier for a prefix the
  regexes never admit stayed green as dead code that reads as coverage.

TWO OF MY OWN TESTS WERE VACUOUS, and the controls are what said so. The
leading-zero test asserted exit 0 and a consistent verdict — both true before
the fix. The clean-review test drove the node script directly, which never
executes the shell guard at all: it passed unchanged with the guard made
unconditional. Both are rewritten to test the layer the defect lives on, and
both now fail when their fix is reverted.

Every fix in this commit fails a named test on reversion, each mutation
verified to have applied before its verdict was read.

* fix(#3829): the carried marker can no longer outlive the carry

A third adversarial pass. Its most important finding is a defect the SECOND
pass talked me into, which is worth recording as plainly as the fix.

THE MARKER BECAME A LIE. Pass 2 objected that bounding the carried-marker strip
still altered a human-written reason, and proposed storing the cell verbatim
instead. That objection was a preference, not a defect — its own driven output
showed exactly one marker, which is correct — and adopting it created a real
one: once the generated marker is stored it can never leave, so a carried
finding that REAPPEARS in a later review still renders "not in the current
review". The ledger then contradicts its own contents. Driven both runs.

The strip is back, bounded to one occurrence and unconditional. The residual
ambiguity is irreducible — a reason ending in exactly that phrase is
indistinguishable from the marker — and it costs nothing real: on a carried row
the render puts the phrase straight back, and on a current row the phrase was
self-contradictory to begin with. The unbounded quantifier is what had to go,
not the strip. The property now states that contract rather than asserting a
verbatim survival the code deliberately does not provide.

Also:

- An ABSENT REVIEW.md abandoned the ledger it was meant to reconcile. The guard
  proceeds when a ledger exists, then the script read the review unconditionally,
  threw, and the trailing fallback swallowed it — the freeze the reconciliation
  path exists to prevent, reached through the door the guard opened.
- Counts are length-bounded as well as digit-only. Bash integers wrap at 2^64,
  so a 20-digit count arrived at the sum as 0 and an inconsistent breakdown
  passed.
- A closing fence must carry only whitespace after its marker; a line with an
  info string is an opener's shape and ended the fence early.
- parseGateCounts matched [ \t\n\v\f\r] where the shipped grep uses
  [[:space:]], which under this UTF-8 locale matches EM SPACE. `\s` is the
  faithful model. Fourth counterexample of that class, and a fixture.
- The agent-domain scan required [A-Z]{2,}, so a one-letter prefix like `C-01`
  — explicit and parseable, not prose — was invisible to it.

AND THE FIDELITY GUARD PAID FOR ITSELF INSIDE ONE SESSION: writing this round's
first draft I put backticks around a token in a code comment again, in the very
commit whose subject is that mistake. The probe failed, named it, and no test
of behaviour could have. Two of my own tests also had to be rewritten: one
asserted things true before its fix, and one drove the node script directly
where the defect lived in the shell.

383 pass across the touched files and the two size ceilings; ten lint gates
green; every fix fails a named test on reversion, each mutation verified to have
applied before its verdict was read.

* docs(#3829): the Source reason is preserved, but not verbatim — say so

Found by claim-auditing the response comment before posting it, which is the
one place this would have been caught: the doc and the code were written in
different commits and only a reader holding both notices they disagree.

The feature reference said the hand-written reason is "preserved verbatim
across re-runs". It is not, and deliberately so — a reason ending in the
literal phrase "(not in the current review)" loses that trailing phrase,
because it is indistinguishable from the carried marker the gate appends.

The exception is stated rather than dropped, with the reason it is the better
trade: storing the marker instead means it never leaves, and a carried finding
that later reappears goes on claiming it is absent from the very review that
reports it. A ledger wrong about its own contents beats losing a duplicated
phrase, but only if the doc admits which one it chose.

FEATURES.md regenerated; gen-features --check and lint:docs green.

* fix(#3829): the gate now emits the counts it computes (B1a/B1b)

Block 1 computed REVIEW_STATUS and the four counts and printed none of
them, then the prose below asked the agent to display four of them. The
shell exits at the closing fence and the agent sees only stdout, so those
values were unobtainable: REQ-REVIEW-08 was unreachable in every shipped
path and the fence was decorative.

The rule was already stated one block down -- "a prose-only gate on a
value no later block can see is not a gate" -- and applied only to block
2. It now governs the block that is this step's primary deliverable.

Both arms emit, and the status gate is mechanical rather than prose:
a clean/skipped/absent review prints nothing, an inconsistent or partial
breakdown prints the countless form, and the full breakdown prints
otherwise. Driven against the review's own case (critical: 1, warning: 9,
info: 8, total: 18) with no appended emitter:

  Code review: 18 findings - 1 critical, 9 warning, 8 info.
  Consider running: /gsd:code-review 1 --fix

* test(#3829): the counts harness stops manufacturing the output it asserts on (B2)

runShippedGateCounts extracted the shipped fence and then APPENDED its own
printf of the six internal variables before running it. Every counts
assertion was green against a script that existed only inside the test
process: the shipped fence emitted nothing, the tested fence emitted six
lines because the test added them. That is why B1a shipped past a suite
that looks like it covers exactly that surface -- the green was
structurally incapable of turning red for it.

The emitter now lives in the fence, so the harness reads the fence's own
stdout and synthesizes nothing. Parity with the mirror moved up a level
with it: renderGateMessage() renders both arms from the mirror's parsed
counts and the assertion compares the WHOLE emitted message, so a drift
in any parsed value changes the string or the arm it selects. Asserting
on the observable is strictly stronger than asserting on five
intermediates, and it can express what the old probe could not -- an
absent review now reports NOTHING, which is a different fact from
reporting a countless review.

A fifth src.includes() assertion converted with it (round 1 retired
four). It pinned the PROSE stating the countless condition, so it went
red when the emitter moved into the fence while the behaviour it named
was untouched -- the pin arguing for its own conversion.

Negative control: reverting the shipped echo now turns 16 tests red.
Before this commit the same reversion turned zero red, which is the
finding.

* fix(#3829): the disposition column is an enum, not any lowercase token (B3)

ADR-227 requires a trust boundary to validate semantic SHAPE and to coerce
a failure to the contract's safe default. The ledger is a trust boundary by
construction -- the rendered instruction tells a human to hand-edit it --
and the prior-row parser captured column 3 as ([a-z]+), checked against
nothing.

One transposed character was enough. `| CR-01 | critical | opne | - |` is
not the literal 'open', so it beat the default, was excluded from the
`open:` headline count, and was carried forward forever. The ledger then
reported the phase fully triaged off a typo.

The asymmetry is what made this a correctness bug rather than a style
point: a typo OUTSIDE [a-z] ('Deferred') already failed to match, lost the
decision and reset the row to open -- safe. A typo INSIDE [a-z] was unsafe.
The parser failed open in the one direction that matters. A row that fails
the enum now yields no prior entry and the row falls back to 'open', by the
same path the capital-D case already took.

The property test could not have caught this: DECIDED is drawn from the
vocabulary, so no property built on it can present an out-of-vocabulary
token. Added JUNK, the arbitrary for the complement, deliberately
lowercase so it stays inside the old capture's own character set -- the
unsafe half is the token that LOOKS like a decision and is not. The new
property also asserts the headline count agrees with the row it renders,
which is the half the defect actually reported wrongly.

Negative control: the new property fails against the ([a-z]+) capture and
passes against the enum.

* fix(#3829): a finding the heading parser cannot match is surfaced, not dropped (B4)

Two independent parsers produce two numbers one paragraph apart -- the
counts from REVIEW.md's frontmatter, the rows from `### <ID>:` heading
matches against a closed CR|BL|WR|IN alternation -- and nothing reconciled
them. A finding the alternation could not reach contributed no row, no note
and no diagnostic, and the ledger then declared `open: 3 of 3` over a set
strictly smaller than the console line had reported one paragraph earlier.
Two findings recorded nowhere, and neither artifact said so.

The PR's own argument for the closed alternation -- that an unlisted prefix
produces no row rather than a MIS-CLASSIFIED one -- is the wrong trade under
this repo's fail-safe rule. A dropped finding is demoted below every finding
that parsed, and an unparseable finding is precisely the one a human most
needs to see.

Block 2 now derives the frontmatter total (anchored inside the findings:
mapping, digit-and-length-bounded like block 1's) and hands it to the
script, which reconciles it against the CURRENT review's matched findings --
order.length, never rows.length, which also counts carried rows and would
either understate the shortfall or invent one. Surfaced exactly as the
stale fix-report case already is: a non-blocking `unparsed: N` key plus the
console line, both naming the two numbers so the claim is checkable.

  Code review disposition recorded: 3 of 3 finding(s) open (2 finding(s)
  recorded NOWHERE: the review reports 5, but only 3 matched the expected
  heading shape `### <CR|BL|WR|IN>-NN: <title>`)

The key is emitted only when there IS a shortfall, so an ordinary ledger
gains no noise key and the unchanged-run check is unaffected.

Four tests, including three negative controls the round owed itself: a
clean review gains no key, an absent/non-numeric total reconciles nothing
rather than fabricating a shortfall, and a total SMALLER than the row count
cannot render `unparsed: -1`. Reversion control: dropping the key turns the
first red.

* fix(#3829): pass --raw to the commit_docs config-get (#3763)

Not from the review -- from a gate the base range added after it. #3763
lands `tests/config-get-raw-guard.test.cjs`, and this branch was its sole
offender: a config-get command substitution without --raw feeds
JSON.stringify output into a bash string comparison, where it silently
never matches for string values. The consumer here is exactly that:

  if [ "$COMMIT_DOCS" = "true" ]

Every other shipped call site in the tree already passes --raw
(spike.md, fast.md, new-milestone.md, sketch-wrap-up.md, ...), so this is
sibling convention, not a new posture.

Worth recording because the two readings are both correct and they
disagree: round 2's review cleared this exact line under ADR-3409 as "the
safe member of that family", since `query config-get <key>` with no --pick
exits 1 on absence and the fallback arm is reachable. That is still true --
--raw does not change it. The base then moved and added a gate that reads
the same line for a different property.

* fix(#3829): scope the count reads to the findings: mapping, not just the frontmatter (m1)

`^[[:space:]]*total:` matches any indented key anywhere in the block, so a
top-level key later named `total:`, `info:` or `critical:` was picked up
ahead of the nested one. The block's own extensive comment is about scoping
the FRONTMATTER, and the scoping stopped one level short of the mapping the
values actually belong to. `status:` was never exposed -- it is anchored to
column 0 because it IS top-level.

The reads now run over the `findings:` block alone, selected by awk and cut
at the next column-0 key. Block 2's REVIEW_TOTAL derivation (added with B4)
already used that filter; this brings block 1 to it, so the two agree by
construction rather than by coincidence.

The mirror models the same scoping, and two fixtures drive it: a top-level
`total: 999` ahead of a nested `total: 1`, and top-level `critical:`/`info:`
ahead of theirs. Reversion control: unanchoring the shipped reads turns them
red.

* fix(#3829): severity comes from the section heading, not just the id prefix (M3)

gsd-code-reviewer.md emits findings under '## Critical Issues' /
'## Warnings' / '## Info', and that heading is the reviewer's own statement
of a finding's severity. The walker already visits every line -- the
fix-report path tracks '## ' sections -- so the signal was in hand and
discarded in favour of the id prefix alone.

A reviewer who mis-numbers a Critical as WR-04 while filing it under
'## Critical Issues' produced a row reading 'warning'. The ledger's Severity
column is the whole basis for triaging it, and it then disagreed both with
the review it summarizes and with the frontmatter count line block 1 prints
from findings.critical.

Section first, prefix as fallback: a finding under no recognized section --
a review that does not use the documented headings, and every row carried
from an earlier review -- keeps the prefix mapping, BL- included. Sections
are matched WHOLE, exactly as the fix-report sections are, so
'## Critical Issues Verification' does not re-tier what sits under it, and
a heading inside a fenced example does not govern.

Five tests: both mis-numbering directions, the prefix fallback across all
four prefixes, the lookalike heading, and the fenced-example case.
Reversion control: prefix-only turns the first two red.

Sixth src.includes() assertion converted with it -- it pinned the exact
source LINE of the enumeration loop, so it went red when that loop was
reformatted while the property it names was strictly widened. It now
asserts the property: every finding id, in order, once each.

* fix(#3829): an untriaged row is carried too, not silently deleted (M1)

The carry-forward kept a prior row only when its disposition was not
'open', so an untriaged row for a finding the current review no longer
reports was dropped entirely. Combined with the reconciliation gap that
left EVERY row open, a re-review deleted the whole ledger.

The re-review loop rewrites REVIEW.md on every iteration, so REVIEW.md does
not retain it either: run 1 records CR-01 open, the re-review renumbers it
to CR-02, run 2's ledger contains neither. That is #3829's complaint
verbatim -- "no trace of what happened to them" -- reproduced by the
artifact built to prevent it. The old justification, "nothing was decided
about it", is exactly the state #3829 says must leave a trace.

Every prior row is now carried, and the carried marker is what keeps it
honest: the row does not claim the finding is live, it records that it was
seen and never triaged. Two costs, stated rather than discovered: a
renumbered finding shows twice until the old row is triaged, and a carried
untriaged row persists until decided. Both are bounded by the phase's own
findings, both are legible from the marker, and both beat a silent delete.

Five tests updated -- they encoded the dropped-untriaged behaviour as the
contract -- plus one new test for the renumbering case M1 names. Reversion
control: restoring the guard turns six red.

Two self-inflicted defects caught while writing this, both by probes round
1 built:

  - Four unescaped backticks in a comment inside the double-quoted node -e
    argument, which bash ran as command substitution. The extractor-parity
    probe fired ("--auto: command not found"). Third time that trap has
    been sprung in this PR, third time the probe caught it.
  - The reworded ledger footer contained the literal carried-marker phrase,
    and the marker-accumulation assertion counts it across the whole file,
    so a doc line read as a second marker. The assertion was right.

* test(#3829): cover the count-length threshold at limit-1, limit and limit+1 (M2)

The guard is `?????????*` -- nine or more characters -- so the limit is
8 digits accepted, 9 rejected. The only cases were 'x', single digits and a
20-digit value, none of which pins the boundary. RULESET.TESTS
boundary-coverage is a hard rule here and it was unmet.

All three points asserted, with the sum kept consistent at each so the
LENGTH rule is what decides the verdict rather than the sum check
incidentally agreeing.

Reversion control is the off-by-one M2 names: dropping one `?` moves the
limit to 7 digits, which no test could previously notice, and now turns
this one red.

* feat(#3829): wire the disposition ledger into the fix path (B1c/B1d)

REQ-REVIEW-09 was unreachable in every shipped path. execute-phase.md's
code_review_gate invokes review with neither --fix nor --auto, so
<NN>-REVIEW-FIX.md cannot exist when the gate runs and every row it writes
is `open` by construction. The operator then runs /gsd:code-review N --fix
by hand -- the very suggestion the step prints -- which writes REVIEW-FIX.md
and never touched the ledger. A phase with 23 findings, all fixed, ended at
`open: 23 / total: 23`: the artifact that exists to distinguish a triaged
finding from a forgotten one asserted that 23 triaged findings were
forgotten. Worse than recording nothing, because it looks authoritative and
is inverted.

Taking remedy (i), not (ii). Narrowing the docs to say the ledger reflects
the previous phase execution is a legitimate choice, but it ships a feature
whose central artifact is inert and then documents the inertness.

ONE ADAPTATION, because the prescribed site does not exist. The review says
to wire code-review.md's --fix/--auto path. code-review.md is not the writer
(gsd-code-fixer writes the report, code-review-fix.md commits it), and more
decisively it has no point that is AFTER the report exists: it delegates
through code-review/steps/dispatch-fix.md, which calls
Workflow(code-review-fix.md) and then exits the workflow. There is nothing
downstream of that call to wire to.

The site is code-review-fix.md, immediately after commit_fix_report. That
is where the report is on disk and committed, it is the canonical
implementation for all fix logic by dispatch-fix.md's own statement, and it
additionally covers a direct invocation of that workflow -- which a wiring
in code-review.md would have missed.

The same step, not a second copy: it consumes PHASE_DIR and PHASE_NUMBER,
both already parsed from the init JSON, and it is idempotent, so a phase
that reaches the gate and then a fix run ends with one ledger reflecting
both rather than two competing ones.

Driven end to end: the gate writes `open: 2 of 2`, the fix path reconciles
to fixed/skipped and `open: 0`. Two tests -- one pins the wiring and its
ordering relative to commit_fix_report and present_results, one drives the
two call sites in sequence. Reversion control: removing the step turns the
first red; the second covers the reconciliation the wiring makes reachable
rather than the wiring itself.

* fix(#3829): a reflowed fix-report title is the same title (m2)

The stale-fix-report guard compared titles with trim() equality. The strict
instinct is right -- ids are reused across re-reviews, so a stale
REVIEW-FIX.md must not mark a brand-new CR-01 as already fixed -- but
gsd-code-fixer.md writes '### {finding_id}: {title}' under no contract that
the title is copied byte-for-byte from REVIEW.md. A fixer that reflows a
long title produced a spurious mismatch note, left a genuinely-fixed row
'open', and told the reader the report named a different finding. That
false-positive mode was acknowledged nowhere.

Whitespace is normalized, and only whitespace: a wrapped title is the same
title, and it is the one divergence that carries no information. Case
changes and truncation stay strict on purpose -- they are the shapes a
genuinely DIFFERENT finding takes, and widening to them would trade a
visible false positive for the silent false negative the strict match
exists to prevent. The residual is now stated in the step rather than left
to be rediscovered.

The note's wording changed with it. It asserted the report "names a
different finding"; both causes reach that branch and the step cannot tell
them apart, so it now reports the observation -- "titles its finding
differently from the review ... a stale report, or a re-titled one" --
rather than a conclusion it has not earned.

Three tests: the reflow case reconciles cleanly, the re-cased case still
reports, and the stale case still reports with the new wording. Reversion
control: restoring the strict comparison turns the reflow test red.

Seventh src.includes() converted -- it pinned the comparison EXPRESSION, so
it went red when the comparison gained normalization while the property it
names was unchanged.

* docs(#3829): describe the flow that ships, not the one implied (m3)

Both reference pages said "/gsd-code-review <N> --fix records fixed and
skipped, which the gate reconciles from REVIEW-FIX.md" -- true in the
abstract, materially misleading in practice, because no shipped path
performed that reconciliation. With B1c/B1d wired it is now real, and the
pages say WHERE it happens rather than leaving a reader to assume the
in-phase gate does it: the gate runs before any fix report exists and
writes all-open, and --fix is what records what happened.

The round's other behaviour changes land here too, since a reference page
that lags the artifact is worse than none:

  - the disposition column is a closed vocabulary, and a value outside it
    falls back to open rather than being treated as a decision
  - severity comes from the section heading when the review uses one, and
    from the ID prefix otherwise
  - an unparsed shortfall is stated rather than dropped
  - titles are compared ignoring whitespace, so a reflowed title still
    reconciles, and a mismatch is reported as an observation rather than as
    a claim that the report is stale
  - EVERY row is carried now, triaged or not, with the cost of the
    renumbered-finding double-entry stated rather than left to be found

docs/FEATURES.md regenerated from the fragment; lint:generated-sync and
lint:docs both exit 0.

* chore(#3829): migrate the emitted-drift ack from a fragment to commit trailers

ADR-3942 landed on next in #3954: the acknowledgment is a git commit
trailer now, and tests/emitted-drift-acks/ no longer exists.

Worth noting for anyone reading the rebase: this did NOT surface as the
modify/delete conflict the migration guidance predicts. This branch ADDED
its fragment rather than modifying an existing one, and the base deleted
only the files that were already there, so the replay was clean and the
fragment survived silently into a directory that no longer exists. Quieter
than a conflict, and worse -- the gate is what catches it, not git.

Two Growth keys rather than the fragment's one: round 2 wired the ledger
into code-review-fix.md, so that file grew too. Both key on the bare
filename, per the Growth namespace.

Emitted-Drift-Ack-Growth: code-review-fix.md — #3829 review round 2, blocker 1c/1d: REQ-REVIEW-09 was unreachable in every shipped path because the in-phase gate runs before any REVIEW-FIX.md exists, so every ledger row it wrote was open and nothing ever reconciled them. This file gains one step, record_disposition, that reads and executes the same lazily-read step after commit_fix_report. It is the only point in the fix flow that is after the report is on disk: code-review.md delegates here through steps/dispatch-fix.md and exits, so it has no such point at all. Growth is one step of prose, no logic is duplicated, and the step is idempotent so the two call sites converge on one ledger.

* chore(#3829): the changeset describes the round's behaviour, not round 1's

It renders into CHANGELOG, so it carries the same misleading implication
minor 3 was about: "the gate ... reconciling fixed/skipped from
REVIEW-FIX.md" reads as though the in-phase gate does it, when the gate
runs before any fix report exists. Says where it happens, and picks up the
round's other user-visible changes -- carried untriaged rows, section-based
severity, the disposition vocabulary, and the unparsed shortfall.

* fix(#3829): a dotted phase number no longer aborts the step

Found by this round's own adversarial review, in its MISSED section: no
finding asked about it, and it is the most serious thing in the round after
the two blockers.

Both callers explicitly accept a dotted phase -- code-review.md:60 and
code-review-fix.md:36 both validate ^[0-9]+(\.[0-9]+)?$ and name "03.1" in
their own error text -- and both fences reconstructed the path with
`printf "%02d" "${PHASE_NUMBER}"`, which cannot format one. Driven with
PHASE_NUMBER=3.1: bash prints `invalid number` and exits 1, and under
`set -euo pipefail` that aborts the step on its FIRST line. An advisory
gate that promises never to block took the phase's entire review report
down with it, and the newly wired fix-path call site inherited the same
defect.

Pad the integer part and carry the sub-number verbatim, so 3.1 -> 03.1 and
3 -> 03, with both arms falling back to the raw value rather than aborting.
Driven: 3.1 now reads 03.1-REVIEW.md and writes
03.1-REVIEW-DISPOSITION.md; the integer path is unchanged.

Two other findings from the same review, both about claims rather than code:

MINOR 2's TEST WAS MIS-NAMED, and the reviewer was right to refute the
claim. It called itself the "reflowed" case while substituting triple
spaces, which is not a reflow. Driven: a genuinely WRAPPED heading is still
not reconciled, because a `###` heading is one line by definition and the
continuation is a separate paragraph. Not widened -- absorbing whatever
follows a heading into the title would swallow arbitrary prose and make the
stale-report check meaningless, and the kept failure mode is the safe one
(a visible mismatch note, never a wrong "fixed"). The test is renamed to
what it covers and the bound is now pinned by its own test.

THE SHELL-SHARING GUARD DID NOT GUARD. Negative-controlling it -- rather
than reading it -- showed that deleting block 2's real REVIEW_FILE
derivation left it GREEN, on the exact defect it was written for. Block 2
prefixes its `node -e` with `REVIEW_FILE="${REVIEW_FILE}" ...` to put the
values in the child's environment, and the detector counted that
self-referential pass-through as a derivation. Pass-throughs are now
excluded, and the control fires. Pre-existing, not introduced here: the
original column-0 anchor matched that same line.

Also worth recording: my first attempt at that control silently patched
nothing and reported clean. Same lesson this PR already learned once.

* fix(#3829): validate the phase number before formatting it, and make the shell guard executable

Three findings from the round review's continuation pass, all confirmed by
driving them.

1. MY OWN DOTTED-PHASE FIX WAS WRONG on the fallback path. `printf "%02d"
   abc` writes `00` to stdout BEFORE it fails, so
   `$(printf ... || printf %s ...)` CONCATENATES the two: `abc` became
   `00abc`, empty became `00`, and a legitimate `08.1` became `0008.1`
   because bash reads the leading zero as octal. An unset PHASE_NUMBER also
   aborted under `set -u` -- in the step that promises never to abort.

   Validate, then format: never format and fall back on failure. Driven
   across every edge the review named -- 3.1 -> 03.1, 3 -> 03, 08.1 -> 08.1,
   09 -> 09, 1.2.3 -> 01.2.3, and abc / empty / -1 / unset carried verbatim
   with exit 0.

2. THE SHELL-SHARING GUARD STILL DID NOT GUARD. Excluding pass-throughs was
   not enough: a structural predicate recognises assignment TOKENS, never
   assignments that derive a usable value, so `REVIEW_FILE=`,
   `REVIEW_FILE=$REVIEW_FILE` and a commented-out assignment all evaded it.
   No regex closes that class.

   The authority moves to execution -- the third time this PR has learned
   that lesson. The real second fence now runs in a fresh shell with nothing
   but the step's two declared inputs and must write the ledger at the
   correct derived path. All four mutations are caught: empty assignment,
   self-reference, commented-out, and deletion. The textual check stays as a
   cheap fast-fail and is labelled as one.

3. THE TITLE-BOUND CORRECTION HAD NOT REACHED THE DOCS. The step comment and
   both docs pages still said a reflowed title reconciles, contradicting the
   bound pinned one commit earlier. Superseded prose left standing reads as
   current to anyone arriving cold, so all three surfaces are rewritten
   rather than annotated, and FEATURES.md regenerated.

Also hoisted `HAS_BASH` to the file's other top-level constants. `const` is
in the temporal dead zone until its declaration runs, and a
`{ skip: !HAS_BASH }` option object is evaluated eagerly, so a bash-gated
test added above the old mid-file declaration threw a ReferenceError that
aborted its whole describe and CANCELLED its siblings -- while the summary
line still read `fail 0`. It caught three separate additions in this round
before I stopped moving tests and moved the constant.

* fix(#3829): refuse an out-of-shape phase number instead of carrying it into a path

Self-found while writing the prompt for the next review pass, which is the
honest provenance: I asked the reviewer whether a path traversal was
reachable through PHASE_NUMBER, then checked before dispatching.

It was, and I had introduced it. The previous commit's fallback carried an
unusable phase number VERBATIM, and PHASE_NUMBER is interpolated into a file
path:

  PHASE_NUMBER='../../etc/passwd'
  -> REVIEW_FILE=/tmp/phase/../../etc/passwd-REVIEW.md

The `printf "%02d"` it replaced had at least mangled that to `00`. A fix
that makes a path more reachable than the bug it replaced is a regression,
whatever it does for the case it was written for.

Both callers already validate ^[0-9]+(\.[0-9]+)?$ (code-review.md:60,
code-review-fix.md:36), so this is defense in depth rather than a live
exploit -- but the step has two call sites now and should not take either
caller's word for its own inputs. It validates the WHOLE value and, on
failure, builds no path at all: PADDED is empty and each fence refuses by
name rather than coercing. Block 1 declines to report counts read from a
path made out of the bad value; block 2 declines to write, which also keeps
it clear of the bare-name ledger defect round 1 closed.

Driven across the shape boundary: 3.1 / 3 / 08.1 / 09 accepted; abc, empty,
unset, 1.2.3, -1, 3., .1, +1, "3 1" and ../../etc/passwd all refused with
exit 0 and a named diagnostic. Reversion control: restoring carry-verbatim
turns the traversal test red.

* fix(#3829): bound the phase number's length, and make the shell guard prove derivation

Third adversarial pass. Two of its three refutations were already closed by
the previous commit (the ../escape and 1/../../escape traversals, and the
unset-input abort); these two were not.

1. A 54-DIGIT PHASE NUMBER WRAPPED SILENTLY. The validator accepted any
   all-digit value, so `$((10#$_int))` overflowed 64 bits and PADDED became
   `-7908320945662590977`. Length-bounded now at 8 digits, exactly as the
   counts already are and for the identical reason -- and the counts guard
   sitting twenty lines away is why this one is embarrassing rather than
   subtle. Driven at the boundary: 8 digits accepted, 9 rejected.

   The bare `${PHASE_NUMBER}` in the suggestion line is hardened to
   `${PHASE_NUMBER:-}` while here. The empty-PADDED guard makes it
   unreachable today, but it is one refactor away from an unbound-variable
   abort under `set -u`, in the step that promises not to abort.

2. THE EXECUTED SHELL GUARD PROVED THE FENCE WORKS, NOT THAT IT DERIVES.
   A single-phase probe is satisfied by a hardcode, and the review
   demonstrated exactly that: replacing the derivation with
   `case ... in 1) PADDED=01 ;; 7) PADDED=07 ;; *) PADDED=07 ;; esac`
   breaks every real phase and passed the entire suite. It now runs two
   distinct phases, 7 and 3.1 -- a hardcode cannot satisfy both, and the
   dotted one additionally pins the integer-part split.

   The claim "given only the declared inputs" was also overstated: the test
   spreads `...process.env` (it needs PATH and HOME). The DERIVED names are
   now explicitly deleted from that environment, so the claim is true rather
   than merely intended.

Also rewrote a comment that had become false: it pinned a describe to the
end of the file because of the HAS_BASH temporal-dead-zone constraint, which
the hoist removed. Superseded prose left standing reads as current to
anyone arriving cold.

The changeset's "the gate stays advisory and never blocks" is now verified
rather than asserted: both fences exit 0 under an unset PHASE_NUMBER and a
traversal-shaped one.

* fix(#3829): validate both inputs, refuse before building a path, and never write through a symlink

Fourth adversarial pass. Four findings, all confirmed by driving them.

1. PHASE_DIR WAS NOT VALIDATED AT ALL. Unset, both fences died with
   `PHASE_DIR: unbound variable` under `set -u` -- the same class as
   PHASE_NUMBER, which I had just spent two commits fixing while its sibling
   input sat one line away. The step declares two inputs; it now validates
   two.

2. THE LENGTH BOUND WAS ON THE WRONG THING. The nine-character glob applied
   to the WHOLE value rather than the integer part, so it falsely rejected
   `12345678.1` (a legal 8-digit phase) while accepting `1.123456`. Each
   component is bounded on its own now; the sub-number is bounded too, since
   it is likewise interpolated into a filename.

3. REJECTED VALUES STILL HAD PATHS BUILT FROM THEM. The refusal guard sat
   AFTER the assignments, so an unusable input still assembled
   `${PHASE_DIR}/-REVIEW.md` and stat'ed it before refusing. The guard is
   now the first thing after validation, and both fences construct paths
   from validated locals rather than from the raw environment.

4. THE LEDGER WRITE FOLLOWED SYMLINKS. From the review's MISSED section, and
   the sharpest thing in it: `fs.writeFileSync` follows a symlink, so a
   pre-existing symlink at the ledger path replaced the contents of whatever
   it pointed at -- outside the phase directory, with the link left intact
   so nothing looked wrong. Driven, and the target's contents were gone.
   This PR introduces the artifact, so it owns the check: an existing ledger
   that is not a regular file is not a ledger, and the advisory gate says so
   and steps over.

The executed shell guard now draws its phases AT RUN TIME. Fixed fixtures
cannot establish derivation -- the review defeated the one-phase version
with a hardcode, then defeated the two-phase version by adding one more arm
to the same case. Any finite sample loses that race. A phase picked per run
cannot be enumerated in advance; the drawn values print in every assertion
message so a failure stays reproducible. Control: the review's three-value
hardcode now fails on three consecutive runs.

Eighth src.includes() converted -- it pinned the literal `${PHASE_DIR}`
interpolation and went red when construction moved to a validated local,
while "writes a REVIEW-DISPOSITION sibling" was untouched. It now asserts
that property, and that REVIEW.md is not written.

The changeset's "stays advisory and never blocks" is verified rather than
asserted: 8 of 8 hostile-input cases across both fences exit 0 -- both
inputs unset, PHASE_DIR unset, a traversal-shaped phase, and a missing
phase directory.

* fix(#3829): check the ledger path before reading it, and pin the write-safety behaviour

Fifth adversarial pass, and the last one this round. Three fixes, three
disclosed residuals.

FIXED

1. A FIFO AT THE LEDGER PATH BLOCKED FOREVER. readFileSync on a FIFO never
   returns, so the step documented as "advisory, never blocks" blocked
   indefinitely -- the literal counterexample to its own headline claim. The
   non-regular-file check ran after that read.

2. THE UNCHANGED-RUN FAST PATH BYPASSED THE CHECK. A symlink whose target
   already matched the rendered ledger read through the link, reported
   `unchanged`, and never reached the refusal.

   Both fixed by the same move: the check is now the FIRST thing the script
   does, before any read or write of that path. Ordering was the defect, not
   the predicate.

3. THE COMMIT TEST FOLLOWED THE LINK the script had just refused. `[ -f ]`
   resolves symlinks, so the guard and its consumer disagreed about the same
   path and the helper could still be handed one. `[ ! -L ]` added.

   And the behaviour shipped with NO regression control -- I hand-drove it
   last commit and did not pin it, which the review caught by grepping for
   the words. Five tests now: symlink, symlink-with-matching-target, FIFO,
   directory, and an ordinary ledger as the negative control so the refusal
   is not a blanket one. mkfifo goes through the process seam like every
   other spawn here.

DISCLOSED, NOT FIXED -- these are stated in the step rather than carried
silently:

- TOCTOU between the lstat and the write. Node exposes no portable
  O_NOFOLLOW write, and an attacker who can write into the phase directory
  mid-run already has what the check would protect. It narrows a real
  accident; it is not a security boundary and the docs claim none.
- A hard link passes isFile() by construction.
- The REVIEW.md and REVIEW-FIX.md reads still resolve symlinks. They are
  reads of files the operator owns, in their own phase directory.

Also narrowed a comment that overclaimed. The randomized guard's domain is
FINITE -- 88 integer and 792 dotted values -- so a mutation enumerating all
880 passes forever, and Math.random() is unseeded, so "reproducible" means
only that the drawn values are printed on failure. Raising the bar is what
it buys; proving derivation is not, and nothing short of reading the fence
is. The previous comment claimed otherwise and was refuted.

* test(#3829): make the write-safety controls portable to the Windows lane

CI caught what neither the local suite nor five adversarial review passes
could: every one of those ran on Linux.

The FIFO test gated on `mkfifo`'s exit code. On the Windows lane mkfifo
EXISTS and exits 0 while producing something that is not a FIFO, so the
guard passed, the test ran against an ordinary path, the ledger wrote
normally, and the assertion failed for a reason unrelated to the behaviour
under test. It now gates on `lstatSync().isFIFO()` -- what was actually
created, not what the command claimed. Control: with the shipped guard
disabled the test still goes red on Linux, where the FIFO is real.

The two symlink tests are skipped on win32, following this repo's existing
convention for symlink-planting tests (tests/settings-jsonc.test.cjs:389
skips the same class; tests/unreachable-guard-drift.test.cjs:726 records the
reason -- symlink creation requires elevated privileges on Windows CI). The
privilege happened to be available on the lane this round, which is exactly
why the convention is not "try it and see".

* fix(#3829): a bare `|` in a deferral reason is prose, not a parse failure

Review round 3, the one blocker. The Source cell is the one field this ledger asks a
human to hand-edit, and "waiting on team A | team B to align" is an ordinary thing to
type there. The prior-row capture admitted a pipe only when escaped, so a bare one
failed the WHOLE line: prior.get() was undefined, the row fell through to `open` with
an empty Source, and the console line read "1 of 1 finding(s) open" — a Critical a
human explicitly deferred, with a documented reason, rendered indistinguishable from
one never triaged, and the reason gone. The exact ambiguity #3829 exists to remove,
reachable by one missing backslash.

The Source cell is the LAST column, so it is now captured through to the end of the
line, less an optional trailing pipe; a bare `|` inside it is prose. The render
escapes a bare pipe on the next write so the table stays a table, and the escaped
form re-parses to itself, so the second run reports `unchanged` — the fixed point
holds. The ledger's own instruction line says so instead of asking the human to
escape.

Why the property never caught it: SOURCE_CELL only ever appended a PRE-ESCAPED pipe,
so the arbitrary built to stress this cell could not reach the one input that broke
it. It now also emits a bare pipe, and the round-trip expectation is the escaped
form of what the human wrote. A fixed regression case drives the reviewer's exact
input through two runs and asserts the decision, the reason, the headline count and
convergence. Negative-controlled: both new tests fail against the previous capture.

The src.includes() pin on the old capture text is retired for the behavioural case —
it was pinning the defect.

* fix(#3829): escape every bare pipe in one write, whatever precedes it

Round 3, found by the adversarial pass over the round's own fix rather than by the
review. The first escape used /(^|[^\\])\|/g, which CONSUMES the character before the
pipe: adjacent bare pipes were escaped one per run (A||B -> A\||B -> A\|\|B, a third
run to converge, breaking the advertised second-run fixed point), and an escaped
backslash before a pipe (A\\|B) hid the pipe behind the wrong parity and left it bare
in the rendered table. The property generator emits at most one bare pipe, which is
the one case the old form got right, so no property reached either.

Scan as pairs instead: an escaped pair (backslash + anything) is kept verbatim and only
a pipe outside one is escaped. One write, then a fixed point. Regression case drives
`A||B and C\\|D` through two runs; it fails against the previous escape.

* fix(#3829): the script leaves by return, so an explicit exit cannot drop its verdict line

Round 3, from the adversarial pass over the round's own fix. The embedded node script
printed its verdict and then called process.exit(0) -- on the 'unchanged' branch only;
the 'recorded' branch fell off the end. Node's "A note on process I/O" documents
process.stdout writes to pipes and sockets as asynchronous on POSIX, and process.exit()
as forcing exit before pending asynchronous stdout writes complete -- so on a POSIX
lane the caller can see exit 0 with no verdict line. This is a hardening against that
documented hazard, not a reproduced defect: the reviewer's empty-second-run stdout,
which first pointed here, turned out to be its own sandbox -- a bare console.log child
printed nothing there either -- and that attribution is withdrawn.

The script now runs inside main() and leaves by return on all four early-exit paths,
so the event loop drains stdout before the process ends. Same exit status either way,
and the || echo fallback is unaffected. A structural test pins the absence of the call
(comment-stripped; dotted, bracketed and whitespace-split spellings). The empty-review
docs-parity pin that asserted the literal process.exit(0) line is retired -- the
round-2 describe drives that property behaviourally.

Also widens the property generator: SOURCE_CELL now reaches adjacent pipes and a
backslash of either parity before a pipe, the two shapes the first render escape got
wrong while passing every input the generator could then produce -- checked against
an independent parity-walk oracle rather than a copy of the render's own scan.

* fix(#3829): record what an --auto iteration fixed, instead of reporting it open

Round 5's major. `record_disposition` runs once, after the whole capped-at-3
`--auto` loop converges — but this workflow keeps ONE final version of REVIEW.md
and REVIEW-FIX.md rather than per-iteration copies, and deletes the .iterN.md
backups on convergence. A finding fixed in iteration 1 was therefore absent from
the final review (it was fixed, so the re-review stopped reporting it) AND from
the final fix report (overwritten by the last iteration), so the row fell back to
the gate's `open` and rendered `open ... (not in the current review)` — the same
bytes a finding that vanished for an unrelated reason produces. That is the one
distinction #3829 exists to make, undone by the artifact built to make it.

The precise site was the two-arm `applied` construction: for an id the current
review does not report, `sameTitle(undefined, h.title)` is false and
`title.has(id)` is false too, so the entry entered NEITHER `applied` NOR
`staleFix`. It was dropped in silence.

Four changes, one defect:

- A third arm. When the review does not report an id at all there is no title to
  disagree with, so this is not the stale-report case — it is what a finding
  looks like once it has been acted on. Record it. The id-reuse hazard stays
  closed by the arm below it: when the review DOES report the id, a title
  mismatch still goes to `staleFix` and is never applied, so a renumbered
  finding cannot inherit an earlier iteration's `fixed`.
- Rows for decided ids the review no longer reports, carried and marked. A
  decision the ledger cannot render is a decision lost — the same silent drop
  the carry-forward loop already refuses for prior rows, one source over.
- The .iterN.md fix-report backups are read alongside the final report, newest
  first, so the most recent statement about an id wins — the precedence a
  duplicate id already gets within one report.
- The shell guard proceeds on a fix report, not only on an existing ledger. A
  direct `/gsd-code-review N --auto` writes no gate ledger, and a converged loop
  leaves `status: clean`, so a fully successful multi-iteration run recorded
  nothing at all.

And the backups now go in `cleanup_iteration_backups`, after the ledger has read
them. #3190's rule is untouched — spent scratch on convergence, retained on
degradation — only the timing moved; deleting them inside the loop erased every
early fix before anything read it. `CONVERGED` does not survive the loop's shell
and is re-derived from the final review's status, which is exactly how the loop
sets it; anything but a proven-clean review retains.

Seven new regression tests plus an ordering test, all eight reversion-controlled
against pre-fix code — every one fires. One is the negative control that matters:
a reused id whose title differs must stay `open`, never inherit `fixed`.

Residual, stated: an id appearing only in an iteration fix report takes its
severity from the id prefix rather than a section heading, because `sectionSev`
is built from the current review. That is the documented fallback for carried
rows, not a new gap.

* test(#3829): pin the two PADDED derivations against a silent desync

Round 5's minor 1. Each fenced block runs in a fresh shell and must derive what
it reads, so the PADDED derivation — the traversal fence between an
attacker-influenceable phase number and a file path, plus the per-component
length bound — is duplicated verbatim. Both copies were independently tested and
nothing asserted they stay in step, which is the shared-parallel-surface shape
CLAUDE.md requires a parity test for, on security-relevant validation logic
rather than incidental repetition.

Compared line by line rather than through a normalizing rewrite: a normalizer
has to be told what may differ, and whatever it is told to tolerate stops being
asserted. Exactly one line may differ — each block refuses by its own name — and
the test names both forms. It also asserts the slice is substantial, since a
parity test over an empty slice passes vacuously.

Control: dropping one `?` from block 2's length bound, which moves that copy's
limit to 7 digits while block 1 keeps 8, turns it red. That is the exact silent
divergence the finding describes.

One correction to the finding's own statement, since it is worth recording: the
cited lines are :324 and ~:480, which are node-script lines; the derivations are
at :48-83 and :211-246. And they are 35-of-36 identical rather than
byte-identical — the refusal message differs, deliberately.

* docs(#3829): state the PHASE_DIR trust boundary instead of carrying it

Round 5's minor 2 asked that the assumption behind PHASE_DIR's validation be
confirmed rather than silently carried forward at the two new call sites. It is
confirmed, and the comment that stood here was wrong about it: "PHASE_DIR is the
step's other declared input and gets the same treatment" describes something the
code does not do.

Both inputs have the SAME provenance — each caller binds them from
`gsd_run query init.phase-op` (code-review-fix.md:7,17; execute-phase.md the
same) — so neither is raw user input and neither is more trusted. The asymmetry
is not about trust. It is that only one of them has a shape: PHASE_NUMBER
carries a documented contract, `^[0-9]+(\.[0-9]+)?$`, asserted by both callers,
so a value outside it is provably wrong and is refused. PHASE_DIR's contract is
"a filesystem path", which admits `..`, absolute and relative forms and
symlinked parents alike; no predicate separates a legitimate planning directory
from an illegitimate one, so a shape check would reject working setups while
proving nothing.

So the emptiness check is adopted as what it actually is — the guard against
`PHASE_DIR: unbound variable` aborting a step that promises never to block — and
the shape check is declined, with the reason written where the next reader meets
it rather than left to be re-derived.

The residual is restated in place rather than left in a PR comment: PHASE_DIR
may itself be a symlink and the ledger is then written through it, outside the
phase directory, deterministically. Left alone deliberately — the write goes
where the caller pointed. Not a security boundary, and nothing here claims one.

* docs(#3829): record why HAS_BASH is a platform assumption, not a probe

Round 5's minor 3 is DECLINED, and the reason is the repo's own contract rather
than a judgement call — written at the constant so the next reader does not
"fix" it and re-enable what the rule exists to prevent.

The gap is real and confirmed: 22 tests carry `{ skip: !HAS_BASH }`, so block
1's bash severity-reporting path has no Windows-lane coverage. But
`local/no-unguarded-nonportable-exec`
(eslint-rules/no-unguarded-nonportable-exec.cjs, DEFECT.WINDOWS-TEST-PORTABILITY)
REQUIRES this guard around `sh -c` / `bash -c` in tests, and its own remedy text
names `if (process.platform !== 'win32')` as the sanctioned form, because these
constructs fail under Windows Git Bash. So the constant is the repo's answer to
this question, not an oversight in this PR.

Swapping it for a runtime `bash` probe would light 22 tests up on a lane the
rule has already determined they cannot pass — trading a legible, rule-encoded
skip for a red matrix. Reversing that is the rule's decision; a change here
belongs with a change there.

* docs(#3829): describe how --auto's iterations reach the disposition ledger

The reconciliation section described the `--fix` path accurately and said
nothing about `--auto`, which is where round 5's major lived. It now states
that the loop overwrites its fix report each pass, that the re-review drops a
finding once it is fixed, that the gate therefore reads the per-iteration
backups newest-first, and that the backups are removed after the ledger has read
them rather than before. It also states the converged-with-no-ledger case: a fix
report on disk is reason enough to record.

FEATURES.md regenerated (176 features / 21 groups). Changeset extended to name
the shipped behaviour rather than only the `--fix` half.

* fix(#3829): clear lint-workflow-shellcheck, a gate the base range added

Not from the review. The rebase onto `next` brought in `lint-workflow-shellcheck`
(#4109), whose baseline was generated before this PR's new step file existed — so
that file's findings are new by construction and `lint:ci` exited 1 on the
rebased head before this round touched anything. The last green CI run predates
the gate. Caught locally rather than by a red push.

Three fixes and one baseline entry, split by whether the finding is real:

- STRUCTURAL (not ShellCheck, not baselineable): the guard's
  `for _f in "…${PADDED}-REVIEW-FIX.iter"*.md` is the bare `for x in $VAR` shape
  that word-splits differently under bash and zsh. Wrapped in
  `$(printf '%s' "$PADDED")`, the linter's own prescribed remedy.

- SC2097/SC2098, and this one was a genuine latent bug rather than a lint nit:
  `FIX_REPORT_FILE="${_pd}/${PADDED}-REVIEW-FIX.md"` sat in the same env-prefix
  list that sets `PADDED`, so its `${PADDED}` expanded the OUTER variable, not
  the one two entries earlier. Both happen to hold the same value here, which is
  exactly why it would have kept being wrong quietly. Built before the command
  now.

- SC2317 ×3 is baselined, not fixed. It fires on
  `return 0 2>/dev/null || exit 0` — the deliberate idiom that lets a fence
  refuse whether it is sourced or executed — and the verdict is a false
  positive: the `exit 0` is reached precisely in the executed case. Rewriting a
  dual-mode refusal to satisfy a wrong unreachability claim trades a real
  behaviour for a clean report. Baseline 207 -> 210.

`lint:ci` exits 0. 173 tests pass across the two touched files.

* fix(#3829): a reused finding id no longer inherits the old finding's decision

Found by this round's own adversarial review, which drove it rather than
reasoned about it — and it refuted the arm I had named as my strongest
suspicion, so it is recorded as a correction, not a discovery.

Finding ids are reused across re-reviews: the --auto loop renumbers. `row()`
inherited a prior decision on an id MATCH ALONE, with nothing checking it was
the same finding. Driven: a prior `CR-01 fixed` row against a review reporting a
brand-new CR-01 rendered the NEW finding `fixed`. A false decision in the
artifact whose entire purpose is telling triaged from forgotten — the same
failure mode round 4's blocker was, reached by the other door.

I had argued this was closed by the stale-report arm. It is not: that arm guards
the FIX-REPORT path only. The PRIOR-LEDGER path had no title check at all.

- The ledger now records each finding's title, in the FRONTMATTER rather than a
  fifth table column: the Source cell is the field a human hand-edits and the one
  that must escape pipes, and a second free-text column doubles that surface for
  no reader benefit.
- A prior decision is inherited only when the recorded title still matches. An
  ABSENT prior title inherits, deliberately — a ledger written before titles were
  recorded carries none, and refusing there would reset every decision in it,
  which is the loss this guard exists to prevent, caused by the guard.
- A decision whose id has been reused is PRESERVED under a `superseded:` key
  rather than dropped. The review's driven refutation was precisely that the
  mismatch was surfaced while the decision was lost. It cannot keep a row — the
  id is taken, and two rows under one id is an ambiguity, not a record — so it is
  carried in the frontmatter, re-emitted every run, deduped by id+title, and
  named on the console.
- And an iteration-derived decision now cites the report it actually came from.
  The Source cell hard-coded the unsuffixed `<NN>-REVIEW-FIX.md`, so a decision
  read out of an iteration backup cited a file that may not exist. A citation the
  reader cannot follow is worse than none. Also the review's finding.

Five new tests. Four fail against the pre-fix step; the fifth — that a ledger
with no recorded title still inherits — is a BACK-COMPAT guard and passes both
ways by construction. It is not a reversion control and is not counted as one.

* fix(#3829): follow the cleanup move through, and stop miscalling a converged run

Three loose ends the earlier cleanup relocation left, two of them found by the
round's own review and one by the suite.

**#3190's own test still pinned the old placement.** T6 asserted the `.iterN.md`
removal lives inside `auto_iteration_loop` — exactly what moving it broke. Its
SEMANTICS are unchanged and still asserted: removed on convergence, retained on
degradation, creation intact. What it now pins additionally is the ordering that
forced the move — the ledger reads the backups BEFORE they are removed — and that
the loop no longer removes what it just wrote. Rewritten rather than deleted: the
assertion was superseded, the guarantee was not.

**`CONVERGED` had become a decoy.** With the removal gone from the loop, the flag
was set in two places and read in none. Deleted, and the prose that still said
"the loop sets it" rewritten to what is true: the loop breaks on exactly one
condition, a clean re-review, which leaves REVIEW.md at `status: clean` — and that
is what `cleanup_iteration_backups` re-derives from.

**A converged final iteration reported the opposite of what happened.** The
post-loop message keyed on the iteration COUNTER alone, so a run that converged ON
iteration 3 exited with `ITERATION == MAX_ITERATIONS` and printed "Reached maximum
iterations. Remaining issues documented in REVIEW-FIX.md" over a run in which
every finding was fixed. Convergence is re-derived from the review the loop left
behind — the same signal the cleanup step reads, so the two cannot disagree.

* docs(#3829): retract two claims this round made and could not support

Both were caught by the round's own adversarial review, both were driven, and
both would have reached the maintainer. Recording the retraction where the claim
was made, rather than only in a PR comment.

**The env-prefix "latent bug" does not exist.** An earlier commit in this round
claimed that `FIX_REPORT_FILE="${_pd}/${PADDED}-REVIEW-FIX.md"`, sitting in the
same `node -e` env-prefix list that sets `PADDED`, expanded the OUTER variable
rather than the one two entries earlier — reading ShellCheck's SC2097/SC2098 as
a defect report. Driven in bash and in dash: assignments in one prefix list take
effect left to right, and the later entry DOES see the earlier one. The warning
is a false positive here. The split is kept, but for readability only; the
comment no longer describes it as a fix.

**The HAS_BASH decline rested on a rule that does not govern these call sites.**
It cited `local/no-unguarded-nonportable-exec` as REQUIRING the
`process.platform !== 'win32'` guard. Checked, and wrong on both halves: the rule
fires only on a file that also chmods an exec bit with an octal literal, and this
file has none — so it never runs here — while
`eslint-rules/lib/platform-guard.cjs` accepts four guard shapes plus
`os.platform()`, not one. A constraint that exists is not a constraint that
applies, and I did not check which.

The decline stands on narrower and honest grounds: whether these fences PASS on
the Windows lane is UNVERIFIED. What evidence there is points at divergence
rather than absence — the rule's subject line is that `bash -c` constructs "fail
on Windows Git Bash", and this PR already measured `mkfifo` existing on that
runner, exiting 0, and creating no FIFO. So a probe would not be a clean win; it
would light 22 tests on a lane whose shell semantics are known to differ and
unknown in detail. That is a measurement to make deliberately, not a change to
make in passing. The gap is real and is now stated as a gap.

* fix(#3829): close four defects the review drove out of the first title fix

The round's own adversarial review re-ran against the reworked tree and refuted
two more claims. Every item below is its finding, verified before acting.

**An iteration-only decision recorded no title, so the reuse guard leaked.**
`applied` stored `{d, src}` and the row took its title from the current review —
which does not report the finding at all. The row shipped with no title, and the
next review reusing that id hit the title-ABSENT back-compat exception and
inherited the old `fixed`. The exact defect the title machinery exists to close,
surviving through the hole opened for legacy ledgers. `applied` now carries the
title it was decided under.

**A changed decision was dropped in favour of the obsolete one.** The dedupe was
a has()-guard, so re-superseding a finding whose decision had since changed left
the older record standing. It now replaces.

**Re-spaced titles double-recorded.** The dedupe keyed on the raw title while
`sameTitle()` collapses whitespace; the key now agrees with the comparison.

**And the frontmatter was not valid YAML.** `title: Parser: loses data` is
rejected outright by a real reader, and the `superseded:` line format was not
YAML at all. Values are emitted as JSON scalars — YAML 1.2 is a JSON superset —
and superseded records are properly nested. Round-tripped through js-yaml in the
tests.

One more, self-inflicted while fixing the above: the parse registered each
carried superseded record TWICE, once at `- id:` under an empty-title key and
again at `title:`. Records doubled on every run. They are collected during the
walk and registered once, complete.

**T6 was vacuous.** The review flipped `= "clean"` to `!=` in the cleanup and the
rewritten T6 still passed — it greps for `FINAL_STATUS`, `rm` and "retained"
occurring somewhere, never wiring them to a branch. T6b now EXECUTES the fence in
both directions against real files. It fails on that exact mutation.

**And a converged final iteration printed two success messages** — the loop's
break already reported it. This branch now stays silent and exists only to
withhold the degradation warning.

Three CI gates the base range brought in, all tripped by this round's own text:

- `/gsd-code-review` in a comment — runtime workflow artifacts take the colon
  form. Now `/gsd:code-review`.
- The preamble-ordering parity test: my PHASE_DIR comment wrote the literal
  `gsd_run` before the shim preamble. Reworded.
- Prompt-stuffing: the file passed 50K. I trimmed 5.8K of my own commentary
  first; even removing every added comment leaves the added CODE over the line,
  and the file entered this round at 44,523 — 89% of the budget. Added to
  SIZE_ONLY_WORKFLOWS with the same reasoning the two existing entries carry, and
  the same acknowledgement: splitting is the real fix.

* test(#3829): extract the cleanup fence without an ad-hoc markdown regex

T6b's helper used `/```bash\n([\s\S]*?)\n```/`, which trips two of the repo's own
rules: `local/no-adhoc-markdown-parsing` (use the sectionizer, not a hand-rolled
fence regex) and `local/no-crlf-fragile-split` (a bare `\n` against readFileSync
content is wrong under Windows autocrlf).

Line-scanned now, CRLF-normalized first — the same shape `bashFences()` in
tests/code-review-pipeline-regression.test.cjs already uses, which solved this
first. `npm run lint` is clean and T6b still fails on the inverted-branch
mutation it exists to catch.

* fix(#3829): withdraw the superseded-decision store; keep the identity guard

Three adversarial passes over this round each found real defects, and passes 2
and 3 were entirely inside the `superseded:` block added in pass 1 — a second
identity scheme, keyed on (id, title), living beside the row store keyed on id.
Pass 3 refuted it on three separate counts: a legacy title that merely looked
like JSON lost its quotes and fabricated a record; a finding that was deferred,
superseded, then returned and fixed left an active row and an obsolete
superseded record standing together, reporting `unchanged` forever; and my own
test for the replacement path never passed the earlier ledger in, so it guarded
nothing.

The construct had no terminal state. It is withdrawn.

**What survives is the safety property.** The ledger records each finding's
title, and a recorded decision is carried forward only while the id still names
the same finding. That is what stops a renumbered `CR-01` inheriting an earlier
`CR-01`'s `fixed` — a false decision in the artifact whose purpose is telling
triaged from forgotten, and the same class as round 4's blocker.

**What is given up, and it is disclosed rather than hidden.** On a detected
reuse the earlier decision loses its row. The drop is reported on the console
naming the id and what had been decided, the previous ledger is committed so the
row remains in git, and docs/features/code-review-pipeline.md states the
limitation.

Two defects from pass 3 are fixed rather than deleted, because they are in the
guard and not the store:

- **Known-empty and NOT-KNOWN were conflated.** `### CR-01:` yields an empty
  title; that is a title. While it emitted no `title:` key it read back as a
  pre-format ledger and inherited across a reused id — the same leak, three
  passes running. Emitted whenever the title is known, empty included; a carried
  row no source knows stays absent, which is the legacy-compatible read.
  Underneath it was a falsy fallback: `(act && act.t) || priorTitle.get(id)`
  discards `''`. Now a typeof check.
- **JSON.parse ran on legacy values.** A pre-format ledger whose bare title was
  written `"quoted"` was parsed and lost its quotes, so the decision stopped
  matching. The frontmatter now declares `titles: json` and the parse is gated on
  it; a ledger without the marker keeps its scalars.

One defect from pass 3 is NOT mine and is not fixed here: a converged run prints
a success message from the loop break AND another from `present_results`. Both
predate this round. My earlier claim that "the duplicate is gone" was true only
of the pair I introduced; the pre-existing pair stands, and widening this round
into `present_results` is not warranted.

188 tests pass. The three new tests fire against the pre-simplification step.
`lint:ci` exits 0. The step file is 55,590 chars, down from a 62,220 peak.

* docs(#3829): stop the ledger promising a preservation it no longer makes

Fourth review pass. No machinery defects this time — both findings are claims in
text this step SHIPS, which is the class this whole stack exists to prevent.

**The rendered ledger still said "Re-running the gate preserves every row and
every disposition."** That was true until the same round gave the step an
intentional drop for a reused finding id, and then it was false in the artifact's
own user-facing footer. It now states what the step does, including the one
exception, where a reader actually meets it.

**And the console asserted "the previous ledger is in git."** Committing the
ledger is gated on `commit_docs`, and a failed commit is swallowed — so under
`commit_docs=false` the overwritten decision may exist nowhere. The note reports
the drop and stops there; asserting a recovery path that may not be there is the
same overclaim in a smaller font.

Two residuals from the same pass are DECLINED and documented rather than fixed,
because both would need the second identity scheme just withdrawn:

- A pre-titles ledger carries no titles, so its decisions inherit on the id
  alone. Refusing there resets every decision in every existing ledger, which is
  the loss the guard exists to prevent.
- Two genuinely distinct findings sharing both an id and a title are
  indistinguishable to an (id, title) key.

The pass also refuted the `titles: json` marker on a ledger written by
`b86ea6065^`, which emitted JSON titles before the marker existed. Declined:
that revision is an intermediate commit on this unpushed branch and has never
been released. The PR's published head writes no titles at all, so a real ledger
is either pre-titles (unmarked, bare — handled) or written by the shipped version
(marked). The unmarked-JSON state cannot reach a user.

Test pinned, and it fails against the pre-correction step.

* docs(#3829): fix four wrong citations and one false size justification

All four came out of a claim-audit of this round's own response comment — an
audit of the text, not the code, which is where the remaining errors were.

- **The caller citation was wrong.** The in-code note said both inputs bind from
  `gsd_run query init.phase-op`. `execute-phase.md:85` uses `init.execute-phase`;
  only `code-review-fix.md:22` uses `init.phase-op`. The substantive point is
  unchanged — both are orchestrator-derived, neither is raw user input — but the
  citation was not checked.
- **A leftover "the prior row is in git."** Removed from the console note last
  commit, left standing in the comment two lines above it.
- **The docs still carried the promise the ledger had just dropped.** The
  rendered footer was corrected; the same sentence in
  `docs/features/code-review-pipeline.md` was not.
- **The SIZE_ONLY_WORKFLOWS justification was false.** It claimed the added CODE
  alone exceeded the threshold. Removing every round-added comment leaves 47,148
  chars against a 50,000 limit, so the file CAN fit — the claim was wrong, and an
  exemption defended on a wrong premise is worse than no exemption.

So the entry is re-justified on what is actually true, and earned first: another
**10,188 chars** of this round's own commentary are cut (62,220 → 52,032, from a
44,523 baseline that was already 89% of the budget). Fitting under is possible
only by stripping essentially all remaining explanation from logic three review
passes found defects in. That is the wrong trade in a file whose house style is
heavy in-fence documentation, and the entry says so rather than implying the
file had no choice.

One measurement corrected while checking: the Windows-lane skip count is **37**,
not the 22 the review cited nor the 26 I first counted. Twenty-two and 26 count
`{ skip: !HAS_BASH }` CALL SITES; a skip on a `describe` cancels its subtests.
Forced the constant false and counted what actually skips.

236 tests pass. `lint:ci` exits 0.

* fix(#3829): the drop report is conditional, and two published claims were not

A fifth adversarial pass, run against the two commits that went out AFTER the
fourth pass and were never reviewed, refuted three claims this round published.

1. The drop is NOT reported unconditionally. `row()` reports only a RECORDED
   decision (`was.d !== 'open'`); a prior row still at `open` is replaced in
   silence. The behaviour is right — `open` records no decision to lose — but
   the shipped ledger legend and BOTH feature docs asserted the report happens
   every time. Text corrected in all three places, which is the same defect
   class this round already corrected once for the preservation promise.

2. The test guarding that console wording was VACUOUS: it ran with no prior
   ledger, so no reuse occurred and its `is in git` assertion could not have
   failed however the console was worded. Driven through a real drop now, with
   the drop asserted as a precondition. A new test covers the `open` arm and
   fails on the pre-fix legend.

3. The HAS_BASH gap is now MEASURED rather than assumed, on native Windows with
   Git Bash 5.2.37 / MINGW64 first on PATH, node v25.2.1:

       HAS_BASH left alone:  179 tests, 127 pass,  0 fail, 52 skipped
       HAS_BASH forced true: 179 tests, 140 pass, 24 fail, 15 skipped

   So 37 of the skips are this guard's, confirming the count the round
   published — and unskipping is NOT a clean win: 24 fail, clustered on
   `bash -c` quoting and spawn failures, exactly the divergence the eslint
   rule's subject line names. The guard stays; it now documents a measured gap.
   The stale "the count is 22" comment is gone.

4. The size-exemption justification was wrong a second time. The overshoot is
   ~2.3K normalized chars, not "essentially all remaining explanation": the
   round's committed peak was 59,246 chars (not 62,220, which was never
   committed), and it entered at 44,466 chars, not 44,523 — both earlier
   figures mixed bytes into a character measurement. Rewritten to the numbers
   the scanner actually produces.

Also: the shipped comment said both callers validate the phase shape without
naming that they validate PADDED_PHASE, not the raw PHASE_NUMBER this step is
handed.

* fix(#3829): renumber this PR's two REQs, which #3661 took while the branch sat

The rebase onto current `next` surfaced a REQ-number collision, not a text
conflict. #3661 landed `REQ-REVIEW-08` (`workflow.code_review_point`) on
`docs/features/code-review-pipeline.md` while this branch also claimed 08 and
09 for severity surfacing and the per-finding disposition. Two different
requirements under one identifier is the kind of thing that reads as correct in
both diffs and is wrong in the merged tree.

Base numbering wins, because it shipped: `REQ-REVIEW-08` stays #3661's. This
PR's two become **REQ-REVIEW-09** (severity surfacing) and **REQ-REVIEW-10**
(per-finding disposition). Swept the whole tree rather than the conflict hunk —
two references sat in files git merged cleanly and never flagged:

- `gsd-core/workflows/code-review-fix.md:450`, the prose stating why
  `record_disposition` is the step's only reachable call site.
- `tests/code-review-pipeline-regression.test.cjs:1782`, the comment on the
  test that pins that call site.

`docs/FEATURES.md` is regenerated from the fragment rather than hand-edited;
`node scripts/gen-features.cjs --check` is green (178 features, 21 groups) and
`lint:generated-sync` exits 0.

Two things stated rather than quietly carried. The `Emitted-Drift-Ack-Growth`
trailer on the round-2 commit still reads `REQ-REVIEW-09` for what is now
REQ-REVIEW-10 — it is a historical acknowledgment of that commit's growth, and
its purpose is unaffected, so it is left rather than rewritten across 52
replayed commits. And `docs/INVENTORY-MANIFEST.json` appeared stale immediately
after the replay, reporting two missing `cli_modules/` entries; that was the
lane's pre-rebase build output, not manifest drift. Rebuilding in the replayed
lane and re-checking shows it in sync and unmodified. Regenerating before the
build would have committed the deletion of two base-added entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): reach the title round-trip with a generator that can break it

Round 6's only finding. The round-5 title tracking introduced a fresh parser
(the `titles: json` / `  - id:` / `    title:` frontmatter walk) and a fresh
bijective contract (`JSON.stringify(oneLine(t))` out, `/^    title: (.*)$/`
plus `JSON.parse` back in), and `tests/code-review-disposition.property.test.cjs`
was untouched since round 4 with no reference to `title` at all. Every heading
the generator built was `'### <id>: finding number <i>'` — never a colon, a
quote, a backslash, or the empty string.

You were right that this is the round-3 shape again, and I would rather
demonstrate that than assert it. Two mutations to the shipped step, each a
plausible edit rather than a contrived one:

  A. render `titles: raw` instead of `titles: json`, so the re-parser never
     JSON.parses and stores the quoted scalar as the title;
  B. `yv = (t) => oneLine(t)` — the bare scalar, no JSON at all.

    mutation A — new generator: FAIL     old generator: pass (3/3)
    mutation B — new generator: FAIL     old generator: pass (3/3)

Both ship past the pre-round suite. The gap was reachable, not theoretical.

What changed:

- `TITLE`, a new arbitrary drawn from the class the render's own comments say
  the escaping is for — `:` (why `yv()` exists), `"` and `\` (what
  stringify/parse must round-trip), the empty string (the known-empty vs
  not-known distinction the render draws explicitly) — plus scalars that MIMIC
  the ledger's own frontmatter grammar (`findings:`, `titles: json`, a nested
  `    title: ` line, `  - id: CR-99`), unicode, surrounding whitespace, and one
  title long enough to outrun a scanner assuming short scalars.
- `FINDINGS` now carries a title per id, so all four properties run the cycle
  over the title contract instead of over a constant. `IDS` keeps the old
  id-only shape it is built from.
- A fourth property asserting the round trip in the two places it is observable:
  the stored scalar must `JSON.parse` back to the trimmed heading title, and a
  hand-recorded decision must survive the next run.

The second half is the one that matters, and its construction is the point.
The decision is made by EDITING THE RENDERED LEDGER IN PLACE, never by writing
a bare row the way the existing properties do. A bare row carries no
frontmatter, so `priorTitle` is empty, `sameFinding()` returns true through its
`!priorTitle.has(id)` back-compat arm, and the title contract is never
consulted — the property would pass over a completely broken round-trip. Both
mutations above go green against the bare-row form. That collapse is why the
property is written this way, and the comment says so in place.

So the assertion is the consequence, not the JSON: a lossy round-trip does not
corrupt a title, it makes `sameFinding()` false and resets a human's `deferred`
to `open` with the reason gone — this PR's own founding failure mode, reached
through the field the round-5 work added.

BOUND, stated rather than quietly omitted: the generator emits no CR or LF. A
`###` heading is one line by definition, so a newline is not an input the
heading parser can be handed; `oneLine()` guards the value's other producers,
not this one.

Two things found while writing it, both corrected here rather than left:

- `runOnce` now returns stdout. The reuse report is a CONSOLE note, not a
  ledger key, so my first draft's `assert.doesNotMatch(ledger, /^reused:/m)`
  was vacuously true forever — a test that cannot fail.
- `expectedTitle` is a TRIM, not a `\s+` collapse. Collapsing is `sameTitle`'s
  COMPARISON rule; `oneLine()` is the STORAGE rule and preserves internal
  whitespace. The collapse form fails on an internal tab against entirely
  correct code, which is how a test gets weakened instead of believed the first
  time it goes red.

The file header claimed "two properties" while three were running; it now
states four, one line each.

239 tests pass across the four pipeline files, 0 skipped. `lint:ci` exits 0
(`lint-workflow-shellcheck`: 203 baseline findings, 0 new).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): the prefix census guard now says four sites, because round 5 added one

Self-found, from re-deriving the round-1 finding-id census this round rather
than carrying the round-1 verdict forward.

The census guard's comment says the prefix set is "written out three times —
the heading matcher, the ledger re-parser, and (by its keys) the severity map".
That was true when it was written. Round 5's title tracking added a fourth
copy: the frontmatter `- id: ((?:CR|BL|WR|IN)-\d+)` matcher that rebuilds
`priorTitle`.

The guard itself did not fall behind, and the reason is worth keeping visible:
`idAlternations()` scans the extracted script by PATTERN rather than walking a
fixed list of sites, so the new alternation was absorbed with no edit. Verified
by running the extractor at this head — three alternations found, one distinct
set, severity map keys `CR,BL,WR` with `IN` on the documented `info` default,
0 domain members not reached.

Only the prose fell behind. Corrected, with the pattern-scan rationale stated
in place so the next reader does not helpfully convert it into the hand-listed
enumeration it deliberately is not — which would be exactly the defect this
guard exists to catch, in the guard.

Census discharge for this round: re-derived at the rebased head over the
extracted shipped script, 3 enumeration sites reached, 0 not reached; the
domain (the prefixes `gsd-code-reviewer.md` can emit, walked across both its
heading template and its prose Label-equivalence paragraph) is unchanged since
round 1 at 4 of 4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): catch a duplicate REQ id in a fragment, since nothing did

Not from your review — this is the test the round owed itself, and I would
rather say why than let it look like scope creep.

The renumber commit earlier in this round has no reversion control without it.
I reverted that fix to check, and the first attempt LOOKED controlled: reverting
only the fragment turned `gen-features --check` red. That is the generated-sync
gate noticing the projection went stale, not anything noticing the collision.
Reverting CONSISTENTLY — fragment plus a regenerated `docs/FEATURES.md` — is
silent:

    gen-features --check   rc=0
    lint:ci                rc=0
    pipeline suite         rc=0

with two `REQ-REVIEW-08` entries standing in one requirement list. Nothing in
the repo reads REQ ids at all, so there was no second place for it to be caught.

The failure this guards is a MERGE, not an edit, which is why review does not
see it: two PRs open at once each append "the next" REQ number to the same list,
and whichever lands second is rebased onto a list that already used it. git
merges them as different lines of one file and reports nothing. Neither PR's
diff shows a collision — each is correct against the tree it was written on.
That is exactly how #3661 and this PR both ended up claiming REQ-REVIEW-08.

Scope, stated because it is the part that could be wrong: the check is WITHIN a
fragment, never across the corpus. Two different features legitimately both
carry `REQ-REVIEW-01..07` — the cross-AI review feature and the code-review
pipeline — so corpus-wide uniqueness would be false on the committed tree and
would have to be weakened the day it first ran. A requirement list belongs to
its feature; that is the scope of the identifier.

It lives in `describe('the committed docs/features/ corpus')` because it is an
invariant over the committed corpus, which is that block's stated job, and it
pins no count — the file's own header rules out counts as shared mutable cells
that every feature PR would have to edit.

Control: green on the committed tree (no fragment carries a duplicate today);
red on the restored collision, naming the file and the id. 85 tests pass in
this file.

Happy to drop this if you would rather the round stayed inside the review's
four corners — but then the renumber ships uncontrolled, and I would rather put
that choice in front of you than make it quietly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): finish the census comment correction, which stopped one line short

Found by this round's own pre-push adversarial review, which refuted the claim
the previous commit made about itself.

`7f019d985` said the census comment correction was complete. It corrected one
site and left two, both in the helper block twelve lines above the test it
belongs to:

- `severityMapKeys`' header still read "The THIRD copy: the severity map's
  keys". With three alternations the map is the FOURTH copy, and has been since
  round 5.
- `idAlternations`' header said "adding a prefix to only two of them is silent",
  written when there were two alternations and never updated to three.

This is the defect the original correction was ABOUT, committed inside the
correction: a fragment of prose carries no supersession marker, so a reader
landing on line 2810 gets the dead count stated as current fact, and the fixed
comment eighty lines down does not reach them. Fixing one surface and leaving
its neighbour is not a partial fix, it is the same fix not done.

The region is now consistent end to end, and both headers say the thing that
actually matters — the scan is by PATTERN, not a fixed list of sites, which is
why round 5's new matcher needed no edit here and why converting it to an
enumeration would reintroduce exactly the drift it guards.

WHILE HERE, a disclosure that was narrower than the truth. `7a6680e8f` said the
`Emitted-Drift-Ack-Growth` trailer still names REQ-REVIEW-09 for what is now
REQ-REVIEW-10, and left it deliberately rather than rewrite 52 replayed
commits. That is right, but it is not the whole set: the message BODIES of
`c94106568` ("wire the disposition ledger into the fix path") and `06282f668`
("migrate the emitted-drift ack") both state "REQ-REVIEW-09 was unreachable in
every shipped path", meaning the disposition requirement, which is now
REQ-REVIEW-10.

Same decision, stated at its real size: three historical references, not one.
They are commit history rather than living documentation — git is the record of
what was believed when — and rewriting the branch to correct a number in a
message would cost every review round its correspondence to the commits it
reviewed. The TREE carries no stale reference; `docs/`, the workflows and the
tests all read REQ-REVIEW-09 for severity surfacing and REQ-REVIEW-10 for the
disposition.

Regression file: 181 tests pass, 0 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): the third stale count, and a disclosure that over-counted itself

Both found by re-running this round's pre-push review after the last fix. It
refuted the commit that claimed the region was consistent — for the second time
in a row — and it was right again.

**The third site.** `:2917` said "And the third copy, which is not an
alternation" and `:2919` said "Without this, both regexes can gain a prefix".
Written when there were two alternations; there are three, so the map is the
fourth copy and it is three regexes that can drift.

Worth saying how it survived two passes, because the mechanism is the point and
it is the same one this PR keeps re-learning. Both earlier passes VERIFIED with
a grep built from the strings I had just fixed — `THIRD copy`, case-sensitive,
plus a handful of phrasings I expected. `the third copy` in lowercase matched
none of them, and `both regexes` was not a phrasing I thought to look for. A
grep returns what you already thought of; that is not a verification of prose,
it is a re-statement of your own assumption. The region is now checked by
reading it end to end, and all four count statements agree: three alternations
(heading matcher, ledger row re-parser, frontmatter `- id:` matcher), with the
severity map as the fourth copy.

**And the disclosure over-counted.** The previous commit widened the historical
REQ-REVIEW-09 references from one to three. Three is wrong. There are TWO
underlying statements:

  - `c94106568`'s message body, and
  - the `Emitted-Drift-Ack-Growth` trailer on `06282f668`.

I counted `06282f668` twice — once as "the trailer" and once as "a body" — when
its only mention IS that trailer (`git show -s --format=%B 06282f668 |
grep -c REQ-REVIEW-09` outside the trailer line: 0). Over-counting is the safe
direction and it is still a wrong number in a message, which is the thing this
round has been correcting all along.

The decision is unchanged: both are commit history rather than living
documentation, and rewriting the branch to fix a number in a message would cost
every review round its correspondence to the commits it reviewed. The TREE
carries no stale reference — 08 is #3661's `workflow.code_review_point`, 09 is
severity surfacing, 10 is the per-finding disposition.

Comment-only in one test file; no assertion, regex or extracted-script
expectation moved. Regression file: 181 tests pass, 0 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* fix(#3829): join the disposition-step dispatch so REQ-LANG-04 inheritance is provable

`lint-response-language-coverage` (#2529, which landed on `next` after this PR was
approved) reported `execute-phase/steps/code-review-disposition.md` as having no
response-language coverage. The step does inherit it: `execute-phase.md` imports
`references/execute-phase-response-language.md` and dispatches the step with
`Read and execute`. The dispatch stub wrapped, leaving the verb at the end of one
line and the path at the start of the next, and `namesFragmentAsEntryPoint` matches
within a single line — so a genuine inheritance was unprovable to the linter.

Rejoining the verb and the path restores it: `namesFragmentAsEntryPoint` goes
false -> true and the lint reports `OK (165 workflows covered)`. Only line breaks
move — the word stream is identical to the previous revision, and the file is
unchanged at 93,390 bytes, so no growth acknowledgment is owed.

This takes the third coverage form the lint documents — inheritance — rather than
the inline directive the CI message names first. Where inheritance is provable the
lint's own comments say a second copy "buys no coverage and adds a sentence that
can drift", and the step file already sits over the prompt-stuffing threshold.

Swept all 76 fragments in the catalog: this is the only one whose parent's previous
line ends with a dispatch verb. The 17 others that are mentioned without a provable
entry point are table-routed or bare prose references carrying no dispatch verb at
all, and correctly hold the pinned inline directive instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JgX6QQmygeZnQqbc3o8RNC

* chore(#3829): regenerate derived artifacts after rebase onto next

The rebase onto current `next` conflicted on the 19 install-tree goldens and
`docs/FEATURES.md`. Those are generated, so the conflicts were resolved
arbitrarily and the generators re-run (`npm run regen:derived`) rather than
hand-merged — a clean textual merge of a generated file attests the merge, never
the content.

Reconciled per artifact against the base's own committed copy rather than against
the pre-regen tree, because the pre-regen tree is the arbitrary resolution:

  - all 19 `tests/fixtures/install-tree/*.json` now differ from
    `upstream/next` by exactly one key,
    `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`;
  - `docs/FEATURES.md` differs by exactly REQ-REVIEW-09/10 and this PR's own
    reference section;
  - `docs/INVENTORY-MANIFEST.json` differs by exactly the same one step file,
    and needed no regeneration to get there.

Nothing the base added was dropped by the arbitrary resolution: the restored
entries (the `gsd-core/agents/` and `gsd-core/commands/gsd/` families, the
compact templates, the `detail/elaboration.md` files, `gsd-secret-read-guard.js`)
are all base-owned and came back through the generator, which is what the
resolve-arbitrarily-then-regenerate discipline is for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* fix(#3829): repair the rebase's conflict resolution in the regression suite

The rebase onto current `next` hit one add/add conflict in this file: #4209's
external-reviewer-evidence describe and this PR's #3829 block were added at the
same insertion point. Resolving it by keeping both sides was correct in
substance and wrong in mechanics — the conflict boundary cuts through two open
blocks that the SHARED trailing `  });\n});` closes, so each side carries +2
unbalanced braces on its own and concatenating them left the file with 683 `{`
against 680 `}`.

`node --check` fails outright, so the whole file deregistered rather than
failing a test — 188 tests silently stopped existing. Rebuilt the region as a
real three-way merge (ancestor a262ad6b6, ours upstream/next, theirs 77ee739c3)
and closed the first side explicitly before the second begins.

Both feature blocks are present exactly once, braces balance 683/683, and the
file runs 188/188 locally. The sibling markdown file resolved the same way is
unaffected and was checked rather than assumed: prose has no block structure to
unbalance, and it differs from the base by 112 added lines with zero removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* test(#3829): inject the unreadable-review failure in a way root cannot bypass

Round 9 finding 1. `runShippedGateCounts({ mode: 0o000 })` does not simulate an
unreadable review under root: root bypasses POSIX read permission bits, so the
fence's `[ -f "$REVIEW_FILE" ] && [ -r "$REVIEW_FILE" ]` guard stays true, the
fixture is read, and the assertion sees a real breakdown where it expects
silence. Reproduced as reported — `node:24-slim`, euid 0:

    not ok 18 - an unreadable REVIEW.md leaves the counts empty and does not abort
    actual: 'Code review: 4 findings — 1 critical, 2 warning, 1 info.\n...'

**The prescribed remedy does not reach this site, so this adapts it rather than
applying it.** Stubbing `fs.readFileSync` to throw EACCES is the right fix where
the read happens in-process; here the read is performed by a spawned `bash`, so
node's `fs` is not on the code path and the stub would change nothing.

What the guard actually has is two legs, and only `-r` is defeated by root:

  - the `-r` leg keeps the mode-bit fixture and declares the lanes it cannot
    bind on (`win32`, `euid 0`), which is exactly what
    tests/plan-review-convergence.test.cjs:2326 does for its own shell-side
    `-r` arm — the repo's existing precedent for this shape;
  - the `-f` leg is new and root-immune: a DIRECTORY at the review path fails
    `-f` for every euid, reaching the same non-reporting arm with the same
    observable. It binds on the bench lane where the first test is skipped.

Skipping the first without adding the second would have traded a false failure
for lost coverage on the only lane that found this.

Reversion control, run as root: reverting this commit fails exactly
`an unreadable REVIEW.md leaves the counts empty and does not abort` and its
enclosing describe `#3861 round 1 — the counts mirror is asserted against the
shipped shell`, with no other change to the failure set. **That answers the
round's open question** — the review flagged the describe as possibly a second
root cause; it is the first one's rollup, and there is no second.

Local (euid 1000): 189/189, both tests run.
Root: 179 pass / 1 skip, the skip naming its reason, the directory test running.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* chore(#3829): refresh the compact-content benchmark baseline for the gate's edit

Self-found in this round; not raised in review. The base range landed #4139's
compact-content benchmark, whose committed baseline records per-workflow token
counts. This PR replaces a 12-line bash block in `execute-phase.md`'s
code_review_gate with a 5-line dispatch paragraph, which moves that workflow's
measured counts by 17 tokens — so the baseline the base just added drifts
against a tree it was measured before.

    DRIFT: split "execute-phase": off 25631 -> 25614 (-17), on 23380 -> 23363 (-17)
    DRIFT: aggregate: off 106923 -> 106906, on 90275 -> 90258

Refreshed with the remedy the script itself names
(`node scripts/benchmark-compact-content.cjs --write`). The regenerated diff
touches only the `execute-phase` entry and the aggregate — every other
workflow's numbers are byte-identical, which is the reconcile this PR's edit
predicts.

Attributed rather than assumed: `tests/benchmark-compact-content.test.cjs` is
27/27 at `upstream/next` with no PR content, and was 26/27 on this head. So the
drift is this PR's, not base noise — and it is invisible to a diff-scoped sweep,
because the PR never touches the baseline file and the base range is what
created it. The sibling `benchmark:compact-content-variants` was checked in the
same pass and reports up to date, so this is the only one of the pair affected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* test(#3829): cover an EMPTY REVIEW.md, a case the body claimed and no test reached

Found by this round's own adversarial audit of the PR body, not by review. The body
has said since round 1 that the suite covers "a REVIEW.md that is missing, empty,
or a directory". Two of those three were true. The empty one was not.

Every `reviewText: ''` call in this file also passes `writeReview: false`, which
makes the file MISSING, not empty — so the arm the body named had no test at all.
They are genuinely different paths through the shipped fence: a missing file never
gets past `[ -f ]`, while an empty one passes both `[ -f ]` and `[ -r ]` and is
actually opened and read.

Probed the shipped fence directly against a real empty file before asserting
anything: exit 0, empty stdout. So the behaviour was already correct and only the
coverage claim was false — which is the same "documented as covered, not covered"
shape this PR exists to make visible in the review gate, found in its own body.

The explanatory comment is deliberately precise about WHY the scan yields nothing,
because the plausible reading is wrong and a later reader would inherit it: it is
not the `NR==1{if($0!="---") exit}` guard. A zero-byte file gives awk no record, so
that action never runs (NR stays 0); the output is empty because `closed` is never
set. The comment also states what the test does not prove on its own — its
observable is identical to the missing-file case, so "the file was read" rests on
the harness and the fence, not on the assertions.

189 -> 190 tests in this file, all passing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* chore(#3829): regenerate derived artifacts after rebase onto next

The rebase onto current `next` conflicted in the 19 install-tree goldens,
`docs/INVENTORY-MANIFEST.json` and the compact-content benchmark baseline.
Those are generated, so they were resolved arbitrarily and regenerated
with their own producers rather than hand-merged: every golden now differs
from `next`'s committed copy by exactly the one PR-owned entry
(`gsd-core/workflows/execute-phase/steps/code-review-disposition.md`), the
manifest by the same entry, and the benchmark baseline by the
`execute-phase` split plus the aggregate.

* chore(#3829): regenerate the platform-conformance tier for the added property test

`next` gained the conformance-tier classifier (#4591) and its CI gate after
this branch was cut. The branch adds `tests/code-review-disposition.property.test.cjs`,
so the generated tier list was one file short (546 != 547). Regenerated
with `node scripts/gen-platform-conformance-tier.cjs --write`; the macOS
tier (`--target macos --check`) already matched.

* chore(#3829): regenerate the two conflicted generated artifacts after the rebase

`next` moved 6 commits past the previous base and conflicted in exactly two
files, both generated:

- `tests/fixtures/compact-content-benchmark-baseline.json` — #4208 (`4cc2a466b`)
  and #4619 (`db4d8a9ba`) both moved the measured token counts, and this branch
  moves the `execute-phase` split too.
- `scripts/lib/platform-conformance-tier.generated.cjs` — #4641 (`4d65c248e`)
  made test-conformance the sole Windows selector and narrowed the tier to
  28.5%, and #4568/#4619 re-ran it after.

Both were resolved arbitrarily during the replay and then regenerated with
their own producers rather than hand-merged, per the generated-artifact rule:
`node scripts/benchmark-compact-content.cjs --write` and
`npm run regen:derived` (which runs `gen-platform-conformance-tier.cjs
--write` for both the default and the macOS target).

Reconciled against `next`'s own committed copies rather than the pre-regen
tree:

- the benchmark baseline differs from `next` by exactly the `execute-phase`
  split (`offTokens` 25827 -> 25810, `onTokens` 23576 -> 23559 — the 17-token
  delta this PR's step-file extraction has carried since round 5) plus the
  `aggregate` that sums it;
- the conformance tier differs from `next` by exactly one added entry,
  `tests/code-review-fix-pipeline-regression.test.cjs`. Under the narrowed
  28.5% selector that is the file the classifier now picks from this PR's
  test set. The arbitrary resolution had carried 282 stale lines computed
  under the pre-#4641 selector (`--numstat` on this commit: 3 insertions,
  282 deletions), and regeneration collapsed them.

The full derived sweep was run, not just the two named producers: all 19
install-tree goldens, `docs/INVENTORY-MANIFEST.json`, `docs/FEATURES.md`,
the macOS conformance tier and the exit-code registries regenerated
byte-identical, so nothing else drifted under the new base.

* fix(#3829): accept N-segment phase ids, and bound length per component

The base range added `scanMarkdownSingleSegmentPhaseRegex` (#4568,
`a2331c01f`), which refuses the single-optional-segment phase regex on
phase-carrying markdown lines under three roots — `gsd-core/workflows/`,
`gsd-core/references/` and `agents/` (`lint-phase-id-drift.cjs:301`,
`:373-387`). It flagged two lines in this step file. Chasing the flag
turned up two real defects behind it, so this commit is those rather than
the comment edit the flag literally asked for.

## Defect 1 — the step refused ids both its callers accept

The comments asserted that both callers validate `^[0-9]+(\.[0-9]+)?$`.
#4568 had widened those two call sites to `^[0-9]+(\.[0-9]+)*$`, so the
prose was stale. Correcting only the prose would have shipped a comment
promising N-segment support over code that refused it, because the step
carried a third `case` arm:

    *.*.*)        _ok=0 ;;   # more than one dot: not the documented shape

`23.1.2` took the refusal arm, `PADDED` came back empty, and the step
printed `Code review reporting skipped (unusable phase number ...)` and
wrote **no ledger** — for a phase id both of its callers accept.

It degraded loudly rather than silently; there is a diagnostic on stdout.
The traversal-fence test asserted that refusal as *correct*, listing
`1.2.3` among the values that must be rejected, so an arity bound and a
shape bound sat folded into one `case` arm with a test pinning the pair.

The arity arm is gone. Deleting it alone would have left the step
**wider** than its callers in one direction — `1..2` has an empty
segment, which `^[0-9]+(\.[0-9]+)*$` refuses and the retired arm had been
masking — so a third arm replaces it:

    *..*)         _ok=0 ;;   # EMPTY SEGMENT

## Defect 2 — the length bound was not per-component, though its comment said so

Removing the arity arm made a second defect reachable. The bound read:

    case "$_pn" in *.*) case "${_pn#*.}" in ?????????*) _ok=0 ;; esac ;; esac

`${_pn#*.}` is the whole tail after the first dot — one component only
while an id has at most two. With N-segment ids accepted, that form
rejects `1.1234567.1`, whose every component is a legal 7 digits, purely
because the tail measures 9 characters. The comment directly above it has
read **"LENGTH-BOUND EACH COMPONENT SEPARATELY"** since before this PR,
and had itself named the composite bound as "too strict" — the same
mistake, surviving one level up.

Both fences now walk the segments and bound each:

    _rest="$_pn"
    while [ -n "$_rest" ]; do
      case "$_rest" in
        *.*) _seg="${_rest%%.*}"; _rest="${_rest#*.}" ;;
        *)   _seg="$_rest";       _rest="" ;;
      esac
      case "$_seg" in ?????????*) _ok=0 ;; esac
    done

The `$((10#...))` overflow guard is preserved, per component: bash
integers wrap at 2^64, so an unbounded integer segment would silently
become a negative padded phase.

**The remaining divergence from the callers is a CLASS, not a list:** any
id carrying a component of nine or more characters is caller-accepted and
fence-refused — `123456789`, `1.999999999`, `1.123456789.1`,
`123456789.1`, `1.1.123456789` and so on. That narrowing is deliberate
and is the overflow guard. An earlier draft named two examples as though
they were exhaustive; that wording is withdrawn.

## What is NOT claimed

- The canonical grammar is `PHASE_NUMBER_TOKEN_SOURCE` in
  `src/phase-id.cts:65`, `\d+[A-Z]?(?:\.\d+)*`, added by **#2128**
  (`09be501eb`, 2026-07-10). An earlier draft dated it to #865
  (2026-06-08); that was the first commit to touch the *file*, not the
  one that added the constant, and it is withdrawn.
- #4568 gave the six shell sites **segment-count** parity with that
  grammar, not textual parity: the canonical source permits an optional
  `[A-Z]`, and the shell literals remain digit-only. Driven: this step
  and both callers all refuse `23A.1`, so they agree with each other and
  are jointly narrower than `src/phase-id.cts`. That is a question about
  the six sites rather than about this step, and it is not touched here.
- **#4619 does not produce N-segment ids.** It only transforms an
  already-supplied `{phase_number}` so `$((10#...))` does not abort on
  one. An earlier draft cited it as the producer; that is withdrawn, and
  is stated rather than silently swapped so a reader can see it was
  corrected.
- That the folded `case` arm is *why* nothing caught this is an
  observation about the test's shape, not an established cause.

## Tests

- `an N-SEGMENT phase number reports counts, exactly as its callers
  accept it` — drives `23.1.2` **and** `1.2.3.4`: the retired guard was
  arity-shaped, so a bound merely moved from two dots to three would pass
  a three-segment-only test.
- `the length bound is PER COMPONENT, not over the whole tail after the
  first dot` — drives an 8-char and a 9-char **middle** segment, the
  position the old form got wrong.
- `the fence agrees with its callers across a probed set spanning both
  boundaries` — example-based, and says so: a finite probe cannot prove
  congruence over an infinite language, and one review pass demonstrated
  that by injecting a `2) _ok=0` arm this test still passed. It is a
  regression pin over the values that actually broke.
- The traversal list loses `1.2.3` (legal at this base) and gains `1..2`
  and `1.2.` — the malformed-dot cases the arity guard had masked.

**Negative controls, re-measured against reconstructed fences:**

    fence state              N-seg   probed   per-comp
    fully pre-fix            FAIL    FAIL     FAIL
    shape-fix only           PASS    FAIL     FAIL
    this tree                PASS    PASS     PASS

Two tests, not one, catch Defect 2: the caller-agreement probe includes
`1.1234567.1`, so the whole-tail bound breaks it too. An earlier draft
claimed the per-component test failed alone — that table was written
before `1.1234567.1` was added to the probe and was not re-measured
afterwards. It is corrected here from a fresh run. The N-segment test
correctly does not fire on Defect 2; it predates the bound work and is
insensitive to it.

Found by this round's own adversarial review passes.

* docs(#3829): name this step's two dispatchers correctly, in code as well as in the PR body

`code-review.md` is not a call site of this step. It carries an identical
`^[0-9]+(\.[0-9]+)*$` validator, which is why it kept getting cited as one, but
it never dispatches `code-review-disposition.md`. The two dispatchers are
`execute-phase.md` (`code_review_gate`) and `code-review-fix.md`
(`record_disposition`) -- and only the second validates anything.

This round corrected that in the PR body and simultaneously wrote the old
conflation into the shipped comments, so the file asserted at line 24 what the
body denied in public, and contradicted its own line 48. Four false assertions,
each duplicated because the fenced block is emitted twice:

- "Both callers explicitly accept ... (code-review.md:63, code-review-fix.md:39)"
- "#4568 widened both of this step's callers" -- it widened the one that
  validates; the other has no validator to widen
- "Both callers already validate ... (code-review.md:63, code-review-fix.md:39)"
- "a SHAPE (..., asserted by both callers)" -- asserted by one

The same conflation had propagated into three comments in
`code-review-pipeline-regression.test.cjs`; corrected there too.

Adds the one fact that follows from naming the dispatchers correctly and that
nothing else in the tree records: `execute-phase.md` applies NO shape gate, so
this fence is not mirroring an upstream guarantee -- it IS the guarantee. A
later reader who believes the caller validates will "simplify" it away.

Comments only. The executable shell is byte-identical to 03adf9474 (verified by
stripping comment lines and diffing). 200/200 regression + property, 47/47
prompt-injection security scan, eslint and lint:generated-sync clean.

The wider [A-Z]-axis divergence between these sites and the canonical
`PHASE_NUMBER_TOKEN_SOURCE` is tracked separately as #4660 and deliberately not
restated here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xdw628PpLYveWfJu7kDyvZ

* chore(#3829): regenerate the two conflicted generated artifacts after the rebase onto next

Both conflicted during the replay onto `eb49ff98d` and were resolved arbitrarily, then
regenerated with their own producers (`npm run regen:derived`,
`node scripts/benchmark-compact-content.cjs --write`) rather than hand-merged. Reconciled
against next's committed copies: the conformance tier differs by the one entry this PR
adds, the benchmark baseline by the `execute-phase` split (the same 17-token delta this
PR's step-file extraction has carried since round 5) plus the aggregate that sums it. The
rest of the derived sweep — 19 install-tree goldens, INVENTORY-MANIFEST, FEATURES, the
macOS tier, the exit-code registries — regenerated byte-identical.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* fix(#3829): a carried row keeps the severity the ledger recorded, instead of re-inferring it from the prefix

The ledger always wrote a severity for every row (table cell and frontmatter key) and
nothing read either back: the prior-row regex skipped the cell as `[^|]*`, the frontmatter
walk collected only titles, and a carried row was rebuilt through sev() from the id prefix,
because sectionSev holds only the findings the CURRENT review reports. So a WR-04 the
reviewer filed under `## Critical Issues` was recorded critical, a human deferred it, and the
next run -- the review no longer reporting it -- silently re-recorded it warning. The one
artifact whose purpose is remembering a finding's severity lost it on the second run, in the
unsafe direction (round 11, reproduced by executing the shipped script twice).

Both persisted copies are now read back, enum-validated (ADR-227, as the disposition column
already is): the table cell first, the frontmatter `severity:` as the fallback for a
hand-mangled cell. Severity precedence is the current review's SECTION, then the RECORDED
value, then the id PREFIX, and the recorded value is inherited only while the id still names
the same finding -- the identity rule the disposition already obeys -- so a reused id starts
from its own review. sev() moves below sameFinding() because it now depends on it.

Tests: a new describe drives the reviewer's exact case (WR-04 under `## Critical Issues`,
deferred by hand, dropped by the next review -> stays critical) plus five controls: recorded
outranks prefix under no recognized section; the current section still outranks recorded; a
REUSED id does not inherit; a mangled cell falls back to the frontmatter and a mangled pair
to the prefix; a bare pre-severity row still infers. A fast-check property assigns each
finding a section independent of its prefix, carries every row through an empty review, and
asserts the section severity survives and the third run reports unchanged.

Negative control, measured against the pre-fix step: the two carry tests, the mangled-cell
test and the property fail; the three precedence/back-compat controls pass at both ends, as
they pin behaviour that predates the fix. Every prior carried-row test used CR-01/IN-01,
whose prefix already matched, so the lossy path had returned the right answer by coincidence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* fix(#3829): a malformed REVIEW.md is reported as unparsed, not passed over as clean

A REVIEW.md with three criticals and an unterminated frontmatter yielded REVIEW_STATUS='',
and the counting arm then printed nothing -- byte-identical to a clean review. The guard
that scopes the frontmatter scan was right to yield no values from an unterminated block; the
reporting arm was wrong to treat 'no status' as 'no review'. Block 2 said `status: none`
rather than `clean`, which is why a careful reader could still separate them (round 11,
Minor).

Both fences now record whether the file was actually READ, separately from what it yielded.
A read file with no parseable status -- unterminated frontmatter, no frontmatter, no
`status:` key, a zero-byte file -- prints `Code review status unparsed: ...` with no
breakdown (there is none to trust) and no --fix suggestion (nothing proves there are
findings). Absent, directory and unreadable stay silent: nothing was read, so nothing is
described. Block 2's skip line names the same distinction, `status: unparsed` vs `none`.

The counts mirror follows the shell: a mirror is always handed a text, so its empty-status
arm is the unparsed one, and the existing 'unterminated frontmatter' and 'no frontmatter at
all' parity fixtures now bind the new message on both sides. The EMPTY-file test from round 9
changes its assertion deliberately: its observable is no longer identical to the missing-file
case, which is the point. Five new tests drive the arm, its three shapes, the three shapes
that stay silent, and block 2's wording.

Negative control, against the previous step: the unterminated, no-status, no-frontmatter and
empty-file tests fail, both parity fixtures fail (the mirror moved and the shell had not),
and block 2's `unparsed` assertion fails; the stays-silent controls pass at both ends.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* docs(#3829): name the unlocked read-modify-write and reference #3780 rather than solving it

The ledger is rendered whole from a prior read with nothing serializing two writers, and
this step has two dispatchers plus an invited hand-edit, so the window is real. It is the
shape #3780 reported for WINDOWS.md under parallel executors, which #4681 closed with a
cross-process lock in src/broken-windows.cts. Not taken here, deliberately: the step is a
shell-embedded script with no dependency on the compiled tree, and adopting the lock module
is its own change. Stated at the write site and as a residual in the feature doc; no lost
update has been reproduced (round 11, Minor).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* docs(#3829): cross-reference the two "review disposition" ledgers in both directions

ADR-3806 canonizes a `## Review Dispositions Ledger` section inside PLAN.md for reviews-mode
planning: append-only per round, over REVIEWS.md findings. This PR's
`<NN>-REVIEW-DISPOSITION.md` is a sibling file beside REVIEW.md for the code-review pipeline,
rewritten idempotently with rows carried. Adjacent names, opposite durability rules, and
neither document mentioned the other -- the round-11 review checked the ADR gate against
3806, cleared it, and flagged exactly that mis-read hazard.

An in-place dated amendment section on ADR-3806 (contributor-standards "Amending an accepted
ADR", pattern 1) and a paragraph in the pipeline feature doc, each naming the other and the
axis on which they differ. docs/FEATURES.md regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* chore(#3829): regenerate the compact-content benchmark baseline after the rebase onto next

`next` moved three commits while the round was in flight and the baseline conflicted again;
resolved arbitrarily during the replay and regenerated with its own producer. It differs from
next's copy by the `execute-phase` split this PR has carried since round 5, plus the aggregate.
The rest of the derived sweep regenerated byte-identical.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* fix(#3829): accept letter-variant phase ids, matching the canonical grammar #4744 widened the dispatchers to

#4744 (#4660) landed on `next` while this round was in flight: it widened the six
shell/markdown phase-number mirrors -- `code-review-fix.md:39`, this step's validating
dispatcher, among them -- to the canonical grammar's letter axis (`12A`, `3A`, `23A.1.2`),
and added a `lint-phase-id-drift` ratchet that flags any digit-only mirror left in the
workflow tree. Rebased onto that base, this step was the one it flagged (two fences, two
sites): `12A` was refused by name and wrote no ledger, for a phase id its own dispatcher
now accepts -- the round-10 class ("the step refused phase ids its validating dispatcher
accepts") re-opened by the base. Found by running the base range's modified gates against
the rebased tree, not by the review.

Both fences now admit an uppercase letter in the character class and pin WHERE it may sit --
only as the last character of the integer part, at most once -- so `23a`, `A23`, `2A3`,
`23AB` and `23.1A` stay refused. The per-component length bound is on the DIGITS (the letter
is one character the `$((10#...))` overflow guard has no stake in, so `12345678A` is within
it exactly as `12345678` is), and the letter is carried verbatim after the padded digits,
`3A` -> `03A`, as `src/phase-id.cts` pads it. The two fences stay line-identical except for
their refusal message (the parity test holds), and every comment literal of the old shape
reads the canonical one.

Tests: a new fixture drives `12A`, `3A`, `23A.1.2` and `12345678A` through the shipped
fence to the padded path; the traversal-fence list gains the five wrong placements; the
caller-agreement probe's regex gains the letter axis with both-direction cases, and
`123456789A` joins the deliberate over-bound narrowing. Negative control against the
pre-widening step: the letter-variant test, the caller-agreement probe and the base's
`scanMarkdownLetterlessPhaseMirror` gate all fail; the traversal-fence test passes at both
ends (the five new placements were already refused, by the narrower class).

Also re-anchors the fence and test comments' `code-review-fix.md` / `code-review.md` citations by
content (the validator, not a line number): the line numbers had drifted by one against the rebased
base, and drift again on every rebase.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* chore(#3829): regenerate the two conflicted generated artifacts after the rebase onto next

Both files conflicted during the replay and were resolved arbitrarily rather than
hand-merged, then regenerated with their own producers — `npm run regen:derived`
and `node scripts/benchmark-compact-content.cjs --write`.

Reconciled against next's own committed copies rather than against the pre-regen
tree, because a clean textual merge of a pinned-number file attests the merge and
not the numbers:

- `tests/fixtures/compact-content-benchmark-baseline.json` differs from next by
  exactly the `execute-phase` split (offTokens 26264 -> 26200, onTokens 24013 ->
  23949) and the `aggregate` that sums it. That is the step-file extraction this
  PR has carried since round 5, re-measured against the new base; no other entry
  moved.
- `scripts/lib/platform-conformance-tier.generated.cjs` differs from next by
  exactly one added entry, `tests/code-review-fix-pipeline-regression.test.cjs` —
  the file the classifier picks from this PR's test set.

The full derived sweep was run, not just the two named producers: all 19
install-tree goldens, docs/FEATURES.md, docs/INVENTORY-MANIFEST.json, the macOS
tier and the exit-code registries regenerate byte-identical under the new base,
so nothing else drifted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): stop block 2 computing an unparsed shortfall from a self-contradicting findings block

Round 12, Minor. Confirmed, and the premise is slightly stronger than stated: block
2 does not merely skip block 1's `critical + warning + info == total` cross-check —
it never derived the three severity counts at all, so the check's inputs were
absent. It bounded `total` for digits and length only and handed it to the
`unparsed:` reconciliation.

So a REVIEW.md whose `findings:` block disagrees with itself (`total: 10` beside
`critical: 1, warning: 1, info: 1`) made block 1 print the countless form —
breakdown suppressed as untrustworthy — while block 2 still computed a shortfall
from that same untrusted number. Two trust models for one field, one fence apart,
with the weaker one downstream. It fails in the safe direction, which is why the
review did not raise it as a blocker; it is still a real inconsistency.

Block 2 now derives `critical`/`blocker`, `warning` and `info` through the same
`findings:`-anchored filter block 1 uses, with the same digit-and-length bound and
the same `10#` on every operand, and blanks `total` when the three disagree with it.

**Adapted, not applied verbatim — and the divergence is the point.** The finding
says to re-apply block 1's cross-check. Block 1's gate is `REVIEW_COUNTS_OK`, which
demands all four counts be numeric, because block 1 DISPLAYS all four and
`6 findings —  critical` is the half-filled line that rule exists to prevent. Block
2 displays none of them; it uses `total` alone, against the number of headings the
row parser matched. Applied verbatim, the all-four rule blanks a perfectly usable
`total: 5` on a review carrying no severity keys and SILENTLY DROPS an `unparsed:`
shortfall this step reports correctly today — trading a safe-direction over-report
for a silent under-report, which is the wrong way round and is the exact failure
class the `unparsed:` key was added to close. Only the CONTRADICTION ports: absent
counts are not a disagreement, because there is nothing to disagree with.

Driven against the shipped fence, not a mirror:

  consistent 1+1+1=3          -> total 3       CONTRADICTION total:10 -> withheld
  blocker: alternation        -> total 2       counts absent          -> total 5 (kept)
  leading zeros 01+01+01=03   -> total 03      one count absent       -> total 5 (kept)
  no findings block           -> withheld      non-numeric count      -> total 5 (kept)

Six regression tests drive the second markdown fence end to end through `runHook`
under bash, asserting on the rendered ledger. Negative controls fire in OPPOSITE
directions, which is what pins the narrowing rather than only the fix:

- revert the fence fix          -> `a contradicting findings block yields no
                                   shortfall` and the `blocker:` twin go red
- apply the VERBATIM all-four   -> `a total with NO severity keys still reconciles`
  prescription instead             and the partial/non-numeric case go red

Restored tree: 214/214.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* docs(#3829): state the one input that produces no unparsed shortfall

The reference page said the shortfall "is stated" whenever `total:` exceeds the
parsed headings. After the round-12 fix that is conditional, and a doc asserting
the unconditional form describes behaviour the step no longer has.

Names the boundary in both directions, because the narrowing is the part a reader
would otherwise get wrong: a `findings:` block whose three severities are all
present, numeric and do not sum to `total` produces no key — the same input on
which the console line already withholds the breakdown — while counts that are
merely absent, partial or non-numeric are not a disagreement and still reconcile
from `total` alone.

`docs/FEATURES.md` regenerated; `gen-features --check` green (182 features, 21
groups).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): close two holes the round's own adversarial pass found in its first attempt

Neither is from the maintainer's review. Both were found by the pre-push adversarial
pass over this round's own claims, which refuted them by execution.

**1. A malformed severity could suppress a real shortfall.** The sibling frontmatter
reads use `cut -d: -f2 | tr -d ' '`, and `tr -d` deletes INTERNAL spaces, so
`critical: 1 0` arrives as the perfectly numeric `10`. That is long-standing in those
reads — its mirror is pinned as a fixture from round 1 — and it was INERT in block 2
until this round made that block read the severities at all. At that point a repaired
number could satisfy the new sum test and suppress an `unparsed:` shortfall that is
genuinely owed. Driven, pre-fix: `critical: 1 0 / warning: 0 / info: 0 / total: 5`
against three parsed headings emitted no `unparsed:` key where `unparsed: 2` was
correct.

The three severity reads now trim the ends only, so an internal space survives into
the digit check and fails it — `_sum_ok=0`, nothing is suppressed. Fail-safe in the
only direction that matters: when the frontmatter is malformed the step declines to
suppress rather than trusting a repaired number.

**Scope, stated:** only the SUPPRESSION inputs are strict. `REVIEW_TOTAL`'s own read
still uses `tr -d ' '`, unchanged and identical to block 1's — narrowing it would
change the `unparsed:` computation itself, which is pre-existing behaviour and wider
than this round. So `total: 1 0` is still read as `10` by both blocks, as before.

**2. The repointed #4748 gate could not see a later rebinding.** Its derivation slices
stop AT the first anchored assignment, so inserting the canonical lookup and then
overriding it with `REVIEW_FILE="${_pd}/WRONG-REVIEW.md"` left every assertion green —
the slice pins a line, not the path the fence actually consumes. A new test pins the
whole file instead: the only `REVIEW_FILE=` bindings permitted are the canonical
lookup (exactly twice, once per fence, each being a fresh shell) and the identity
pass-through that hands it to the embedded node script as an env prefix.

Negative controls, both the adversarial pass's own mutations, against the restored
tree at 215/215 and 164/164:
- restore `tr -d ' '` on the severity reads -> the internal-space test reds
- insert the WRONG-REVIEW override after the lookup -> the rebinding test reds

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): one parser for every count block 2 reads, and widen the rebinding guard

A second adversarial pass, run against the first pass's own fixes, refuted three of
them by execution. Fixes to review findings are the class most likely to carry a new
defect, which is why that pass exists; all three were real.

**1. `cut -d: -f2` takes the SECOND FIELD, not the scalar.** So `critical: 1: junk`
arrived as the perfectly numeric `1`, and `1+0+0 != 5` was read as a contradiction
that SUPPRESSED a shortfall genuinely owed. The previous fix trimmed the ends but
still cut at the wrong place, so it closed the internal-space shape and left this one
open. `-f2-` keeps everything after the first colon; the malformed scalar stays
malformed and `total: 5` still reconciles.

**2. The deliberate asymmetry was wrong, and it FABRICATED.** The previous fix parsed
the severities strictly and left `total` lenient, on the reasoning that narrowing
`total` was out of scope. Driven: `critical: 5 0` with `total: 1 0` repaired only the
total to `10`, rejected the severity, skipped the contradiction check, and invented
`unparsed: 7` against three parsed headings. Both uniform policies behave sanely —
strict rejects the malformed total, lenient detects `50 != 10`. A field is either
trustworthy or it is not; parsing one leniently and its sibling strictly is the shape
that fabricates. Every count this block reads now goes through one parser.

**Scope, restated because it moved:** the previous commit said `REVIEW_TOTAL`'s read
was deliberately unchanged. That is no longer true and the reasoning behind it did not
survive contact — the asymmetry it protected is what produced the fabrication. Block
1's reads are still untouched; its own all-four gate runs over consistently-parsed
values, so it has no equivalent split.

**3. The rebinding guard missed an indented or exported assignment.** `^REVIEW_FILE=`
let both `  REVIEW_FILE=...` and `export REVIEW_FILE=...` through, and each executes
exactly like a bare one. The predicate now absorbs leading whitespace and an optional
`export` before the accept-list decides.

Driven after the fix, against three parsed headings:

  critical: '1: junk'  total: 5      -> total 5 kept, unparsed: 2 reported
  critical: '5 0'      total: '1 0'  -> total rejected, no unparsed key
  critical: '1 0'      total: 5      -> total 5 kept (unchanged)
  consistent / contradiction / blocker / leading-zero / absent — all unchanged

Negative controls, each the adversarial pass's own mutation, against 381/381:
- `-f2-` back to `-f2`            -> the second-colon test reds
- `total` back to lenient `tr -d` -> the fabricated-shortfall test reds
- an INDENTED rebinding           -> the rebinding guard reds
- an `export` rebinding           -> the rebinding guard reds

Residual, disclosed: a duplicate `critical:`/`blocker:` key is still resolved by
`grep -m1` taking the first match. Duplicate keys are invalid YAML and the same
first-match rule is long-standing in the sibling reads; detecting them is a wider
change than this round.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): one parser for the whole step, and prove a contradiction from a partial sum

A third adversarial pass, run against the second pass's fixes. Three findings, plus
one this round's own negative control caught afterwards.

**1. An absent severity still bounds the sum from below.** Counts are non-negative,
so a missing one can only ADD: when the severities that ARE present already sum to
MORE than `total`, the block disagrees with itself whatever the absent value is.
Requiring all three before comparing missed that — driven: `critical: 4`,
`warning: 4`, no `info:`, `total: 5` reconciled against a total the present counts had
already refuted. The comparison is two-armed now: EQUALITY when all three are known, a
LOWER BOUND when they are not. An UNDERshoot stays reconcilable, because that is
exactly what the absent count explains.

**2. Block 1 now uses the same parser, so the console and the ledger cannot
contradict each other.** Tightening block 2 first left the two fences disagreeing about
the same bytes. Driven: `critical: 1 0` with `total: 1 0` repairs to 10 and 10, which
SUM — so block 1 reported `10 findings — 10 critical, 0 warning, 0 info.` from a
`findings:` block containing no such numbers, while the ledger recorded three rows and
no shortfall. Block 1's reads move to `cut -d: -f2-` plus an end-trim; both fences now
take the countless arm on that input. The counts mirror moves with them — its whole job
is modelling the shipped pipeline, and it modelled the retired one.

This is wider than the review's finding and I want that visible: the finding was about
block 2 alone. But a disclosed divergence between a console line and a ledger is the
confusion this PR exists to remove, so it is fixed rather than documented.

**3. `REVIEW_FILE+=-wrong` executes and was missed.** The rebinding guard matched only
`=`; `+=` appends (driven: `REVIEW_FILE=good; REVIEW_FILE+=-wrong` prints `good-wrong`).

**4. My first negative control for (2) was VACUOUS, and that is the reason for the new
`BLOCK 1 withholds a breakdown built from REPAIRED counts` test.** Reverting block 1's
parser left the suite green: on every fixture that existed both parsers landed on the
same arm, so parity could not see the difference. A SELF-CONSISTENT repaired breakdown
separates them, and the test pins it directly rather than through parity.

**Correction to the previous commit's claim.** It said moving `total` to a strict read
"changes nothing for a well-formed review". That is false: `total:\t5\t` is valid YAML
(`yaml.parse` returns 5) which the old `tr -d ' '` rejected and the new trim accepts.
The change is an improvement, not a no-op, and the claim was the wrong shape.

Also from the third pass's MISSED: the fixtures exercised malformed `critical` and
`total` only, so they did not pin the four-field symmetry the fix claims. Every field
now gets every malformed shape.

Negative controls, against the restored tree at 383/383 (220 in the pipeline file):
- revert block 1's parser        -> the repaired-counts test reds (was vacuous; now fires)
- revert the overshoot arm       -> the absent-severity contradiction test reds
- a `+=` rebinding               -> the rebinding guard reds

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): finish the mirror update, and give status the same parser as the counts

A fourth adversarial pass. The important finding is that the PREVIOUS commit's mirror
update was HALF APPLIED, and the suite could not see it.

**1. The counts mirror was still on the retired parser.** That file carries TWO
helpers: `first`, used for `status:`, and `firstIn`, used for ALL FOUR counts. The
previous commit updated `first` and the comment above it, and left `firstIn` on
`split(':')[1].replace(/ /g,'')` — so the shipped block had moved to `-f2-` + end-trim
and the mirror had not, while the parity assertion stayed green.

It stayed green for the same reason this round's earlier negative control was vacuous:
on every fixture that existed, both parsers reach the COUNTLESS arm, so the rendered
message is identical and parity cannot see the divergence. Two fixtures now separate
them — `a self-consistent repaired breakdown` (10 == 10+0+0, so the retired parser
renders a full breakdown from a `findings:` block containing no such numbers) and
`tab-separated counts` (valid YAML the retired `tr -d ' '` made non-numeric). Reverting
`firstIn` reds both.

This is the same shape this PR's round-3 reply already recorded about itself: a fix
verified with a grep built from the strings just fixed. The region is checked by
reading it end to end now.

**2. `status:` kept the retired parser after the counts moved off it, and it is the
read where truncation costs most.** `cut -d: -f2` turned the valid YAML scalar
`status: clean:junk` into the bare `clean`, so an unusable status took the CLEAN arm
and suppressed BOTH the console report and the ledger. Driven, both parsers side by
side. The whole scalar matches no arm now, so the step reports. One parser for every
scalar this step reads, in both fences.

**Two claim corrections, no code change:**

- The previous commit implied block 1's console output was preserved for every
  well-formed review. It is not: `critical:\t1` is valid YAML that the retired
  `tr -d ' '` left non-numeric (countless form) and the trim now reads (full
  breakdown). That is an improvement, and the claim was the wrong shape. The
  `tab-separated counts` fixture pins it.
- "Block 1 and block 2 can no longer contradict each other about counts" was
  overstated. It is true of the PARSER, which is what changed. They can still differ
  when the body carries MORE findings than `total:` declares: the reconciliation
  reports a shortfall only, and the excess direction is deliberately clamped so a
  review under-declaring its own total cannot render `unparsed: -1` — pinned by the
  round-2 test `a total SMALLER than the rows is not reported as a negative shortfall`.
  That is pre-existing and out of this round's scope; stating it rather than widening
  scope again.

Negative controls, against the restored tree at 223/223 (387 across both files):
- revert the mirror's `firstIn`  -> both new parity fixtures red
- revert the `status:` parser    -> the clean-arm suppression test reds

`npm run lint:ci` exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): pin the reads to LC_ALL=C, so the parser cannot depend on the machine

A fifth adversarial pass. It refuted the claim that the shipped reads and their JS
mirror are equivalent, and the counterexample is a locale.

**The POSIX character classes are locale-defined, and glibc's C.UTF-8 disagrees with
both C and en_US.UTF-8.** Driven, same sed, same input, three locales:

    LC_ALL=C          clean<U+2003>  ->  clean<U+2003>   (kept)
    LC_ALL=C.UTF-8    clean<U+2003>  ->  clean           (trimmed)
    LC_ALL=en_US.UTF-8 clean<U+2003> ->  clean<U+2003>   (kept)

C.UTF-8 classifies U+2003 — and U+1680, U+2000-U+200A, U+205F, U+3000 — as BOTH
[[:space:]] and [[:blank:]]. So `status: clean<U+2003>` trimmed to the bare `clean`,
took the CLEAN arm, and silently suppressed both the console report and the ledger —
but only on machines whose locale said so. That is the same suppression the previous
commit fixed for `clean:junk`, with a machine-dependent trigger instead of a parse one.

`[[:blank:]]` is NOT the fix — it is locale-defined too, and C.UTF-8 puts U+2003 in it
as well. Nor is `[ \t]`: POSIX bracket expressions provide no escape at all, so `\t` there is a
backslash and a `t`. GNU sed's default reading of it as TAB is an extension, and the
same GNU sed asked for conformance shows the other reading on this host:

    sed -E          's/[ \t]+$//'  draft  ->  draft
    sed --posix -E  's/[ \t]+$//'  draft  ->  draf

A POSIX-conforming sed is therefore expected to truncate `status: draft` to `draf`.
That expectation is derived from POSIX plus the `--posix` demonstration above; it was
NOT driven against a BSD/macOS sed, because this host has none. The portable
fix is to pin the locale: under C the class is exactly {space, tab, NL, VT, FF, CR},
which is precisely what the mirror already spells out literally. The two now agree by
construction rather than by coincidence of the machine.

All 20 read sites (10 `grep`, 10 `sed`, both fences) are pinned. The mirror's four
anchors move from JS `\s` to the same literal class, closing the divergence in the
other direction — `\s` matches a U+2003 indent that the pinned `grep` does not.

**A second gap, found by this round's own control rather than by the reviewer.** The
mirror has two helpers, and the previous commit proved `firstIn` (the counts) was
pinned by a fixture. `first` (the status) was NOT: reverting it left the suite green.
Every pre-existing status fixture left both parsers on the SAME arm — `issues:found`
truncates to `issues`, which is no more `clean` than `issues:found` is — so the status
mirror could drift unseen, exactly as `firstIn` had. The fixture that separates them
is one where truncation FLIPS the arm: `status: clean:junk`.

That is the third time this round a mirror edit was invisible to the fixtures that
existed, and the question that finds it every time is: what input actually separates
the two versions?

**Two prose corrections in the step**, which had gone stale rather than wrong-headed:
the block-2 comment still said "the sibling reads use `tr -d`" after they had all been
moved off it, and the mirror's class comment claimed an equivalence it did not yet have.

Negative controls, each driven against the committed tree:
- drop LC_ALL=C from the shipped seds  -> 5 red, incl. both locale-invariance tests
- revert the `first` status mirror     -> `a status whose truncation would flip the arm` reds
- revert the mirror anchors to `\s`    -> `locale-invariant on a unicode-space indented count key` reds

The locale-invariance tests are the durable guard: the parity fixtures only run under
whatever locale the suite inherits, so they can catch this only on a machine that
already has the bug. These drive the same input under both locales and assert the
shipped fence does not care.

238/238 across the two pipeline files. `npm run lint:ci` exits 0, with
lint-workflow-shellcheck reporting 212 pre-existing findings and 0 new.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): pin the awk selectors too, and assert the invariant instead of claiming it

A sixth adversarial pass, and it refuted the previous commit's central claim. That commit
pinned all 10 `grep` and all 10 `sed` reads to LC_ALL=C and then said the parser no longer
depends on the machine. It does: the `findings:` MAPPING SELECTOR is an `awk`, and both
copies of it were left unpinned.

`awk '/^findings:[[:space:]]*$/{f=1; next} f&&/^[^[:space:]]/{exit} f'` resolves its two
character classes through the ambient locale exactly as grep's and sed's did. Driven, on
`findings:<U+2003>`:

    fence 1, LC_ALL=C        Code review found issues.
    fence 1, LC_ALL=C.UTF-8  Code review: 1 findings — 1 critical, 0 warning, 0 info.
    fence 2, LC_ALL=C        TOTAL=''
    fence 2, LC_ALL=C.UTF-8  TOTAL=1

Under C.UTF-8 the opener matched and the mapping opened; under C it did not. The same
review rendered a breakdown on one machine and the countless message on another, with
every grep and sed already pinned.

**Why it was missed is the more useful part.** The previous commit's census counted
`grep` and `sed` sites and reported zero unpinned — because it SEARCHED FOR THE TOOLS IT
HAD JUST EDITED rather than for the tools that were there. That is the same shape as this
round's other three misses: a check built from the thing just changed cannot see what the
change forgot. So this commit does not just add the fourth and fifth pins; it replaces the
claim with an assertion the next edit cannot fool:

  `every locale-sensitive tool in the step is pinned to LC_ALL=C` walks the step file and
  fails on ANY unpinned `grep`/`sed`/`awk`, naming line and call. `cut -d: -f2-` and
  `tr -d '\r'` stay exempt, and the exemption is principled rather than residual: neither
  resolves a character class or a collation — one splits on a single ASCII byte, the other
  deletes one literal byte.

The mirror's block boundary moves to the same literal classes, for the same reason the
anchors did last commit — `/^findings:\s*$/` and `/^\S/` model neither pinned side.

**A correction to the previous commit's message, made in place.** It asserted that BSD sed
reads `[ \t]` as a literal backslash and `t`, stated as driven fact. It was not driven —
this host has no BSD sed. The claim is now stated as what it is: POSIX bracket expressions
provide no escape, GNU's TAB reading is an extension, and GNU sed asked for conformance
demonstrates the other reading here (`sed --posix -E 's/[ \t]+$//'` turns `draft` into
`draf`). The conclusion is unchanged; the evidence class was overstated.

Also measured while establishing that LC_ALL=C is safe for non-ASCII, and worth recording
because it makes the pin a strict improvement rather than a wash: on a REVIEW.md carrying a
single invalid UTF-8 byte, GNU grep under C.UTF-8 reports `binary file matches` and emits
nothing, blanking EVERY read; under C the value parses and is rejected on its merits. Valid
UTF-8 is untouched either way — the C space class is entirely bytes < 0x80, which no UTF-8
multibyte sequence contains, so the trim cannot split a character.

Negative controls, each driven and restored:
- unpin the four `awk` selectors  -> 3 red, incl. the new invariant test naming both lines
- revert the mirror block boundary to `\s`/`\S` -> both `findings:` opener tests red

241/241 across the two pipeline files. `npm run lint:ci` exits 0 — after it caught a real
defect in the new test itself: `split('\n')` on readFileSync content is banned here
(DEFECT.WINDOWS-CRLF-TEST-PORTABILITY), and it now uses `splitLines()`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): make the pin invariant see calls that are not piped

The invariant added in the previous commit keyed on `| grep|sed|awk`, which is true of
every call in the step today and is exactly the wrong thing to rely on. A guard written
around the shapes that happen to exist cannot see the shape a later edit introduces —
`awk '...' < "$f"` or `$(grep ...)` would have walked straight past it, which is the same
property that let the two awk selectors sit unpinned through a commit claiming the parser
was locale-independent.

It now blanks the PINNED calls and treats anything still naming one of the three tools on
a non-comment line as an offender, so the check is "every call is pinned" rather than
"every piped call is pinned". Comment lines stay exempt: the step's prose names unpinned
forms while explaining why they were retired.

Driven both ways against the invariant alone:
- inject a NON-PIPED unpinned `awk '...' < "$REVIEW_FILE"` -> reds (the old form did not)
- unpin the four piped `awk` selectors                     -> still reds (no regression)
- unmodified tree                                          -> green

241/241 across the two pipeline files, `npm run lint:ci` exits 0. No shipped behaviour
changes in this commit; it only widens what the test can see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): close the invariant's path-qualified hole, and state the limit it keeps

A seventh adversarial pass. It confirmed the shipped fix — census clean at 24 pinned calls,
whole-fence output identical under C and C.UTF-8 on every input built to separate them, and
both fences byte-identical on a well-formed review carrying `café 東京` and `naïve résumé` —
and then refuted the claim I made about the GUARD, not the feature.

`/usr/bin/awk '/[[:space:]]/{exit}' < "$REVIEW_FILE"` passed the invariant. The preceding-
character class shielded any match preceded by `/` or `.`, so a path-qualified call was
invisible. A path-qualified call is still a call; the class no longer shields either. The
substrings that motivated the exclusion are unaffected — `parsed`, `passed` and `awkward`
have a word character on one side or the other, so the boundary still rejects them.

**The rest of that finding is disclosed rather than fixed, deliberately.** `$AWK "$f"` and a
command name computed inside the embedded `node -e` block also evade the check, and they are
not closable by this mechanism: it scans text, not shell or JavaScript command structure. The
same pass that found them showed that widening the regex further only trades those false
negatives for false positives on quoted strings and awk program text. So the test now STATES
its boundary instead of implying it has none — the previous comment claimed a census "a future
edit cannot fool", which was exactly the kind of overclaim this round has been correcting.
What it catches is enumerated there, driven, along with which direction its false answers go.

Driven against the invariant alone, each mutation reporting its own substitution count:
- `/usr/bin/awk ... < "$f"` (the exact evasion) -> reds; before this commit it did NOT
- `env awk ... < "$f"`                          -> reds
- `LC_ALL=C.UTF-8 awk ... < "$f"` (wrong pin)   -> reds
- unmodified tree                               -> green

241/241 across the two pipeline files, `npm run lint:ci` exits 0. No shipped behaviour changes
in this commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): stop the guard's comment promising a closed list of what it misses

An eighth adversarial pass. It confirmed the delta was test-only and that both fences are
byte-identical across it — exit, stdout, stderr and rendered ledger — and then refuted the
comment again, with two more evasions: a command name fragmented in shell (`a''wk`), and an
executable command substitution on a physical line starting with `#` inside a multiline
quoted argument, which the comment exemption skips.

Both are real. Neither is the point. Three passes running have each found one more evasion of
a TEXT scan, which is the actual finding: **the list cannot be closed.** A comment that
enumerates residuals is false the moment someone is cleverer than the enumeration, and fixing
it by appending the newest example just resets the clock.

So the comment no longer claims an inventory. It says what this is — a regression guard against
the accident that has now happened twice in this round, a read added or edited without its pin
in a file where every other read has one — and what it is not: a proof. The examples are marked
as illustrations. The operative instruction is the one that survives any future evasion: treat
anything it reports as real, and never treat its silence as proof a new read is pinned.

No logic changed; the guard catches exactly what it caught before. 241/241 across the two
pipeline files.

`npm run lint:ci` exits 0 — after it caught this commit twice over, which is worth recording
because both were in prose I had just written to be careful:
- the previous message's example path `docs/grep.md` read as a genuine docs reference from this
  file, and lint-docs-guard-registration demanded a baseline entry for a path that does not
  exist. The illustration is now `bin/grep-wrapper`, and the reason is stated inline.
- an earlier commit's `split('\n')` on readFileSync content tripped the CRLF-portability rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): the guard's two directions are not symmetric, so stop saying they are

A ninth adversarial pass. It confirmed the delta before it was comment-only with the guard's
logic byte-identical, and that all six advertised forms are still caught — and then found the
one absolute the rewrite left behind.

The comment said "treat anything this guard reports as real" three lines above admitting the
guard produces loud false positives. Driven: `echo ok # grep is discussed` is reported, and
labelled `(unpinned)`, though it is prose and no unpinned read exists. Both sentences were
mine, in the same comment, written in the same edit that was supposed to remove overclaiming.

The instruction is now the accurate one, which is that the two directions are NOT symmetric:
a REPORT is cheap to adjudicate — read the line, a trailing comment or a path is obvious — while
SILENCE proves nothing, because the known evasions are silent and so is any evasion nobody has
thought of yet. Investigate every report; never read silence as proof a new read is pinned.

Comment-only. Guard logic untouched, `gsd-core/` byte-identical to cac648aab — three consecutive
passes have now confirmed the shipped behaviour unchanged. 241/241 across the two pipeline files,
`npm run lint:ci` exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): a report needs its context to adjudicate, not just its own line

A tenth adversarial pass, and the last one this round runs. It confirmed `gsd-core/` is the
SAME TREE OBJECT as at cac648aab (105b5ed9b) with the guard's caught and silent sets unchanged,
and refuted one more sentence of the same comment.

"A report is cheap to adjudicate (read the line)" is false, driven: the identical reported
physical line `  grep` is a COMMAND after `:` and an ARGUMENT after `printf '%s\n' \`. The guard
reports both, and the reported line alone does not distinguish them — the preceding line is what
settles it. The comment now says so, with that counterexample in it.

This is the fourth consecutive pass to find a defect in this one comment and none in the shipped
code, which is itself the result worth recording: the shipped fix has been frozen since cac648aab
and confirmed byte-identical by four passes, while the prose describing a best-effort text scanner
took four attempts to stop overclaiming. Writing an accurate description of what a heuristic does
NOT do turns out to be harder than the heuristic.

Comment-only; guard logic untouched. 241/241 across the two pipeline files, `npm run lint:ci`
exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): re-home the lookup's phase-id gate beside the step, not in a block upstream removed

#4781 (#4628) removed #4748's letter-axis work from `tests/nsegment-phase-grammar.test.cjs`,
including the `#4748 — the REVIEW.md lookup` describe block. This PR had four tests living in
that block, because that is where the gate was when #3829 moved the lookup out of
`execute-phase.md` and into the lazily-read step file.

Those four tests assert properties of THIS PR's step file, not of #4748's sites. Rebasing onto
the removal would have deleted them silently — the branch would still be green, with its own
coverage quietly gone. They move here instead, unchanged in substance, beside the step they
guard: an unrelated upstream revert can no longer take this PR's coverage with it.

One assertion did NOT come along. The old block also checked that `execute-phase.md`'s init
parse list names `padded_phase`; #4781 removed that field from the list, and the assertion is a
property of #4748's site rather than of this step. Carrying it here would only have pinned
someone else's revert to this PR.

The step never depended on that field in the first place — it computes PADDED itself, validating
PHASE_NUMBER for shape and traversal and padding the digit run through `10#` while carrying an
optional letter verbatim. That self-containment is why the removal costs this PR nothing but the
tests' address.

5 pass, 0 fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* chore(#3829): regenerate the derived artifacts the round-12 rebase invalidated

The base moved from 092d9256b to 003d982c8 (six merges) while this round was in flight, so the
branch was rebased and the derived state had to be re-derived rather than hand-merged.

`platform-conformance-tier.generated.cjs` regains `tests/configured-entrypoint-validation.test.cjs`,
added upstream by #4249. Resolving the conflict hunk-by-hunk in favour of this branch had dropped
that entry; `lint:generated-sync` caught it, which is what that check is for.

`compact-content-benchmark-baseline.json` carries the measured values at the new base rather than
this branch's stale pair: split "execute-phase" off 26200 -> 26242, on 23949 -> 23952
(8.59% -> 8.73%); aggregate off 108243 -> 108285, on 91595 -> 91598 (15.38% -> 15.41%). The
benchmark reports drift and exits 0 either way, so a stale baseline does not announce itself here
— it announces itself in CI.

**The growth acknowledgment is back, and the reason is worth stating.** Against the previous base
this PR left `execute-phase.md` a net -103 bytes: the extraction removed more than the dispatch
paragraph added, so the file ended up smaller than the base's copy and the
`Emitted-Drift-Ack-Growth` trailer became false and was dropped. #4781 then rewrote that file
upstream, and against the new base the same extraction nets +36 bytes (93421 -> 93457) — which is
the figure this PR originally reported at round 2. The size delta was never a property of this
change alone; it is a property of this change against whichever base it sits on, and it has now
been both signs in one round. The trailer was restored then. Round 14 rebased onto `029acd915`, where #4830's
re-land moved `execute-phase.md` again and the same extraction is -103 once more
(93564 -> 93461), so the trailer is false a second time and this commit no longer
carries it. Third sign flip, same reason each time.

`lint:generated-sync` and `benchmark-compact-content --check` both clean afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* chore(#3829): regenerate derived artifacts after rebase onto next

The rebase onto `c9a5cc3e1` conflicted in the 19 install-tree goldens.
Those are generated, so they were resolved arbitrarily and regenerated
with their own producers (`npm run regen:derived`, then
`benchmark-compact-content.cjs --write`) rather than hand-merged — a clean
textual merge of a generated file attests the merge, never the content.

Reconciled per artifact against the base's own committed copy rather than
against the pre-regen tree, because the pre-regen tree is the arbitrary
resolution:

- every install-tree golden now differs from `c9a5cc3e1`'s copy by exactly
  one entry, `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`,
  which is this PR's own new step file;
- the compact-content benchmark baseline by the `execute-phase` entry
  (offTokens 26184 -> 26167) and the aggregate that sums it.

`docs/INVENTORY-MANIFEST.json` was in the at-risk set but regenerated
byte-identical, so it carries no change here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RJPv9LNaPV3Cd3NCpfFKGb

* chore(#3829): regenerate derived artifacts after rebase onto 029acd915

The rebase onto current `next` conflicted in two generated files — the
compact-content benchmark baseline and the platform conformance tier. Both
were resolved arbitrarily and regenerated with their own producers
(`npm run regen:derived`, then `benchmark-compact-content.cjs --write`)
rather than hand-merged: a clean textual merge of a generated file attests
the merge, never the content.

Reconciled per artifact against the base's own committed copy rather than
against the pre-regen tree, because the pre-regen tree is the arbitrary
resolution. Every differing key belongs to a file this PR actually touches:

- `tests/fixtures/install-tree/claude.json` differs from `029acd915`'s copy
  by exactly one entry,
  `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`, this
  PR's own new step file;
- `scripts/lib/platform-conformance-tier.generated.cjs` by exactly one
  entry, `tests/code-review-fix-pipeline-regression.test.cjs`, a test this
  PR adds;
- `tests/fixtures/compact-content-benchmark-baseline.json` by the
  `execute-phase` entry (offTokens 26344 -> 26280) and the aggregate that
  sums it, which this PR moves by editing `execute-phase.md`.

The other 18 install-tree goldens, `docs/FEATURES.md` and
`docs/INVENTORY-MANIFEST.json` were regenerated too and came back
byte-identical, so they carry no change here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* test(#3829): follow #4748's REVIEW.md-lookup gate to the step that now owns the lookup

Self-found while rebasing onto `029acd915`, not raised in review.

#4830 (`8a5166598c`) re-landed #4768's letter-suffix work on `next`, restoring the
`#4748 — execute-phase.md resolves the REVIEW.md path from init's padded_phase`
block in `tests/nsegment-phase-grammar.test.cjs`. That block had been removed by
#4781, which is why round 13 re-homed this PR's own four tests out of it. The
restored block anchors on a line this PR deletes:

    expected 1 line(s) containing "REVIEW_FILE=\"${PHASE_DIR}/${PADDED}-REVIEW.md\"", found 0

It fails at the describe level, so all four of its tests go with it. It was green
before this rebase only because the base did not carry the block yet.

Putting the line back is not available. #3829 moved the lookup into the lazily-read
step file because `execute-phase.md` did not fit under ADR-857's frozen pre-phase-6
ceiling (93600); the parent is at 93461, and the three lines this gate anchors on
cost 183 (measured, not computed: `git show 029acd915:... | sed -n '1168,1170p' | wc -c`). The block would also be dead code — the step performs the lookup.

So the gate follows the lookup. Two of its four assertions are properties of the
lookup and are re-pointed at the step file: that no fence hands PHASE_NUMBER to
`printf "%02d"`, and that both lookups are preceded by a PADDED binding that pads
the digit run through `10#` and carries the letter verbatim. The regression control
on the lookup line itself comes along, now over both fences. The init-parse-list
assertion stays on `execute-phase.md`, which still names `padded_phase`.

Two assertions do NOT come along, and they are the two that were properties of the
INLINE site rather than of the lookup: the `PADDED="{padded_phase}"` literal binding
(the step derives PADDED itself, validating PHASE_NUMBER for shape and traversal
first), and the composition run over the three live lines. The step's executable
coverage — a composition run plus a padding-agrees-with-the-canonical-normalizer
matrix over letter ids — already exists in `tests/code-review-pipeline-regression.test.cjs`
under "#3829 — the step's REVIEW.md lookup resolves a letter-suffixed phase without
a shell re-pad". Mirroring it here would be a second implementation of one grammar.

That coverage is NOT equivalent, and the difference is worth stating rather than glossing. At its
original site #4748's gate was a DATAFLOW pin: `execute-phase.md` bound init's own
`{padded_phase}`, so the lookup could not disagree with the canonical normalizer because it never
computed anything. The step reconstructs the value in shell, so that pin is not available at this
address and AGREEMENT with the normalizer is what replaces it. A follow-on commit adds a
fast-check property asserting that agreement over generated ids, because the existing 13-shape
matrix samples 2 of 26 letters and cannot see a divergence outside its own points.

One consequence is disclosed rather than absorbed: `padded_phase` is now parsed but unused in
`execute-phase.md` (`:95`). Removing it from that parse list is #4830's call on its own site, not
this PR's, so the assertion that it is still named stays.

#4748's property is unchanged: a letter-suffixed phase resolves its own REVIEW.md,
and an already-padded `08` does not read as octal.

Negative-controlled rather than asserted. Against the shipped step: 162 pass, 0 fail.
Dropping `$_let` from both PADDED bindings reds "every lookup is preceded by a PADDED
binding that carries the letter run"; dropping `10#` reds it too. The step file was
restored byte-identical after each control.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* test(#3829): assert the step's padding against the canonical normalizer by property, not by 13 points

From this round's own pre-push adversarial review, not from the maintainer's.

The prior commit re-points #4748's REVIEW.md-lookup gate at the step that now owns the lookup. The
review's finding was that this is not coverage-equivalent, and it is right: at the original site the
gate was a DATAFLOW pin — `execute-phase.md` bound init's own `{padded_phase}`, so the lookup could
not disagree with the canonical normalizer because it never computed anything. The step reconstructs
the value in shell, so agreement with the normalizer is what has to replace the pin.

That agreement was already asserted, but over a 13-shape matrix. Its words: "future canonical
grammar changes could therefore diverge without this gate detecting them." Correct — the matrix
samples 2 of 26 letters and a bounded set of segment shapes, and this PR has twice been told that a
generator which cannot reach the interesting input is a fixture with extra steps (round 3's
SOURCE_CELL, round 5's title generator). Same defect, third address.

So the agreement is now a property over generated ids: a digit run inside the step's own 8-digit
bound, an optional single A-Z, and up to two dot segments, asserted equal to
`normalizePhaseName(id)` through the shipped shell derivation of BOTH fences. Milestone `N-N` forms
are outside the step's accepted domain and are asserted nowhere here rather than silently passed.

Negative-controlled on two mutations, and the second is the one that justifies the property rather
than the matrix:

  %02d -> %03d                      matrix RED,   property RED
  drop the letter when outside {A,B} matrix GREEN, property RED

The second is the added coverage, demonstrated rather than argued: the matrix is structurally
unable to reach a letter it does not enumerate. The step file was restored byte-identical after
each control.

numRuns is 25 — each case spawns bash twice through the process seam, once per fence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* fix(#3829): pad the phase id as a string, so the step agrees with the canonical normalizer

Found by the property the previous commit added, on its second adversarial pass. This is a real
divergence in shipped behaviour, not a test-only correction.

`normalizePhaseName` (src/init.cts, via gsd-core/bin/lib/phase-id.cjs) left-pads a phase id's digit
run to a MINIMUM of two and otherwise PRESERVES it — `padStart(2, '0')`. The step re-derived the same
value arithmetically, `printf "%02d" "$((10#$_dig))"`, which does not preserve: it collapses every
leading-zero run longer than two.

  id          normalizePhaseName   the step (before)
  8           08                   08
  08          08                   08
  008         008                  08     <-- diverges
  0008        0008                 08     <-- diverges
  00000008    00000008             08     <-- diverges
  0008A       0008A                08A    <-- diverges

Consequence: for such an id the gate resolves `08-REVIEW.md` while init emits `008-REVIEW.md`, so it
finds no review and says so — advisory, and therefore silent. That is the failure class #4748 exists
to close, reached by a different road: not a letter this time, but a leading-zero run.

The fix is a string pad that implements `padStart(2, '0')` exactly, at both fences:

  case "${#_dig}" in 1) PADDED="0${_dig}${_let}${_sub}" ;; *) PADDED="${_dig}${_let}${_sub}" ;; esac

It is strictly less machinery than what it replaces. `10#` existed only to stop bash reading a
leading zero as octal inside `$(( ))`; with no arithmetic there is no octal hazard to guard, so the
remedy is retired rather than kept. `scripts/lint-phase-id-drift.cjs` — which exists to flag
unsanctioned `printf "%02d"` re-pads of phase-carrying variables — is green, and now has one less
re-pad to tolerate.

Two dependent assertions move with it, and both are now stated as the PROPERTY rather than as one
spelling of the remedy: the round-14 gate in `tests/nsegment-phase-grammar.test.cjs` asserts the
binding does no arithmetic and carries both `${_dig}` and `${_let}`, and the regression pin in
`tests/code-review-pipeline-regression.test.cjs` follows the new form.

The property's generator is widened in the same commit, because its first cut could not have found
this: it built the digit run with `String(fc.integer(...))`, which can never produce a leading zero,
so it had silently LOST the `08`/`09` coverage the 13-shape matrix beside it already had. The run is
now generated as a digit string, and segment depth goes to four (the repo exercises `1.2.3.4`). The
step's grammar is unbounded in depth; four is a stated bound, and it is this property's residual.

Negative-controlled, and the matrix is the control's control — it stays GREEN on both:

  revert to the arithmetic pad          matrix GREEN, property RED
  drop the letter when outside {A,B}    matrix GREEN, property RED

The step file was restored byte-identical after each. 407 pass / 0 fail across both test files;
lint:ci, gen-features --check, gen-platform-conformance-tier --check, benchmark-compact-content
--check and gen-install-tree-fixtures all clean, with no regenerated artifact moving.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* test(#3829): hold #4748's gate by execution, and retire the comments the string pad made false

Third adversarial pass on this round. Two findings, both fair, neither a correctness defect.

FIRST — the gate pinned a SPELLING, not the property. Its objection was concrete: an equivalent
multi-line string pad would have failed a regex that matches one `case ... esac` line. That is a
false-positive generator, and a guard that false-fires is a guard that gets deleted.

The split is now honest about what each layer can hold. A STATIC gate can hold the DEFECT SHAPE —
no arithmetic in the binding above each lookup — and that is all it asserts. Correctness is held by
EXECUTION: the base gate regains a composition test that runs the shipped derivation slice of both
fences against `normalizePhaseName`, over `3A 8 9 08 008 0008A 23A.1.2`.

That composition test is the one I removed two commits ago, and removing it was the weaker call. At
#4748's ORIGINAL site the gate could be static because the property was a literal binding of init's
own `{padded_phase}`; nothing could disagree, because nothing computed. At this site the step
derives the value, so the property is behavioural and only execution holds it. `008` is in the list
because it is the case the arithmetic pad got wrong and no prior fixture covered.

Controlled three ways, and the middle one is the finding being answered:

  arithmetic pad restored            4 fail   caught
  EQUIVALENT multi-line string pad   0 fail   no false positive
  drop the letter outside {A,B}      1 fail   semantic drift caught

SECOND — the step carried comments the fix had made false, in four places. Two explained `10#` as
part of the live phase derivation; two justified the eight-digit bound by bash integer overflow.
Neither described the code any more. They are rewritten to current truth rather than annotated,
because a fragment carries no supersession marker and overturned prose reads as canon:

  - the octal rationale now says the pad performs no arithmetic and needs no `10#`, and notes that
    `10#` survives in this step only on the severity COUNTS, which really are numbers being added;
  - the length bound now states that overflow is unreachable since the pad stopped converting, and
    that the bound stays for the reason it always also had — every component is interpolated into a
    filename, and filesystem components are finite.

408 pass / 0 fail across both files. lint:ci, lint-phase-id-drift, gen-features --check,
gen-platform-conformance-tier --check, benchmark-compact-content --check and the
prompt-injection-scan security suite are all clean, and no regenerated artifact moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* docs(#3829): correct four comments, including one this round's own rewrite got wrong

Fourth adversarial pass. Comment-only; no code, no test logic, no regenerated artifact moves.

ONE OF THESE IS MY OWN ERROR, introduced two commits ago. Rewriting the length-bound rationale, I
replaced the dead integer-overflow justification with "the bound stays because every component is
interpolated into a filename and filesystem components are finite". That is false, and it was driven
false: the bound is PER SEGMENT, every segment is joined into ONE filename component, and depth is
unbounded. Thirty 8-digit segments yield a 278-character PADDED and a 294-character name against a
NAME_MAX of 255. Replacing a dead rationale with a wrong one is worse than leaving the dead one, so
the comment now states what the bound actually does and names the composite-length gap as a residual
of this validator that predates the pad change. It is not fixed here; it is stated.

The other three are stale rather than wrong:

- `overflow guard` named the length check in two places. Nothing overflows any more -- the pad does
  no arithmetic -- so it is the digit bound, and is called that.
- The `#4748` block header in `tests/nsegment-phase-grammar.test.cjs` still said the step pads
  through `10#`, still said the composition run did not survive the move, and still said the
  executable coverage was "cited rather than copied" -- while the composition test sat twenty lines
  below it. All three were true when written and none survived this round. The header now records
  why a STATIC assertion cannot hold a BEHAVIOURAL property, and that the PR's fast-check property
  is a different instrument over the same contract rather than the same test twice.
- The severity-count comment said a base-inference failure "takes the whole advisory step down under
  `set -e`". It does not: the arithmetic sits inside an `if` condition, a TESTED context, where
  `set -e` is inert, and the consistency check is SKIPPED instead -- which the regression suite
  already records. `10#` stays; only the account of what it prevents is corrected. This one predates
  the round and is corrected because it is adjacent and factually wrong, not because it blocked
  anything.

408 pass / 0 fail across both files; lint:ci, lint-phase-id-drift, gen-features --check,
benchmark-compact-content --check and the prompt-injection-scan security suite all clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* chore(#3829): regenerate the compact-content benchmark baseline after the rebase onto fac0e9de8

The rebase onto current `next` conflicted on this generated fixture, as it has in every
recent round: the base regenerates it for its own token deltas and this branch regenerates
it for `execute-phase.md`'s, so both sides rewrite the same keys. Resolved arbitrarily
during the replay and regenerated with its own producer
(`scripts/benchmark-compact-content.cjs --write`), never hand-merged.

Key-level drift against the base's committed copy is exactly two entries and both are this
PR's own: `splits.execute-phase` (the workflow this PR edits) and `aggregate`, which is the
sum over the splits and therefore moves whenever any split does. Zero foreign keys moved.

The full generator sweep was re-run after the replay -- gen-features, gen-inventory-manifest,
gen-platform-conformance-tier (both targets), gen-install-tree-fixtures and the benchmark --
and this fixture is the only artifact that moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V5xnzzipRHPHq8Z3doQvn3

* chore(#3829): regenerate the compact-content benchmark baseline after the rebase onto b956bb7c6

`next` moved again while this round was running -- #4902 landed at 06:27Z and touches this same
generated fixture -- so the branch went back to CONFLICTING within minutes of the previous push.
This is the second rebase of the round, not a correction of the first.

Resolved arbitrarily during the replay and regenerated with its own producer
(`scripts/benchmark-compact-content.cjs --write`), never hand-merged. Key-level drift against the
new base's committed copy is again exactly `splits.execute-phase` and `aggregate` (the sum over
splits) -- zero foreign keys.

The full generator sweep was re-run after this replay as well; this fixture is the only artifact
that moved. Build inputs were untouched by the base range, so the lane's existing build stands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V5xnzzipRHPHq8Z3doQvn3

* fix(#3829): refresh the launcher preamble in place, without hoisting it above the guard

Found by this round's own post-rebase validity gate, not by review: a reader flip-test over the 85
tests that read the two at-risk workflow files was clean before the replay and failed after it, on
`runtime-launcher-parity (#373)` invariant (B).

The cause is a real interaction. #4902 landed on `next` mid-round and rewrote the canonical launcher
preamble; invariant (B) counts occurrences of that exact snippet, so this step's older copy matched
zero times even though it sat in the right place. The remedy the invariant names --
`node scripts/sync-runtime-launcher.cjs` -- fixes the count but also HOISTS the preamble to the top
of the block, and that is wrong here: it moved the shim ahead of the status guard, and six of this
PR's own tests exist to pin that ordering (`runDispositionGuard` asserts the block opens with its
guard, then the shim). Running the tool verbatim turned one red into seven.

So the preamble text is refreshed to the current canonical snippet IN PLACE, at the offset it
already occupied. Both constraints hold at once, verified by execution rather than by reading:
invariant (B) sees exactly one canonical occurrence and it precedes the first `gsd_run` call, while
the guard still opens the block (shim at offset 19897 of the second fence, and the test wants > 0).
`runtime-launcher-parity` + `code-review-pipeline-regression` together: 282 tests, 281 pass.

Not fixed here, and not ours: `(K2) end-to-end: the resolved local tool honors
git.allow_default_branch_commits (#4834)` -- the one remaining failure -- fails identically on a
detached worktree at pristine `b956bb7c6` carrying none of this PR's content (37 tests, 36 pass,
same single failure). Reported rather than chased.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V5xnzzipRHPHq8Z3doQvn3

* fix(#3829): record the shortfall when NO finding in a review parses

Found by this round's own pre-push adversarial review, not by the maintainer -- and it is the
same defect class round 2 raised as Blocker 4, surviving in the one corner that round's fix did
not reach.

The unparsed reconciliation exists so a finding the CR|BL|WR|IN heading parser cannot match is
SURFACED rather than dropped. It reported faithfully whenever SOME findings parsed. It reported
nothing at all when NONE did, on a phase with no prior ledger and no fix report: a REVIEW.md
declaring two Criticals, both written under a prefix the alternation does not carry, produced no
ledger, no console line and no diagnostic. That is precisely the silent drop this reconciliation
was added to close, reachable exactly where the evidence is weakest -- the run in which not one
finding was understood.

Two exits discarded it, and the first one is the one that actually fired. The shortfall was
derived beside the render, while the exit that stands down for "nothing to record" keys on
order.length and sits ~200 lines earlier; it returned before the value existed. The later
rows.length exit had the same hole but was unreachable for this input. So the derivation moves
above the earlier exit -- order is final from the heading walk and never grows again, so the value
is unchanged -- and both exits now decline to fire while a shortfall is outstanding. The result is
a zero-row ledger carrying an unparsed key: an honest record that the review declared findings and
none of them were understood, which is strictly better than the file not existing.

Scoped, not removed. A genuinely clean review is untouched: a declared total of 0 is not greater
than order.length, so unparsed is 0 and both returns still fire exactly as before. The new test's
companion pins that, and it is why the exit was relaxed conditionally rather than deleted.

Driven at every step rather than reasoned about. Before the fix, three cases through the shipped
script: two unmatched findings on a first run wrote NOTHING; two unmatched plus one matched
reported unparsed: 1; two unmatched against an existing ledger reported unparsed: 2. Only the
first was silent, which is why the mechanism read as covered. After the fix the first renders a
ledger with unparsed: 2 and names it on the console, and the other two are byte-unchanged.

The new regression test was negative-controlled against the pre-fix step file and goes red there
(1 pass / 1 fail over the pair); post-fix both pass. Its companion clean-review control is green
on both sides, so the pair is not passing by accident.

The PR's six test files: 554 tests, 554 pass, 0 fail, 0 skipped. lint:ci exits 0. The generator
sweep still produces no drift.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): pin the SECOND exit that discarded the shortfall, and prove the pair kills it

Self-found by the round review of the previous commit, not by the maintainer: that commit added
the !unparsed conjunct to BOTH exits but tested only one of them. Deleting the later one left
every new test green -- a surviving mutant, which is coverage in name only.

The two exits are reached by different inputs, which is why one fixture cannot pin both. The
earlier exit stands down the moment a fix report exists, so an input carrying one sails past it
and lands on the later rows.length return. The fixture therefore needs a fix report that
contributes NO row: an id the alternation CAN match becomes a carried row, rows.length is 1, and
the later guard never decides. The first draft of this test used CR-99 and was vacuous for exactly
that reason -- it passed with the guard deleted. It names SEC-03 now.

Mutation-controlled in all three directions, since a test that kills no mutant pins nothing:

  earlier exit loses !unparsed   -> test 1 RED,  test 2 green, test 3 green
  later exit loses !unparsed     -> test 1 RED,  test 2 RED,   test 3 green
  both exits neutered (over-fire)-> test 1 green,test 2 green, test 3 RED
  unmutated                      -> all three green

Every test kills at least one mutant and no mutant survives all three, so the pair covers both
roads to the drop and the clean-review control covers the over-fire the relaxation could have
introduced.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): close the fourth cell — the later exit's own over-fire

Self-found again by the round review, which drove the mutant set rather than trusting the matrix
the previous commit asserted: neutering ONLY the later exit survived all three tests. The previous
commit's matrix was accurate and incomplete, which is the more dangerous shape -- it reads as a
closed argument.

The guards form a 2x2 and only three cells were pinned. T1 pins the earlier exit's under-fire, T2
the later exit's under-fire, T3 the earlier exit's over-fire. Nothing pinned the LATER exit's own
over-fire, and it is reachable: a fix report -- even one naming no matchable id -- makes the
earlier exit stand down, so control arrives at the later exit carrying a genuinely clean review and
no shortfall. Neuter that exit and a phase with nothing to report grows a zero-row ledger reading
"0 of 0 finding(s) open", with all three earlier tests green through it. T4 is that case.

Five mutants driven over the four tests:

  earlier exit loses !unparsed        -> T1 RED
  later exit loses !unparsed          -> T1 RED, T2 RED
  both exits -> if(false)             -> T3 RED, T4 RED
  ONLY later exit -> if(false)        -> T4 RED          (the survivor this commit kills)
  ONLY earlier exit -> if(false)      -> all four green

The last one is reported as an EQUIVALENT mutant rather than an open cell, and the distinction is
the point of stating it: the earlier exit's over-fire condition is a strict subset of the later
exit's -- it adds only fixReports.length === 0 -- so removing it is subsumed and changes no
observable behaviour. A test cannot kill a mutant that does not alter output, and pretending
otherwise would mean writing one that asserts on internals.

The PR's six test files: 556 tests, 556 pass, 0 fail, 0 skipped. lint:ci exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* fix(#3829): keep the embedded record-builder inside a Windows command line

Block 2 runs the disposition record-builder as `node -e "<script>"`, so the entire
script is a single argv entry. Windows caps a command line at 32767 characters
(CreateProcess) and Node surfaces the overflow as ENAMETOOLONG from spawn -- the
process never starts. Linux's ~2 MB ARG_MAX cannot see that cliff at all.

The previous fix (bd83bc88e, hoisting the shortfall derivation above the earlier
exit) grew the extracted script from 31941 to 33353 characters. There were 826
characters of headroom; it spent 1412. On the next CI run ubuntu and macOS stayed
green and `conformance test (windows-latest, 24, shard 1/3)` went red with 83
failures -- 79 reporting `spawn_failed` out of runShippedDisposition, the other 4
asserting on a ledger that was never written. That same shard was SUCCESS at
596968aa0, the head before that commit.

Measured on native Windows (node v25.2.1), bisected: the largest `-e` argument that
still spawns is 32728 characters; 32729 fails. Both payloads driven directly:

    pre-fix   33353 chars  ->  SPAWN-FAIL ENAMETOOLONG
    post-fix  15916 chars  ->  SPAWN-OK

Note the shape of it: the fix that makes this gate report a silently-dropped finding
was itself silently dropped on Windows, because the whole script stopped launching.

What changes here is placement -- not content, not behaviour. 17 long rationale
comment blocks move out of the quoted payload into a new "Design notes for the
embedded record-builder" section in this file's prose, each anchored to the code line
it preceded so the pairing survives the move. Comment runs shorter than five lines
stay inline, where adjacency is cheap. Nothing is deleted.

    extracted script  33353 -> 15916 chars (16812 under the measured limit)
    executable code   byte-identical at 10853 bytes, verified by diffing the payload
                      with all comment lines stripped from both sides
    this file         78542 -> 79538 bytes -- the prose moved, it did not grow

Verified green: this PR's own test file at 250 tests across all 36 suites, 0 fail;
changeset-lint, docs-lint, default-flip-documentation, lint:ci, gen-emitted-baseline,
workflow-size-budget (133 tests), and lint-workflow-shellcheck (204 pre-existing
findings, 0 new).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): pin the embedded script under the Windows command-line budget

Nothing guarded the size of the `node -e` payload, so the regression the previous
commit fixes was invisible to every Linux gate and surfaced only as 83 Windows
failures that said `spawn_failed` and never mentioned length. Without a guard the
next addition to the script re-breaks Windows exactly the same way, and finds out
the same expensive way.

The test asserts the extracted script stays under a 24 KiB budget -- the 32767
CreateProcess cap less roughly 8 KiB of deliberate headroom, so the script has
somewhere to grow before this fires.

It is bounded from BELOW as well, and that half is the point: a pure length
assertion passes when the extractor returns '', which is exactly what a moved fence
or a renamed delimiter would produce. A guard that reports a comfortable 0 bytes is
the vacuous-oracle shape. The lower bound makes a broken extractor fail loudly here
instead of reporting success.

Negative-controlled rather than assumed. Against the PRE-fix step file the test goes
red on the real payload (33353 > 24576); against the fixed file it passes at 15916.
A new test that has only ever been run against fixed code can be green because it
hit a branch the bug never lived on.

This PR's own test file: 250 tests, 250 pass, 0 fail, 0 skipped, across all 36
suites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): say what the command-line guard does not prove

The round review's MISSED, adopted. The guard counts the extracted JavaScript, but the
32767 cap applies to the whole serialized command line -- executable path, quoting and
backslash escaping included -- so a quote-heavy payload expands on the way out, and the
32728 figure it cites is one host's measured threshold rather than the CI runner's.

The budget is unchanged and still correct; only the claim around it moves. 24576 leaves
roughly 8 KiB for both effects, which is a practical margin, not a proof that everything
the guard admits will spawn. Stating that in the test is cheaper than having a future
reader infer a guarantee the assertion cannot make.

Comment-only. No assertion, no budget and no behaviour changes.

This PR's own test file: 221 tests, 0 fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-22 21:04:09 -04:00

4603 lines
284 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# GSD Feature Reference
> Feature index and reference for GSD Core. For architecture details, see [Architecture](ARCHITECTURE.md). For command syntax, see [Command Reference](COMMANDS.md). Return to [docs index](README.md).
---
<!-- FEATURES:START — generated by scripts/gen-features.cjs; do not edit by hand -->
## Table of Contents
- [Core Features](#core-features)
- [Project Initialization](#1-project-initialization)
- [Phase Discussion](#2-phase-discussion)
- [UI Design Contract](#3-ui-design-contract)
- [Phase Planning](#4-phase-planning)
- [Phase Execution](#5-phase-execution)
- [Work Verification](#6-work-verification)
- [Ship](#65-ship)
- [UI Review](#7-ui-review)
- [Milestone Management](#8-milestone-management)
- [Planning Features](#planning-features)
- [Phase Management](#9-phase-management)
- [Quick Mode](#10-quick-mode)
- [Autonomous Mode](#11-autonomous-mode)
- [Freeform Routing](#12-freeform-routing)
- [Note Capture](#13-note-capture)
- [Auto-Advance (Next)](#14-auto-advance-next)
- [Review Dispositions Ledger](#3806-review-dispositions-ledger)
- [Quick Batch Mode](#4015-quick-batch-mode)
- [Quality Assurance Features](#quality-assurance-features)
- [Nyquist Validation](#15-nyquist-validation)
- [Plan Checking](#16-plan-checking)
- [Post-Execution Verification](#17-post-execution-verification)
- [Node Repair](#18-node-repair)
- [Health Validation](#19-health-validation)
- [Cross-Phase Regression Gate](#20-cross-phase-regression-gate)
- [Requirements Coverage Gate](#21-requirements-coverage-gate)
- [TDD-Applicability Predicate](#4273-tdd-applicability-predicate)
- [Context Engineering Features](#context-engineering-features)
- [Context Window Monitoring](#22-context-window-monitoring)
- [Session Management](#23-session-management)
- [Session Reporting](#24-session-reporting)
- [Multi-Agent Orchestration](#25-multi-agent-orchestration)
- [Model Profiles](#26-model-profiles)
- [Compact Content Mode](#4139-compact-content-mode)
- [Brownfield Features](#brownfield-features)
- [Codebase Mapping](#27-codebase-mapping)
- [Existing Codebase Onboarding](#27b-existing-codebase-onboarding)
- [Post-Execute Codebase Drift Detection](#27a-post-execute-codebase-drift-detection)
- [Utility Features](#utility-features)
- [Debug System](#28-debug-system)
- [Todo Management](#29-todo-management)
- [Statistics Dashboard](#30-statistics-dashboard)
- [Update System](#31-update-system)
- [Settings Management](#32-settings-management)
- [Test Generation](#33-test-generation)
- [Infrastructure Features](#infrastructure-features)
- [Git Integration](#34-git-integration)
- [CLI Tools](#35-cli-tools)
- [Multi-Runtime Support](#36-multi-runtime-support)
- [Hook System](#37-hook-system)
- [Developer Profiling](#38-developer-profiling)
- [Execution Hardening](#39-execution-hardening)
- [Verification Debt Tracking](#40-verification-debt-tracking)
- [v1.27 Features](#v127-features)
- [Fast Mode](#41-fast-mode)
- [Cross-AI Peer Review](#42-cross-ai-peer-review)
- [Backlog Parking Lot](#43-backlog-parking-lot)
- [Persistent Context Threads](#44-persistent-context-threads)
- [PR Branch Filtering](#45-pr-branch-filtering)
- [Security Hardening](#46-security-hardening)
- [Multi-Repo Workspace Support](#47-multi-repo-workspace-support)
- [Discussion Audit Trail](#48-discussion-audit-trail)
- [v1.28 Features](#v128-features)
- [Forensics](#49-forensics)
- [Milestone Summary](#50-milestone-summary)
- [Workstream Namespacing](#51-workstream-namespacing)
- [Manager Dashboard](#52-manager-dashboard)
- [Assumptions Discussion Mode](#53-assumptions-discussion-mode)
- [UI Phase Auto-Detection](#54-ui-phase-auto-detection)
- [Multi-Runtime Installer Selection](#55-multi-runtime-installer-selection)
- [v1.29 Features](#v129-features)
- [Windsurf Runtime Support](#56-windsurf-runtime-support)
- [Internationalized Documentation](#57-internationalized-documentation)
- [v1.31 Features](#v131-features)
- [Schema Drift Detection](#59-schema-drift-detection)
- [Security Enforcement](#60-security-enforcement)
- [Documentation Generation](#61-documentation-generation)
- [Discuss Chain Mode](#62-discuss-chain-mode)
- [Single-Phase Autonomous](#63-single-phase-autonomous)
- [Scope Reduction Detection](#64-scope-reduction-detection)
- [Claim Provenance Tagging](#65-claim-provenance-tagging)
- [Worktree Toggle](#66-worktree-toggle)
- [Project Code Prefixing](#67-project-code-prefixing)
- [Claude Code Skills Migration](#68-claude-code-skills-migration)
- [v1.32 Features](#v132-features)
- [STATE.md Consistency Gates](#69-statemd-consistency-gates)
- [Autonomous `--to N` Flag](#70-autonomous---to-n-flag)
- [Research Gate](#71-research-gate)
- [Verifier Milestone Scope Filtering](#72-verifier-milestone-scope-filtering)
- [Read-Before-Edit Guard Hook](#73-read-before-edit-guard-hook)
- [Context Reduction](#74-context-reduction)
- [Discuss-Phase `--power` Flag](#75-discuss-phase---power-flag)
- [Debug `--diagnose` Flag](#76-debug---diagnose-flag)
- [Phase Dependency Analysis](#77-phase-dependency-analysis)
- [Anti-Pattern Severity Levels](#78-anti-pattern-severity-levels)
- [Methodology Artifact Type](#79-methodology-artifact-type)
- [Planner Reachability Check](#80-planner-reachability-check)
- [Playwright-MCP UI Verification](#81-playwright-mcp-ui-verification)
- [Pause-Work Expansion](#82-pause-work-expansion)
- [Response Language Config](#83-response-language-config)
- [Manual Update Procedure](#84-manual-update-procedure)
- [New Runtime Support (Trae, Cline, Augment Code)](#85-new-runtime-support-trae-cline-augment-code)
- [Autonomous `--interactive` Flag](#86-autonomous---interactive-flag)
- [Commit-Docs Guard Hook](#87-commit-docs-guard-hook)
- [Community Hooks Opt-In](#88-community-hooks-opt-in)
- [v1.34.0 Features](#v1340-features)
- [Global Learnings Store](#89-global-learnings-store)
- [Queryable Codebase Intelligence](#90-queryable-codebase-intelligence)
- [Execution Context Profiles](#91-execution-context-profiles)
- [Gates Taxonomy](#92-gates-taxonomy)
- [Code Review Pipeline](#93-code-review-pipeline)
- [Socratic Exploration](#94-socratic-exploration)
- [Safe Undo](#95-safe-undo)
- [Plan Import](#96-plan-import)
- [Rapid Codebase Scan](#97-rapid-codebase-scan)
- [Autonomous Audit-to-Fix](#98-autonomous-audit-to-fix)
- [Improved Prompt Injection Scanner](#99-improved-prompt-injection-scanner)
- [Stall Detection in Plan-Phase](#100-stall-detection-in-plan-phase)
- [Hard Stop Safety Gates in /gsd-progress --next](#101-hard-stop-safety-gates-in-gsd-progress---next)
- [Adaptive Model Preset](#102-adaptive-model-preset)
- [Post-Merge Hunk Verification](#103-post-merge-hunk-verification)
- [v1.35.0 Features](#v1350-features)
- [New Runtime Support (Cline, CodeBuddy, Qwen Code)](#104-new-runtime-support-cline-codebuddy-qwen-code)
- [GSD-2 Reverse Migration](#105-gsd-2-reverse-migration)
- [AI Integration Phase Wizard](#106-ai-integration-phase-wizard)
- [AI Eval Review](#107-ai-eval-review)
- [v1.36.0 Features](#v1360-features)
- [Plan Bounce](#108-plan-bounce)
- [External Code Review Command](#109-external-code-review-command)
- [Cross-AI Execution Delegation](#110-cross-ai-execution-delegation)
- [Architectural Responsibility Mapping](#111-architectural-responsibility-mapping)
- [Extract Learnings](#112-extract-learnings)
- [Context-Window-Aware Prompt Thinning](#114-context-window-aware-prompt-thinning)
- [Configurable CLAUDE.md Path](#115-configurable-claudemd-path)
- [TDD Pipeline Mode](#116-tdd-pipeline-mode)
- [v1.37.0 Features](#v1370-features)
- [Spike Command](#117-spike-command)
- [Sketch Command](#118-sketch-command)
- [Agent Size-Budget Enforcement](#119-agent-size-budget-enforcement)
- [Shared Boilerplate Extraction](#120-shared-boilerplate-extraction)
- [Knowledge Graph Integration](#121-knowledge-graph-integration)
- [v1.40.0 Features](#v1400-features)
- [Skill Surface Consolidation](#122-skill-surface-consolidation)
- [Namespace Meta-Skills (Two-Stage Routing)](#123-namespace-meta-skills-two-stage-routing)
- [Context-Window Utilization Guard](#124-context-window-utilization-guard)
- [Phase-Lifecycle Status-Line Read-Side](#125-phase-lifecycle-status-line-read-side)
- [v1.41.0 Features](#v1410-features)
- [Per-Phase-Type Model Selection](#126-per-phase-type-model-selection)
- [Dynamic Routing with Failure-Tier Escalation](#127-dynamic-routing-with-failure-tier-escalation)
- [Update Banner Opt-In](#128-update-banner-opt-in)
- [Issue-Driven Orchestration Guide](#129-issue-driven-orchestration-guide)
- [Graphify Commit-Based Staleness](#130-graphify-commit-based-staleness)
- [v1.42.1 Features](#v1421-features)
- [Package Legitimacy Gate](#132-package-legitimacy-gate)
- [Skill Surface Budgeting](#133-skill-surface-budgeting)
- [Installer Migrations](#134-installer-migrations)
- [Custom Ship PR Body Sections](#135-custom-ship-pr-body-sections)
- [Review Default Reviewers](#136-review-default-reviewers)
- [Fallow Structural Review Pre-Pass](#137-fallow-structural-review-pre-pass)
- [End-of-Phase Human Verification Mode](#138-end-of-phase-human-verification-mode)
- [Quota and Rate-Limit Failure Classification](#139-quota-and-rate-limit-failure-classification)
- [Statusline Context Position](#140-statusline-context-position)
- [Milestone Tag Creation Toggle](#141-milestone-tag-creation-toggle)
- [Structured JSON Error Mode](#142-structured-json-error-mode)
- [UAT-Passed Predicate](#143-uat-passed-predicate)
- [Spec-Phase Edge-Completeness Probe](#144-spec-phase-edge-completeness-probe)
- [v1.43.0 Features](#v1430-features)
- [MemPalace Memory Capability](#145-mempalace-memory-capability)
- [Spec-Phase Prohibition Probe](#146-spec-phase-prohibition-probe)
- [Capability Management Command](#147-capability-management-command)
- [Smart Entry Launcher](#148-smart-entry-launcher)
- [v1.7.0 Features](#v170-features)
- [Embeddable Orchestration System (Host-Integration Interface)](#149-embeddable-orchestration-system-host-integration-interface)
- [Discoverability Registries](#150-discoverability-registries)
- [Companion MCP Server](#151-companion-mcp-server)
- [Statusline Token Count & Git Segment](#152-statusline-token-count--git-segment)
- [Model Catalog Advances](#153-model-catalog-advances)
- [Claude Orchestration Capability (BETA)](#154-claude-orchestration-capability-beta)
- [External-Job Capability](#155-external-job-capability)
- [API-Coverage Gate](#156-api-coverage-gate)
- [State Rebuild & Configurable Graph Path](#157-state-rebuild--configurable-graph-path)
- [Broken-Windows Ledger](#158-broken-windows-ledger)
- [Complexity-Triggered Refactor](#159-complexity-triggered-refactor)
- [Archive Quick Tasks at Milestone Close](#160-archive-quick-tasks-at-milestone-close)
- [Verify-Command Path Grounding](#161-verify-command-path-grounding)
- [Statusline STATE.md Freshness Marker](#162-statusline-statemd-freshness-marker)
- [Read-Only Planning Snapshot (`planning inspect`)](#163-read-only-planning-snapshot-planning-inspect)
- [Live-DOM UAT Capability](#164-live-dom-uat-capability)
- [Opt-In Parallel Reviewer Lanes](#165-opt-in-parallel-reviewer-lanes)
- [Machine-Readable State Contract (`.planning/state.json`)](#166-machine-readable-state-contract-planningstatejson)
- [Stated Failing Direction](#167-stated-failing-direction)
- [Runtime Identity](#168-runtime-identity)
- [Context Drift Gate](#3348-context-drift-gate)
- ["Failure Is a Value" — Strict Argv Rejection and the `--pick` Absence Contract](#3884-failure-is-a-value--strict-argv-rejection-and-the---pick-absence-contract)
- [No Silent Swallow, No Verdict From Dropped Data](#3885-no-silent-swallow-no-verdict-from-dropped-data)
- [Runtime Marker Resolution, Derived Codex Sandbox, and In-Phase Short-Form Dependencies](#3897-runtime-marker-resolution-derived-codex-sandbox-and-in-phase-short-form-dependencies)
- [The Raw Terminator Is Banned by Construction](#3910-the-raw-terminator-is-banned-by-construction)
- [Hooks Declare Their Crash Policy](#3911-hooks-declare-their-crash-policy)
- [gsd-tools Declares Outcomes, Pinned at v1](#3912-gsd-tools-declares-outcomes-pinned-at-v1)
- [Reachable Lint Rules and a Non-Destructive Quick-Task Append](#3951-reachable-lint-rules-and-a-non-destructive-quick-task-append)
- [Per-Task External-Tracker Content-Resolution Seam](#3970-per-task-external-tracker-content-resolution-seam)
- [Unreadable-Directory Scope Signal](#4014-unreadable-directory-scope-signal)
- [Graphify CLI Preferred for Planner and Researcher Graph Queries](#4836-graphify-cli-preferred-for-planner-and-researcher-graph-queries)
---
## Core Features
### 1. Project Initialization
**Command:** `/gsd-new-project [--auto @file.md]`
**Purpose:** Transform a user's idea into a fully structured project with research, scoped requirements, and a phased roadmap.
**Requirements:**
- REQ-INIT-01: System MUST conduct adaptive questioning until project scope is fully understood
- REQ-INIT-02: System MUST spawn parallel research agents to investigate the domain ecosystem
- REQ-INIT-03: System MUST extract requirements into v1 (must-have), v2 (future), and out-of-scope categories
- REQ-INIT-04: System MUST generate a phased roadmap with requirement traceability
- REQ-INIT-05: System MUST require user approval of the roadmap before proceeding
- REQ-INIT-06: System MUST prevent re-initialization when `.planning/PROJECT.md` already exists
- REQ-INIT-07: System MUST support `--auto @file.md` flag to skip interactive questions and extract from a document
**Produces:**
| Artifact | Description |
|----------|-------------|
| `PROJECT.md` | Project vision, constraints, technical decisions, evolution rules |
| `REQUIREMENTS.md` | Scoped requirements with unique IDs (REQ-XX) |
| `ROADMAP.md` | Phase breakdown with status tracking and requirement mapping |
| `STATE.md` | Initial project state with position, decisions, metrics |
| `config.json` | Workflow configuration |
| `research/SUMMARY.md` | Synthesized domain research |
| `research/STACK.md` | Technology stack investigation |
| `research/FEATURES.md` | Feature implementation patterns |
| `research/ARCHITECTURE.md` | Architecture patterns and trade-offs |
| `research/PITFALLS.md` | Common failure modes and mitigations |
**Process:**
1. **Questions** — Adaptive questioning guided by the "dream extraction" philosophy (not requirements gathering)
2. **Research** — 4 parallel researcher agents investigate stack, features, architecture, and pitfalls
3. **Synthesis** — Research synthesizer combines findings into SUMMARY.md
4. **Requirements** — Extracted from user responses + research, categorized by scope
5. **Roadmap** — Phase breakdown mapped to requirements, with granularity setting controlling phase count
**Functional Requirements:**
- Questions adapt based on detected project type (web app, CLI, mobile, API, etc.)
- Research agents have web search capability for current ecosystem information
- Granularity setting controls phase count: `coarse` (2-4), `standard` (4-6), `fine` (6-10)
- `--auto` mode extracts all information from the provided document without interactive questioning
- Existing codebase context (from `/gsd-map-codebase`) is loaded if present
---
### 2. Phase Discussion
**Command:** `/gsd-discuss-phase [N] [--auto] [--batch]`
**Purpose:** Capture user's implementation preferences and decisions before research and planning begin. Eliminates the gray areas that cause AI to guess.
**Requirements:**
- REQ-DISC-01: System MUST analyze the phase scope and identify decision areas (gray areas)
- REQ-DISC-02: System MUST categorize gray areas by type (visual, API, content, organization, etc.)
- REQ-DISC-03: System MUST ask only questions not already answered in prior CONTEXT.md files
- REQ-DISC-04: System MUST persist decisions in `{phase}-CONTEXT.md` with canonical references
- REQ-DISC-05: System MUST support `--auto` flag to auto-select recommended defaults
- REQ-DISC-06: System MUST support `--batch` flag for grouped question intake
- REQ-DISC-07: System MUST scout relevant source files before identifying gray areas (code-aware discussion)
- REQ-DISC-08: System MUST adapt gray area language to product-outcome terms when USER-PROFILE.md indicates a non-technical owner (learning_style: guided, jargon in frustration_triggers, or high-level explanation depth)
- REQ-DISC-09: When REQ-DISC-08 applies, advisor_research rationale paragraphs MUST be rewritten in plain language — same decisions, translated framing
**Produces:** `{padded_phase}-CONTEXT.md` — User preferences that feed into research and planning
**Gray Area Categories:**
| Category | Example Decisions |
|----------|-------------------|
| Visual features | Layout, density, interactions, empty states |
| APIs/CLIs | Response format, flags, error handling, verbosity |
| Content systems | Structure, tone, depth, flow |
| Organization | Grouping criteria, naming, duplicates, exceptions |
---
### 3. UI Design Contract
**Command:** `/gsd-ui-phase [N]`
**Purpose:** Lock design decisions before planning so that all components in a phase share consistent visual standards.
**Requirements:**
- REQ-UI-01: System MUST detect existing design system state (shadcn components.json, Tailwind config, tokens)
- REQ-UI-02: System MUST ask only unanswered design contract questions
- REQ-UI-03: System MUST validate against 7 dimensions (Copywriting, Visuals, Color, Typography, Spacing, Registry Safety, Inventory Provenance)
- REQ-UI-04: System MUST enter revision loop if validation returns BLOCKED (max 2 iterations)
- REQ-UI-05: System MUST offer shadcn initialization for React/Next.js/Vite projects without `components.json`
- REQ-UI-06: System MUST enforce registry safety gate for third-party shadcn registries
**Produces:** `{padded_phase}-UI-SPEC.md` — Design contract consumed by executors
**7 Validation Dimensions:**
1. **Copywriting** — CTA labels, empty states, error messages
2. **Visuals** — Focal points, visual hierarchy, icon accessibility
3. **Color** — Accent usage discipline, 60/30/10 compliance
4. **Typography** — Font size/weight constraint adherence
5. **Spacing** — Grid alignment, token consistency
6. **Registry Safety** — Third-party component inspection requirements
7. **Inventory Provenance** — Component inventory enumerated from the installed design system, not recalled
**shadcn Integration:**
- Detects missing `components.json` in React/Next.js/Vite projects
- Guides user through `ui.shadcn.com/create` preset configuration
- Preset string becomes a planning artifact reproducible across phases
- Safety gate requires `npx shadcn view` and `npx shadcn diff` before third-party components
---
### 4. Phase Planning
**Command:** `/gsd-plan-phase [N] [--auto] [--skip-research] [--skip-verify]`
**Purpose:** Research the implementation domain and produce verified, atomic execution plans.
**Requirements:**
- REQ-PLAN-01: System MUST spawn a phase researcher to investigate implementation approaches
- REQ-PLAN-02: System MUST produce plans with 2-3 tasks each, sized for a single context window
- REQ-PLAN-03: System MUST structure plans as XML with `<task>` elements containing `name`, `files`, `action`, `verify`, and `done` fields
- REQ-PLAN-04: System MUST include `read_first` and `acceptance_criteria` sections in every plan
- REQ-PLAN-05: System MUST run plan checker verification loop (up to 3 iterations) unless `--skip-verify` is set
- REQ-PLAN-06: System MUST support `--skip-research` flag to bypass research phase
- REQ-PLAN-07: System MUST prompt user to run `/gsd-ui-phase` if frontend phase detected and no UI-SPEC.md exists (UI safety gate)
- REQ-PLAN-08: System MUST include Nyquist validation mapping when `workflow.nyquist_validation` is enabled
- REQ-PLAN-09: System MUST verify all phase requirements are covered by at least one plan before planning completes (requirements coverage gate)
- REQ-PLAN-10: System MUST support an optional `<reversibility rating="reversible|costly|one-way">` element recording how costly a decision would be to undo, and MUST insert a `checkpoint:decision` before the task implementing a `one-way` decision unless `--no-reversibility-gates` is set (`costly` is flagged without blocking; `reversible` and unrated flow normally)
**Produces:**
| Artifact | Description |
|----------|-------------|
| `{phase}-RESEARCH.md` | Ecosystem research findings |
| `{phase}-{N}-PLAN.md` | Atomic execution plans (2-3 tasks each) |
| `{phase}-VALIDATION.md` | Test coverage mapping (Nyquist layer) |
**Plan Structure (XML):**
```xml
<task type="auto">
<name>Create login endpoint</name>
<files>src/app/api/auth/login/route.ts</files>
<action>
Use jose for JWT. Validate credentials against users table.
Return httpOnly cookie on success.
</action>
<verify>curl -X POST localhost:3000/api/auth/login returns 200 + Set-Cookie</verify>
<done>Valid credentials return cookie, invalid return 401</done>
</task>
```
**Plan Checker Verification (8 Dimensions):**
1. Requirement coverage — Plans address all phase requirements
2. Task atomicity — Each task is independently committable
3. Dependency ordering — Tasks sequence correctly
4. File scope — No excessive file overlap between plans
5. Verification commands — Each task has testable done criteria
6. Context fit — Tasks fit within a single context window
7. Gap detection — No missing implementation steps
8. Nyquist compliance — Tasks have automated verify commands (when enabled)
---
### 5. Phase Execution
**Command:** `/gsd-execute-phase <N>`
**Purpose:** Execute all plans in a phase using wave-based parallelization with fresh context windows per executor.
**Requirements:**
- REQ-EXEC-01: System MUST analyze plan dependencies and group into execution waves
- REQ-EXEC-02: System MUST spawn independent plans in parallel within each wave
- REQ-EXEC-03: System MUST give each executor a fresh context window (200K tokens)
- REQ-EXEC-04: System MUST produce atomic git commits per task
- REQ-EXEC-05: System MUST produce a SUMMARY.md for each completed plan
- REQ-EXEC-06: System MUST run post-execution verifier to check phase goals were met
- REQ-EXEC-07: System MUST support git branching strategies (`none`, `phase`, `milestone`)
- REQ-EXEC-08: System MUST invoke node repair operator on task verification failure (when enabled)
- REQ-EXEC-09: System MUST run prior phases' test suites before verification to catch cross-phase regressions
**Produces:**
| Artifact | Description |
|----------|-------------|
| `{phase}-{N}-SUMMARY.md` | Execution outcomes per plan |
| `{phase}-VERIFICATION.md` | Post-execution verification report |
| Git commits | Atomic commits per task |
**Wave Execution:**
- Plans with no dependencies → Wave 1 (parallel)
- Plans depending on Wave 1 → Wave 2 (parallel, waits for Wave 1)
- Continues until all plans complete
- File conflicts force sequential execution within same wave
**Executor Capabilities:**
- Reads PLAN.md with full task instructions
- Has access to PROJECT.md, STATE.md, CONTEXT.md, RESEARCH.md
- Commits each task atomically with structured commit messages
- Uses `--no-verify` on commits during parallel execution to avoid build lock contention
- Handles checkpoint types: `auto`, `checkpoint:human-verify`, `checkpoint:decision`, `checkpoint:human-action`
- Reports deviations from plan in SUMMARY.md
**Parallel Safety:**
- **Pre-commit hooks**: Skipped by parallel agents (`--no-verify`), run once by orchestrator after each wave
- **STATE.md locking**: File-level lockfile prevents concurrent write corruption across agents
---
### 6. Work Verification
**Command:** `/gsd-verify-work [N]`
**Purpose:** User acceptance testing — walk the user through testing each deliverable and auto-diagnose failures.
**Requirements:**
- REQ-VERIFY-01: System MUST extract testable deliverables from the phase
- REQ-VERIFY-02: System MUST present deliverables one at a time for user confirmation
- REQ-VERIFY-03: System MUST spawn debug agents to diagnose failures automatically
- REQ-VERIFY-04: System MUST create fix plans for identified issues
- REQ-VERIFY-05: System MUST inject cold-start smoke test for phases modifying server/database/seed/startup files
- REQ-VERIFY-06: System MUST produce UAT.md with pass/fail results
**Produces:** `{phase}-UAT.md` — User acceptance test results, plus fix plans if issues found
---
### 6.5. Ship
**Command:** `/gsd-ship [N] [--draft]`
**Purpose:** Bridge local completion → merged PR. After verification passes, push branch, create PR with auto-generated body from planning artifacts, optionally trigger review, and track in STATE.md.
**Requirements:**
- REQ-SHIP-01: System MUST verify phase has passed verification before shipping
- REQ-SHIP-02: System MUST push branch and create PR via `gh` CLI
- REQ-SHIP-03: System MUST auto-generate PR body from SUMMARY.md, VERIFICATION.md, and REQUIREMENTS.md
- REQ-SHIP-04: System MUST update STATE.md with shipping status and PR number
- REQ-SHIP-05: System MUST support `--draft` flag for draft PRs
- REQ-SHIP-06: System MUST support append-only project PR body sections configured with `ship.pr_body_sections`
**Prerequisites:** Phase verified, `gh` CLI installed and authenticated, work on feature branch
**Produces:** GitHub PR with rich body, optional configured PRD-style sections, STATE.md updated
**User documentation:** [Custom PR Body Sections](ship-pr-body-sections.md)
---
### 7. UI Review
**Command:** `/gsd-ui-review [N]`
**Purpose:** Retroactive 6-pillar visual audit of implemented frontend code. Works standalone on any project.
**Requirements:**
- REQ-UIREVIEW-01: System MUST score each of the 6 pillars on a 1-4 scale
- REQ-UIREVIEW-02: System MUST capture screenshots via Playwright CLI to `.planning/ui-reviews/`
- REQ-UIREVIEW-03: System MUST create `.gitignore` for screenshot directory
- REQ-UIREVIEW-04: System MUST identify top 3 priority fixes
- REQ-UIREVIEW-05: System MUST work standalone (without UI-SPEC.md) using abstract quality standards
**6 Audit Pillars (scored 1-4):**
1. **Copywriting** — CTA labels, empty states, error states
2. **Visuals** — Focal points, visual hierarchy, icon accessibility
3. **Color** — Accent usage discipline, 60/30/10 compliance
4. **Typography** — Font size/weight constraint adherence
5. **Spacing** — Grid alignment, token consistency
6. **Experience Design** — Loading/error/empty state coverage
**Produces:** `{padded_phase}-UI-REVIEW.md` — Scores and prioritized fixes
---
### 8. Milestone Management
**Commands:** `/gsd-audit-milestone`, `/gsd-complete-milestone`, `/gsd-new-milestone [name]`
**Purpose:** Verify milestone completion, archive, tag release, and start the next development cycle.
**Requirements:**
- REQ-MILE-01: Audit MUST verify all milestone requirements are met
- REQ-MILE-02: Audit MUST detect stubs, placeholder implementations, and untested code
- REQ-MILE-03: Audit MUST check Nyquist validation compliance across phases
- REQ-MILE-04: Complete MUST archive milestone data to MILESTONES.md
- REQ-MILE-05: Complete MUST offer git tag creation for the release
- REQ-MILE-06: Complete MUST offer squash merge or merge with history for branching strategies
- REQ-MILE-07: Complete MUST clean up UI review screenshots
- REQ-MILE-08: New milestone MUST follow same flow as new-project (questions → research → requirements → roadmap)
- REQ-MILE-09: New milestone MUST NOT reset existing workflow configuration
---
## Planning Features
### 9. Phase Management
**Commands:** `/gsd-phase`, `/gsd-phase --insert [N]`, `/gsd-phase --remove [N]`
**Purpose:** Dynamic roadmap modification during development.
**Requirements:**
- REQ-PHASE-01: Add MUST append a new phase to the end of the current roadmap
- REQ-PHASE-02: Insert MUST use decimal numbering (e.g., 3.1) between existing phases
- REQ-PHASE-03: Remove MUST renumber all subsequent phases
- REQ-PHASE-04: Remove MUST prevent removing phases that have been executed
- REQ-PHASE-05: All operations MUST update ROADMAP.md and create/remove phase directories
- REQ-PHASE-06: Bare-number phase lookup MUST resolve digit-leading slug names consistently across phase verbs, preserve project-code-prefixed result shaping, and fail loudly when multiple directories match
---
### 10. Quick Mode
**Command:** `/gsd-quick [--full] [--discuss] [--research]`
**Purpose:** Ad-hoc task execution with GSD guarantees but a faster path.
**Requirements:**
- REQ-QUICK-01: System MUST accept freeform task description
- REQ-QUICK-02: System MUST use same planner + executor agents as full workflow
- REQ-QUICK-03: System MUST skip research, plan checker, and verifier by default
- REQ-QUICK-04: `--full` flag MUST enable plan checking (max 2 iterations) and post-execution verification
- REQ-QUICK-05: `--discuss` flag MUST run lightweight pre-planning discussion
- REQ-QUICK-06: `--research` flag MUST spawn focused research agent before planning
- REQ-QUICK-07: Flags MUST be composable (`--discuss --research --full`)
- REQ-QUICK-08: System MUST track quick tasks in `.planning/quick/YYMMDD-xxx-slug/`
- REQ-QUICK-09: System MUST produce atomic commits for quick task execution
---
### 11. Autonomous Mode
**Command:** `/gsd-autonomous [--from N]`
**Purpose:** Run all remaining phases autonomously — discuss → plan → execute per phase.
**Requirements:**
- REQ-AUTO-01: System MUST iterate through all incomplete phases in roadmap order
- REQ-AUTO-02: System MUST run discuss → plan → execute for each phase
- REQ-AUTO-03: System MUST pause for explicit user decisions (gray area acceptance, blockers, validation)
- REQ-AUTO-04: System MUST re-read ROADMAP.md after each phase to catch dynamically inserted phases
- REQ-AUTO-05: `--from N` flag MUST start from a specific phase number
---
### 12. Freeform Routing
**Command:** `/gsd-progress --do` (see also `/gsd-manager` for interactive routing)
**Purpose:** Analyze freeform text and route to the appropriate GSD command.
**Requirements:**
- REQ-DO-01: System MUST parse user intent from natural language input
- REQ-DO-02: System MUST map intent to the best matching GSD command
- REQ-DO-03: System MUST confirm the routing with the user before executing
- REQ-DO-04: System MUST handle project-exists vs no-project contexts differently
- REQ-DO-05: Routing rules MUST order specific operations before the generic keyword rules they shadow (specific-before-generic)
- REQ-DO-06: Dispatch MUST forward only arguments the selected command accepts; the freeform sentence is forwarded only when that command explicitly accepts a freeform task description
---
### 13. Note Capture
**Command:** `/gsd-capture`
**Purpose:** Zero-friction idea capture without interrupting workflow. Append timestamped notes, list all notes, or promote notes to structured todos.
**Requirements:**
- REQ-NOTE-01: System MUST save timestamped note files with a single Write call
- REQ-NOTE-02: System MUST support `list` subcommand to show all notes from project and global scopes
- REQ-NOTE-03: System MUST support `promote N` subcommand to convert a note into a structured todo
- REQ-NOTE-04: System MUST support `--global` flag for global scope operations
- REQ-NOTE-05: System MUST NOT use Task, AskUserQuestion, or Bash — runs inline only
---
### 14. Auto-Advance (Next)
**Command:** `/gsd-progress --next`
**Purpose:** Automatically detect current project state and advance to the next logical workflow step, eliminating the need to remember which phase/step you're on.
**Requirements:**
- REQ-NEXT-01: System MUST read STATE.md, ROADMAP.md, and phase directories to determine current position
- REQ-NEXT-02: System MUST detect whether discuss, plan, execute, or verify is needed
- REQ-NEXT-03: System MUST invoke the correct command automatically
- REQ-NEXT-04: System MUST suggest `/gsd-new-project` if no project exists
- REQ-NEXT-05: System MUST suggest `/gsd-complete-milestone` when all phases are complete
**State Detection Logic:**
| State | Action |
|-------|--------|
| No `.planning/` directory | Suggest `/gsd-new-project` |
| Phase has no CONTEXT.md | Run `/gsd-discuss-phase` |
| Phase has no PLAN.md files | Run `/gsd-plan-phase` |
| Phase has plans but no SUMMARY.md | Run `/gsd-execute-phase` |
| Phase executed but no VERIFICATION.md | Run `/gsd-verify-work` |
| All phases complete | Suggest `/gsd-complete-milestone` |
---
### 3806. Review Dispositions Ledger
**Purpose:** Reviews-mode planning (`/gsd-plan-phase {N} --reviews`) has required every current
actionable REVIEWS.md finding to be incorporated into PLAN.md or explicitly deferred/rejected
there since v1.5.0 (#724/#728). Nothing canonized *where* in PLAN.md, *what shape*, or how a
REVIEWS.md line reference survives the next round rewriting the file wholesale. Two
independently-invented, mutually incompatible disposition formats were observed across two
consecutive rounds of the same phase, each written by a different planner subagent instance
improvising from prose alone.
**Behavior:** The existing return-payload tables from `references/planner-reviews.md` Step 4 —
`### Review Feedback Addressed` / `### Review Feedback Deferred` — are now the canonical
**Review Dispositions Ledger**, promoted verbatim in shape into the affected PLAN.md itself under
a `## Review Dispositions Ledger` heading. Each reviews-mode round gets its own
`### Round {N} — {REVIEWS_sha}` subsection, where `{REVIEWS_sha}` is the commit that wrote that
round's REVIEWS.md snapshot (`workflows/review.md` already commits REVIEWS.md as its own commit).
A REVIEWS.md line reference cites `L##@{REVIEWS_sha}`; a bare line number is non-conforming. The
ledger is append-only — a later round adds a new row naming what it supersedes rather than editing
or deleting an earlier round's tables.
The contract is stated once, in `references/planner-reviews.md`; `workflows/plan-phase.md`'s
`<review_incorporation_contract>` and `agents/gsd-plan-checker.md`'s Review Incorporation dimension
both reference it by name rather than restating it, guarded by a parity test
(`tests/plan-review-convergence.test.cjs`) that fails if the three drift apart.
`{Concern}`/`{Reason}` stay free text — the reviewer roster is capability-owned and open to
third-party additions, so no closed reviewer/severity enum is introduced.
**Known limits:** No lint or check verb enforces this shape yet — a follow-up (tracked as part 2
of #3806) will add deterministic enforcement once a migration story for the two pre-existing ad-hoc
formats already in the wild is decided. Legacy PLAN.md content written before this convention is
not migrated or flagged.
**Reference:** [ADR-3806](adr/3806-review-dispositions-ledger.md) · [Cross-AI Peer Review](#42-cross-ai-peer-review)
---
### 4015. Quick Batch Mode
**Command:** `/gsd-quick-batch [--file <path>] [--jobs auto|N] [--validate] [--research] [--resume <batch-id>]`
**Purpose:** Batch several `/gsd-quick`-shaped tasks together — one coordinator plans, dispatches, and merges them as a single run, with per-item leaves and deterministic merge ordering (ADR-1239 "Quick-batch binding").
**Requirements:**
- REQ-QB-01: System MUST accept an inline task list (≥2 items) or `--file <path>`
- REQ-QB-02: System MUST reject `--discuss` and `--full` with a usage error before any dispatch
- REQ-QB-03: System MUST reject a malformed `--jobs` value before any dispatch
- REQ-QB-04: System MUST resolve effective concurrency as `min(task count, jobsN, capacity)` for `--jobs N`, or `capacity` alone for `--jobs auto`
- REQ-QB-05: System MUST force a mutating (worktree/executor) wave's concurrency to 1 when isolation is `none`, without capping a non-mutating (research/planning-only) wave
- REQ-QB-06: System MUST dispatch a planner per eligible item per DAG layer, providing the full batch task catalog and always requiring `depends_on`/`files_modified` frontmatter
- REQ-QB-07: System MUST recompute execution waves after each planning layer from the planners' declared dependencies/files
- REQ-QB-08: System MUST serialize worktree create/merge/cleanup while allowing already-created worktrees to run concurrently
- REQ-QB-09: System MUST merge items strictly in the deterministic wave order, never completion order
- REQ-QB-10: System MUST NOT call the STATE.md completion primitive for an item routed to `human_needed`
- REQ-QB-11: System MUST fail an item routed to `gaps_found`/`merge_failed`/`scope_violation` without rollback, without an automatic retry, and with its worktree preserved
- REQ-QB-12: System MUST support `--resume <batch-id>` to re-derive eligibility and dispatch only still-runnable items, refusing closed on an unknown batch id or a diverged base revision
---
## Quality Assurance Features
### 15. Nyquist Validation
**Purpose:** Map automated test coverage to phase requirements before any code is written. Named after the Nyquist sampling theorem — ensures a feedback signal exists for every requirement.
**Requirements:**
- REQ-NYQ-01: System MUST detect existing test infrastructure during plan-phase research
- REQ-NYQ-02: System MUST map each requirement to a specific test command
- REQ-NYQ-03: System MUST identify Wave 0 tasks (test scaffolding needed before implementation)
- REQ-NYQ-04: Plan checker MUST enforce Nyquist compliance as 8th verification dimension
- REQ-NYQ-05: System MUST support retroactive validation via `/gsd-validate-phase`
- REQ-NYQ-06: System MUST be disableable via `workflow.nyquist_validation: false`
**Produces:** `{phase}-VALIDATION.md` — Test coverage contract
**Retroactive Validation (`/gsd-validate-phase [N]`):**
- Scans implementation and maps requirements to tests
- Identifies gaps where requirements lack automated verification
- Spawns auditor to generate tests (max 3 attempts)
- Never modifies implementation code — only test files and VALIDATION.md
- Flags implementation bugs as escalations for user to address
---
### 16. Plan Checking
**Purpose:** Goal-backward verification that plans will achieve phase objectives before execution.
**Requirements:**
- REQ-PLANCK-01: System MUST verify plans against 8 quality dimensions
- REQ-PLANCK-02: System MUST loop up to 3 iterations until plans pass
- REQ-PLANCK-03: System MUST produce specific, actionable feedback on failures
- REQ-PLANCK-04: System MUST be disableable via `workflow.plan_check: false`
---
### 17. Post-Execution Verification
**Purpose:** Automated check that the codebase delivers what the phase promised.
**Requirements:**
- REQ-POSTVER-01: System MUST check against phase goals, not just task completion
- REQ-POSTVER-02: System MUST produce VERIFICATION.md with pass/fail analysis
- REQ-POSTVER-03: System MUST log issues for `/gsd-verify-work` to address
- REQ-POSTVER-04: System MUST be disableable via `workflow.verifier: false`
---
### 18. Node Repair
**Purpose:** Autonomous recovery when task verification fails during execution.
**Requirements:**
- REQ-REPAIR-01: System MUST analyze failure and choose one strategy: RETRY, DECOMPOSE, or PRUNE
- REQ-REPAIR-02: RETRY MUST attempt with a concrete adjustment
- REQ-REPAIR-03: DECOMPOSE MUST break task into smaller verifiable sub-steps
- REQ-REPAIR-04: PRUNE MUST remove unachievable tasks and escalate to user
- REQ-REPAIR-05: System MUST respect repair budget (default: 2 attempts per task)
- REQ-REPAIR-06: System MUST be configurable via `workflow.node_repair_budget` and `workflow.node_repair`
---
### 19. Health Validation
**Command:** `/gsd-health [--repair] [--backfill]`
**Purpose:** Validate `.planning/` directory integrity and auto-repair issues.
**Requirements:**
- REQ-HEALTH-01: System MUST check for missing required files
- REQ-HEALTH-02: System MUST validate configuration consistency
- REQ-HEALTH-03: System MUST detect orphaned plans without summaries
- REQ-HEALTH-04: System MUST check phase numbering and roadmap sync
- REQ-HEALTH-05: `--repair` flag MUST auto-fix recoverable issues except DESTRUCTIVE-risk ones, which it MUST report but never auto-apply
- REQ-HEALTH-06: `--backfill` flag MUST synthesize missing MILESTONES.md entries from archived milestone snapshots
---
### 20. Cross-Phase Regression Gate
**Purpose:** Prevent regressions from compounding across phases by running prior phases' test suites after execution.
**Requirements:**
- REQ-REGR-01: System MUST run test suites from all completed prior phases after phase execution
- REQ-REGR-02: System MUST report any test failures as cross-phase regressions
- REQ-REGR-03: Regressions MUST be surfaced before post-execution verification
- REQ-REGR-04: System MUST identify which prior phase's tests were broken
**When:** Runs automatically during `/gsd-execute-phase` before the verifier step.
---
### 21. Requirements Coverage Gate
**Purpose:** Ensure all phase requirements are covered by at least one plan before planning completes.
**Requirements:**
- REQ-COVGATE-01: System MUST extract all requirement IDs assigned to the phase from ROADMAP.md
- REQ-COVGATE-02: System MUST verify each requirement appears in at least one PLAN.md
- REQ-COVGATE-03: Uncovered requirements MUST block planning completion
- REQ-COVGATE-04: System MUST report which specific requirements lack plan coverage
**When:** Runs automatically at the end of `/gsd-plan-phase` after the plan checker loop.
---
### 4273. TDD-Applicability Predicate
**Purpose:** Give the workflow engine one code-owned computation for whether TDD's RED/GREEN/REFACTOR
procedure applies to a given plan, instead of restating the same precedence logic as hand-written prose in
each dispatch backend — a restatement that had already drifted between two backends (#4264, #4265). This
is Phase 1 of epic #4272 (ADR-3473's fourth application of the single-owner-predicate pattern): it ships
the isolated `phase.tdd-applicable` query verb only. Wiring `execute-phase.md` and its
executor-isolation-dispatch step to consume the verb instead of their own inline predicates is a later
phase of the same epic.
**Command:** `gsd-tools query phase.tdd-applicable <plan-file> [--cli-flag]`
**Requirements:**
- REQ-TDDA-01: System MUST resolve applicability via a fixed precedence: `--cli-flag` (explicit override) >
plan frontmatter `type: tdd` > any task in the plan carrying `tdd="true"` > project config
`workflow.tdd_mode`
- REQ-TDDA-02: System MUST report which precedence tier decided the outcome (`cli_flag`, `plan_frontmatter`,
`task_attribute`, `config`, or `none`) alongside the boolean result
- REQ-TDDA-03: System MUST emit JSON (`applicable`, `source`, `plan_type`, `config_tdd_mode`,
`cli_flag_present`) so callers can consume the decision without re-deriving it
---
## Context Engineering Features
### 22. Context Window Monitoring
**Purpose:** Prevent context rot by alerting both user and agent when context is running low.
**Requirements:**
- REQ-CTX-01: Statusline MUST display context usage percentage to user
- REQ-CTX-02: Context monitor MUST inject agent-facing warnings at the WARNING fire-point — ≤35% remaining by default, overridable per project via `hooks.context_warning_threshold`
- REQ-CTX-03: Context monitor MUST inject agent-facing warnings at the CRITICAL fire-point — ≤25% remaining by default, overridable per project via `hooks.context_critical_threshold`
- REQ-CTX-04: Warnings MUST debounce (5 tool uses between repeated warnings)
- REQ-CTX-05: Severity escalation (WARNING→CRITICAL) MUST bypass debounce
- REQ-CTX-06: Context monitor MUST differentiate GSD-active vs non-GSD-active projects
- REQ-CTX-07: Warnings MUST be advisory, never imperative commands that override user preferences
- REQ-CTX-08: All hooks MUST fail silently and never block tool execution
**Architecture:** Two-part bridge system:
1. Statusline writes metrics to `/tmp/claude-ctx-{session}.json`
2. Context monitor reads metrics and injects `additionalContext` warnings
---
### 23. Session Management
**Commands:** `/gsd-pause-work`, `/gsd-resume-work`, `/gsd-progress`
**Purpose:** Maintain project continuity across context resets and sessions.
**Requirements:**
- REQ-SESSION-01: Pause MUST save current position and next steps to `continue-here.md` and structured `HANDOFF.json`
- REQ-SESSION-02: Resume MUST restore full project context from HANDOFF.json (preferred) or state files (fallback)
- REQ-SESSION-03: Progress MUST show current position, next action, and overall completion
- REQ-SESSION-04: Progress MUST read all state files (STATE.md, ROADMAP.md, phase directories)
- REQ-SESSION-05: All session operations MUST work after `/clear` (context reset)
- REQ-SESSION-06: HANDOFF.json MUST include blockers, human actions pending, and in-progress task state
- REQ-SESSION-07: Resume MUST surface human actions and blockers immediately on session start
---
### 24. Session Reporting
**Command:** `/gsd-pause-work --report`
**Purpose:** Generate a structured post-session summary document capturing work performed, outcomes achieved, and estimated resource usage.
**Requirements:**
- REQ-REPORT-01: System MUST gather data from STATE.md, git log, and plan/summary files
- REQ-REPORT-02: System MUST include commits made, plans executed, and phases progressed
- REQ-REPORT-03: System MUST estimate token usage and cost based on session activity
- REQ-REPORT-04: System MUST include active blockers and decisions made
- REQ-REPORT-05: System MUST recommend next steps
**Produces:** `.planning/reports/SESSION_REPORT.md`
**Report Sections:**
- Session overview (duration, milestone, phase)
- Work performed (commits, plans, phases)
- Outcomes and deliverables
- Blockers and decisions
- Resource estimates (tokens, cost)
- Next steps recommendation
---
### 25. Multi-Agent Orchestration
**Purpose:** Coordinate specialized agents with fresh context windows for each task.
**Requirements:**
- REQ-ORCH-01: Each agent MUST receive a fresh context window
- REQ-ORCH-02: Orchestrators MUST be thin — spawn agents, collect results, route next
- REQ-ORCH-03: Context payload MUST include all relevant project artifacts
- REQ-ORCH-04: Parallel agents MUST be truly independent (no shared mutable state)
- REQ-ORCH-05: Agent results MUST be written to disk before orchestrator processes them
- REQ-ORCH-06: Failed agents MUST be detected (spot-check actual output vs reported failure)
---
### 26. Model Profiles
**Command:** `/gsd-config --profile <quality|balanced|budget|adaptive|inherit>`
**Purpose:** Control which AI model each agent uses, balancing quality vs cost.
**Requirements:**
- REQ-MODEL-01: System MUST support 4 profiles: `quality`, `balanced`, `budget`, `inherit`
- REQ-MODEL-02: Each profile MUST define model tier per agent (see profile table)
- REQ-MODEL-03: Per-agent overrides MUST take precedence over profile
- REQ-MODEL-04: `inherit` profile MUST defer to runtime's current model selection
- REQ-MODEL-04a: `inherit` profile MUST be used when running non-Anthropic providers (OpenRouter, local models) to avoid unexpected API costs
- REQ-MODEL-05: Profile switch MUST be programmatic (script, not LLM-driven)
- REQ-MODEL-06: Model resolution MUST happen once per orchestration, not per spawn
**Profile Assignments:**
| Agent | `quality` | `balanced` | `budget` | `inherit` |
|-------|-----------|------------|----------|-----------|
| gsd-planner | Opus | Opus | Sonnet | Inherit |
| gsd-roadmapper | Opus | Sonnet | Sonnet | Inherit |
| gsd-executor | Opus | Sonnet | Sonnet | Inherit |
| gsd-phase-researcher | Opus | Sonnet | Haiku | Inherit |
| gsd-project-researcher | Opus | Sonnet | Haiku | Inherit |
| gsd-research-synthesizer | Sonnet | Sonnet | Haiku | Inherit |
| gsd-debugger | Opus | Sonnet | Sonnet | Inherit |
| gsd-codebase-mapper | Sonnet | Haiku | Haiku | Inherit |
| gsd-verifier | Sonnet | Sonnet | Haiku | Inherit |
| gsd-plan-checker | Sonnet | Sonnet | Haiku | Inherit |
| gsd-integration-checker | Sonnet | Sonnet | Haiku | Inherit |
| gsd-nyquist-auditor | Sonnet | Sonnet | Haiku | Inherit |
---
### 4139. Compact Content Mode
**Config:** `workflow.compact_content: false`
**Purpose:** Per-project opt-in to token-minimized variants of GSD's own shipped prompt
content — workflow instructions, planning-artifact templates, and non-Claude agent-persona
payloads — so the always-loaded instruction window leaves more of the model's attention on
the developer's own code (ADR-4139 Decision 2: finite attention, not per-invocation price,
since prompt caching already discounts the latter).
Nothing is compressed at runtime. Compact variants are hand-authored, reviewed files sitting
beside their canonical siblings; the config key only chooses which one gets read. With the
key off (the default), every covered workflow, template, and agent persona behaves exactly as
it did before this feature existed.
**Requirements:**
- REQ-COMPACT-01: System MUST default `workflow.compact_content` to `false` — off costs
nothing and changes no existing behavior
- REQ-COMPACT-02: Eagerly `@`-included workflow files MUST keep their host-guaranteed load;
compactness on this stream comes from a spine + deferred `detail/*.md` elaboration, never
from converting the `@`-include itself
- REQ-COMPACT-03: A missed runtime `Read` of a deferred elaboration or compact variant MUST
degrade to a complete, correct, terser state — never to a state with no instructions
- REQ-COMPACT-04: No compact variant MAY weaken or remove protected content (guardrails,
output-format contracts, few-shot examples, security language, structural headings)
- REQ-COMPACT-05: An agent with no compact persona variant registered MUST fall back to its
canonical persona and disclose the fallback inside the served payload, never fail or serve
nothing
- REQ-COMPACT-06: `/gsd-new-project` MUST ask the question and persist the answer;
`/gsd-settings` and `/gsd-config` MUST toggle it on an already-initialized project
**Config:**
| Setting | Type | Default | Description |
|---------|------|---------|-------------|
| `workflow.compact_content` | boolean | `false` | When `true`, loads token-minimized instruction/template/agent-persona variants wherever one is registered; falls back to canonical content everywhere else |
**See also:** [ADR-4139](../adr/4139-compact-content-seam.md), [CONFIGURATION.md](../CONFIGURATION.md#workflow-toggles), [USER-GUIDE.md](../USER-GUIDE.md)
---
## Brownfield Features
### 27. Codebase Mapping
**Command:** `/gsd-map-codebase [area]`
**Purpose:** Analyze an existing codebase before starting a new project or as the mapping handoff from `/gsd-onboard`, so GSD understands what exists.
**Requirements:**
- REQ-MAP-01: System MUST spawn parallel mapper agents for each analysis area
- REQ-MAP-02: System MUST produce structured documents in `.planning/codebase/`
- REQ-MAP-03: System MUST detect: tech stack, architecture patterns, coding conventions, concerns
- REQ-MAP-04: Subsequent `/gsd-new-project` MUST load codebase mapping and focus questions on what's being added
- REQ-MAP-05: Optional `[area]` argument MUST scope mapping to a specific area
**Produces:**
| Document | Content |
|----------|---------|
| `STACK.md` | Languages, frameworks, databases, infrastructure |
| `ARCHITECTURE.md` | Patterns, layers, data flow, boundaries |
| `CONVENTIONS.md` | Naming, file organization, code style, testing patterns |
| `CONCERNS.md` | Technical debt, security issues, performance bottlenecks |
| `STRUCTURE.md` | Directory layout and file organization |
| `TESTING.md` | Test infrastructure, coverage, patterns |
| `INTEGRATIONS.md` | External services, APIs, third-party dependencies |
**Incremental remap — `--paths` (#2003):** The mapper accepts an optional
`--paths <p1,p2,...>` scope hint. When provided, it restricts exploration
to the listed repo-relative prefixes instead of scanning the whole tree.
This is the pathway used by the post-execute codebase-drift gate to refresh
only the subtrees the phase actually changed. Each produced document carries
`last_mapped_commit` in its YAML frontmatter so drift can be measured
against the mapping point, not HEAD.
---
### 27b. Existing Codebase Onboarding
**Command:** `/gsd-onboard [--fast] [--text]`
**Purpose:** Guide first-time setup for an existing repository by checking brownfield state, routing through codebase mapping and docs ingest, then handing off to project initialization without silently overwriting planning artifacts.
**Requirements:**
- REQ-ONBOARD-01: System MUST detect existing code, package manifests, planning documents, partial `.planning/` state, and complete or missing codebase-map files.
- REQ-ONBOARD-02: System MUST hand off to `/gsd-map-codebase` or `/gsd-map-codebase --fast` when brownfield code lacks the required `.planning/codebase/` map files; fast-map readiness is partial and MUST NOT be treated as sufficient for `/gsd-new-project`.
- REQ-ONBOARD-03: System MUST offer `/gsd-ingest-docs` before `/gsd-new-project` when ADR/PRD/SPEC/RFC candidates exist and no project exists.
- REQ-ONBOARD-04: System MUST refuse to report onboarding complete until `PROJECT.md`, `REQUIREMENTS.md`, `ROADMAP.md`, and `STATE.md` all exist.
- REQ-ONBOARD-05: System MUST create or confirm `.planning/onboarding/SUMMARY.md` only after project setup exists.
- REQ-ONBOARD-06: System MUST support `--text` for numbered plain-text gates on runtimes without interactive menus.
**Produces:**
| Artifact | Description |
|----------|-------------|
| `.planning/codebase/` | Codebase map produced by the `/gsd-map-codebase` handoff |
| `.planning/PROJECT.md`, `REQUIREMENTS.md`, `ROADMAP.md`, `STATE.md` | Planning setup produced by `/gsd-new-project` or `/gsd-ingest-docs` |
| `.planning/onboarding/SUMMARY.md` | Onboarding status, artifact index, and next-command summary |
---
### 27a. Post-Execute Codebase Drift Detection
**Introduced by:** #2003
**Trigger:** Runs automatically at the end of every `/gsd-execute-phase`
**Configuration:**
- `workflow.drift_threshold` (integer, default `3`) — minimum new
structural elements before the gate acts.
- `workflow.drift_action` (`warn` | `auto-remap`, default `warn`) —
warn-only or spawn `gsd-codebase-mapper` with `--paths` scoped to
affected subtrees.
**What counts as drift:**
- New directory outside mapped paths
- New barrel export at `(packages|apps)/*/src/index.*`
- New migration file (supabase/prisma/drizzle/src/migrations/…)
- New route module under `routes/` or `api/`
**Non-blocking guarantee:** any internal failure (missing STRUCTURE.md,
git errors, mapper spawn failure) logs a single line and the phase
continues. Drift detection cannot fail verification.
**Requirements:**
- REQ-DRIFT-01: System MUST detect the four drift categories from `git diff
--name-status last_mapped_commit..HEAD`
- REQ-DRIFT-02: Action fires only when element count ≥ `workflow.drift_threshold`
- REQ-DRIFT-03: `warn` action MUST NOT spawn any agent
- REQ-DRIFT-04: `auto-remap` action MUST pass sanitized `--paths` to the mapper
- REQ-DRIFT-05: Detection/remap failure MUST be non-blocking for `/gsd-execute-phase`
- REQ-DRIFT-06: `last_mapped_commit` round-trip through YAML frontmatter
on each `.planning/codebase/*.md` file
---
## Utility Features
### 28. Debug System
**Command:** `/gsd-debug [description]`
**Purpose:** Systematic debugging with persistent state across context resets.
**Requirements:**
- REQ-DEBUG-01: System MUST create debug session file in `.planning/debug/`
- REQ-DEBUG-02: System MUST track hypotheses, evidence, and eliminated theories
- REQ-DEBUG-03: System MUST persist state so debugging survives context resets
- REQ-DEBUG-04: System MUST require human verification before marking resolved
- REQ-DEBUG-05: Resolved sessions MUST append to `.planning/debug/knowledge-base.md`
- REQ-DEBUG-06: Knowledge base MUST be consulted on new debug sessions to prevent re-investigation
**Debug Session States:** `gathering` → `investigating` → `fixing` → `verifying` → `awaiting_human_verify` → `resolved`
---
### 29. Todo Management
**Commands:** `/gsd-capture [desc]`, `/gsd-capture --list`
**Purpose:** Capture ideas and tasks during sessions for later work.
**Requirements:**
- REQ-TODO-01: System MUST capture todo from current conversation context
- REQ-TODO-02: Todos MUST be stored in `.planning/todos/pending/`
- REQ-TODO-03: Completed todos MUST move to `.planning/todos/completed/`
- REQ-TODO-04: Check-todos MUST list all pending items with selection to work on one
---
### 30. Statistics Dashboard
**Command:** `/gsd-stats`
**Purpose:** Display project metrics — phases, plans, requirements, git history, and timeline.
**Requirements:**
- REQ-STATS-01: System MUST show phase/plan completion counts
- REQ-STATS-02: System MUST show requirement coverage
- REQ-STATS-03: System MUST show git commit metrics
- REQ-STATS-04: System MUST support multiple output formats (json, table, bar)
---
### 31. Update System
**Command:** `/gsd-update`
**Purpose:** Update GSD to the latest version with changelog preview.
**Requirements:**
- REQ-UPDATE-01: System MUST check for new versions via npm
- REQ-UPDATE-02: System MUST display changelog for new version before updating
- REQ-UPDATE-03: System MUST be runtime-aware and target the correct directory
- REQ-UPDATE-04: System MUST back up locally modified files to `gsd-local-patches/`
- REQ-UPDATE-05: `/gsd-update --reapply` MUST restore local modifications after update
- REQ-UPDATE-06: `/gsd-update --next` (alias `--rc`) MUST target the `@next` RC dist-tag for version check and install; omitting the flag MUST keep `@latest` behavior unchanged (ADR #660)
- REQ-UPDATE-07: System MUST back up user-added files found inside GSD-managed directories to `gsd-user-files-backup/` before the clean install
- REQ-UPDATE-08: When that backup is non-empty, the update MUST offer an explicit restore choice before finishing, and MUST leave the backup intact whichever way the user answers
- REQ-UPDATE-09: A restore MUST NOT overwrite a path the newly installed release ships, MUST NOT overwrite a different file already on disk, and MUST report best-effort compatibility warnings for restored files without blocking on them
---
### 32. Settings Management
**Command:** `/gsd-settings`
**Purpose:** Interactive configuration of workflow toggles and model profile.
**Requirements:**
- REQ-SETTINGS-01: System MUST present current settings with toggle options
- REQ-SETTINGS-02: System MUST update `.planning/config.json`
- REQ-SETTINGS-03: System MUST support saving as global defaults (`~/.gsd/defaults.json`)
**Configurable Settings:**
| Setting | Type | Default | Description |
|---------|------|---------|-------------|
| `mode` | enum | `interactive` | `interactive` or `yolo` (auto-approve) |
| `granularity` | enum | `standard` | `coarse`, `standard`, or `fine` |
| `model_profile` | enum | `balanced` | `quality`, `balanced`, `budget`, or `inherit` |
| `models.<phase_type>` | enum | (none) | Per-phase-type tier override (`planning`, `discuss`, `research`, `execution`, `verification`, `completion`). Values: `opus`, `sonnet`, `haiku`, `inherit`. Coarse phase-level tuning that wins over `model_profile` but loses to per-agent `model_overrides`. See [CONFIGURATION.md](CONFIGURATION.md#per-phase-type-models-models--added-in-v140). Added in v1.40 |
| `granularities.<phase_type>` | enum | (none) | Per-phase-type granularity override (`planning`, `discuss`, `research`, `execution`, `verification`, `completion`). Values: `coarse`, `standard`, `fine`. Mirrors `models.<phase_type>` for granularity. See [CONFIGURATION.md](CONFIGURATION.md#core-settings). Added in v1.43 ([#68](https://github.com/open-gsd/gsd-core/issues/68)). `/gsd-plan-phase --granularity <coarse\|standard\|fine>` overrides all config-based granularity for a single invocation (takes precedence over `granularities.planning`, top-level `granularity`, and `planning.granularity`). ([#703](https://github.com/open-gsd/gsd-core/issues/703)) |
| `dynamic_routing.enabled` | boolean | `false` | Master switch for failure-tier escalation. When `true`, agents resolve to `tier_models[default_tier]` and escalate one tier on orchestrator-detected soft failure. Capped by `max_escalations`. See [CONFIGURATION.md](CONFIGURATION.md#dynamic-routing-with-failure-tier-escalation-dynamic_routing--added-in-v140). Added in v1.40 |
| `workflow.research` | boolean | `true` | Domain research before planning |
| `workflow.plan_check` | boolean | `true` | Plan verification loop |
| `workflow.verifier` | boolean | `true` | Post-execution verification |
| `workflow.auto_advance` | boolean | `false` | Auto-chain discuss→plan→execute |
| `workflow.nyquist_validation` | boolean | `true` | Nyquist test coverage mapping |
| `workflow.ui_phase` | boolean | `true` | UI design contract generation |
| `workflow.ui_safety_gate` | boolean | `true` | Prompt for ui-phase on frontend phases |
| `workflow.node_repair` | boolean | `true` | Autonomous task repair |
| `workflow.node_repair_budget` | number | `2` | Max repair attempts per task |
| `planning.commit_docs` | boolean | `true` | Commit `.planning/` files to git |
| `planning.search_gitignored` | boolean | `false` | Include gitignored files in searches |
| `parallelization.enabled` | boolean | `true` | Run independent plans simultaneously |
| `git.branching_strategy` | enum | `none` | `none`, `phase`, or `milestone` |
---
### 33. Test Generation
**Command:** `/gsd-add-tests [N]`
**Purpose:** Generate tests for a completed phase based on UAT criteria and implementation.
**Requirements:**
- REQ-TEST-01: System MUST analyze completed phase implementation
- REQ-TEST-02: System MUST generate tests based on UAT criteria and acceptance criteria
- REQ-TEST-03: System MUST use existing test infrastructure patterns
---
## Infrastructure Features
> **Looking for a third-party add-on instead?** See the [GSD Community Capability Registry & EoS Registry](registries/README.md) — non-endorsing discoverability catalogs for community-contributed Capabilities and EoS host integrations.
### 34. Git Integration
**Purpose:** Atomic commits, branching strategies, and clean history management.
**Requirements:**
- REQ-GIT-01: Each task MUST get its own atomic commit
- REQ-GIT-02: Commit messages MUST follow structured format: `type(scope): description`
- REQ-GIT-03: System MUST support 3 branching strategies: `none`, `phase`, `milestone`
- REQ-GIT-04: Phase strategy MUST create one branch per phase
- REQ-GIT-05: Milestone strategy MUST create one branch per milestone
- REQ-GIT-06: Complete-milestone MUST offer squash merge (recommended) or merge with history
- REQ-GIT-07: System MUST respect `commit_docs` setting for `.planning/` files
- REQ-GIT-08: System MUST auto-detect `.planning/` in `.gitignore` and skip commits
**Commit Format:**
```
type(phase-plan): description
# Examples:
docs(08-02): complete user registration plan
feat(08-02): add email confirmation flow
fix(03-01): correct auth token expiry
```
---
### 35. CLI Tools
**Purpose:** Programmatic utilities for workflows and agents, replacing repetitive inline bash patterns.
**Requirements:**
- REQ-CLI-01: System MUST provide atomic commands for state, config, phase, roadmap operations
- REQ-CLI-02: System MUST provide compound `init` commands that load all context for each workflow
- REQ-CLI-03: System MUST support `--raw` flag for machine-readable output
- REQ-CLI-04: System MUST support `--cwd` flag for sandboxed subagent operation
- REQ-CLI-05: All operations MUST use forward-slash paths on Windows
**Command Categories:** State (11 subcommands), Phase (5), Roadmap (3), Verify (8), Template (2), Frontmatter (4), Scaffold (4), Init (12), Validate (2), Progress, Stats, Todo
---
### 36. Multi-Runtime Support
**Purpose:** Run GSD across multiple AI coding agent runtimes.
**Requirements:**
- REQ-RUNTIME-01: System MUST support Claude Code, OpenCode, Kilo, Codex, Copilot, Antigravity, Trae, Cline, Augment Code, CodeBuddy, Qwen Code
- REQ-RUNTIME-02: Installer MUST transform content per runtime (tool names, paths, frontmatter)
- REQ-RUNTIME-03: Installer MUST support interactive and non-interactive (`--claude --global`) modes
- REQ-RUNTIME-04: Installer MUST support both global and local installation
- REQ-RUNTIME-05: Uninstall MUST cleanly remove all GSD files without affecting other configurations
- REQ-RUNTIME-06: Installer MUST handle platform differences (Windows, macOS, Linux, WSL, Docker)
- REQ-RUNTIME-07: Runtimes with lifecycle hook support MUST register per-turn context-headroom tracking events at install time
- REQ-RUNTIME-08: Native packaging manifests MUST be version-stamped and enable runtime-native install/update/uninstall flows
**Runtime Transformations:**
| Aspect | Claude Code | OpenCode | Kilo | Codex | Copilot | Antigravity | Cursor | Trae | Cline | Augment | CodeBuddy | Qwen Code |
|--------|------------|----------|-------|-------|---------|-------------|--------|------|-------|---------|-----------|-----------|
| Commands | Slash commands | Slash commands | Slash commands | Skills (TOML) | Slash commands | Skills | Skills + Slash commands | Skills | Rules | Skills + Slash commands | Slash commands | Skills |
| Agent format | Claude native | `mode: subagent` | `mode: subagent` | Skills | Tool mapping | Skills | Skills | Skills | Rules | Skills | Skills | Skills |
| Skills emission | N/A | On-demand SKILL.md (1.4.0) | On-demand SKILL.md (1.4.0) | `/skills` picker (1.4.0) | N/A | N/A | SKILL.md | N/A | On-demand SKILL.md (1.4.0) | N/A | N/A | N/A |
| Hook events | `SessionStart`, `PreToolUse`, `PostToolUse`, `SubagentStop`, `Stop`, `PreCompact`, `FileChanged` | N/A | N/A | `SessionStart`, `SubagentStart`, `Stop`, `PostToolUse` | `sessionStart` | N/A | `sessionStart`, `postToolUse` | N/A | `PreToolUse` | N/A | N/A | `SessionStart`, `PreToolUse`, `PostToolUse`, `SubagentStop`, `Stop`, `PreCompact` |
| Config | `settings.json` | `opencode.json(c)` | `kilo.json(c)` | TOML | Instructions | Config | Config | Config | `.clinerules` | Config | Config | Config |
**Cursor artifact surfaces:** `gsd install --cursor` writes two artifact kinds:
- `~/.cursor/skills/gsd-<name>/SKILL.md` — rich skills with YAML frontmatter, Cursor tool-name mapping, and adapter context header (existing surface)
- `~/.cursor/commands/gsd-<name>.md` — plain markdown slash commands (no frontmatter) invocable via `/` in the Agent input (Cursor 1.6+)
**Native skills emission (1.4.0):** Three runtimes now emit GSD as on-demand native skills (`skills/<name>/SKILL.md`) at install time, in addition to their existing command and agent surfaces. Skills respect the active install profile and are removed on uninstall.
- **Cline** (global installs, Cline >= v3.48.0) — emits skills alongside the existing `.clinerules/` directory
- **Kilo** — emits skills alongside `command/` and `agents/`
- **OpenCode** — emits skills alongside its existing surfaces
**New slash-command surfaces (1.4.0):**
- **CodeBuddy** — `/gsd-*` slash commands written to `~/.codebuddy/commands/`
- **Augment** — `commands/gsd-<name>.md` written to `~/.augment/commands/`
- **Cursor** (Cursor >= 1.6) — `.cursor/commands/gsd-<name>.md` so GSD appears in the `/` command menu
**Cross-runtime lifecycle hooks (1.4.0):** Each supported runtime registers lifecycle hook events for per-turn context-headroom tracking and workflow state management. Notable registrations:
- **Claude Code:** `SubagentStop`, `Stop`, `PreCompact` (context-headroom warnings), `FileChanged` (hot-reloads `.planning/config.json` mid-session)
- **Qwen Code:** `SubagentStop`, `Stop`, `PreCompact`
- **Codex:** `SubagentStart`, `Stop`, `PostToolUse` (new in 1.4.0); on Windows the `SessionStart` hook entry gains a `commandWindows` field so the `.cmd` shim is used for native execution
- **Cline:** `PreToolUse`
- **Cursor:** `sessionStart` (injects workflow state), `postToolUse` (nudges `.planning` updates)
- **Copilot:** `sessionStart`
**Runtime-specific enrichments (1.4.0):**
- Codex emits `service_tier: flex` for light-tier agents; GSD skills appear in the Codex `/skills` picker via `SKILL.md` (no `agents/openai.yaml` sidecar is emitted — doing so caused duplicate autocomplete entries, #1326)
**Native packaging:**
- **Claude Code:** GSD Core ships a `.claude-plugin/plugin.json` manifest, enabling installation and lifecycle management via `claude plugin install|enable|disable|update gsd-core`. Commands load under the `/gsd-core:` namespace (e.g. `/gsd-core:plan-phase`), avoiding slash-command collisions with the classic npm installer which uses `/gsd:`. Always-on guard and update hooks are wired automatically via `hooks/hooks.json`. The plugin path is additive — the npm installer (`npx @opengsd/gsd-core`) remains fully supported.
---
### 37. Hook System
**Purpose:** Runtime event hooks for context monitoring, status display, and update checking.
**Requirements:**
- REQ-HOOK-01: Statusline MUST display model, current task, directory, and context usage
- REQ-HOOK-02: Context monitor MUST inject agent-facing warnings at threshold levels
- REQ-HOOK-03: Update checker MUST run in background on session start
- REQ-HOOK-04: All hooks MUST respect `CLAUDE_CONFIG_DIR` env var
- REQ-HOOK-05: All hooks MUST include 3-second stdin timeout guard
- REQ-HOOK-06: All hooks MUST fail silently on any error
- REQ-HOOK-07: Context usage MUST normalize for autocompact buffer (16.5% reserved)
- REQ-HOOK-08: Update banner MUST be opt-in and silent unless an update is available (PR #2795)
**Statusline Display:**
```text
[⬆ /gsd-update │] model │ [current task │] directory [█████░░░░░ 50%]
```
Color coding: <50% green, <65% yellow, <80% orange, ≥80% red with skull emoji
**Update Banner (opt-in, when GSD statusline isn't used):**
When the user declines (or keeps a non-GSD) statusline, the installer offers a SessionStart banner that surfaces update availability without occupying statusline real estate. The banner reads `~/.cache/gsd/gsd-update-check.json` (written by `gsd-check-update-worker.js`) and emits one line only when an update is available:
```text
GSD update available: 1.39.0 → 1.40.0. Run /gsd-update.
```
The banner is silent when up-to-date and rate-limits "check failed" diagnostics to once per 24 hours. Removed cleanly by `npx @opengsd/gsd-core --uninstall` or by deleting the SessionStart entry that references `gsd-update-banner.js`.
---
### 38. Developer Profiling
**Command:** `/gsd-profile-user [--questionnaire] [--refresh]`
**Purpose:** Analyze Claude Code session history to build behavioral profiles across 8 dimensions, generating artifacts that personalize Claude's responses to the developer's style.
**Dimensions:**
1. Communication style (terse vs verbose, formal vs casual)
2. Decision patterns (rapid vs deliberate, risk tolerance)
3. Debugging approach (systematic vs intuitive, log preference)
4. UX preferences (design sensibility, accessibility awareness)
5. Vendor/technology choices (framework preferences, ecosystem familiarity)
6. Frustration triggers (what causes friction in workflows)
7. Learning style (documentation vs examples, depth preference)
8. Explanation depth (high-level vs implementation detail)
**Generated Artifacts:**
- `USER-PROFILE.md` — Full behavioral profile with evidence citations
- `CLAUDE.md` profile section — Auto-discovered by Claude Code
**Flags:**
- `--questionnaire` — Interactive questionnaire fallback when session history is unavailable
- `--refresh` — Re-analyze sessions and regenerate profile
**Pipeline Modules:**
- `profile-pipeline.cjs` — Session scanning, message extraction, sampling
- `profile-output.cjs` — Profile rendering, questionnaire, artifact generation
- `gsd-user-profiler` agent — Behavioral analysis from session data
**Requirements:**
- REQ-PROF-01: Session analysis MUST cover at least 8 behavioral dimensions
- REQ-PROF-02: Profile MUST cite evidence from actual session messages
- REQ-PROF-03: Questionnaire MUST be available as fallback when no session history exists
- REQ-PROF-04: Generated artifacts MUST be discoverable by Claude Code (CLAUDE.md integration)
---
### 39. Execution Hardening
**Purpose:** Three additive quality improvements to the execution pipeline that catch cross-plan failures before they cascade.
**Components:**
**1. Pre-Wave Dependency Check** (execute-phase)
Before spawning wave N+1, verify key-links from prior wave artifacts exist and are wired correctly. Catches cross-plan dependency gaps before they cascade into downstream failures.
**2. Cross-Plan Data Contracts — Dimension 9** (plan-checker)
New analysis dimension that checks plans sharing data pipelines have compatible transformations. Flags when one plan strips data that another plan needs in its original form.
**3. Export-Level Spot Check** (verify-phase)
After Level 3 wiring verification passes, spot-check individual exports for actual usage. Catches dead stores that exist in wired files but are never called.
**Requirements:**
- REQ-HARD-01: Pre-wave check MUST verify key-links from all prior wave artifacts before spawning next wave
- REQ-HARD-02: Cross-plan contract check MUST detect incompatible data transformations between plans
- REQ-HARD-03: Export spot-check MUST identify dead stores in wired files
---
### 40. Verification Debt Tracking
**Command:** `/gsd-audit-uat`
**Purpose:** Prevent silent loss of UAT/verification items when projects advance past phases with outstanding tests. Surfaces verification debt across all prior phases so items are never forgotten.
**Components:**
**1. Cross-Phase Health Check** (progress.md Step 1.6)
Every `/gsd-progress` call scans ALL phases in the current milestone for outstanding items (pending, skipped, blocked, human_needed, gaps_found). Displays a non-blocking warning section with actionable links.
A verification report counts as outstanding under EITHER terminal non-passing status: `human_needed` contributes its `human_verification:` entries, and `gaps_found` contributes both its `human_verification:` and its `gaps:` entries, excluding any already closed. What counts as closed is per key: a `gaps:` entry closes on `status: resolved` and nothing else — the same rule the `## Gaps` markdown reader applies, so one authored entry cannot read closed in one reader and open in the other — while a `human_verification:` entry also closes on a bare `resolution:` field, provided no `status:` contradicts it (#3850).
**2. `status: partial`** (verify-work.md, UAT.md)
New UAT status that distinguishes between "session ended" and "all tests resolved". Prevents `status: complete` when tests are still pending, blocked, or skipped without reason.
**3. `result: blocked` with `blocked_by` tag** (verify-work.md, UAT.md)
New test result type for tests blocked by external dependencies (server, physical device, release build, third-party services). Categorized separately from skipped tests.
**4. HUMAN-UAT.md Persistence** (execute-phase.md)
When verification returns `human_needed`, items are persisted as a trackable HUMAN-UAT.md file with `status: partial`. Feeds into the cross-phase health check and audit systems.
**5. Phase Completion Warnings** (phase.cjs, transition.md)
`phase complete` CLI returns verification debt warnings in its JSON output. Transition workflow surfaces outstanding items before confirmation.
**Requirements:**
- REQ-DEBT-01: System MUST surface outstanding UAT/verification items from ALL prior phases in `/gsd-progress`
- REQ-DEBT-02: System MUST distinguish incomplete testing (partial) from completed testing (complete)
- REQ-DEBT-03: System MUST categorize blocked tests with `blocked_by` tags
- REQ-DEBT-04: System MUST persist human_needed verification items as trackable UAT files
- REQ-DEBT-05: System MUST warn (non-blocking) during phase completion and transition when verification debt exists
- REQ-DEBT-06: `/gsd-audit-uat` MUST scan all phases, categorize items by testability, and produce a human test plan
---
## v1.27 Features
### 41. Fast Mode
**Command:** `/gsd-fast [task description]`
**Purpose:** Execute trivial tasks inline without spawning subagents or generating PLAN.md files. For tasks too small to justify planning overhead: typo fixes, config changes, small refactors, forgotten commits, simple additions.
**Requirements:**
- REQ-FAST-01: System MUST execute the task directly in the current context without subagents
- REQ-FAST-02: System MUST produce an atomic git commit for the change
- REQ-FAST-03: System MUST track the task in `.planning/quick/` for state consistency
- REQ-FAST-04: System MUST NOT be used for tasks requiring research, multi-step planning, or verification
**When to use vs `/gsd-quick`:**
- `/gsd-fast` — One-sentence tasks executable in under 2 minutes (typo, config change, small addition)
- `/gsd-quick` — Anything needing research, multi-step planning, or verification
---
### 42. Cross-AI Peer Review
**Command:** `/gsd-review --phase N [--claude] [--codex] [--coderabbit] [--opencode] [--qwen] [--cursor] [--agy] [--antigravity] [--ollama] [--lm-studio] [--llama-cpp] [--kimi-code] [--all]`
**Purpose:** Invoke external AI CLIs (Claude, Codex, CodeRabbit, OpenCode, Qwen Code, Cursor, Antigravity, Kimi Code) and local OpenAI-compatible servers (Ollama, LM Studio, llama.cpp) to independently review phase plans. Produces structured REVIEWS.md with per-reviewer feedback.
Each reviewer is a **declared lane**: its binary, prompt and output channels, timeout, availability probe, and empty-output policy come from a capability manifest rather than hand-written per-CLI logic, so a reviewer can be shipped as an installable capability instead of a core change.
**Requirements:**
- REQ-REVIEW-01: System MUST detect available AI CLIs on the system
- REQ-REVIEW-02: System MUST build a structured review prompt from phase plans
- REQ-REVIEW-03: System MUST invoke each selected CLI independently
- REQ-REVIEW-04: System MUST collect responses and produce `REVIEWS.md`
- REQ-REVIEW-05: Reviews MUST be consumable by `/gsd-plan-phase --reviews`
- REQ-REVIEW-06: System MUST support project-level no-flag defaults via `review.default_reviewers`
- REQ-REVIEW-07: Reviewer precedence MUST be explicit flags > `--all` > `review.default_reviewers` > all detected reviewers
**Produces:** `{phase}-REVIEWS.md` — Per-reviewer structured feedback
**User configuration note:**
- Set `review.default_reviewers` in `.planning/config.json` (or via `gsd config-set`) to control no-flag `/gsd-review` fan-out.
- `review.default_reviewers` may include configured `review.reviewer_instances` names; each instance runs as an independent reviewer identity backed by its configured adapter/model. Instance names are not CLI flags.
- Use `--all` for a full pre-merge sweep without changing project defaults.
- For local model servers with small context windows, set `review.max_prompt_tokens_per_reviewer` to auto-trim prompts per reviewer — see [Prompt budgets for small-context reviewers](../docs/CONFIGURATION.md#prompt-budgets-for-small-context-reviewers) in CONFIGURATION.md.
**Why record which model produced a review (#2295):** `reviewers:` in the frontmatter recorded which CLIs ran, but not which model each one resolved to. Without a pin, the model is whatever the CLI's own config or internal default happens to pick, so a "Codex vs Antigravity" comparison could quietly be a frontier model against a cheap-tier default with nothing in the record to say so — and a CLI update, or an unrelated config edit, could silently make past and future reviews incomparable.
The fix records the model *and its provenance*. Provenance is what makes the value trustworthy: `pinned` (from `review.models.<slug>`) is certain, while `banner` and `transcript` are recovered from third-party CLI output this project does not own — a startup banner or an undocumented session log.
That third-party dependence is a real trade-off, held honestly rather than papered over: the `banner` and `transcript` arms read formats GSD does not control, so they are best-effort by design and degrade to `unknown` rather than guessing or failing the run. A recorded `unknown` is a real answer — a wrong model name attributed to a review would be worse than none.
---
### 43. Backlog Parking Lot
**Commands:** `/gsd-capture --backlog <description>`, `/gsd-review-backlog`, `/gsd-capture --seed <idea>`, `/gsd-capture --list-seeds [status]`
**Purpose:** Capture ideas that aren't ready for active planning. Backlog items use 999.x numbering to stay outside the active phase sequence. Seeds are forward-looking ideas with trigger conditions that surface automatically at the right milestone. `--list-seeds` provides a read-only audit of all parked seeds (with optional status filter) without waiting for the next milestone.
**Requirements:**
- REQ-BACKLOG-01: Backlog items MUST use 999.x numbering to stay outside active phase sequence
- REQ-BACKLOG-02: Phase directories MUST be created immediately so `/gsd-discuss-phase` and `/gsd-plan-phase` work on them
- REQ-BACKLOG-03: `/gsd-review-backlog` MUST support promote, keep, and remove actions per item
- REQ-BACKLOG-04: Promoted items MUST be renumbered into the active milestone sequence
- REQ-SEED-01: Seeds MUST capture the full WHY and WHEN to surface conditions
- REQ-SEED-02: `/gsd-new-milestone` MUST scan seeds and present matches
- REQ-SEED-03: `/gsd-capture --list-seeds` MUST list seeds with status, scope, and trigger for audit, with optional status filtering
**Produces:**
| Artifact | Description |
|----------|-------------|
| `.planning/phases/999.x-slug/` | Backlog item directory |
| `.planning/seeds/SEED-YYMMDD-xxx-slug.md` | Seed with trigger conditions |
---
### 44. Persistent Context Threads
**Command:** `/gsd-thread [name | description]`
**Purpose:** Lightweight cross-session knowledge stores for work that spans multiple sessions but doesn't belong to any specific phase. Lighter weight than `/gsd-pause-work` — no phase state, no plan context.
**Requirements:**
- REQ-THREAD-01: System MUST support create, list, and resume modes
- REQ-THREAD-02: Threads MUST be stored in `.planning/threads/` as markdown files
- REQ-THREAD-03: Thread files MUST include Goal, Context, References, and Next Steps sections
- REQ-THREAD-04: Resuming a thread MUST load its full context into the current session
- REQ-THREAD-05: Threads MUST be promotable to phases or backlog items
**Produces:** `.planning/threads/{slug}.md` — Persistent context thread
---
### 45. PR Branch Filtering
**Command:** `/gsd-pr-branch [target branch]`
**Purpose:** Create a clean branch suitable for pull requests by filtering out `.planning/` commits. Reviewers see only code changes, not GSD planning artifacts.
**Requirements:**
- REQ-PRBRANCH-01: System MUST identify commits that only modify `.planning/` files
- REQ-PRBRANCH-02: System MUST create a new branch with planning commits filtered out
- REQ-PRBRANCH-03: Code changes MUST be preserved exactly as committed
- REQ-PRBRANCH-04: System MUST NOT delete a `.planning/` path the target branch already tracks
- REQ-PRBRANCH-05: Verification MUST assert against the active filter mode's contract, not an unconditional zero
**Filter modes.** `planning.pr_strict` selects what "filtered" means. The default mode treats `.planning/` as two populations: structural state that belongs in review (`STATE.md`, `ROADMAP.md`, `MILESTONES.md`, `PROJECT.md`, `REQUIREMENTS.md`, `milestones/**`) and transient per-phase artifacts that do not (`phases/`, `quick/`, `research/`, `threads/`, `todos/`, `debug/`, `seeds/`, `codebase/`, `ui-reviews/`). Strict mode collapses that distinction: nothing under `.planning/` reaches the PR branch, and a commit is carried over only when it touches at least one file outside `.planning/`.
Strict mode exists because the two ways to keep planning private are not equivalent. Turning off `planning.commit_docs` keeps `.planning/` out of git, which also takes parallel executor worktrees with it — a worktree is checked out from a commit, so an untracked planning tree is simply absent inside it and the executor has no `PLAN.md` to read. Strict mode leaves planning committed, so worktrees and revert paths keep working, and moves the guarantee to the publication boundary instead. See [Publish PRs without planning artifacts](how-to/publish-prs-without-planning-artifacts.md).
Both modes filter by forcing the excluded paths back to whatever the target branch already tracks, in the index *and* the working tree. Un-staging alone would record a deletion of any planning file the target branch carries, and would leave the picked file untracked on disk, where it makes a later cherry-pick of the same path abort.
---
### 46. Security Hardening
**Purpose:** Defense-in-depth security for GSD's planning artifacts. Because GSD generates markdown files that become LLM system prompts, user-controlled text flowing into these files is a potential indirect prompt injection vector.
**Components:**
**1. Centralized Security Module** (`security.cjs`)
- Path traversal prevention — validates file paths resolve within the project directory
- Prompt injection detection — scans for known injection patterns in user-supplied text
- Safe JSON parsing — catches malformed input before state corruption
- Field name validation — prevents injection through config field names
- Shell argument validation — sanitizes user text before shell interpolation
**2. Prompt Injection Guard Hook** (`gsd-prompt-guard.js`)
PreToolUse hook that scans Write/Edit calls targeting `.planning/` for injection patterns. Advisory-only — logs detection for awareness without blocking legitimate operations.
**3. Workflow Guard Hook** (`gsd-workflow-guard.js`)
PreToolUse hook that detects when Claude attempts file edits outside a GSD workflow context. Advises using `/gsd-quick` or `/gsd-fast` instead of direct edits. Configurable via `hooks.workflow_guard` (default: false).
**4. CI-Ready Injection Scanner** (`prompt-injection-scan.security.test.cjs`)
Test suite that scans all agent, workflow, and command files for embedded injection vectors.
**Requirements:**
- REQ-SEC-01: All user-supplied file paths MUST be validated against the project directory
- REQ-SEC-02: Prompt injection patterns MUST be detected before text enters planning artifacts
- REQ-SEC-03: Security hooks MUST be advisory-only (never block legitimate operations)
- REQ-SEC-04: JSON parsing of user input MUST catch malformed data gracefully
- REQ-SEC-05: macOS `/var` → `/private/var` symlink resolution MUST be handled in path validation
---
### 47. Multi-Repo Workspace Support
**Purpose:** Auto-detection and project root resolution for monorepos and multi-repo setups. Supports workspaces where `.planning/` may need to resolve across repository boundaries.
**Requirements:**
- REQ-MULTIREPO-01: System MUST auto-detect multi-repo workspace configuration
- REQ-MULTIREPO-02: System MUST resolve project root across repository boundaries
- REQ-MULTIREPO-03: Executor MUST record per-repo commit hashes in multi-repo mode
---
### 48. Discussion Audit Trail
**Purpose:** Auto-generate `DISCUSSION-LOG.md` during `/gsd-discuss-phase` for full audit trail of decisions made during discussion.
**Requirements:**
- REQ-DISCLOG-01: System MUST auto-generate DISCUSSION-LOG.md during discuss-phase
- REQ-DISCLOG-02: Log MUST capture questions asked, options presented, and decisions made
- REQ-DISCLOG-03: Decision IDs MUST enable traceability from discuss-phase to plan-phase
---
## v1.28 Features
### 49. Forensics
**Command:** `/gsd-forensics [description]`
**Purpose:** Post-mortem investigation of failed or stuck GSD workflows.
**Requirements:**
- REQ-FORENSICS-01: System MUST analyze git history for anomalies (stuck loops, long gaps, repeated commits)
- REQ-FORENSICS-02: System MUST check artifact integrity (completed phases have expected files)
- REQ-FORENSICS-03: System MUST generate a markdown report saved to `.planning/forensics/`
- REQ-FORENSICS-04: System MUST offer to create a GitHub issue with findings
- REQ-FORENSICS-05: System MUST NOT modify project files (read-only investigation)
**Produces:**
| Artifact | Description |
|----------|-------------|
| `.planning/forensics/report-{timestamp}.md` | Post-mortem investigation report |
**Process:**
1. **Scan** — Analyze git history for anomalies: stuck loops, long gaps between commits, repeated identical commits
2. **Integrity Check** — Verify completed phases have expected artifact files
3. **Report** — Generate markdown report with findings, saved to `.planning/forensics/`
4. **Issue** — Offer to create a GitHub issue with findings for team visibility
---
### 50. Milestone Summary
**Command:** `/gsd-milestone-summary [version]`
**Purpose:** Generate comprehensive project summary from milestone artifacts for team onboarding.
**Requirements:**
- REQ-SUMMARY-01: System MUST aggregate phase plans, summaries, and verification results
- REQ-SUMMARY-02: System MUST work for both current and archived milestones
- REQ-SUMMARY-03: System MUST produce a single navigable document
**Produces:**
| Artifact | Description |
|----------|-------------|
| `MILESTONE-SUMMARY.md` | Comprehensive navigable summary of milestone artifacts |
**Process:**
1. **Collect** — Aggregate phase plans, summaries, and verification results from the target milestone
2. **Synthesize** — Combine artifacts into a single navigable document with cross-references
3. **Output** — Write `MILESTONE-SUMMARY.md` suitable for team onboarding and stakeholder review
---
### 51. Workstream Namespacing
**Command:** `/gsd-workstreams`
**Purpose:** Parallel workstreams for concurrent work on different milestone areas.
**Requirements:**
- REQ-WS-01: System MUST isolate workstream state in separate `.planning/workstreams/{name}/` directories
- REQ-WS-02: System MUST validate workstream names (alphanumeric + hyphens only, no path traversal)
- REQ-WS-03: System MUST support list, create, switch, status, progress, complete, resume subcommands
**Produces:**
| Artifact | Description |
|----------|-------------|
| `.planning/workstreams/{name}/` | Isolated workstream directory structure |
**Process:**
1. **Create** — Initialize a named workstream with isolated `.planning/workstreams/{name}/` directory
2. **Switch** — Change active workstream context for subsequent GSD commands
3. **Manage** — List, check status, track progress, complete, or resume workstreams
---
### 52. Manager Dashboard
**Command:** `/gsd-manager`
**Purpose:** Interactive command center for managing multiple phases from one terminal.
**Requirements:**
- REQ-MGR-01: System MUST show overview of all phases with status
- REQ-MGR-02: System MUST filter to current milestone scope
- REQ-MGR-03: System MUST show phase dependencies and conflicts
**Produces:** Interactive terminal output
**Process:**
1. **Scan** — Load all phases in the current milestone with their statuses
2. **Display** — Render overview showing phase dependencies, conflicts, and progress
3. **Interact** — Accept commands to navigate, inspect, or act on individual phases
---
### 53. Assumptions Discussion Mode
**Command:** `/gsd-discuss-phase` with `workflow.discuss_mode: 'assumptions'`
**Purpose:** Replace interview-style questioning with codebase-first assumption analysis.
**Requirements:**
- REQ-ASSUME-01: System MUST analyze codebase to generate structured assumptions before asking questions
- REQ-ASSUME-02: System MUST classify assumptions by confidence level (Confident/Likely/Unclear)
- REQ-ASSUME-03: System MUST produce identical CONTEXT.md format as default discuss mode
- REQ-ASSUME-04: System MUST support confidence-based skip gate (all HIGH = no questions)
**Produces:**
| Artifact | Description |
|----------|-------------|
| `{phase}-CONTEXT.md` | Same format as default discuss mode |
**Process:**
1. **Analyze** — Scan codebase to generate structured assumptions about implementation approach
2. **Classify** — Categorize assumptions by confidence level: Confident, Likely, Unclear
3. **Gate** — If all assumptions are HIGH confidence, skip questioning entirely
4. **Confirm** — Present unclear assumptions as targeted questions to the user
5. **Output** — Produce `{phase}-CONTEXT.md` in identical format to default discuss mode
---
### 54. UI Phase Auto-Detection
**Part of:** `/gsd-new-project` and `/gsd-progress`
**Purpose:** Automatically detect UI-heavy projects and surface `/gsd-ui-phase` recommendation.
**Requirements:**
- REQ-UI-DETECT-01: System MUST detect UI signals in project description (keywords, framework references)
- REQ-UI-DETECT-02: System MUST annotate ROADMAP.md phases with `ui_hint` when applicable
- REQ-UI-DETECT-03: System MUST suggest `/gsd-ui-phase` in next steps for UI-heavy phases
- REQ-UI-DETECT-04: System MUST NOT make `/gsd-ui-phase` mandatory
**Process:**
1. **Detect** — Scan project description and tech stack for UI signals (keywords, framework references)
2. **Annotate** — Add `ui_hint` markers to applicable phases in ROADMAP.md
3. **Surface** — Include `/gsd-ui-phase` recommendation in next steps for UI-heavy phases
---
### 55. Multi-Runtime Installer Selection
**Part of:** `npx @opengsd/gsd-core`
**Purpose:** Select multiple runtimes in a single interactive install session.
**Requirements:**
- REQ-MULTI-RT-01: Interactive prompt MUST support multi-select (e.g., Claude Code + Antigravity)
- REQ-MULTI-RT-02: CLI flags MUST continue to work for non-interactive installs
**Process:**
1. **Detect** — Identify available AI CLI runtimes on the system
2. **Prompt** — Present multi-select interface for runtime selection
3. **Install** — Configure GSD for all selected runtimes in a single session
---
## v1.29 Features
### 56. Windsurf Runtime Support
**Part of:** `npx @opengsd/gsd-core`
**Purpose:** Add Windsurf as a supported AI CLI runtime for GSD installation and execution.
**Requirements:**
- REQ-WINDSURF-01: Installer MUST detect Windsurf runtime and offer it as a target
- REQ-WINDSURF-02: GSD commands MUST function correctly within Windsurf sessions
**Process:**
1. **Detect** — Identify Windsurf runtime availability on the system
2. **Install** — Configure GSD skills and hooks for the Windsurf environment
---
### 57. Internationalized Documentation
**Part of:** `docs/`
**Purpose:** Provide GSD documentation in Portuguese, Korean, and Japanese.
**Requirements:**
- REQ-I18N-01: Documentation MUST be available in Portuguese (pt), Korean (ko), and Japanese (ja)
- REQ-I18N-02: Translations MUST stay synchronized with English source documents
**Process:**
1. **Translate** — Convert core documentation into target languages
2. **Publish** — Make translated documentation accessible alongside English originals
---
## v1.31 Features
### 59. Schema Drift Detection
**Command:** Automatic during `/gsd-execute-phase`
**Purpose:** Detect when ORM schema files are modified without corresponding migration or push commands, preventing false-positive verification.
**Requirements:**
- REQ-SCHEMA-01: System MUST detect modifications to ORM schema files (Prisma, Drizzle, Payload, Sanity, Mongoose)
- REQ-SCHEMA-02: System MUST verify corresponding migration/push commands exist when schema changes are detected
- REQ-SCHEMA-03: System MUST implement two-layer defense: plan-time injection and execute-time gate
- REQ-SCHEMA-04: System MUST support `GSD_SKIP_SCHEMA_CHECK` env var to override detection
- REQ-SCHEMA-05: System MUST prevent false-positive verification when schema is modified without migration
**Process:**
1. **Detect** — Monitor ORM schema file modifications during plan execution
2. **Verify** — Check that corresponding migration/push commands are present in the plan
3. **Gate** — Block execution if schema drift is detected without migration (execute-time gate)
4. **Inject** — Add migration reminders during plan generation (plan-time injection)
**Config:** `GSD_SKIP_SCHEMA_CHECK` environment variable to bypass detection.
---
### 60. Security Enforcement
**Command:** `/gsd-secure-phase <N>`
**Purpose:** Threat-model-anchored security verification for phase implementations.
**Requirements:**
- REQ-SEC-01: System MUST perform threat-model-anchored verification (not blind scanning)
- REQ-SEC-02: System MUST support configurable OWASP ASVS verification levels (1-3)
- REQ-SEC-03: System MUST block phase advancement based on configurable severity threshold
- REQ-SEC-04: System MUST spawn `gsd-security-auditor` agent for analysis
**Produces:**
| Artifact | Description |
|----------|-------------|
| Security audit report | Threat-model-anchored findings with severity classification |
**Process:**
1. **Model** — Build threat model from phase implementation context
2. **Audit** — Spawn `gsd-security-auditor` to verify against threat model
3. **Gate** — Block phase advancement if findings meet or exceed `security_block_on` severity
**Config:**
| Setting | Type | Default | Description |
|---------|------|---------|-------------|
| `security_enforcement` | boolean | `true` | Enable threat-model security verification |
| `security_asvs_level` | number (1-3) | `1` | OWASP ASVS verification level |
| `security_block_on` | string | `"high"` | Minimum severity to block phase advancement |
---
### 61. Documentation Generation
**Command:** `/gsd-docs-update`
**Purpose:** Generate and verify project documentation with accuracy checks.
**Requirements:**
- REQ-DOCS-01: System MUST spawn `gsd-doc-writer` agent to generate documentation
- REQ-DOCS-02: System MUST spawn `gsd-doc-verifier` agent to check accuracy
- REQ-DOCS-03: System MUST verify generated documentation against actual implementation
**Produces:**
| Artifact | Description |
|----------|-------------|
| Updated project documentation | Generated and verified documentation files |
**Process:**
1. **Generate** — Spawn `gsd-doc-writer` to create or update documentation from implementation
2. **Verify** — Spawn `gsd-doc-verifier` to check documentation accuracy against codebase
3. **Output** — Produce verified documentation with accuracy annotations
---
### 62. Discuss Chain Mode
**Flag:** `/gsd-discuss-phase <N> --chain`
**Purpose:** Auto-chain discuss, plan, and execute phases in one flow to reduce manual command sequencing.
**Requirements:**
- REQ-CHAIN-01: System MUST auto-chain discuss → plan → execute when `--chain` flag is provided
- REQ-CHAIN-02: System MUST respect all gate settings between chained phases
- REQ-CHAIN-03: System MUST halt the chain if any phase fails
**Process:**
1. **Discuss** — Run discuss-phase to gather context
2. **Plan** — Automatically invoke plan-phase with gathered context
3. **Execute** — Automatically invoke execute-phase with generated plan
---
### 63. Single-Phase Autonomous
**Flag:** `/gsd-autonomous --only N`
**Purpose:** Execute just one phase autonomously instead of all remaining phases.
**Requirements:**
- REQ-ONLY-01: System MUST execute only the specified phase number when `--only N` is provided
- REQ-ONLY-02: System MUST follow the same discuss → plan → execute flow as full autonomous mode
- REQ-ONLY-03: System MUST stop after the specified phase completes
**Process:**
1. **Select** — Identify the target phase from `--only N` argument
2. **Execute** — Run full autonomous flow (discuss → plan → execute) for that single phase
3. **Stop** — Halt after the phase completes instead of advancing to the next
---
### 64. Scope Reduction Detection
**Part of:** `/gsd-plan-phase`
**Purpose:** Prevent silent requirement dropping during plan generation with three-layer defense.
**Requirements:**
- REQ-SCOPE-01: System MUST prohibit planners from reducing scope without explicit justification
- REQ-SCOPE-02: System MUST have plan-checker verify requirement dimension coverage
- REQ-SCOPE-03: System MUST have orchestrator recover dropped requirements and re-inject them
- REQ-SCOPE-04: System MUST implement three-layer defense: planner prohibition, checker dimension, orchestrator recovery
**Process:**
1. **Prohibit** — Planner instructions explicitly forbid scope reduction
2. **Check** — Plan-checker verifies all phase requirements are covered in the plan
3. **Recover** — Orchestrator detects dropped requirements and re-injects them into the planning loop
---
### 65. Claim Provenance Tagging
**Part of:** `/gsd-plan-phase --research-phase <N>`
**Purpose:** Ensure research claims are tagged with source evidence and assumptions are logged separately.
**Requirements:**
- REQ-PROVENANCE-01: Researcher MUST mark claims with source evidence references
- REQ-PROVENANCE-02: Assumptions MUST be logged separately from sourced claims
- REQ-PROVENANCE-03: System MUST distinguish between evidenced facts and inferred assumptions
**Process:**
1. **Research** — Researcher gathers information from codebase and domain sources
2. **Tag** — Each claim is annotated with its source (file path, documentation, API response)
3. **Separate** — Assumptions without direct evidence are logged in a distinct section
---
### 66. Worktree Toggle
**Config:** `workflow.use_worktrees: false`
**Purpose:** Disable git worktree isolation for users who prefer sequential execution.
**Requirements:**
- REQ-WORKTREE-01: System MUST respect `workflow.use_worktrees` setting when deciding isolation strategy
- REQ-WORKTREE-02: System MUST default to `true` (worktrees enabled) for backward compatibility
- REQ-WORKTREE-03: System MUST fall back to sequential execution when worktrees are disabled
**Config:**
| Setting | Type | Default | Description |
|---------|------|---------|-------------|
| `workflow.use_worktrees` | boolean | `true` | When `false`, disables git worktree isolation |
---
### 67. Project Code Prefixing
**Config:** `project_code: "ABC"`
**Purpose:** Prefix phase directory names with a project code for multi-project disambiguation.
**Requirements:**
- REQ-PREFIX-01: System MUST prefix phase directories with project code when configured (e.g., `ABC-01-setup/`)
- REQ-PREFIX-02: System MUST use standard naming when `project_code` is not set
- REQ-PREFIX-03: System MUST apply prefix consistently across all phase operations
**Config:**
| Setting | Type | Default | Description |
|---------|------|---------|-------------|
| `project_code` | string | (none) | Prefix for phase directory names |
---
### 68. Claude Code Skills Migration
**Part of:** `npx @opengsd/gsd-core`
**Purpose:** Migrate GSD commands to Claude Code 2.1.88+ skills format with backward compatibility.
**Requirements:**
- REQ-SKILLS-01: Installer MUST write `skills/gsd-*/SKILL.md` for Claude Code 2.1.88+
- REQ-SKILLS-02: Installer MUST auto-clean legacy `commands/gsd/` directory
- REQ-SKILLS-03: Installer MUST maintain backward compatibility with older Claude Code versions via the legacy `commands/gsd/` path
**Process:**
1. **Detect** — Check Claude Code version to determine skills support
2. **Migrate** — Write `skills/gsd-*/SKILL.md` files for each GSD command
3. **Clean** — Remove legacy `commands/gsd/` directory if skills are installed
4. **Fallback** — Maintain legacy `commands/gsd/` path compatibility for older Claude Code versions
---
## v1.32 Features
### 69. STATE.md Consistency Gates
**Commands:** `state validate [--strict]`, `state sync [--verify]`, `state planned-phase --phase N --plans N`
**Purpose:** Detect and repair drift between STATE.md and the actual filesystem, preventing cascading errors from stale state.
**Requirements:**
- REQ-STATE-01: `state validate` MUST detect drift between STATE.md fields and filesystem reality
- REQ-STATE-02: `state sync` MUST reconstruct STATE.md from actual project state on disk
- REQ-STATE-03: `state sync --verify` MUST perform a dry-run showing proposed changes without writing
- REQ-STATE-04: `state planned-phase` MUST record the state transition after plan-phase completes (Planned/Ready to execute)
- REQ-STATE-05: `state validate` MUST report a `Last activity` value that no reader can parse, rather than validating clean
- REQ-STATE-06: `state validate --strict` MUST reflect `valid` in the process exit status, leaving the default exit status unchanged
**Produces:**
| Artifact | Description |
|----------|-------------|
| Updated `STATE.md` | Corrected state reflecting filesystem reality |
**Process:**
1. **Validate** — Compare STATE.md fields against filesystem (phase directories, plan files, summaries)
2. **Sync** — Reconstruct STATE.md from disk when drift is detected
3. **Transition** — Record post-planning state with plan count for execute-phase readiness
---
### 70. Autonomous `--to N` Flag
**Flag:** `/gsd-autonomous --to N`
**Purpose:** Stop autonomous execution after completing a specific phase, allowing partial autonomous runs.
**Requirements:**
- REQ-TO-01: System MUST stop execution after the specified phase number completes
- REQ-TO-02: System MUST follow the same discuss -> plan -> execute flow for each phase up to N
- REQ-TO-03: `--to N` MUST be combinable with `--from N` for bounded autonomous ranges
**Process:**
1. **Bound** — Set the upper phase limit from `--to N` argument
2. **Execute** — Run autonomous flow for each phase up to and including phase N
3. **Stop** — Halt after phase N completes
---
### 71. Research Gate
**Part of:** `/gsd-plan-phase`
**Purpose:** Block planning when RESEARCH.md has unresolved open questions, preventing plans built on incomplete information.
**Requirements:**
- REQ-RESGATE-01: System MUST scan RESEARCH.md for unresolved open questions before planning begins
- REQ-RESGATE-02: System MUST block plan-phase entry when open questions exist
- REQ-RESGATE-03: System MUST surface the specific unresolved questions to the user
**Process:**
1. **Scan** — Check RESEARCH.md for open questions section with unresolved items
2. **Gate** — Block planning if unresolved questions are found
3. **Surface** — Display the specific open questions requiring resolution
---
### 72. Verifier Milestone Scope Filtering
**Part of:** `/gsd-execute-phase` (verifier step)
**Purpose:** Distinguish between genuine gaps and items deferred to later phases, reducing false negatives in verification.
**Requirements:**
- REQ-VSCOPE-01: Verifier MUST check whether a gap is addressed in a later milestone phase
- REQ-VSCOPE-02: Gaps addressed in later phases MUST be marked as "deferred", not "gap"
- REQ-VSCOPE-03: Only genuine gaps (not covered by any future phase) MUST be reported as failures
**Process:**
1. **Verify** — Run standard goal-backward verification
2. **Filter** — Cross-reference detected gaps against later milestone phases
3. **Classify** — Mark deferred items separately from genuine gaps
---
### 73. Read-Before-Edit Guard Hook
**Part of:** Hooks (`PreToolUse`)
**Purpose:** Prevent infinite retry loops in non-Claude runtimes by ensuring files are read before editing.
**Requirements:**
- REQ-RBE-01: Hook MUST detect Edit/Write tool calls that target files not previously read in the session
- REQ-RBE-02: Hook MUST advise reading the file first (advisory, non-blocking)
- REQ-RBE-03: Hook MUST prevent infinite retry loops common in runtimes without built-in read-before-edit enforcement
---
### 74. Context Reduction
**Part of:** prompt assembly pipeline
**Purpose:** Reduce context prompt sizes through markdown truncation and cache-friendly prompt ordering.
**Requirements:**
- REQ-CTXRED-01: System MUST truncate oversized markdown artifacts to fit within context budgets
- REQ-CTXRED-02: System MUST order prompts for cache-friendly assembly (stable prefixes first)
- REQ-CTXRED-03: Reduction MUST preserve essential information (headings, requirements, task structure)
- REQ-CTXRED-04: Skill `description:` fields MUST be ≤ 100 chars; enforced by `npm run lint:descriptions` (see `scripts/lint-descriptions.cjs` and `tests/skill-frontmatter-contract.test.cjs`)
**Process:**
1. **Measure** — Calculate total prompt size for the workflow
2. **Truncate** — Apply markdown-aware truncation to oversized artifacts
3. **Order** — Arrange prompt sections for optimal KV-cache reuse
---
### 75. Discuss-Phase `--power` Flag
**Flag:** `/gsd-discuss-phase --power`
**Purpose:** File-based bulk question answering for discuss-phase, enabling batch input from a prepared answers file.
**Requirements:**
- REQ-POWER-01: System MUST accept a file containing pre-written answers to discussion questions
- REQ-POWER-02: System MUST map answers to the corresponding gray area questions
- REQ-POWER-03: System MUST produce CONTEXT.md identical to interactive discuss-phase
---
### 76. Debug `--diagnose` Flag
**Flag:** `/gsd-debug --diagnose`
**Purpose:** Diagnosis-only mode that investigates without attempting fixes.
**Requirements:**
- REQ-DIAG-01: System MUST perform full debug investigation (hypotheses, evidence, root cause)
- REQ-DIAG-02: System MUST NOT attempt any code modifications
- REQ-DIAG-03: System MUST produce a diagnostic report with findings and recommended fixes
---
### 77. Phase Dependency Analysis
**Command:** `/gsd-manager --analyze-deps`
**Purpose:** Detect phase dependencies and suggest `Depends on` entries for ROADMAP.md before running `/gsd-manager`.
**Requirements:**
- REQ-DEP-01: System MUST detect file overlap between phases
- REQ-DEP-02: System MUST detect semantic dependencies (API/schema producers and consumers)
- REQ-DEP-03: System MUST detect data flow dependencies (output producers and readers)
- REQ-DEP-04: System MUST suggest dependency entries with user confirmation before writing
**Produces:** Dependency suggestion table; optionally updates ROADMAP.md `Depends on` fields
---
### 78. Anti-Pattern Severity Levels
**Part of:** `/gsd-resume-work`
**Purpose:** Mandatory understanding checks at resume with severity-based anti-pattern enforcement.
**Requirements:**
- REQ-ANTI-01: System MUST classify anti-patterns by severity level
- REQ-ANTI-02: System MUST enforce mandatory understanding checks at session resume
- REQ-ANTI-03: Higher severity anti-patterns MUST block workflow progression until acknowledged
---
### 79. Methodology Artifact Type
**Part of:** Planning artifacts
**Purpose:** Define consumption mechanisms for methodology documents, ensuring they are consumed correctly by agents.
**Requirements:**
- REQ-METHOD-01: System MUST support methodology as a distinct artifact type
- REQ-METHOD-02: Methodology artifacts MUST have defined consumption mechanisms for agents
---
### 80. Planner Reachability Check
**Part of:** `/gsd-plan-phase`
**Purpose:** Validate that plan steps are achievable before committing to execution.
**Requirements:**
- REQ-REACH-01: Planner MUST validate that each plan step references reachable files and APIs
- REQ-REACH-02: Unreachable steps MUST be flagged during planning, not discovered during execution
---
### 81. Playwright-MCP UI Verification
**Part of:** `/gsd-verify-work` (optional)
**Purpose:** Automated visual verification using Playwright-MCP during verify-phase.
**Requirements:**
- REQ-PLAY-01: System MUST support optional Playwright-MCP visual verification during verify-phase
- REQ-PLAY-02: Visual verification MUST be opt-in, not mandatory
- REQ-PLAY-03: System MUST capture and compare visual state against UI-SPEC.md expectations
---
### 82. Pause-Work Expansion
**Part of:** `/gsd-pause-work`
**Purpose:** Support non-phase contexts with richer handoff data for broader pause-work applicability.
**Requirements:**
- REQ-PAUSE-01: System MUST support pausing in non-phase contexts (quick tasks, debug sessions, threads)
- REQ-PAUSE-02: Handoff data MUST include richer context appropriate to the current work type
---
### 83. Response Language Config
**Config:** `response_language`
**Purpose:** Cross-phase language consistency for non-English users.
**Requirements:**
- REQ-LANG-01: System MUST respect `response_language` setting across all phases and agents
- REQ-LANG-02: Setting MUST propagate to all spawned agents for consistent language output
- REQ-LANG-03: Every workflow MUST carry response-language coverage — through an exact inline directive, a shared `@`-referenced directive (`gsd-core/references/response-language-directive.md`), or inheritance from the parent workflow that dispatches it; enforced in CI by `scripts/lint-response-language-coverage.cjs` (#2529)
- REQ-LANG-04: A covering directive MUST name inter-tool narration, not only the question/prompt surface. A directive names it by using the word "narration" or the phrase "between tool calls"; the class it denotes is the model's running commentary between tool calls, status updates, progress notes and findings included, and enumerating those items without naming the class does not satisfy the rule. A directive worded around questions and prompts alone leaves the model's running commentary in English beside translated answers, which is the defect #2529 reports; `scripts/lint-response-language-coverage.cjs` rejects it (#2529)
**Config:**
| Setting | Type | Default | Description |
|---------|------|---------|-------------|
| `response_language` | string | (none) | Language code for agent responses (e.g., `"pt"`, `"ko"`, `"ja"`) |
---
### 84. Manual Update Procedure
**Part of:** `docs/manual-update.md`
**Purpose:** Document a manual update path for environments where `npx` is unavailable or npm publish is experiencing outages.
**Requirements:**
- REQ-MANUAL-01: Documentation MUST describe step-by-step manual update procedure
- REQ-MANUAL-02: Procedure MUST work without npm access
---
### 85. New Runtime Support (Trae, Cline, Augment Code)
**Part of:** `npx @opengsd/gsd-core`
**Purpose:** Extend GSD installation to Trae IDE, Cline, and Augment Code runtimes.
**Requirements:**
- REQ-TRAE-01: Installer MUST support `--trae` flag for Trae IDE installation
- REQ-CLINE-01: Installer MUST support Cline via `.clinerules` configuration
- REQ-AUGMENT-01: Installer MUST support Augment Code with skill conversion and config management
---
### 86. Autonomous `--interactive` Flag
**Flag:** `/gsd-autonomous --interactive`
**Purpose:** Lean-context autonomous mode that keeps discuss-phase interactive (user answers questions) while dispatching plan and execute as background agents on runtimes that support nested background dispatch; on Claude Code, plan and execute run inline to preserve worktree isolation and independent verification.
**Requirements:**
- REQ-INTERACT-01: `--interactive` MUST run discuss-phase inline with interactive questions (not auto-answered)
- REQ-INTERACT-02: `--interactive` MUST dispatch plan-phase and execute-phase as background agents for context isolation on runtimes where a backgrounded agent can spawn subagents; on Claude Code, plan and execute run inline
- REQ-INTERACT-03: `--interactive` MUST enable pipeline parallelism — discuss Phase N+1 while Phase N builds (applies on runtimes that support nested background dispatch; on Claude Code, discuss does not overlap planning/execution)
- REQ-INTERACT-04: Main context MUST only accumulate discuss conversations (lean context) on runtimes that support nested background dispatch; on Claude Code, inline plan/execute also accumulate in the main context
**Process:**
1. **Discuss inline** — Run discuss-phase in the main context with user interaction
2. **Dispatch** — On runtimes that support nested background dispatch: send plan and execute to background agents with fresh context windows. On Claude Code: run plan and execute inline.
3. **Pipeline** — On runtimes with background dispatch: while background agents build Phase N, begin discussing Phase N+1. On Claude Code: phases run sequentially.
---
### 87. Commit-Docs Guard Hook
**Hook:** `gsd-commit-docs.js`
**Purpose:** PreToolUse hook that enforces the `commit_docs` configuration, preventing `.planning/` files from being committed when `planning.commit_docs` is `false`.
**Requirements:**
- REQ-COMMITDOCS-01: Hook MUST intercept git commit commands that stage `.planning/` files
- REQ-COMMITDOCS-02: Hook MUST block commits containing `.planning/` files when `commit_docs` is `false`
- REQ-COMMITDOCS-03: Hook MUST be advisory — does not block when `commit_docs` is `true` or absent
---
### 88. Community Hooks Opt-In
**Hooks:** `gsd-validate-commit.sh`, `gsd-session-state.sh`, `gsd-phase-boundary.sh`
**Purpose:** Optional git and session hooks for GSD projects, gated behind `hooks.community: true` in config.
**Requirements:**
- REQ-COMMUNITY-01: All community hooks MUST be no-ops unless `hooks.community` is `true` in `.planning/config.json`
- REQ-COMMUNITY-02: `gsd-validate-commit.sh` MUST enforce Conventional Commits format on git commit messages
- REQ-COMMUNITY-03: `gsd-session-state.sh` MUST track session state transitions
- REQ-COMMUNITY-04: `gsd-phase-boundary.sh` MUST enforce phase boundary checks
**Config:**
| Setting | Type | Default | Description |
|---------|------|---------|-------------|
| `hooks.community` | boolean | `false` | Enable optional community hooks for commit validation, session state, and phase boundaries |
| `hooks.commit_types` | array of strings | `[]` | Extra Conventional Commits types `gsd-validate-commit.sh` accepts, in addition to the built-in `feat, fix, docs, style, refactor, perf, test, build, ci, chore` — never replaces them. Each entry must match `^[a-z][a-z0-9-]*$` (lowercase letters, digits, hyphens); non-conforming or non-string entries are dropped. Example: `{ "hooks": { "community": true, "commit_types": ["enhance", "enh", "revert"] } }`. |
---
## v1.34.0 Features
### 89. Global Learnings Store
**Commands:** Auto-triggered at phase completion; consumed by planner
**Config:** `features.global_learnings`
**Purpose:** Persist cross-session, cross-project learnings in a global store so the planner agent can learn from patterns across the entire project history — not just the current session.
**Requirements:**
- REQ-LEARN-01: Learnings MUST be auto-copied from `.planning/` to the global store at phase completion
- REQ-LEARN-02: The planner agent MUST receive relevant learnings at spawn time via injection
- REQ-LEARN-03: Injection MUST be capped by `learnings.max_inject` to avoid context bloat
- REQ-LEARN-04: Feature MUST be opt-in via `features.global_learnings: true`
**Config:**
| Setting | Type | Default | Description |
|---------|------|---------|-------------|
| `features.global_learnings` | boolean | `false` | Enable cross-project learnings pipeline |
| `learnings.max_inject` | number | (system default) | Maximum learnings entries injected into planner |
---
### 90. Queryable Codebase Intelligence
**Command:** `/gsd-map-codebase --query [<term>|status|diff|refresh]`
**Config:** `intel.enabled`
**Purpose:** Maintain a queryable JSON index of codebase structure, API surface, dependency graph, file roles, and architecture decisions in `.planning/intel/`. Enables targeted lookups without reading the entire codebase.
**Requirements:**
- REQ-INTEL-01: Intel files MUST be stored as JSON in `.planning/intel/`
- REQ-INTEL-02: `query` mode MUST search across all intel files for a term and group results by file
- REQ-INTEL-03: `status` mode MUST report freshness (FRESH/STALE, stale threshold: 24 hours)
- REQ-INTEL-04: `diff` mode MUST compare current intel state to the last snapshot
- REQ-INTEL-05: `refresh` mode MUST spawn the intel-updater agent to rebuild all files
- REQ-INTEL-06: Feature MUST be opt-in via `intel.enabled: true`
**Intel files produced:**
| File | Contents |
|------|----------|
| `stack.json` | Technology stack and dependencies |
| `api-map.json` | Exported functions and API surface |
| `dependency-graph.json` | Inter-module dependency relationships |
| `file-roles.json` | Role classification for each source file |
| `arch-decisions.json` | Detected architecture decisions |
---
### 91. Execution Context Profiles
**Config:** `context_profile`
**Purpose:** Select a pre-configured execution context (mode, model, workflow settings) tuned for a specific type of work without manually adjusting individual settings.
**Requirements:**
- REQ-CTX-01: `dev` profile MUST optimize for iterative development (balanced model, plan_check enabled)
- REQ-CTX-02: `research` profile MUST optimize for research-heavy work (higher model tier, research enabled)
- REQ-CTX-03: `review` profile MUST optimize for code review work (verifier and code_review enabled)
**Available profiles:** `dev`, `research`, `review`
**Config:**
| Setting | Type | Default | Description |
|---------|------|---------|-------------|
| `context_profile` | string | (none) | Execution context preset: `dev`, `research`, or `review` |
---
### 92. Gates Taxonomy
**References:** `gsd-core/references/gates.md`
**Agents:** plan-checker, verifier
**Purpose:** Define 4 canonical gate types that structure all workflow decision points, enabling plan-checker and verifier agents to apply consistent gate logic.
**Gate types:**
| Type | Description |
|------|-------------|
| **Confirm** | User approves before proceeding (e.g., roadmap review) |
| **Quality** | Automated quality check must pass (e.g., plan verification loop) |
| **Safety** | Hard stop on detected risk or policy violation |
| **Transition** | Phase or milestone boundary acknowledgment |
**Requirements:**
- REQ-GATES-01: plan-checker MUST classify each checkpoint as one of the 4 gate types
- REQ-GATES-02: verifier MUST apply gate logic appropriate to the gate type
- REQ-GATES-03: Hard stop safety gates MUST never be bypassed by `--auto` flags
---
### 93. Code Review Pipeline
**Commands:** `/gsd-code-review`, `/gsd-code-review --fix`
**Purpose:** Structured review of source files changed during a phase, with a separate auto-fix pass that commits each fix atomically.
**Requirements:**
- REQ-REVIEW-01: `gsd-code-review` MUST scope files to the phase using SUMMARY.md and git diff fallback
- REQ-REVIEW-02: Review MUST support three depth levels: `quick`, `standard`, `deep`
- REQ-REVIEW-03: Findings MUST be severity-classified: Critical, Warning, Info
- REQ-REVIEW-04: `gsd-code-review --fix` MUST read REVIEW.md and fix Critical + Warning findings by default
- REQ-REVIEW-05: Each fix MUST be committed atomically with a descriptive message
- REQ-REVIEW-06: `--auto` flag MUST enable fix + re-review iteration loop, capped at 3 iterations
- REQ-REVIEW-07: Feature MUST be gated by `workflow.code_review` config flag
- REQ-REVIEW-08: `workflow.code_review_point` MUST select which loop point the automatic review step registers at (`execute:post` default, or `execute:wave:post`), independent of the `workflow.code_review` on/off gate and of manual `/gsd-code-review` invocation (#3661)
- REQ-REVIEW-09: The in-phase `code_review_gate` MUST report the per-severity counts it parses from REVIEW.md, so a review with one `info` finding is distinguishable from a review with a Critical
- REQ-REVIEW-10: Each finding MUST carry a recorded disposition, so a triaged finding is distinguishable from a forgotten one
**Config:**
| Setting | Type | Default | Description |
|---------|------|---------|-------------|
| `workflow.code_review` | boolean | `true` | Enable code review commands |
| `workflow.code_review_point` | string | `execute:post` | Loop point for the automatic review: `execute:post` (once per phase, default) or `execute:wave:post` (once per completed wave, scoped to what changed since the phase's prior review). See below. |
| `workflow.code_review_depth` | string | `standard` | Default review depth: `quick`, `standard`, or `deep` |
| `workflow.code_review_depth_overrides` | array | `[]` | Ordered `{ paths, depth }` rules that escalate depth for directories matched by path prefix against the changed-file set (#2554). See below. |
**Reviewing per wave instead of per phase (#3661)**
Setting `workflow.code_review_point` to `execute:wave:post` moves the automatic review from
"once, after the whole phase's waves have all landed" to "once per completed wave." Each
wave's review scopes to what changed since the phase's *previous* review — the whole phase's
diff on the first wave, then just that wave's diff on every wave after — so review batches
stay small instead of growing with the phase. A finding introduced early is caught after the
wave that introduced it, not after the last wave of the phase.
This only affects the *automatic* dispatch inside a wave-based phase execution. Manual
`/gsd-code-review <phase>` runs are gated by `workflow.code_review` alone and are unaffected
by this key. `/gsd-autonomous` and `/gsd-quick` have no wave granularity of their own, so
setting this to `execute:wave:post` means automatic review does not run inside those two
flows — the same way every other wave-scoped capability step already behaves for them.
**Path-scoped code review depth overrides**
`workflow.code_review_depth_overrides` matches rules against the review's changed-file set by whole-segment directory-path prefix — `src/auth` matches `src/auth/token.ts` and `src/auth` itself, never `src/authfoo/x.ts` or `docs/src/auth/x.ts` — and is case-sensitive, following git.
Escalation is **whole-review, not per-file**: depth is a single scalar handed to the reviewer agent, not a per-file setting, so the strongest matching tier across the whole rule set applies to every file in the review — a sensitive file is never reviewed shallowly because it shared a review with an unrelated one.
v1 supports **directory-prefix matching only, not glob syntax**: no glob engine (`minimatch`, `picomatch`, `fast-glob`) exists in this project and none was added for this feature. A path containing `*` or `?` (e.g. `src/auth/**`) is a configuration error rather than a silent near-miss, because accepting it as sugar for a prefix would make unsupported patterns look armed when they match nothing. Every use case in the issue is expressible as a directory prefix. See [Scope code review depth by path](how-to/scope-code-review-depth-by-path.md) for the resolution order, error table, and a worked example.
**Optional external reviewer lanes (#4209):** `/gsd-code-review` accepts the same reviewer-lane flags as `/gsd-review` — any flag the roster declares (run `gsd_run review-lane flags` to list them for your installation, e.g. `--codex`, `--agy`). No reviewer-lane flag is the default and is byte-for-byte unchanged from before #4209: zero lane selection, plan, or invoke calls, and only the internal `gsd-code-reviewer` agent runs. Passing one or more flags asks those lanes to independently review the same already-resolved file scope alongside the internal agent, through the same shared capability-trait interpreter and `review-lane plan`/`invoke` machinery `/gsd-review` uses — no second implementation. Each lane's prompt carries only the repository root, canonical file paths, review depth, and base SHA, never source file contents, under four fixed prohibitions (no source mutation, no test execution, no background processes, no polling). External findings are unverified corroborating evidence: `gsd-code-reviewer` independently re-verifies every claim against the actual source before writing it to `REVIEW.md`, so there remains exactly one `REVIEW.md` schema regardless of how many lanes ran. An explicitly requested lane that is unavailable or fails is reported as a warning, never silently dropped and never a raw-CLI fallback. This is separate from `/gsd-review`, which reviews `PLAN.md` files before execution, not source code.
**In-phase review reporting and disposition**
`/gsd-execute-phase`'s `code_review_gate` runs code review, then reports what it found:
```
Code review: 23 findings — 1 critical, 9 warning, 8 info.
Consider running: /gsd-code-review 1 --fix
```
Both `critical:` and its documented tier-equivalent `blocker:` are accepted. A REVIEW.md written
without a `findings:` block has no counts to report, and the gate falls back to the countless form
rather than printing a half-filled line.
The gate then writes `<NN>-REVIEW-DISPOSITION.md` beside the review — one row per finding ID,
defaulting to `open`:
| Finding | Severity | Disposition | Source |
|---------|----------|-------------|--------|
| CR-01 | critical | open | - |
| WR-01 | warning | fixed | 01-REVIEW-FIX.md |
`open` means recorded but not yet triaged, and it is the only value the gate assigns on its own.
**Where the reconciliation happens, which is not where you might expect.** The in-phase gate runs
immediately after review, and at that moment `<NN>-REVIEW-FIX.md` does not exist — the gate invokes
review with neither `--fix` nor `--auto` — so every row it writes is `open`. The `fixed` and
`skipped` outcomes are reconciled by `/gsd-code-review <N> --fix`, which records them once the fix
report is on disk. Running the gate alone therefore tells you what was found; running `--fix` is
what records what happened to it.
**`--auto`'s iterations are reconciled too, and that takes reading more than the final report.**
The loop overwrites `REVIEW-FIX.md` on every pass and the re-review stops reporting a finding once
it is fixed, so a finding closed in iteration 1 appears in neither final artifact. The gate
therefore also reads the per-iteration backups the loop writes (`<NN>-REVIEW-FIX.iterN.md`), newest
first, so the most recent statement about an ID wins; the backups are removed after the ledger has
read them, not before. Without this a fully successful multi-iteration run recorded its early fixes
as `open (not in the current review)` — indistinguishable from a finding that vanished for an
unrelated reason, which is the one distinction this ledger exists to make. A finding a fix report
decided but the current review no longer reports gets a row of its own, carrying that decision and
marked *(not in the current review)*.
A converged `--auto` run also reaches the gate with a clean review and, on a direct
`/gsd-code-review` invocation, no ledger from the in-phase gate. A fix report on disk is reason
enough to record: without that, a run in which every finding was fixed and committed produced no
disposition record at all.
**A decision belongs to a finding, not to an ID.** IDs are reused across re-reviews — `--auto`
renumbers — so the ledger records each finding's title alongside its disposition, and a recorded
decision is carried forward only while the ID still names the same finding. Without that, a prior
`CR-01 fixed` would be inherited by a brand-new `CR-01`, which is a false decision in the one
artifact whose purpose is telling triaged from forgotten.
**The limitation that follows, stated rather than hidden.** When an ID *is* reused, the new finding
is recorded `open` (nobody decided anything about it) and the earlier decision **loses its row** —
the ledger keys rows on the finding ID, and two rows under one ID is an ambiguity, not a record. The
drop is reported on the console, naming the ID and what had been decided — for a *recorded* decision.
A prior row still sitting at `open` is replaced silently, and deliberately: `open` means nobody had
decided anything, so there is no decision to lose. Preserving a recorded decision in the file
was tried and withdrawn: it needed a second identity scheme and produced a fresh defect on each of
three review passes. Recovery is the ledger's own git history where `commit_docs` is on — which is
why the console reports the drop rather than pointing at a commit that may not exist.
**Two residuals, since (ID, title) is not proof of identity.** A ledger written before titles were
recorded carries none, so its decisions are inherited on the ID alone — refusing there would reset
every decision in every existing ledger, which is the loss the guard exists to prevent. And two
genuinely distinct findings that share both an ID and a title are indistinguishable to this key.
Separating them needs a second identity scheme, which is the thing that was just withdrawn.
Reconciliation applies an outcome only when the fix report names the **same** finding, because
finding IDs are reused across re-reviews and a stale report would otherwise declare a brand-new
`CR-01` already fixed. Titles are compared with runs of whitespace collapsed, so a fixer that
re-spaces a title still reconciles. A title that differs otherwise — including one **wrapped across
lines**, since a `###` heading is one line and the continuation is a separate paragraph — leaves the
row `open` and is reported, naming the ids it could not reconcile, rather than passing silently. The report does not claim to
know whether such a report is stale or merely re-titled, because it cannot tell.
The disposition column is a closed vocabulary — `open`, `fixed`, `skipped`, `deferred`. A
hand-edited value outside it is not treated as a decision: the row falls back to `open`, so a
typo cannot quietly mark a phase as triaged.
Severity comes from the section a finding sits under (`## Critical Issues`, `## Warnings`,
`## Info`) when the review uses those headings, and from the ID prefix otherwise. The section is
the reviewer's own statement of severity, so a Critical filed under `## Critical Issues` is
recorded critical even if its ID was mis-numbered `WR-04`. That severity is then **remembered**:
a row the current review no longer reports, or reports under no recognized heading, keeps the
severity the ledger recorded rather than having it re-inferred from the ID prefix — so the
mis-numbered `WR-04` above stays `critical` on the second run, deferred or not. The precedence is
the current review's section, then the recorded value, then the prefix; the recorded value is
inherited only while the ID still names the same finding, by the same title check the disposition
uses, so a reused ID starts from its own review. That check has the same compatibility arm the
disposition has: a ledger written before titles were recorded carries no title to compare, so its
severities — like its decisions — inherit on the ID alone.
If the review's `total:` exceeds the number of findings whose headings the gate could parse, the
shortfall is stated — on the console and as an `unparsed:` key in the ledger's frontmatter. A
finding the gate cannot record is the one a human most needs to see, so it is never dropped
silently. The one input that produces no shortfall is a `findings:` block that disagrees with
itself: where `critical` (or `blocker`), `warning` and `info` are all present, numeric and do not
sum to `total`, the gate has no trustworthy number to reconcile against and withholds the key
rather than reporting a figure derived from one — the same input on which the console line
already withholds the severity breakdown. Counts that are merely absent, partial or non-numeric
are not a disagreement, and the shortfall is still reported from `total` alone. `deferred` is the one disposition the gate never writes: it is recorded by hand, and the
reason recorded beside it in the Source cell is preserved across re-runs, a literal `|` included
once escaped. One exception, because it cannot be resolved: a reason ending in the literal phrase
*(not in the current review)* loses that trailing phrase, since it is indistinguishable from the
carried marker the gate appends. The alternative is worse — a stored marker never leaves, so a
carried finding that later reappears would keep claiming it is absent from the review reporting it.
Re-running the gate keeps every row it can, so a decision recorded here is not
overwritten by a later pass. A finding the current review no longer reports — `--auto` re-reviews
and rewrites REVIEW.md, so this happens routinely — is **carried** rather than dropped, marked
*(not in the current review)*. That holds whether or not it was triaged: losing a decided row would
erase the record that the finding was seen, and losing an *untriaged* one would erase the record
that it was never answered, which is the trace this ledger exists to keep. The cost is that a
renumbered finding shows under both IDs until the old row is decided; the marker makes that legible.
A run that changes nothing rewrites nothing, so a re-executed phase does not produce a docs commit
with no content.
One residual is concurrency: the ledger is read, rebuilt and written whole, with no lock. Two
dispatchers can run this step (`code_review_gate` and `code-review-fix`'s `record_disposition`) and
a human is invited to hand-edit the file, so two writers overlapping would lose one's update. This is
the shape #3780 reported for `WINDOWS.md` under parallel executors, closed there by a cross-process
lock (#4681); the disposition step does not take that lock. No lost update has been reproduced; the
window is stated so it is not mistaken for a guarantee.
Not to be confused with the **Review Dispositions Ledger** of reviews-mode planning
([ADR-3806](../adr/3806-review-dispositions-ledger.md), `docs/features/review-dispositions-ledger.md`):
that one is a `## Review Dispositions Ledger` section inside `PLAN.md`, append-only per round, over
`REVIEWS.md` findings. This artifact is a sibling file beside `REVIEW.md`, rewritten idempotently
with rows carried. Same word, different inputs, writers, files and durability rules; neither governs
the other.
The record is a sibling artifact rather than a section inside REVIEW.md because `--auto`'s
re-review loop rewrites REVIEW.md on every iteration — a ledger kept inside it would not survive
the next pass — and because REVIEW.md has a single writer (`gsd-code-reviewer`) that the gate is
not. The gate remains advisory throughout: it reports and records, and never blocks phase
completion.
---
### 94. Socratic Exploration
**Command:** `/gsd-explore [topic]`
**Purpose:** Guide a developer through exploring an idea via Socratic probing questions before committing to a plan. Routes outputs to the appropriate GSD artifact: notes, todos, seeds, research questions, requirements updates, or a new phase.
**Requirements:**
- REQ-EXPLORE-01: Exploration MUST use Socratic probing — ask questions before proposing solutions
- REQ-EXPLORE-02: Session MUST offer to route outputs to the appropriate GSD artifact
- REQ-EXPLORE-03: An optional topic argument MUST prime the first question
- REQ-EXPLORE-04: Exploration MUST optionally spawn a research agent for technical feasibility
- REQ-EXPLORE-05: A research pass MUST disposition each surfaced claim (admit / refute / abstain) and route every abstention to a visible Unresolved Ledger — never smoothing an ungrounded claim into the narrative as confident prose
---
### 95. Safe Undo
**Command:** `/gsd-undo --last N | --phase NN | --plan NN-MM`
**Purpose:** Roll back GSD phase or plan commits safely using the phase manifest and git log, with dependency checks and a hard confirmation gate before any revert is applied.
**Requirements:**
- REQ-UNDO-01: `--phase` mode MUST identify all commits for the phase via manifest and git log fallback
- REQ-UNDO-02: `--plan` mode MUST identify all commits for a specific plan
- REQ-UNDO-03: `--last N` mode MUST display recent GSD commits for interactive selection
- REQ-UNDO-04: System MUST check for dependent phases/plans before reverting
- REQ-UNDO-05: A confirmation gate MUST be shown before any git revert is executed
---
### 96. Plan Import
**Command:** `/gsd-import --from <filepath>`
**Purpose:** Ingest an external plan file into the GSD planning system with conflict detection against `PROJECT.md` decisions, converting it to a valid GSD PLAN.md and validating it through the plan-checker.
**Requirements:**
- REQ-IMPORT-01: Importer MUST detect conflicts between the external plan and existing PROJECT.md decisions
- REQ-IMPORT-02: All detected conflicts MUST be presented to the user for resolution before writing
- REQ-IMPORT-03: Imported plan MUST be written as a valid GSD PLAN.md format
- REQ-IMPORT-04: Written plan MUST pass `gsd-plan-checker` validation
---
### 97. Rapid Codebase Scan
**Command:** `/gsd-map-codebase --fast [--focus tech|arch|quality|concerns]`
**Purpose:** Lightweight alternative to `/gsd-map-codebase` that spawns a single mapper agent for one or two combined focus areas, producing targeted output in `.planning/codebase/` without the overhead of 4 parallel agents.
**Requirements:**
- REQ-SCAN-01: Scan MUST spawn exactly one mapper agent (not four parallel agents)
- REQ-SCAN-02: Focus area MUST be one of: `tech`, `arch`, `quality`, `concerns`, or the combined `tech+arch` shorthand (default: `tech+arch`); combined focus runs as a single agent covering both areas in one pass
- REQ-SCAN-03: Output MUST be written to `.planning/codebase/` in the same format as `/gsd-map-codebase`
---
### 98. Autonomous Audit-to-Fix
**Command:** `/gsd-audit-fix [--source <audit>] [--severity high|medium|all] [--max N] [--dry-run]`
**Purpose:** End-to-end pipeline that runs an audit, classifies findings as auto-fixable vs. manual-only, then autonomously fixes auto-fixable issues with test verification and atomic commits.
**Requirements:**
- REQ-AUDITFIX-01: Findings MUST be classified as auto-fixable or manual-only before any changes
- REQ-AUDITFIX-02: Each fix MUST be verified with tests before committing
- REQ-AUDITFIX-03: Each fix MUST be committed atomically
- REQ-AUDITFIX-04: `--dry-run` MUST show classification table without applying any fixes
- REQ-AUDITFIX-05: `--max N` MUST limit the number of fixes applied in one run (default: 5)
---
### 99. Improved Prompt Injection Scanner
**Hook:** `gsd-prompt-guard.js`, `gsd-read-injection-scanner.js`
**Script:** `scripts/prompt-injection-scan.sh`, `scripts/base64-scan.sh`
**Purpose:** Defense-in-depth detection of prompt injection attempts in planning artifacts and ingested content. Live hooks inline their own pattern subsets for hook independence (they do not import from `security.cts`). The CI scanner (`scanForInjection` in `security.cts`) provides a centralized engine for codebase-wide scanning in tests.
**Requirements:**
- REQ-SCAN-INJ-01: Live hooks MUST detect invisible Unicode characters (zero-width spaces, soft hyphens, Unicode tag block U+E0000–E007F)
- REQ-SCAN-INJ-02: Live hooks MUST detect known injection patterns (instruction override, role manipulation, system-prompt extraction, fake message boundaries). Base64-decode scanning is a CI-time control (`scripts/base64-scan.sh`), not a live hook — live hooks match a base64-exfiltration phrase regex only, they do not decode.
- REQ-SCAN-INJ-03: ~~Scanner MUST apply entropy analysis~~ — Entropy analysis (`scanEntropyAnomalies`) was removed in #2198 as dead code (zero production callers; live hooks do not perform entropy analysis). This requirement is deferred pending a maintainable live implementation.
- REQ-SCAN-INJ-04: Scanner MUST remain advisory-only — detection is logged, not blocking
- REQ-SCAN-INJ-05: A scanner that could not establish its file list MUST NOT report clean (#3908). The CI scanners (`prompt-injection-scan.sh`, `base64-scan.sh`, `secret-scan.sh`) distinguish four outcomes rather than collapsing them into exit 0:
| Outcome | Exit | Meaning |
|---|---|---|
| scanned, no findings | `0` | files were in scope and none matched |
| findings | `1` | the scan's own verdict |
| nothing in scope | `NO_INPUT` | the diff resolved and was genuinely empty — e.g. a docs-only PR |
| could not scan | `UNAVAILABLE` | the file list was never established: a bad ref, no repository, or a repository with no commits |
Codes come from the exit-code registry ([ADR-3889](adr/3889-process-exit-contract.md)), sourced from `gsd-core/bin/shared/exit-codes.sh`, never written into the scripts. Every one is non-zero, so a caller written `if ! scanner; then` behaves identically for a clean scan and trips for everything else — this can turn a false green red, never a red green. `.github/workflows/security-scan.yml` treats *nothing in scope* as a pass and *could not scan* as a failure; previously the latter passed silently, having scanned nothing.
---
### 100. Stall Detection in Plan-Phase
**Command:** `/gsd-plan-phase`
**Purpose:** Detect when the planner revision loop has stalled — producing the same output across multiple iterations — and break the cycle by escalating to a different strategy or exiting with a clear diagnostic.
**Requirements:**
- REQ-STALL-01: Revision loop MUST detect identical plan output across consecutive iterations
- REQ-STALL-02: On stall detection, system MUST escalate strategy before retrying
- REQ-STALL-03: Maximum stall retries MUST be bounded (capped at the existing max 3 iterations)
---
### 101. Hard Stop Safety Gates in /gsd-progress --next
**Command:** `/gsd-progress --next`
**Purpose:** Prevent `/gsd-progress --next` from entering runaway loops by adding hard stop safety gates and a consecutive-call guard that interrupts autonomous chaining when repeated identical steps are detected.
**Requirements:**
- REQ-NEXT-GATE-01: `/gsd-progress --next` MUST track consecutive same-step calls
- REQ-NEXT-GATE-02: On repeated same-step, system MUST present a hard stop gate to the user
- REQ-NEXT-GATE-03: User MUST explicitly confirm to continue past a hard stop gate
---
### 102. Adaptive Model Preset
**Config:** `model_profile: "adaptive"`
**Purpose:** Role-based model assignment that automatically selects the appropriate model tier based on the current agent's role, rather than applying a single tier to all agents.
**Requirements:**
- REQ-ADAPTIVE-01: `adaptive` preset MUST assign model tiers based on agent role (planner → quality tier, executor → balanced tier, etc.)
- REQ-ADAPTIVE-02: `adaptive` MUST be selectable via `/gsd-config --profile adaptive`
---
### 103. Post-Merge Hunk Verification
**Command:** `/gsd-update --reapply`
**Purpose:** After applying local patches post-update, verify that all hunks were actually applied by comparing the expected patch content against the live filesystem. Surface any dropped or partial hunks immediately rather than silently accepting incomplete merges.
**Requirements:**
- REQ-PATCH-VERIFY-01: Reapply-patches MUST verify each hunk was applied after the merge
- REQ-PATCH-VERIFY-02: Dropped or partial hunks MUST be reported to the user with file and line context
- REQ-PATCH-VERIFY-03: Verification MUST run after all patches are applied, not per-patch
---
## v1.35.0 Features
### 104. New Runtime Support (Cline, CodeBuddy, Qwen Code)
**Part of:** `npx @opengsd/gsd-core`
**Purpose:** Extend GSD installation to Cline, CodeBuddy, and Qwen Code runtimes.
**Requirements:**
- REQ-CLINE-02: Cline install MUST write `.clinerules` to `~/.cline/` (global) or `./.cline/` (local). No custom slash commands — rules-based integration only. Flag: `--cline`.
- REQ-CODEBUDDY-01: CodeBuddy install MUST deploy skills to `~/.codebuddy/skills/gsd-*/SKILL.md` (emitted `user-invocable: false`), `/gsd-*` slash commands to `~/.codebuddy/commands/gsd-*.md`, and subagents to `~/.codebuddy/agents/gsd-*.md`. The commands surface is the sole `/` menu entry point. No `mcp.json` is written (gsd ships no MCP server). Flag: `--codebuddy`.
- REQ-QWEN-01: Qwen Code install MUST deploy skills to `~/.qwen/skills/gsd-*/SKILL.md`, following the open standard used by Claude Code 2.1.88+. `QWEN_CONFIG_DIR` env var overrides the default path. Flag: `--qwen`.
**Runtime summary:**
| Runtime | Install Format | Config Path | Flag |
|---------|---------------|-------------|------|
| Cline | `.clinerules` | `~/.cline/` or `./.cline/` | `--cline` |
| CodeBuddy | Skills (`SKILL.md`) | `~/.codebuddy/skills/` | `--codebuddy` |
| Qwen Code | Skills (`SKILL.md`) | `~/.qwen/skills/` | `--qwen` |
---
### 105. GSD-2 Reverse Migration
**Command:** `/gsd-import --from-gsd2 [--dry-run] [--force] [--path <dir>]`
**Purpose:** Migrate a project from GSD-2 format (`.gsd/` directory with Milestone→Slice→Task hierarchy) back to the v1 `.planning/` format, restoring full compatibility with all GSD v1 commands.
**Requirements:**
- REQ-FROM-GSD2-01: Importer MUST read `.gsd/` from the specified or current directory
- REQ-FROM-GSD2-02: Milestone→Slice hierarchy MUST be flattened to sequential phase numbers (M001/S01→phase 01, M001/S02→phase 02, M002/S01→phase 03, etc.)
- REQ-FROM-GSD2-03: System MUST guard against overwriting an existing `.planning/` directory without `--force`
- REQ-FROM-GSD2-04: `--dry-run` MUST preview all changes without writing any files
- REQ-FROM-GSD2-05: Migration MUST produce `PROJECT.md`, `REQUIREMENTS.md`, `ROADMAP.md`, `STATE.md`, and sequential phase directories
**Flags:**
| Flag | Description |
|------|-------------|
| `--dry-run` | Preview migration output without writing files |
| `--force` | Overwrite an existing `.planning/` directory |
| `--path <dir>` | Specify the GSD-2 root directory |
---
### 106. AI Integration Phase Wizard
**Command:** `/gsd-ai-integration-phase [N]`
**Purpose:** Guide developers through selecting, integrating, and planning evaluation for AI/LLM capabilities in a project phase. Produces a structured `AI-SPEC.md` that feeds into planning and verification.
**Requirements:**
- REQ-AISPEC-01: Wizard MUST present an interactive decision matrix covering framework selection, model choice, and integration approach
- REQ-AISPEC-02: System MUST surface domain-specific failure modes and eval criteria relevant to the project type
- REQ-AISPEC-03: System MUST spawn 3 parallel specialist agents: domain-researcher, framework-selector, and eval-planner
- REQ-AISPEC-04: Output MUST produce `{phase}-AI-SPEC.md` with framework recommendation, implementation guidance, and evaluation strategy
**Produces:** `{phase}-AI-SPEC.md` in the phase directory
---
### 107. AI Eval Review
**Command:** `/gsd-eval-review [N]`
**Purpose:** Retroactively audit an executed AI phase's evaluation coverage against the `AI-SPEC.md` plan. Identifies gaps between planned and implemented evaluation before the phase is closed.
**Requirements:**
- REQ-EVALREVIEW-01: Review MUST read `AI-SPEC.md` from the specified phase
- REQ-EVALREVIEW-02: Each eval dimension MUST be scored as COVERED, PARTIAL, or MISSING
- REQ-EVALREVIEW-03: Output MUST include findings, gap descriptions, and remediation guidance
- REQ-EVALREVIEW-04: `EVAL-REVIEW.md` MUST be written to the phase directory
**Produces:** `{phase}-EVAL-REVIEW.md` with scored eval dimensions, gap analysis, and remediation steps
---
## v1.36.0 Features
### 108. Plan Bounce
**Command:** `/gsd-plan-phase N --bounce`
**Purpose:** After plans pass the checker, optionally refine them through an external script (a second AI, a linter, a custom validator). The bounce step backs up each plan, runs the script, validates YAML frontmatter integrity on the result, re-runs the plan checker, and restores the original if anything fails.
**Requirements:**
- REQ-BOUNCE-01: `--bounce` flag or `workflow.plan_bounce: true` activates the step; `--skip-bounce` always disables it
- REQ-BOUNCE-02: `workflow.plan_bounce_script` must point to a valid executable; missing script produces a warning and skips
- REQ-BOUNCE-03: Each plan is backed up to `*-PLAN.pre-bounce.md` before the script runs
- REQ-BOUNCE-04: Bounced plans with broken YAML frontmatter or that fail the plan checker are restored from backup
- REQ-BOUNCE-05: `workflow.plan_bounce_passes` (default: 2) controls how many refinement passes the script receives
**Configuration:** `workflow.plan_bounce`, `workflow.plan_bounce_script`, `workflow.plan_bounce_passes`
---
### 109. External Code Review Command
**Command:** `/gsd-ship` (enhanced)
**Purpose:** Before the manual review step in `/gsd-ship`, automatically run an external code review command if configured. The command receives the diff and phase context via stdin and returns a JSON verdict (`APPROVED` or `REVISE`). Falls through to the existing manual review flow regardless of outcome.
**Requirements:**
- REQ-EXTREVIEW-01: `workflow.code_review_command` must be set to a command string; null means skip
- REQ-EXTREVIEW-02: Diff is generated against `BASE_BRANCH` with `--stat` summary included
- REQ-EXTREVIEW-03: Review prompt is piped via stdin (never shell-interpolated)
- REQ-EXTREVIEW-04: 120-second timeout; stderr captured on failure
- REQ-EXTREVIEW-05: JSON output parsed for `verdict`, `confidence`, `summary`, `issues` fields
**Configuration:** `workflow.code_review_command`
---
### 110. Cross-AI Execution Delegation
**Command:** `/gsd-execute-phase N --cross-ai`
**Purpose:** Delegate individual plans to an external AI runtime for execution. Plans with `cross_ai: true` in their frontmatter (or all plans when `--cross-ai` is used) are sent to the configured command via stdin. Successfully handled plans are removed from the normal executor queue.
**Requirements:**
- REQ-CROSSAI-01: `--cross-ai` forces all plans through cross-AI; `--no-cross-ai` disables it
- REQ-CROSSAI-02: `workflow.cross_ai_execution: true` and plan frontmatter `cross_ai: true` required for per-plan activation
- REQ-CROSSAI-03: Task prompt is piped via stdin to prevent injection
- REQ-CROSSAI-04: Dirty working tree produces a warning before execution
- REQ-CROSSAI-05: On failure, user chooses: retry, skip (fall back to normal executor), or abort
**Configuration:** `workflow.cross_ai_execution`, `workflow.cross_ai_command`, `workflow.cross_ai_timeout`
---
### 111. Architectural Responsibility Mapping
**Command:** `/gsd-plan-phase` (enhanced research step)
**Purpose:** During phase research, the phase-researcher now maps each capability to its architectural tier owner (browser, frontend server, API, CDN/static, database). The planner cross-references tasks against this map, and the plan-checker enforces tier compliance as Dimension 7c.
**Requirements:**
- REQ-ARM-01: Phase researcher produces an Architectural Responsibility Map table in RESEARCH.md (Step 1.5)
- REQ-ARM-02: Planner sanity-checks task-to-tier assignments against the map
- REQ-ARM-03: Plan checker validates tier compliance as Dimension 7c (WARNING for general mismatches, BLOCKER for security-sensitive ones)
**Produces:** `## Architectural Responsibility Map` section in `{phase}-RESEARCH.md`
---
### 112. Extract Learnings
**Command:** `/gsd-extract-learnings N`
**Purpose:** Extract structured knowledge from completed phase artifacts. Reads PLAN.md and SUMMARY.md (required) plus VERIFICATION.md, UAT.md, and STATE.md (optional) to produce four categories of learnings: decisions, lessons, patterns, and surprises. Optionally captures each item to an external knowledge base via `capture_thought` tool.
**Requirements:**
- REQ-LEARN-01: Requires PLAN.md and SUMMARY.md; exits with clear error if missing
- REQ-LEARN-02: Each extracted item includes source attribution (artifact and section)
- REQ-LEARN-03: If `capture_thought` tool is available, captures items with `source`, `project`, and `phase` metadata
- REQ-LEARN-04: If `capture_thought` is unavailable, completes successfully and logs that external capture was skipped
- REQ-LEARN-05: Running twice overwrites the previous `LEARNINGS.md`
**Produces:** `{phase}-LEARNINGS.md` with YAML frontmatter (phase, project, counts per category, missing_artifacts)
**Optional integration — `capture_thought`:** `capture_thought` is a **convention, not a bundled tool**. GSD does not ship one and does not require one. The workflow checks whether any MCP server in the current session exposes a tool named `capture_thought` and, if so, calls it once per extracted learning with the signature below. If no such tool is present, the step is skipped silently and `LEARNINGS.md` remains the primary output.
Expected tool signature:
```javascript
capture_thought({
category: "decision" | "lesson" | "pattern" | "surprise",
phase: <phase_number>,
content: <learning_text>,
source: <artifact_name>
})
```
Users who run a memory / knowledge-base MCP server (for example, ExoCortex-style servers, `claude-mem`, or `mem0`-style servers) can implement this tool name to have learnings routed into their knowledge base automatically with `project`, `phase`, and `source` metadata. Everyone else can use `/gsd-extract-learnings` without any extra setup — the `LEARNINGS.md` artifact is the feature.
With `features.global_learnings: true`, phase completion runs the extraction for the just-completed phase automatically and copies the artifact to the global store at `~/.gsd/knowledge/` (#3683) — extraction and copy failures never block completion. With the gate off (the default), extraction stays fully manual.
---
### 114. Context-Window-Aware Prompt Thinning
**Purpose:** Reduce static prompt overhead by ~40% for models with context windows under 200K tokens. Extended examples and anti-pattern lists are extracted from agent definitions into reference files loaded on demand via `@` required_reading.
**Requirements:**
- REQ-THIN-01: When `CONTEXT_WINDOW < 200000`, executor and planner agent prompts omit inline examples
- REQ-THIN-02: Extracted content lives in `references/executor-examples.md` and `references/planner-antipatterns.md`
- REQ-THIN-03: Standard (200K-500K) and enriched (500K+) tiers are unaffected
- REQ-THIN-04: Core rules and decision logic remain inline; only verbose examples are extracted
**Reference files:** `executor-examples.md`, `planner-antipatterns.md`
---
### 115. Configurable CLAUDE.md Path
**Purpose:** Allow projects to store their CLAUDE.md in a non-root location. The `claude_md_path` config key controls where `/gsd-profile-user` and related commands write the generated CLAUDE.md file.
**Requirements:**
- REQ-CMDPATH-01: `claude_md_path` defaults to `./.claude/CLAUDE.md` (a valid project-scoped memory location; changed from `./CLAUDE.md` in v1.5 per [#1098](https://github.com/open-gsd/gsd-core/issues/1098) so generated content does not pollute a hand-crafted repo-root `CLAUDE.md`)
- REQ-CMDPATH-02: Profile generation commands read the path from config and write to the specified location
- REQ-CMDPATH-03: Relative paths are resolved from the project root
- REQ-CMDPATH-04: `generate-claude-md` never overwrites an existing instruction file that lacks GSD section markers (a hand-crafted file) unless `--force` is passed
**Configuration:** `claude_md_path`
---
### 116. TDD Pipeline Mode
**Purpose:** Opt-in TDD (red-green-refactor) as a first-class phase execution mode. When enabled, the planner aggressively selects `type: tdd` for eligible tasks and the executor enforces RED/GREEN/REFACTOR gate sequence with fail-fast on unexpected GREEN before RED.
**Requirements:**
- REQ-TDD-01: `workflow.tdd_mode` config key (boolean, default `false`)
- REQ-TDD-02: When enabled, planner applies TDD heuristics from `references/tdd.md` to all eligible tasks (business logic, APIs, validations, algorithms, state machines)
- REQ-TDD-03: Executor enforces gate sequence for `type: tdd` plans — RED commit (`test(...)`) must precede GREEN commit (`feat(...)`)
- REQ-TDD-04: Executor fails fast if tests pass unexpectedly during RED phase (feature already exists or test is wrong)
- REQ-TDD-05: End-of-phase collaborative review checkpoint verifies gate compliance across all TDD plans (advisory, non-blocking)
- REQ-TDD-06: Gate violations surfaced in SUMMARY.md under `## TDD Gate Compliance` section
**Configuration:** `workflow.tdd_mode`
**Reference files:** `tdd.md`, `checkpoints.md`
---
## v1.37.0 Features
### 117. Spike Command
**Command:** `/gsd-spike [idea] [--quick]`
**Purpose:** Run 2–5 focused feasibility experiments before committing to an implementation approach. Each experiment uses Given/When/Then framing, produces executable code, and returns a VALIDATED / INVALIDATED / PARTIAL verdict. Companion `/gsd-spike --wrap-up` packages findings into a project-local skill.
**Requirements:**
- REQ-SPIKE-01: Each experiment MUST produce a Given/When/Then hypothesis before any code is written
- REQ-SPIKE-02: Each experiment MUST include working code or a minimal reproduction
- REQ-SPIKE-03: Each experiment MUST return one of: VALIDATED, INVALIDATED, or PARTIAL verdict with evidence
- REQ-SPIKE-04: Results MUST be stored in `.planning/spikes/NNN-experiment-name/` with a README and MANIFEST.md
- REQ-SPIKE-05: `--quick` flag skips intake conversation and uses the argument text as the experiment direction
- REQ-SPIKE-06: `/gsd-spike --wrap-up` MUST package findings into `.claude/skills/spike-findings-[project]/`
**Produces:**
| Artifact | Description |
|----------|-------------|
| `.planning/spikes/NNN-name/README.md` | Hypothesis, experiment code, verdict, and evidence |
| `.planning/spikes/MANIFEST.md` | Index of all spikes with verdicts |
| `.claude/skills/spike-findings-[project]/` | Packaged findings (via `/gsd-spike --wrap-up`) |
---
### 118. Sketch Command
**Command:** `/gsd-sketch [idea] [--quick] [--text]`
**Purpose:** Explore design directions through throwaway HTML mockups before committing to implementation. Produces 2–3 interactive variants per design question, all viewable directly in a browser with no build step. Companion `/gsd-sketch --wrap-up` packages winning decisions into a project-local skill.
**Requirements:**
- REQ-SKETCH-01: Each sketch MUST answer one specific visual design question
- REQ-SKETCH-02: Each sketch MUST include 2–3 meaningfully different variants in a single `index.html` with tab navigation
- REQ-SKETCH-03: All interactive elements (hover, click, transitions) MUST be functional
- REQ-SKETCH-04: Sketches MUST use real-ish content, not lorem ipsum
- REQ-SKETCH-05: A shared `themes/default.css` MUST provide CSS variables adapted to the agreed aesthetic
- REQ-SKETCH-06: `--quick` flag skips mood intake; `--text` flag replaces `AskUserQuestion` with numbered lists for non-Claude runtimes
- REQ-SKETCH-07: The winning variant MUST be marked in the README frontmatter and with a ★ in the HTML tab
- REQ-SKETCH-08: `/gsd-sketch --wrap-up` MUST package winning decisions into `.claude/skills/sketch-findings-[project]/`
**Produces:**
| Artifact | Description |
|----------|-------------|
| `.planning/sketches/NNN-name/index.html` | 2–3 interactive HTML variants |
| `.planning/sketches/NNN-name/README.md` | Design question, variants, winner, what to look for |
| `.planning/sketches/themes/default.css` | Shared CSS theme variables |
| `.planning/sketches/MANIFEST.md` | Index of all sketches with winners |
| `.claude/skills/sketch-findings-[project]/` | Packaged decisions (via `/gsd-sketch --wrap-up`) |
---
### 119. Agent Size-Budget Enforcement
**Purpose:** Keep agent prompt files lean with tiered line-count limits enforced in CI. Oversized agents are caught before they bloat context windows in production.
**Requirements:**
- REQ-BUDGET-01: `agents/gsd-*.md` files are classified into three tiers: XL (≤ 1 600 lines), Large (≤ 1 000 lines), Default (≤ 500 lines)
- REQ-BUDGET-02: Tier assignment is declared in the file's YAML frontmatter (`size: xl | large | default`)
- REQ-BUDGET-03: `tests/agent-size-budget.test.cjs` enforces limits and fails CI on violation
- REQ-BUDGET-04: Files without a `size` frontmatter key default to the Default (500-line) limit
**Test file:** `tests/agent-size-budget.test.cjs`
---
### 120. Shared Boilerplate Extraction
**Purpose:** Reduce duplication across agents by extracting two common boilerplate blocks into shared reference files loaded on demand. Keeps agent files within size budget and makes boilerplate updates a single-file change.
**Requirements:**
- REQ-BOILER-01: Mandatory-initial-read instructions extracted to `references/mandatory-initial-read.md`
- REQ-BOILER-02: Project-skills-discovery instructions extracted to `references/project-skills-discovery.md`
- REQ-BOILER-03: Agents that previously inlined these blocks MUST now reference them via `@` required_reading
**Reference files:** `references/mandatory-initial-read.md`, `references/project-skills-discovery.md`
---
### 121. Knowledge Graph Integration
**Purpose:** Build, query, and inspect a lightweight knowledge graph of the project in `.planning/graphs/`. Opt-in per project. Exposed as the `/gsd-graphify` user-facing command and the `gsd-tools.cjs graphify …` programmatic verb family. Complements `/gsd-map-codebase --query` (snapshot-oriented) with a graph-oriented view of nodes and edges across commands, agents, workflows, and phases.
**Requirements:**
- REQ-GRAPH-01: Opt-in via `graphify.enabled: true` in `.planning/config.json`. When disabled, `/gsd-graphify` prints an activation hint and stops without writing.
- REQ-GRAPH-02: Slash-command `/gsd-graphify` exposes subcommands `build`, `query <term>`, `status`, `diff`. The programmatic CLI `node gsd-tools.cjs graphify …` additionally exposes `snapshot`, which is also invoked automatically as the final step of `graphify build`.
- REQ-GRAPH-03: Build runs within the configurable `graphify.build_timeout` (seconds); exceeding the timeout aborts cleanly without leaving a partial graph.
- REQ-GRAPH-04: `graphify.cjs` falls back to `graph.links` when `graph.edges` is absent so older graph artifacts keep rendering.
- REQ-GRAPH-05: Graphify is invoked through `gsd-tools.cjs graphify ...` command handlers.
- REQ-GRAPH-06: The knowledge-graph location is configurable via `graphify.graph_path` (issue #1825) so one umbrella-level cross-repo graph can serve multiple sibling projects; `query`/`status`/`diff` read the configured graph (relative to project root), with a byte-identical `.planning/graphs/` default when unset.
- REQ-GRAPH-07: `status` reports the resolved graph location as `graph_path` — the same absolute path `query`/`diff` read, after the `graphify.graph_path` override is applied — so a caller shelling out to the `graphify` CLI passes it as `--graph` instead of re-deriving the default location (issue #4836).
**Configuration:** `graphify.enabled`, `graphify.build_timeout`, `graphify.graph_path`
**Reference files:** `commands/gsd/graphify.md`, `bin/lib/graphify.cjs`
---
## v1.40.0 Features
### 122. Skill Surface Consolidation
**Purpose:** Cut the eager skill-listing overhead by folding 31 micro-skills into 4 new grouped parents and 6 existing parents that absorb sub-operations as flags. Zero functional loss — every removed micro-skill's behavior survives via a flag on a consolidated parent. After consolidation, `commands/gsd/*.md` ships 60 sub-skills (plus 6 namespace meta-skills, see #123).
**Requirements:**
- REQ-CONSOLIDATE-01: Four new grouped skills replace clusters of micro-skills:
- `/gsd-capture` — folds add-todo (default), note (`--note`), add-backlog (`--backlog`), plant-seed (`--seed`), check-todos (`--list`)
- `/gsd-phase` — folds add-phase (default), insert-phase (`--insert`), remove-phase (`--remove`), edit-phase (`--edit`)
- `/gsd-config` — folds settings-advanced (`--advanced`), settings-integrations (`--integrations`), set-profile (`--profile`)
- `/gsd-workspace` — folds new-workspace (`--new`), list-workspaces (`--list`), remove-workspace (`--remove`)
- REQ-CONSOLIDATE-02: Six existing parents absorb wrap-up / sub-operations as flags: `/gsd-update --sync`, `/gsd-update --reapply`, `/gsd-sketch --wrap-up`, `/gsd-spike --wrap-up`, `/gsd-map-codebase --fast`, `/gsd-map-codebase --query`, `/gsd-code-review --fix`, `/gsd-progress --do`, `/gsd-progress --next`.
- REQ-CONSOLIDATE-03: `/gsd-next` is not the retired workflow-advance command; it is reserved for the state-aware smart-entry launcher. Workflow advancement remains under `/gsd-progress --next`.
- REQ-CONSOLIDATE-04: Deleted micro-skill slash forms (the bare `gsd-add-todo`, `gsd-add-backlog`, `gsd-plant-seed`, `gsd-check-todos`, `gsd-add-phase`, `gsd-insert-phase`, `gsd-remove-phase`, `gsd-edit-phase`, `gsd-new-workspace`, `gsd-list-workspaces`, `gsd-remove-workspace`, `gsd-settings-advanced`, `gsd-settings-integrations`, `gsd-set-profile`, `gsd-sketch-wrap-up`, `gsd-spike-wrap-up`, `gsd-reapply-patches`, `gsd-code-review-fix`, …) MUST resolve to "Unknown command" — no shadow stubs.
- REQ-CONSOLIDATE-05: `autonomous.md` invokes `/gsd-code-review --fix` (was previously calling the deleted `gsd-code-review-fix`).
**Reference issue:** [#2790](https://github.com/open-gsd/gsd-core/issues/2790)
---
### 123. Namespace Meta-Skills (Two-Stage Routing)
**Purpose:** Replace the flat eager skill listing with a two-stage hierarchical routing layer. The model sees 6 namespace routers instead of 86 entries, selects a namespace, then routes to the sub-skill. Descriptions use pipe-separated keyword tags (≤ 60 chars) for routing density.
**Commands:**
- `/gsd-workflow` — phase pipeline router (discuss / plan / execute / verify / phase / progress / next)
- `/gsd-project` — project lifecycle (milestones, audits, summary)
- `/gsd-quality` — quality gates (code review, debug, audit, security, eval, ui)
- `/gsd-context` — codebase intelligence (map, graphify, docs, learnings)
- `/gsd-manage` — config / workspace / workstreams / thread / update / ship / inbox
- `/gsd-ideate` — exploration & capture (explore, sketch, spike, spec, capture)
**Token cost:**
| | Entries | Approx tokens |
|---|---|---|
| Pre-1.40 full install | 86 | ~2,150 |
| Namespace meta-skills | 6 | ~120 |
**Requirements:**
- REQ-NS-01: Six `commands/gsd/ns-*.md` namespace routers ship with pipe-separated keyword-tag descriptions (≤ 60 chars).
- REQ-NS-02: Existing sub-skills are unchanged and still invocable directly — namespace skills are additive, not a replacement for direct slash forms.
- REQ-NS-03: The body of each namespace router contains a routing table that maps user intent to the correct concrete sub-skill on the post-#2790 consolidated surface.
- REQ-NS-04: Tests validate namespace files exist, include matching command `requires`, and reference only existing sub-skill files.
**Reference issue:** [#2792](https://github.com/open-gsd/gsd-core/issues/2792)
---
### 124. Context-Window Utilization Guard
**Command:** `/gsd-health --context`
**Purpose:** Quality guard against context-window saturation. Two thresholds: 60 % utilization warns ("consider `/gsd-thread`"), 70 % is critical ("reasoning quality may degrade"; matches the fracture-point per recent context-attention research).
**Requirements:**
- REQ-CTX-GUARD-01: `/gsd-health --context` prints a structured status line with current utilization, threshold tier (`ok` / `warn` / `critical`), and a remediation suggestion.
- REQ-CTX-GUARD-02: The same triage is exposed as `gsd-tools.cjs validate context --tokens-used <int> --context-window <int>` — a structured envelope for status-line and hook callers (#125). Both flags are required; the handler returns the same `{ percent, state }` envelope as the pure classifier in REQ-CTX-GUARD-03.
- REQ-CTX-GUARD-03: The classifier (`bin/lib/context-utilization.cjs`) is pure: input `(tokensUsed, contextWindow)`, output `{ percent, state }`. Easy to unit-test, easy to reuse from any caller.
**Reference issue:** [#2792](https://github.com/open-gsd/gsd-core/issues/2792)
---
### 125. Phase-Lifecycle Status-Line Read-Side
**Purpose:** Surface phase orchestration state on the status-line. `parseStateMd()` reads four new STATE.md frontmatter fields and `formatGsdState()` renders in-flight, idle, and progress scenes. Write-side wiring follows in a later RC.
**Requirements:**
- REQ-LIFECYCLE-01: `parseStateMd()` reads four optional fields:
- `active_phase` — phase number when an orchestrator is in flight
- `next_action` — recommended next command when idle
- `next_phases` — YAML flow array of next phase numbers
- `progress` — nested `total_phases` / `completed_phases` / `percent` block
- REQ-LIFECYCLE-02: `formatGsdState()` checks the lifecycle fields in priority order and emits the first matching scene (Phase active → Idle next-recommended → Milestone complete → Default fallback).
- REQ-LIFECYCLE-03: All four fields default to undefined; existing STATE.md files render byte-for-byte identically.
**Reference issue:** [#2833](https://github.com/open-gsd/gsd-core/issues/2833) — see [`docs/STATE-MD-LIFECYCLE.md`](reference/state-md.md) for the full field reference and rendering rules.
---
## v1.41.0 Features
### 126. Per-Phase-Type Model Selection
**Purpose:** Express model tuning at the phase level (planning, research, execution, verification) without learning the full agent taxonomy. Sits between per-agent `model_overrides` (precise, verbose) and the global `model_profile` tier (coarse, uniform).
**Config key:** `models` in `.planning/config.json`
**Phase-type slots:**
| Slot | Agents assigned |
|------|-----------------|
| `planning` | `gsd-planner`, `gsd-roadmapper`, `gsd-pattern-mapper` |
| `discuss` | `gsd-assumptions-analyzer` |
| `research` | `gsd-phase-researcher`, `gsd-project-researcher`, `gsd-research-synthesizer`, `gsd-codebase-mapper`, `gsd-ui-researcher` |
| `execution` | `gsd-executor`, `gsd-debugger`, `gsd-doc-writer` |
| `verification` | `gsd-verifier`, `gsd-plan-checker`, `gsd-integration-checker`, `gsd-nyquist-auditor`, `gsd-ui-checker`, `gsd-ui-auditor`, `gsd-doc-verifier`, `gsd-code-reviewer` |
| `completion` | (reserved for future subagent) |
**Accepted values:** `"opus"` / `"sonnet"` / `"haiku"` / `"inherit"`
**Resolution precedence (highest → lowest):**
```text
1. model_overrides[<agent>]
2. dynamic_routing.tier_models[<tier>] (when enabled)
3. models[<phase_type>] (this feature)
4. model_profile
5. Runtime default
```
**Requirements:**
- REQ-PHASE-MODELS-01: Six named `models.*` slots accepted by `config-schema.cjs` and `config-schema.ts`; `config-set` rejects unknown phase-types.
- REQ-PHASE-MODELS-02: Configs without a `models` block behave byte-for-byte identically to pre-v1.41 behavior.
- REQ-PHASE-MODELS-03: `discuss` and `completion` are accepted by the schema for forward compatibility; setting them today is a no-op until a subagent maps to each.
**Reference issue:** [#3023](https://github.com/open-gsd/gsd-core/pull/3030)
---
### 127. Dynamic Routing with Failure-Tier Escalation
**Purpose:** Pay for the cheap tier by default; escalate to a more capable model automatically when the orchestrator detects a soft failure (verification inconclusive, plan-check FLAG, etc.).
**Config key:** `dynamic_routing` in `.planning/config.json`
**Behavior:**
- `enabled: false` (default) — feature is off; all agents use the precedence chain unchanged.
- `enabled: true` — the resolver picks `tier_models[default_tier]` for the first spawn and escalates one tier up on orchestrator-detected soft failure, capped by `max_escalations`.
**Composition:** `model_overrides` always wins; `dynamic_routing.tier_models[<tier>]` resolves above `models.<phase_type>` and `model_profile`.
**Requirements:**
- REQ-DYNROUTE-01: `dynamic_routing.enabled` acts as a master switch; when `false` or block is absent, zero behavior change.
- REQ-DYNROUTE-02: New resolver `resolveModelForTier(cwd, agent, attempt)` in `core.cjs` is the single call-site for orchestrator integration.
- REQ-DYNROUTE-03: `max_escalations` caps the escalation chain to prevent runaway cost.
**Reference issue:** [#3024](https://github.com/open-gsd/gsd-core/pull/3031)
---
### 128. Update Banner Opt-In
**Purpose:** Surface update availability to users who have declined or bypassed the GSD statusline, without requiring the statusline.
**Behavior:**
- At install time, if the installer detects no GSD statusline, it offers an opt-in `SessionStart` hook.
- The hook reads the existing `~/.cache/gsd/gsd-update-check.json` cache — the same cache used by the statusline — and prints a banner only when an update is available.
- Silent when up-to-date.
- Failure diagnostics rate-limited to once per 24 h.
- Cleanly removed by `npx @opengsd/gsd-core --uninstall`.
**Requirements:**
- REQ-BANNER-01: Banner does not install without explicit opt-in.
- REQ-BANNER-02: No additional network requests — reuses the existing background update-check cache.
- REQ-BANNER-03: Uninstall path removes the banner hook.
**Reference issue:** [#2795](https://github.com/open-gsd/gsd-core/pull/2795)
---
### 129. Issue-Driven Orchestration Guide
**Purpose:** Document a recipe for driving the full GSD workflow from a GitHub / Linear / Jira issue, mapping tracker-centric concepts onto existing GSD primitives.
**Document:** [`docs/issue-driven-orchestration.md`](issue-driven-orchestration.md)
**Covered workflow:**
1. Create an isolated workspace per issue (`/gsd-workspace --new`)
2. Run the manager dashboard to get oriented (`/gsd-manager`)
3. Execute autonomously (`/gsd-autonomous`)
4. Verify and review (`/gsd-verify-work`, `/gsd-review`)
5. Ship and close the issue (`/gsd-ship`)
No new commands or daemon process — purely a documentation artifact that maps existing primitives onto a tracker-driven workflow.
**Reference issue:** [#2840](https://github.com/open-gsd/gsd-core/pull/2840)
---
### 130. Graphify Commit-Based Staleness
**Purpose:** Surface whether the architecture graph was built from the current commit or an older one, complementing the existing mtime-based stale signal.
**Command:** `/gsd-graphify status`
**New fields returned (graphify v0.7+ graphs):**
| Field | Type | Description |
|-------|------|-------------|
| `built_at_commit` | string | Commit SHA the graph was built from |
| `current_commit` | string | Current `git HEAD` |
| `commits_behind` | number | How many commits behind HEAD the graph is |
| `commit_stale` | boolean \| null | `true`=stale, `false`=current, `null`=unavailable (pre-v0.7, non-git) |
**Rendered output (when signal is available):**
```
Source commit: abc1234 (3 commits behind HEAD)
```
**Security:** `built_at_commit` validated as 4–40 hex chars before reaching `git` — a hostile `graph.json` cannot inject dashed options into argv.
**Fallback:** pre-v0.7 graphs and non-git checkouts return `commit_stale: null`; callers fall back to the existing mtime-based `stale` flag. No behavior change for existing users.
**Reference issue:** [#3170](https://github.com/open-gsd/gsd-core/issues/3170)
---
## v1.42.1 Features
### 132. Package Legitimacy Gate
**Purpose:** Stop hallucinated, suspicious, or slopsquatting package names before they reach a shell install command.
**Behavior:**
- Phase research writes a `## Package Legitimacy Audit` table for recommended packages.
- Packages verified only through search are treated as `[ASSUMED]`, not trusted.
- `[SLOP]` packages are removed from recommendations.
- Plans that need `[ASSUMED]` or suspicious packages add a human verification checkpoint.
- Executor install failures stop for human verification instead of auto-trying similarly named packages.
**Requirements:**
- REQ-PKG-GATE-01: Research MUST record package registry, age, download/source signals, legitimacy verdict, and disposition.
- REQ-PKG-GATE-02: Planner MUST gate unverified or suspicious package installs before execution.
- REQ-PKG-GATE-03: Executor MUST NOT auto-substitute package names after failed package-manager installs.
**Reference:** [v1.42.1 Release Notes](RELEASE-NOTES-LEGACY.md)
---
### 133. Skill Surface Budgeting
**Purpose:** Let users reduce installed skill and agent surface area when context budget matters.
**Install profiles:**
| Profile | Purpose |
|---------|---------|
| `core` | Minimal main-loop surface |
| `standard` | Core plus common phase-management commands |
| `full` | Complete surface; default |
**Runtime control:** `/gsd-surface` lists profile state and enables, disables, or resets skill clusters without reinstalling.
**Requirements:**
- REQ-SURFACE-01: Installer MUST resolve `--profile=<name>` and persist the active profile in `.gsd-profile`.
- REQ-SURFACE-02: `--minimal` and `--core-only` MUST remain aliases for `--profile=core`.
- REQ-SURFACE-03: Runtime surface state MUST persist outside the install profile marker.
**Reference:** [ADR-0011](adr/0011-skill-surface-budget-module.md)
---
### 134. Installer Migrations
**Purpose:** Make runtime config cleanup explicit, auditable, and rollback-aware during installs and updates.
**Capabilities:**
- First-time baseline migration records managed files.
- Legacy stale-file cleanup uses ownership evidence before deleting or rewriting.
- User-owned artifacts are preserved.
- Ambiguous GSD-looking files block with a clear report instead of being silently overwritten.
- Migration plans support dry-run reporting and rollback protection.
**Requirements:**
- REQ-INSTALL-MIGRATION-01: Migration records MUST include metadata, install scope, and ownership evidence.
- REQ-INSTALL-MIGRATION-02: Destructive actions MUST fail closed when ownership is ambiguous.
- REQ-INSTALL-MIGRATION-03: Install failures MUST restore the pre-install state when rollback data exists.
**Reference:** [Installer Migrations](installer-migrations.md)
---
### 135. Custom Ship PR Body Sections
**Command:** `/gsd-ship`
**Config key:** `ship.pr_body_sections`
**Purpose:** Add project-specific PRD-style sections to generated PR bodies without editing GSD workflow files.
**Behavior:** Configured sections append after the required `Summary`, `Changes`, `Requirements Addressed`, `Verification`, and `Key Decisions` sections. They can copy from artifact headings, render templates, or fall back to static text.
**Requirements:**
- REQ-SHIP-SECTIONS-01: Custom sections MUST NOT replace, remove, or reorder required PR sections.
- REQ-SHIP-SECTIONS-02: Unknown template tokens MUST be rejected by config validation.
- REQ-SHIP-SECTIONS-03: Disabled sections MUST stay in config without appearing in PR output.
**Reference:** [Custom PR Body Sections](ship-pr-body-sections.md)
---
### 136. Review Default Reviewers
**Command:** `/gsd-review`
**Config key:** `review.default_reviewers`
**Purpose:** Let teams choose the default reviewer subset for no-flag `/gsd-review` runs.
**Precedence:**
```text
explicit reviewer flags -> --all -> review.default_reviewers -> all detected reviewers
```
**Requirements:**
- REQ-REVIEW-DEFAULTS-01: Missing `review.default_reviewers` MUST preserve the previous all-detected behavior.
- REQ-REVIEW-DEFAULTS-02: Empty arrays MUST be rejected; remove the key to restore all-detected behavior.
- REQ-REVIEW-DEFAULTS-03: Known but unavailable reviewers MUST be skipped with diagnostics rather than hard-failing the run.
**Reference:** [Configuration Reference](CONFIGURATION.md#reviewer-defaults-for-gsd-review)
---
### 137. Fallow Structural Review Pre-Pass
**Command:** `/gsd-code-review`
**Config keys:** `code_quality.fallow.*`
**Purpose:** Add an optional structural analysis pass before the agent review.
**Behavior:** When enabled, GSD resolves a `fallow` binary, runs a bounded audit, writes `FALLOW.json`, and embeds structural findings in `REVIEW.md`.
**Requirements:**
- REQ-FALLOW-01: Fallow MUST be opt-in and disabled by default.
- REQ-FALLOW-02: Missing or failing fallow runs MUST produce clear diagnostics.
- REQ-FALLOW-03: Findings larger than the embed budget MUST be skipped with a warning, preserving the raw JSON artifact.
**Reference:** [Configuration Reference](CONFIGURATION.md#code-quality-settings)
---
### 138. End-of-Phase Human Verification Mode
**Config key:** `workflow.human_verify_mode`
**Purpose:** Reduce mid-flight human checkpoint interruptions while preserving human verification requirements.
**Behavior:** The default `"end-of-phase"` mode embeds human checks into `<verify><human-check>` blocks for phase review. `"mid-flight"` restores blocking `checkpoint:human-verify` tasks.
**Requirements:**
- REQ-HUMAN-VERIFY-01: `checkpoint:decision` and `checkpoint:human-action` MUST remain blocking regardless of mode.
- REQ-HUMAN-VERIFY-02: Human-needed verification MUST remain pending until the end-of-phase review resolves it.
- REQ-HUMAN-VERIFY-03: Configs without the key MUST use `"end-of-phase"`.
**Reference:** [Checkpoints Reference](../gsd-core/references/checkpoints.md)
---
### 139. Quota and Rate-Limit Failure Classification
**Command:** `/gsd-execute-phase`
**Purpose:** Treat provider quota and rate-limit failures as wait-and-resume conditions, not normal executor failures.
**Behavior:** Agent output is classified for signals such as `429`, `rate limit`, `usage limit`, `RESOURCE_EXHAUSTED`, and `usage_limit_reached`. Matching failures present a wait-for-reset recovery path.
**Requirements:**
- REQ-QUOTA-01: Quota failures MUST NOT offer immediate retry as the primary recovery.
- REQ-QUOTA-02: Classification MUST cover Claude, Copilot, Codex, and generic provider sentinels.
- REQ-QUOTA-03: Non-quota failures MUST continue through the normal execution failure path.
**Reference:** [Provider Rate Limit Signals](research/provider-rate-limit-signals.md)
---
### 140. Statusline Context Position
**Config key:** `statusline.context_position`
**Purpose:** Keep the context meter visible in narrow terminals.
**Options:**
| Value | Behavior |
|-------|----------|
| `"end"` | Default; render context meter near the line tail |
| `"front"` | Render context meter immediately after the model name |
**Requirements:**
- REQ-STATUSLINE-POS-01: Invalid values MUST be rejected by config validation.
- REQ-STATUSLINE-POS-02: Missing config MUST preserve existing end-position rendering.
**Reference:** [Configuration Reference](CONFIGURATION.md#statusline-settings)
---
### 141. Milestone Tag Creation Toggle
**Command:** `/gsd-complete-milestone`
**Config key:** `git.create_tag`
**Purpose:** Let projects with external release automation complete milestones without creating local git tags.
**Behavior:** `git.create_tag: false` skips milestone tag creation. The workflow still updates milestone artifacts and state.
**Requirements:**
- REQ-MILESTONE-TAG-01: Missing config MUST preserve automatic tag creation.
- REQ-MILESTONE-TAG-02: Existing tag collisions MUST fail clearly instead of overwriting tags.
- REQ-MILESTONE-TAG-03: Disabling tag creation MUST NOT skip milestone archival.
**Reference:** [Configuration Reference](CONFIGURATION.md#git-branching)
---
### 142. Structured JSON Error Mode
**CLI:** `gsd-tools --json-errors`
**Purpose:** Give automation callers stable machine-readable error envelopes.
**Behavior:** Commands that fail under `--json-errors` return structured `ok: false` payloads with error kind, message, command context, and exit mapping instead of prose-only stderr.
**Requirements:**
- REQ-JSON-ERRORS-01: Unknown commands, validation errors, timeouts, native failures, fallback failures, and internal errors MUST map to canonical error kinds.
- REQ-JSON-ERRORS-02: CLI exit code mapping MUST remain stable for automation callers.
- REQ-JSON-ERRORS-03: Human-readable output MUST remain the default when `--json-errors` is absent.
**Reference:** [JSON Error Mode](json-errors.md)
---
### 143. UAT-Passed Predicate
**CLI:** `node gsd-tools.cjs phase uat-passed <N> [--require-verification]`
**Purpose:** Provide a runtime-neutral, automatable predicate that evaluates HUMAN-UAT results for a phase and returns a structured pass/fail verdict with full diagnostic detail.
**Behavior:** Locates `*-UAT.md` and optionally `*-VERIFICATION.md` files for the given phase, parses UAT test blocks (heading-block parser, column-0 result lines) with a markdown-aware stripper that removes false-positive contexts (YAML frontmatter, fenced code blocks, HTML comments, and blockquotes). Returns `passed: true` only when at least one check exists AND all checks pass AND no blockers — fail-closed, no vacuous pass. The `--require-verification` flag requires at least one `*-VERIFICATION.md` with an allowlisted passing status; the command fails without one.
**Output envelope:** `{ passed, uat_files[], verification_files[], checks[], blockers[], no_uat_artifacts, policy: { require_verification } }`
| Field | Type | Description |
|-------|------|-------------|
| `passed` | `boolean` | `true` only when ≥1 check exists AND all passing AND no blockers |
| `uat_files` | `string[]` | Filenames of `*-UAT.md` files evaluated |
| `verification_files` | `string[]` | Filenames of `*-VERIFICATION.md` files evaluated |
| `checks[]` | `{ file, test, name, result, passing }[]` | Per-item results from heading blocks |
| `blockers[]` | `string[]` | Human-readable failure reasons (frontmatter, failing/missing items, policy, malformed markdown) |
| `no_uat_artifacts` | `boolean` | `true` when no test items were parsed; `passed` is always `false` when `true` |
| `policy.require_verification` | `boolean` | Whether `--require-verification` was active |
**Requirements:**
- REQ-UAT-PRED-01: The predicate MUST ignore result lines inside YAML frontmatter, fenced code blocks, HTML comments, and blockquotes.
- REQ-UAT-PRED-02: `passed: true` MUST require at least one check AND all checks passing AND no blockers (fail-closed, no vacuous pass).
- REQ-UAT-PRED-03: `--require-verification` MUST cause the command to fail when no `*-VERIFICATION.md` file with an allowlisted passing status is found.
- REQ-UAT-PRED-04: `blockers[]` contains all human-readable failure reasons including frontmatter issues, policy violations, and malformed markdown — NOT limited to a subset of `checks[]`.
- REQ-UAT-PRED-05: The module MUST be runtime-neutral (no runtime-specific env checks or exit shortcuts).
- REQ-UAT-PRED-06: A heading block with no column-0 `result:` line emits `result:'missing'` (blocker); test items are never silently dropped.
**Reference:** [Phase Management Commands](COMMANDS.md#phase-uat-passed-n---require-verification)
---
### 144. Spec-Phase Edge-Completeness Probe
**Command:** `/gsd-spec-phase`
**Purpose:** Surface the omitted domain-boundary edges that silently invalidate a requirement — touching intervals, empty inputs, rounding ties, grapheme truncation — before they become production defects. Runs as `Step 5.5` of spec-phase, after the ambiguity gate.
**Behavior:** For each SPEC requirement the probe classifies its data/behavior shape, then raises only the *applicable* categories from a closed 8-category taxonomy (boundary, adjacency, empty, encoding, ordering, precision, idempotency, concurrency) via a relevance filter. Each raised category proposes one concrete candidate edge, which the author resolves to exactly one of four states:
| State | Meaning | Downstream effect |
|-------|---------|-------------------|
| `covered` | An acceptance criterion handles the edge | Pass/fail line written into the SPEC Acceptance Criteria block; lifted into `plan-phase` `must_haves.truths` |
| `dismissed` | The edge cannot occur (requires a non-empty reason) | Recorded with its reason; empty dismissals are rejected |
| `backstop` | Intent recorded, needs a held-out/property-based test | Lifted into `must_haves.truths` as a non-inferable check |
| `unresolved` | Deferred | Soft-gates the spec; row stamped `⚠ Edge unresolved — planner must treat as assumption` |
When a requirement's prose matches **no** shape cue, the probe does not silently drop it (#1110): it emits a single `unclassified — review manually` candidate so the zero-cue requirement is surfaced for the author to resolve like any other (specify / dismiss-with-reason / defer) — a manual-review nudge, not a hard block.
**Non-English projects: the probe reads English via `text_en`, the SPEC does not have to (#2773, durable fix #3717).** The shape cues are English word-boundary patterns, so a project running with [`response_language`](CONFIGURATION.md) set would otherwise have *every* requirement match nothing, classify to zero shapes, and land in `unclassified` — the taxonomy silently contributing nothing to exactly the kind of spec it exists to harden. `spec-phase` Step 5.5 therefore populates an optional `text_en` field alongside each requirement's `text` with a faithful **English translation**: `text_en` is engine input, never user-facing output, so it is translated while `text` keeps the requirement's own wording (the SPEC stays in the original language), requirement ids are left untouched, and any acceptance criteria written back from the resolved edges return to `response_language`. Translation makes the classifier *applicable*; it does not make it omniscient. A requirement carrying no shape cue in **any** language still classifies to zero — that is the classifier's recorded recall gap (ADR-857 §98), not a translation failure — and the remedy there is the same one an English project uses: author an explicit `shapes` array on the requirement instead of relying on prose classification.
The resolved edges populate a `## Edge Coverage` section in `SPEC.md`. Unresolved *applicable* edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under `--auto`, the probe **never auto-dismisses** — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs `[auto] edge coverage: C covered, B backstop, U unresolved`. The one exception is an `unclassified` candidate: `--auto` leaves it **`unresolved`** (surfaced as a flagged assumption), never auto-`backstop` — a missing shape is not evidence an edge exists, so minting a held-out edge obligation would be a false claim.
The load-bearing wire is the `plan-phase` lift: `covered` and `backstop` edges become `must_haves.truths` the verifier can check, so the section is not merely documentation. A `backstop` edge is lifted as a **structured non-inferable marker** (`{ statement, verification: backstop }`, a flat scalar — not a prose note), which the **honest verifier** then consumes (see below) — closing the loop the edge-probe opened.
**Honest verifier — abstention on non-inferable checks (#1154).** A non-inferable (`backstop`) truth is one whose correct behavior is not derivable from the spec alone, so the verifier cannot self-detect the gap and would confidently false-pass it (~100% of the time). At verify time, a `backstop` truth the verifier cannot confirm with **explicit evidence** (a passing wired held-out/property-based test, or a directly-observed behavior) **abstains** → `human_needed` with reason `insufficient_spec` (reported as `unverified — held-out test recommended`), **never a silent `passed`**. This is the verify-time, truth-axis mirror of the prohibition judgment-tier disposition (ADR-550 D4): exogenous (driven by the `backstop` tag, never a self-judged "abstain if unsure"), routing-not-diagnosis (the held-out test carries the omitted rule), and capable-tier dependent (reliable on `sonnet`+; the budget `haiku` tier degrades toward current behavior). An inferable truth is never abstained (the over-abstention guard). Reference: [Honest Verifier](../gsd-core/references/honest-verifier.md).
**Requirements:**
- REQ-EDGE-01: The edge pass MUST run after the ambiguity gate and emit a `## Edge Coverage` SPEC section.
- REQ-EDGE-02: The relevance filter MUST raise only applicable categories; each raised edge resolves to exactly one of covered/dismissed/backstop/unresolved.
- REQ-EDGE-03: A `dismissed` resolution MUST require a non-empty reason.
- REQ-EDGE-04: An unresolved applicable edge MUST trigger the soft gate; write-anyway stamps the row as a planner assumption.
- REQ-EDGE-05: `--auto` MUST never auto-dismiss — auto-cover or auto-backstop only.
- REQ-EDGE-06: `plan-phase` MUST lift `covered` criteria and `backstop` notes into `must_haves.truths`.
- REQ-EDGE-07: A requirement whose prose matches no shape cue MUST surface an `unclassified — review manually` candidate (never silently dropped); `--auto` MUST leave it `unresolved`, never auto-`backstop`.
- REQ-EDGE-08: `plan-phase` MUST lift a `backstop` edge into `must_haves.truths` as a structured flat-scalar marker (`{ statement, verification: backstop }`), never a prose parenthetical.
- REQ-HONEST-01: At verify time a `backstop` truth that cannot be confirmed with explicit evidence MUST abstain → `human_needed` (reason `insufficient_spec`), never `passed`; an inferable truth MUST never be abstained (over-abstention guard); abstention MUST be exogenous (driven by the `backstop` tag, not self-judgment).
**Reference:** [Edge Probe](../gsd-core/references/edge-probe.md)
---
## v1.43.0 Features
### 145. MemPalace Memory Capability
**Purpose:** Opt-in cross-session and cross-project memory via the [MemPalace](https://github.com/MemPalace/mempalace) external service (local-first, MCP + CLI). Wires deliberate recall before discuss/plan and verbatim capture + temporal-KG sync at phase boundaries through the ADR-857 capability mechanism. Default-resilient: disabled by default, every hook is `onError: skip`, and an absent MemPalace installation leaves the loop unchanged.
**Commands:** `/gsd-mempalace-recall`, `/gsd-mempalace-capture`
**Requirements:**
- REQ-MP-01: Opt-in via `mempalace.enabled: true`. Default `false` — the loop is unchanged when unset.
- REQ-MP-02: At `plan:pre`, skill `mempalace-recall` produces `MEMORY-RECALL.md` from prior decisions, patterns, and surprises retrieved via wake-up + semantic search + KG timeline. When MemPalace is unreachable, writes an "unavailable" stub and continues.
- REQ-MP-03: At `discuss:post`, `plan:post`, and `verify:post`, skill `mempalace-capture` files the phase artifact verbatim into the appropriate MemPalace room (`decisions`, `planning`, `milestones`). Capture is idempotent via `mempalace_check_duplicate`.
- REQ-MP-04: At `ship:post`, agent `gsd-mempalace-curator` writes a diary entry, proposes cross-project tunnels (when `mempalace.cross_project_tunnels: true`), and runs wing-scoped sync pruning.
- REQ-MP-05: `mempalace.memory_mode` has three wired values: `augment` (default — palace is an additive recall layer alongside GSD native memory, which stays authoritative), `kg_backend` (knowledge-graph queries resolve against the palace's temporal KG as the primary source, `.planning/graphs/` as fallback; non-KG drawer recall stays additive), `replace` (recall resolves through the palace as the source of truth, native memory as fallback). Every mode is `onError:skip` and default-resilient — an unreachable palace degrades to native memory and GSD keeps writing `.planning/graphs/`, so no mode loses memory. Cross-mode migration of existing `.planning/graphs/` into the palace is out of scope (not yet implemented).
- REQ-MP-06: Every hook is `onError: skip`. No hook carries `blocking: true`. Memory never halts or fails a phase.
- REQ-MP-07: Interactive runs prefer MCP tools; headless/cron runs prefer the MemPalace CLI (`mempalace wake-up`, `mempalace search`, `mempalace mine`, `mempalace sync`).
- REQ-MP-08: `mempalace.auto_capture_hooks` is **forward-declared and not yet functional**. No native Claude Code hooks (`stop`, `precompact`, `session-start`) are installed by this key; the capability's hooks array is empty. This key is reserved for the future "Connected Capability" phase. Default `false`.
**Configuration:** `mempalace.enabled`, `mempalace.memory_mode`, `mempalace.wing`, `mempalace.recall_on_discuss`, `mempalace.recall_on_plan`, `mempalace.capture_artifacts`, `mempalace.mirror_kg`, `mempalace.cross_project_tunnels`, `mempalace.diary_journal`, `mempalace.auto_capture_hooks`
See [Configuration Reference](CONFIGURATION.md#mempalace-settings) for full schema and [How to enable cross-session memory with MemPalace](how-to/enable-cross-session-memory-with-mempalace.md) for a setup walkthrough.
---
### 146. Spec-Phase Prohibition Probe
**Command:** `/gsd-spec-phase`
**Purpose:** Surface the unwritten *must-NOT* constraints — the values/safety/ethics interpretations a feature could silently become that the author would never want but the spec does not forbid — before any code is written. The edge probe reaches data-shape edges; it structurally cannot reach prohibitions. This is the missing instrument, running as `Step 5.6` of spec-phase, after the edge probe.
**Behavior:** A two-stage, prose-orchestrated pass per requirement (no compiled recall engine — recall is inherently model-driven, ADR-550 D7b):
1. **Recall (adversarial probe):** *"What could this feature silently become that the author would NOT want, but the spec does not forbid?"* — model-robust open-vocabulary elicitation across values/safety/ethics.
2. **Precision (one-pass classifier):** drop routine-engineering items, keep genuine values/safety/ethics prohibitions — collapses the raw list to the load-bearing few.
Each surfaced prohibition is resolved to exactly one of three states:
| State | Meaning | Downstream effect |
|-------|---------|-------------------|
| `resolved` | Confirmed a real must-NOT | NEGATIVE acceptance criterion written into the SPEC `## Prohibitions (must-NOT)` section; lifted into `plan-phase` `must_haves.prohibitions` (its own sibling block, never `truths`) |
| `dismissed` | Not a genuine prohibition (requires a non-empty reason) | Recorded with its reason; empty dismissals are rejected |
| `unresolved` | Deferred | Soft-gates the spec; surfaced as a planner assumption |
Each resolved prohibition carries a `verification` tier — `test` (a negative test can enforce it) or `judgment` (only human/LLM judgment can). At verify time, judgment-tier prohibitions route to a never-silent / never-hard-halt soft gate (autonomous emits an `unverified-prohibition — human review recommended` flag); test-tier prohibitions are enforced via the deterministic `check prohibition-enforcement` gate — green when the wired negative test / lint rule passes, hard-gate (flagged, non-green) when missing or failing, in both interactive and autonomous modes (#1259, ADR-550 D5d). Under `--auto`, the probe **never auto-dismisses**. Canon-bound concerns (OWASP / GDPR / fairness) are referred to `/gsd-secure-phase` rather than minting SPEC prohibitions (ADR-550 D6).
The load-bearing wire is the `plan-phase` lift into `must_haves.prohibitions`, so the section is not merely documentation.
**Deterministic prohibition-check descriptor source (#1278).** A resolved `test`-tier prohibition MAY carry an optional **`check` descriptor** — the flat-scalar keys `check_kind` (`node-test` | `lint-rule`), `check_target`, and `check_rule` (lint-rule only) — authored at spec-phase. `projectProhibitions` projects these scalars deterministically and verify-phase reads them back to locate the check handed to `check prohibition-enforcement`, so a wired, passing test closes the gap with **zero manual descriptor authoring** (previously the verify-phase LLM had to invent `{kind, target, rule}` each run, #1259). The descriptor is **optional and backward-compatible** — a descriptor-less prohibition parses and disposes byte-identically to today — and **fail-closed**: a partial, invalid, or absent descriptor falls through to the producer's existing fail-closed locate, never a silent green. `failFirst` stays a verify-time caller attestation (machine-proven fail-first is tracked in #1279).
**Requirements:**
- REQ-PROHIB-01: The prohibition pass MUST run after the edge probe and emit a `## Prohibitions (must-NOT)` SPEC section.
- REQ-PROHIB-02: Stage 1 MUST ask the adversarial recall question; Stage 2 MUST drop routine-engineering items and keep values/safety/ethics prohibitions.
- REQ-PROHIB-03: A `dismissed` resolution MUST require a non-empty reason.
- REQ-PROHIB-04: `--auto` MUST never auto-dismiss.
- REQ-PROHIB-05: `plan-phase` MUST lift resolved prohibitions into `must_haves.prohibitions` (never `truths`).
- REQ-PROHIB-06: A well-formed but unwired `test`-tier prohibition MUST fail closed at verify time — never a silent pass.
- REQ-PROHIB-07: A `test`-tier prohibition with a **machine-proven-fail-first**, genuinely-passing (non-vacuous) wired mechanical check (a `node --test` negative test OR a lint/AST rule) MUST dispose green and be satisfiable; a missing, un-provable, or non-passing check MUST hard-gate (flagged, non-green) in both interactive and autonomous modes. Fail-first is **machine-proven, not caller-attested** (#1279, ADR-550 D5d): before a clean pass greens, the producer independently runs the wired check against a known violation (the descriptor's `violationFixture`) and confirms it goes RED — a lint rule via the violating fixture, a node test via the violating subject injected through the `GSD_PROHIB_SUBJECT` convention; absent a violation source it fails closed, never falling back to attestation. (Enforcement half shipped #1259; deterministic descriptor auto-locate in #1278.)
**Reference:** [Prohibition Probe](../gsd-core/references/prohibition-probe.md)
---
### 147. Capability Management Command
**Command:** `gsd capability install | update | remove | list | outdated | disable | enable`
**Purpose:** The user-facing CLI for the ADR-1244 capability ecosystem — install, upgrade, remove, list, check for updates, and toggle GSD capabilities (first-party and third-party overlays) from a registry / git / npm / tarball / local source. Wires the Phase-3/4 lifecycle library (source resolver, install ledger, trust gate) to a command users actually run.
**Behavior:**
- `install <spec> [--integrity sha512-…] [--scope global|project] [--yes] [--shared-file <rel>]…` — resolve (copy-only) → verify integrity / SHA pin → `engines.gsd` gate → disclose executable surfaces → consent (`--yes` grants; without it an executable install aborts after printing the disclosure and writes nothing) → validate → extract → record the ledger.
- `update [<id> | --all] [--scope] [--yes]` — re-resolve the capability's recorded source and upgrade via atomic stage-then-swap; re-consent when the executable set changed; `--all` reports a per-capability outcome and exits non-zero on any partial failure.
- `remove <id> [--purge-data] [--scope]` — strip the ledger-recorded files + marker-isolated shared edits; first-party capabilities are rejected (use the product uninstaller).
- `list [--json]` — first-party + installed overlay capabilities (both scopes) as a JSON array.
- `outdated [--json] [--scope]` — light remote peek of each installed overlay's recorded source (ADR-1244 D6 per-source matrix: git `ls-remote --tags`, npm `view … version` resolving the highest version matching the recorded range, local re-read; tarball → `manual`, registry → `unknown`) reporting `outdated` / `current` / `pinned` / `manual` / `unknown` per capability. A source pinned to an immutable ref (git `#sha:` or `#tag:`, or an exact npm version) is reported `pinned`. A bare git `#<ref>` is classified at the remote: if it resolves exclusively under `refs/tags/` it is an immutable tag → `pinned`; if it resolves to a mutable branch (or is ambiguous) it is `unknown`. Bounded subprocesses (git ≤30s, npm ≤60s) and a failing peek degrades that row to `unknown` without crashing the command. `--json` for machine output, default for a table.
- `disable | enable <id>` — toggle activation state (equivalent to `gsd capability set <id> --off` / `--on`).
**Trust boundary:** install never executes capability code (copy-only staging); executable surfaces require explicit consent; sources are gated by the **project-scoped** `capabilities.strict_known_registries` policy (fail-closed on a malformed/unparseable value); every shared-config write/delete is realpath-confined to the scope root, and a name collision with a user's `mcpServers` entry is never clobbered.
**Reference:** [`gsd capability` command reference](reference/gsd-capability-command.md) · [ADR-1244](adr/1244-capability-ecosystem.md)
---
### 148. Smart Entry Launcher
**Command:** `/gsd-next`
**Tool:** `gsd-tools smart-entry [--json]`
**Purpose:** Provide a state-aware front door that reads project/workflow state, classifies the user's situation, presents a short menu, and dispatches exactly one existing GSD command.
**Requirements:**
- REQ-SMART-ENTRY-01: Detection MUST be read-only and deterministic; classification lives in `gsd-tools smart-entry`.
- REQ-SMART-ENTRY-02: The launcher MUST never perform project work directly; it only displays a menu and dispatches one command.
- REQ-SMART-ENTRY-03: The workflow MUST fall back to `/gsd-progress` if detection fails.
- REQ-SMART-ENTRY-04: Each classified situation MUST provide exactly one recommended action and valid slash commands.
- REQ-SMART-ENTRY-05: Text-mode runtimes MUST receive a numbered-list fallback instead of being stranded by interactive UI assumptions.
**Situations:** no project, paused, blocked, verify failed, needs first phase, planning, executing, verify pending, idle stranded, complete, unknown.
**Reference:** [Smart Entry Design](superpowers/specs/2026-06-27-gsd-smart-entry-design.md)
---
## v1.7.0 Features
> These are features new to **@opengsd/gsd-core 1.7.0** (the current release line: 1.0.0 → 1.2.0 → … → 1.6.1 → 1.7.0). The preceding `v1.27`–`v1.43.0` sections use the retired get-shit-done-cc / get-shit-done-redux feature numbering and are not gsd-core releases — see [Legacy Release Notes](RELEASE-NOTES-LEGACY.md).
### 149. Embeddable Orchestration System (Host-Integration Interface)
**Purpose:** Express every host integration against one public, versioned contract (ADR-1239 Phase A, #1690) instead of bespoke per-host wiring, so onboarding a new host becomes additive descriptor work.
**Behavior:** The interface exposes six interface points (`command`, `dispatch`, `model`, `hooks`, `state`, `artifact`), eight negotiated axes, and a `PROTOCOL_VERSION` handshake that negotiates down to `min(host, engine)`. In 1.7.0, 14 runtimes were migrated onto the interface via imperative adapters (OpenCode #2087, Cursor #2089, Cline #2090, Hermes #2091, Qwen #2092, Kilo #2093, Trae #2094, Kimi #2095, Antigravity #2096, Augment #2097), a declarative adapter (Codex #2088), plus full lifecycle-hook wiring for CodeBuddy (#2098), GitHub Copilot (#2099), and Windsurf (#2100). Descriptors gained an `extensionEvents` vocabulary (#1946), and `/gsd-surface` now reproduces a runtime's agent output byte-for-byte from the installer's descriptors (#1575).
**New runtimes:** ZCode (Z.ai — Agentic Development Environment for GLM-5.2, #1925), pi (`npx @opengsd/gsd-core --pi`, #2102), and a repo-local VS Code extension driven through the adapter (#2103). The retired Gemini CLI now redirects to Antigravity CLI, its official successor (#1928).
**Reference:** [The Embeddable Orchestration System](explanation/embeddable-orchestration-system.md) · [Host-Integration Interface](reference/host-integration-interface.md) · [Interface versioning policy](explanation/interface-versioning-policy.md)
---
### 150. Discoverability Registries
**Purpose:** Two non-endorsing catalogs for third-party extensions (#2182).
**Behavior:** The **Community Capability Registry** (#2188) lists third-party Feature Capabilities installed with `gsd capability install`; the **EoS Registry** (#2193) lists third-party host integrations built on the ADR-1239 interface. Every entry embeds a live release badge and links to a GitHub Discussion. Registration is a documentation PR, regenerated with `npm run gen:registry`.
**Reference:** [GSD Registries](registries/README.md)
---
### 151. Companion MCP Server
**Command:** `gsd-mcp-server`
**Purpose:** A companion MCP server exposing GSD over stdio JSON-RPC 2.0, covering interface points 1 and 5 (#1681).
**Behavior:** OpenCode installs auto-register it as `mcp.gsd` (#1682). OpenCode also gained the `opencode-subset` hook dialect plus `session.idle` handling (#1682) and now runs GSD's lifecycle safety hooks — prompt-injection guard, read-before-edit guard, and injection scanner (#1923).
---
### 152. Statusline Token Count & Git Segment
**Purpose:** Opt-in statusline additions surfacing more session context.
**Behavior:** An absolute token count on the context meter (#2161) and a git branch + working-state segment (#2163), both opt-in. A companion opt-in **compact GSD-state format** condenses the GSD state segment (#2162).
**Configuration:** `statusline.*`
---
### 153. Model Catalog Advances
**Purpose:** Refresh the default model tiers and how models are surfaced.
**Behavior:** Codex/OpenAI defaults advance to the **GPT-5.6 family (Sol / Terra / Luna)** (#2122); the verbose `(1M context)` model suffix collapses to a compact `(1M)` badge (#2160). GSD warns when model config changes without re-running the installer on static-frontmatter runtimes such as Codex and OpenCode (#1688).
**Reference:** [Configuration](CONFIGURATION.md) · [Configure model profiles](how-to/configure-model-profiles.md)
---
### 154. Claude Orchestration Capability (BETA)
**Purpose:** A default-off, BETA, Claude-only capability that adopts Claude Code's Workflow tool for parallel sub-agent orchestration (#1143).
**Reference:** [The Claude orchestration capability](explanation/claude-orchestration-capability.md)
---
### 155. External-Job Capability
**Purpose:** A default-off capability that externalizes long-running compute as asynchronous external jobs, e.g. SLURM submission (#1165).
**Configuration:** `external_job.submit_timeout_ms`, `external_job.poll_timeout_ms`, `external_job.artifact_dir` (#1164)
---
### 156. API-Coverage Gate
**Command:** `/gsd-verify-work`
**Purpose:** A phase that integrates an external API, SDK, or service can no longer seal verification without a decided coverage matrix (#1562).
**Behavior:** At seal time the gate reads the phase scope — the plan bodies, falling back to this phase's ROADMAP section — and runs the deterministic detector over it. An integration signal without a `COVERAGE.md` matrix blocks the seal; no signal passes.
**Unestablished scope is not a negative verdict (#3909).** A phase with no plan body *and* no roadmap section gives the detector nothing to examine. The gate used to run detection over zero bytes and pass, certifying "no external-API integration" from a probe that never looked. It now holds the seal instead, reporting `scope_unavailable: true`. A phase whose plans are real and simply contain no API vocabulary is unaffected — the discriminator is *bytes examined*, never *signals found*.
**Breaking change:** a phase that previously sealed because its detector could not establish a scope is now correctly held. Add the phase plan, or record a reasoned `No external API integration: <reason>` declaration in `COVERAGE.md`. See [Resolve a skipped capability probe](how-to/resolve-a-skipped-capability-probe.md).
---
### 157. State Rebuild & Configurable Graph Path
**Behavior:** A new `gsd-tools state rebuild` subcommand re-derives `STATE.md` from source (#1830). The new `graphify.graph_path` setting makes the knowledge-graph location configurable, so a single umbrella graph can serve several projects (#1825).
---
### 158. Broken-Windows Ledger
**Behavior:** A cross-phase defect register at `.planning/WINDOWS.md` accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths (#1950). `/gsd-ship` blocks while any entry is `open`; an entry can be `waived` only with a recorded reason (auditable) or marked `fixed` (removed from the blocking set). `/gsd-progress` surfaces the open + waived counts.
**Commands:** `gsd-tools windows status | append | waive | fixed`.
**Config:** `workflow.windows_enforce` (gate active, default `false` — opt-in enforcement). Enable with `gsd config-set workflow.windows_enforce true`. Tracking (the ledger itself, populated by the executor) is always on; only the ship gate is opt-in.
**Backward compatibility:** A project with no `.planning/WINDOWS.md` reports `open_count: 0` and ships cleanly; the gate only activates once windows are recorded.
**Milestone attribution (#4487):** each entry carries a `milestone` field, stamped at record time from the workstream's resolved milestone version (STATE.md `milestone:` frontmatter, or the ROADMAP.md in-progress marker as a fallback). Phase numbers are unique only within one active `phases/` directory — `milestone complete` frees them for reuse — so this is what lets an entry be attributed to the milestone it was actually recorded under, even after that milestone is archived and its phase numbers reused. `null` when no milestone could be resolved, including every entry recorded before this field existed.
**Configuration:** `graphify.graph_path`
---
### 159. Complexity-Triggered Refactor
**Behavior:** An `execute:post` step measures the complexity of the files a phase touched (decision-point counting over comment- and literal-stripped source, no external dependency) and surfaces a scoped refactor proposal at `.planning/phases/<N>/<NN>-REFACTOR.md` when a function's score exceeds `refactor.complexity_threshold` or its growth over its recorded anchor exceeds `refactor.complexity_jump_delta` — whichever trips first, both reported. Trigger semantics are strictly greater (ESLint's `complexity: {max: N}` convention), so a score exactly equal to the threshold does not trigger. The anchor is set the first time a function is observed and moves only when the proposal is dispositioned via `refactor accept` or `refactor decline` — never when the score alone improves — so the jump delta is cumulative growth since the last conscious decision about that function, not the change made in a single phase. Advisory by default: the proposal is informational only, never edits code, and never blocks. Opt-in `refactor.trigger_strict` records an untriaged proposal as an open `deviation` entry in the broken-windows ledger (#1950) instead — it does not block on its own; ship-blocking is broken-windows' existing `ship:pre` gate, enabled separately with `workflow.windows_enforce`. Without broken-windows installed, strict mode still records the proposal locally and says so. Enabling `refactor.trigger_strict` without `workflow.windows_enforce` also on (or with broken-windows absent) surfaces a typed `refactor_strict_not_enforcing` warning on every triggering evaluate, naming the exact remediation, so this enforcement gap is never silent. A declined proposal resolves its ledger entry as `waived` with the recorded reason; an accepted one resolves as `fixed`. The metric is approximate by construction: biased against a flat `switch`, blind to nesting depth, JS/TS-family only, and a renamed function loses its anchor (issue #1953).
**Commands:** `gsd-tools refactor evaluate | status | accept | decline`.
**Config:** `refactor.trigger_enabled` (master gate, default `false`), `refactor.complexity_threshold` (default `15`), `refactor.complexity_jump_delta` (default `5`), `refactor.trigger_strict` (default `false`). See [Configuration Reference](CONFIGURATION.md#refactor-trigger-settings).
**Backward compatibility:** Off by default. When `refactor.trigger_enabled` is `false` the hook never runs and writes nothing; a project that never enables it is completely unaffected.
---
### 160. Archive Quick Tasks at Milestone Close
**Command:** `/gsd-complete-milestone` (forward path), `/gsd-cleanup` (retroactive path), `gsd-tools milestone complete --archive-quick` / `gsd-tools milestone archive-quick <version>` (#2142)
**Behavior:** `.planning/quick/` otherwise accumulates one directory per `/gsd-quick` task forever. `/gsd-complete-milestone` now offers a Yes/Skip prompt — when accepted, it moves every directory under `.planning/quick/` into `.planning/milestones/<version>-quick/`, (re)writes that archive directory's `README.md` (an index built by scanning the archive directory, one entry per task, linked to its `SUMMARY.md` when one exists), and clears the data rows of `STATE.md`'s `### Quick Tasks Completed` table while preserving its header and detected column variant. `/gsd-cleanup` offers the same archival retroactively, for milestones that were already closed before their quick tasks were swept, via the narrower `milestone archive-quick <version>` command — identical move/index/reset behavior, but without touching `ROADMAP.md`, `REQUIREMENTS.md`, `MILESTONES.md`, or milestone-completion guards, so it can be re-run safely against an already-completed milestone.
**Why opt-in.** Phase-directory archival is default-ON (#1871) — omitting a phase directory from an archive would silently leave stale execution history in the way of the next milestone's roadmap. Quick tasks carry no such downstream conflict, so archival here defaults OFF: a user who never passes `--archive-quick` sees zero behavior change. This is a deliberate asymmetry with phase archival, not an oversight.
**Why bucket-all, not per-milestone.** `.planning/quick/` is a flat directory with no on-disk record of which milestone a given task belongs to. Splitting tasks per milestone was considered and rejected — inferring provenance from dates (creation time vs. a milestone's shipped date) is a proxy, not a fact, and a wrong inference on a one-way `mv` is silently irreversible. Archival instead buckets everything currently in `.planning/quick/` into the one milestone being completed (or, on the retroactive path, the one milestone chosen), and says so in the confirmation prompt.
**Why the index is built from disk, not from `STATE.md`'s table.** The `### Quick Tasks Completed` table is a running log a workflow step appends to — it demonstrably drifts from what's actually in `.planning/quick/` (the motivating case: 53 rows against 49 directories, ~22 rows pointing at directories that no longer existed, 18 directories with no row at all). Building the archive's `README.md` index by scanning the archive directory itself, rather than trusting the table, means the index can never inherit that drift; a re-run's index also naturally includes entries a prior run already archived, since it's re-derived from what's physically present.
**Known limits:**
- No per-milestone provenance — bucket-all is the only option (see above).
- A `### Quick Tasks Completed` table whose columns match neither registered variant (with/without a Status column) is left untouched with a warning rather than reset, since clearing it would risk destroying rows under a schema GSD doesn't recognize.
- A `STATE.md` with no `### Quick Tasks Completed` section at all is a normal, silent no-op for the reset step — the section is created lazily by `/gsd-quick`, not present in the project template.
See [Archiving quick tasks](how-to/handle-quick-and-fast-tasks.md#archiving-quick-tasks) for the full walkthrough.
---
### 161. Verify-Command Path Grounding
**Command:** `/gsd-plan-phase` (automatic), `gsd-tools check verify-command-paths <N>` (#2401)
**Behavior:** A planner authoring a per-task `<automated>` verify command has no line of sight to whether the path it just wrote actually resolves, and `gsd-plan-checker` had no deterministic way to check — so it hand-reasoned the filesystem and, in the motivating case, prescribed two successively-wrong replacement paths (the second citing a `package.json` that did not exist). Two changes close that:
1. **Prior-command inheritance.** The nearest prior phase's `<automated>` commands are surfaced to the planner as `prior_verify_commands`, **at every context window**. Cross-phase enrichment was previously gated on `context_window >= 500000`; at 200k the planner re-invented the command and got it wrong. This payload is a handful of one-liners, so it is never gated.
2. **A deterministic probe.** `gsd-tools check verify-command-paths <N>` resolves each `<automated>` command's target directory and reports whether it exists and holds the manifest the command needs. `/gsd-plan-phase` runs it before the plan-check pass and hands the JSON to the checker, which acts on `severity` instead of guessing.
**It never executes command text.** PLAN.md is model-authored, so running it from the checker would be arbitrary code execution — and would trigger the real lint/build as a side effect. The probe only resolves paths and stats directories; a `package.json` it finds is read for script names only.
**Why a recognizer, not a shell parser.** Interpreting shell would mean maintaining a bad shell. Exactly two forms are grounded — a leading `cd <literal>` chain and `npm --prefix <literal>` — and any path carrying a variable, glob, substitution, or `~` returns `unresolvable`, which is a warning and never a blocker. The parser's incompleteness is the specification: it degrades to "cannot prove" rather than growing features. Refusing to guess is the fix, not a limitation of it.
**It reports, it never prescribes.** The payload carries the target that failed and what was missing; there is deliberately no `suggestion` field. Choosing the replacement is the planner's job — and the planner now has the prior phase's proven command to reach for.
**Not findings:** a target an earlier task in this phase creates (`pending_creation`), a command with no `cd`/`--prefix` at all, and the Nyquist `MISSING — Wave 0 …` sentinel, which Dimension 8 owns.
**Known limits:**
- Only `cd <literal>` and `npm --prefix <literal>` are recognized. `pushd`, `make -C`, `yarn --cwd`, `pnpm -C`, and `cargo --manifest-path` report `unresolvable`.
- Verdicts are relative to the *checker's* project root. Under parallel worktree execution the executor's root differs, so a bare ancestor climb (`cd ../..`) is reported `outside_root` as a warning rather than asserted about.
- `script_missing` is advisory only — this phase may be adding the script — so a genuinely mistyped npm script still reaches the executor.
See [Resolve verify-command path findings](how-to/resolve-verify-command-path-findings.md) and [`gsd-tools check verify-command-paths`](COMMANDS.md#gsd-tools-check-verify-command-paths).
---
### 162. Statusline STATE.md Freshness Marker
**Config key:** `statusline.show_state_freshness` (default `false`)
**Purpose:** A solo developer returning to a project after time away reads "Phase 4, executing" in `STATE.md` and acts on it — without noticing the codebase has moved 40 commits since that line was written. `/gsd-health` reports this as `W024`, but only if the user thinks to run it. The statusline is the one surface seen continuously without asking (#2734).
**Behavior:** Renders `state ~N commits back` inside the GSD-state segment when `STATE.md` carries a `state_head` stamp (#2573) and `HEAD` is at least `STATE_HEAD_ADVISORY_COMMITS` (20) commits past it. Both statusline formats carry it — the default renderer and the compact `statusline.state_format` one.
**The threshold is 20, deliberately not 1.** With `commit_docs: true` (the default) the commit carrying a `STATE.md` sync advances `HEAD` by one, so a `> 0` threshold would render `state ~1 commits back` permanently on a project that is by construction fresh — alarm fatigue on the one always-visible surface.
**It degrades to silence rather than to a wrong answer.** The marker is absent — never "fresh" — when the stamp is malformed, when the project root does not own its `.git` (an enclosing unrelated repo would otherwise answer), in a `planning.sub_repos` workspace (the outer `HEAD` never advances when code lands in children), when history was rewound past the stamp, and when git is unavailable or slow. A freshness claim the project cannot substantiate degrades to *unknown*.
**Cost:** exactly one bounded `git rev-list` call per render, and only when enabled *and* a stamp is present — `rev-list --left-right --count` answers ancestry and distance together, and repo pinning is a filesystem check rather than a subprocess. Disabled (the default) it adds none.
**A proxy, never a drift measurement.** The count includes commits that touched nothing `STATE.md` describes, and the stamp restamps on every state write — so a low count means "something wrote STATE recently", not "STATE is accurate". Rendered with a `~`; never gate on it.
**Reference:** [Configuration](CONFIGURATION.md) · [Read the statusline freshness marker](how-to/read-the-statusline-freshness-marker.md) · [ADR-2164](adr/2164-statusline-scope-boundary.md)
---
### 163. Read-Only Planning Snapshot (`planning inspect`)
**Command:** `gsd-tools query planning inspect`
**Purpose:** Give downstream consumers — harness UIs, mission-control surfaces, dashboards, bots — one schema-versioned JSON document describing everything `.planning/` knows, so nothing outside gsd-core has to parse `ROADMAP.md` / `REQUIREMENTS.md` / `*-PLAN.md` / `*-SUMMARY.md` a second time. gsd-core is the single source of `.planning/` truth; a second parser is a second answer.
**Requirements:**
- REQ-INSP-01: `PLANNING_INSPECT_SCHEMA_VERSION = 1` is emitted as `schema_version`. Consumers MUST reject any other value rather than best-effort-parse an unknown shape.
- REQ-INSP-02: Read-only. The command mutates no planning state, and mutates nothing on disk, under any input.
- REQ-INSP-03: Unknown or conflicting evidence serializes as `null` / `"unknown"` with a coded entry in `diagnostics[]` — never inferred, reconciled, or defaulted. Every key is always present; a key is never omitted to signal absence.
- REQ-INSP-04: Argument errors fail loud (non-zero exit, typed `ERROR_REASON`); data gaps do not. v1 takes no arguments, and a stray positional or unknown flag is a usage error rather than a silently-ignored one.
- REQ-INSP-05: Roadmap acceptance, verification status, and UAT items are reported side by side per phase and are never folded into a single verdict. A ROADMAP checkbox carries `authoritative: false` — completion is derived from disk state.
- REQ-INSP-06: `accepted_phases` and `completed_plans` are independent fractions. `percent` is `null` whenever the scope is not `complete`, per the same rule the roadmap and progress surfaces follow.
- REQ-INSP-07: Payloads over ~50 KB use the existing `@file:` spill channel, resolved transparently before stdout.
**Why it does not simply serialize the internal snapshot.** `PlanningSnapshot` (the diagnostic-rule subject introduced by ADR-3180 §8.1) is deliberately additive and still growing — four fields at Phase 10, twenty-plus by Phase 12. Handing that shape to external consumers would freeze an internal contract by accident. `planning inspect` declares its own flat schema and maps into it, so a field added to `PlanningSnapshot` never changes what this command emits.
**Composed, never re-derived.** Milestone identity and phase enumeration arrive via `buildPlanningSnapshot`; completion from `isPhaseComplete` (disk-strict); live-plan counting from `scanPhasePlans`; the percentage arithmetic from `clampPercent`; STATE fields from `stateFieldValue`; plan bodies from the Plan Document Module; requirement IDs from `parseRequirements`; UAT items from `parseUatItems`. Markdown structure is read through the Markdown Sectionizer and Markdown Table Model seams, so the Traceability table is resolved by column name against its registered schema rather than by a position-anchored regex.
**Known limit — task-scoped file provenance.** A `<task>` declares the files it plans to touch, but `SUMMARY.md`'s `## Files Created/Modified` describes the whole plan. Spreading that list across a plan's tasks would be inference, so a task's `changed_files` is populated only where the summary attributes files to that specific task; otherwise it is `null` with `provenance: "plan_scoped"`. Closing this needs a change to the SUMMARY format, not to the reader.
**Reference:** [CLI Tools](CLI-TOOLS.md#planning-inspect) · [Consume the planning snapshot](how-to/consume-the-planning-snapshot.md)
---
### 164. Live-DOM UAT Capability
**Config key:** `workflow.live_dom_uat` (default `false`)
**Purpose:** A phase with a live-UI acceptance criterion could not be finished by the agent that executed it. `gsd-executor` carries no browser tools, so it correctly returned a `checkpoint:human-action` — even though the work was not human-only, just tool-less. Every such phase quietly degraded from *executed by the executor* to *executed, then finished by hand in the orchestrator*, and the plan's `autonomous: false` marker could not distinguish "a human must judge this" from "the executor lacks the tool" (#2856).
**Behavior:** A default-off capability owns one boolean key, one agent, and one additive step. When the key is on, `gsd-dom-verifier` runs at `execute:wave:post` and writes `{phase}-DOM-VERIFY.md`; the orchestrator's `automated_ui_verification` step additionally considers `mcp__chrome-devtools__*` / `mcp__claude-in-chrome__*` when present.
**The executor's tool surface is unchanged in every configuration.** Widening it was the reported proposal and was refused: for a first-party agent the static `tools:` list is the only control that exists — no capability can grant tools to one ([ADR-1244](adr/1244-capability-ecosystem.md) D2), no hook kind grants tool permissions ([ADR-857](adr/857-capability-system.md) D4), and there is no per-dispatch override. Browser reach lives in one purpose-built agent that carries no `Bash`.
**Two independent gates, both fail-closed.** The capability's `activationKey` makes it resolve inactive when the key is off — `resolveLoopHooks` renders a hook only on `state.active === true` — and the step carries its own `when` guard. Tool presence alone never activates it: a browser MCP configured for unrelated work is not driven by default.
**The pre-existing Playwright path is untouched.** `mcp__playwright__*` keeps the gating it already had (presence plus an active UI phase). Pulling it behind a new default-off key would have silently removed working behavior from current users on upgrade; the key gates only the newly added families.
**It tolerates the browser-profile lock rather than coordinating it.** `chrome-devtools-mcp` holds an exclusive lock on its profile, so parallel waves collide. `--isolated` is a flag on the operator's own MCP server registration — GSD neither launches that server nor passes its arguments — so the verifier reports `could_not_look` / `profile_locked`, names the flag, and stops. No retry, no held-up wave.
**`nothing_to_report` is never conflated with `could_not_look`.** A report claiming no issues when it never opened a browser is worse than no report; the artifact carries a closed reason enum so the two are always distinguishable.
**Known limits:** no sandbox — once enabled, nothing constrains which origins are reached ([ADR-1244](adr/1244-capability-ecosystem.md) D5); DOM observation only, no screenshot diffing, accessibility audit, or performance tracing.
**Reference:** [Configuration](CONFIGURATION.md) · [Enable live-DOM verification](how-to/enable-live-dom-verification.md) · [Explanation](explanation/live-dom-uat-capability.md) · [Agents](AGENTS.md)
---
### 165. Opt-In Parallel Reviewer Lanes
**Command:** `/gsd-review`, `/gsd-plan-review-convergence`
**Config key:** `review.parallel_lanes` (default `false`)
**Purpose:** Reviewer lanes within one review pass have no data dependency on each other — they all inspect the same immutable plan snapshot — but were dispatched strictly one at a time, so a pass with Codex, Antigravity and Claude cost roughly the sum of three long reviewer calls. The serialization was a deliberate, unconditional protection against provider rate limits, which made it a global policy imposed on users whose providers could comfortably take concurrent requests, or who run local model servers with no limits at all (#3034).
**Behavior:** With the key enabled, the `invoke_reviewers` step dispatches each selected lane as a background job and joins all of them before `REVIEWS.md` and consensus are rendered. Wall-clock cost falls toward the slowest lane rather than the sum. Default remains `false`, preserving the existing sequential dispatch and its rate-limit protection.
**The guard is strict equality, and it fails safe.** Only the exact value `true` opts in — `"1"`, `"yes"` and `"TRUE"` all stay sequential, so a mistyped config gets the conservative behavior rather than concurrent requests at a rate-limited provider. A failure to read the config falls back to sequential too. This polarity is deliberately the opposite of the `commit_docs` guard, which fails open: there, failing open preserves user intent; here it would fire the very requests the default exists to prevent.
**Result ordering is unchanged in both modes.** Per-lane results are written to slug-scoped files and concatenated in reviewer-selection order after the join, so `gsd-review-lane-results.jsonl` reads identically whether lanes ran sequentially or concurrently. Completion order never reaches the artifact. This also means concurrent lanes never share an append handle — a lane result larger than the pipe-atomicity bound cannot interleave and corrupt the `models:` / `model_sources:` frontmatter that `write_reviews` renders from that file.
**Per-lane semantics are untouched.** Timeouts, prompt budgets, the diagnostic stub for an empty or failed lane, explicit-lane failure ([ADR-2782](adr/2782-reviewer-lane-capability-surface.md) D4), trust/egress checks and result-file layout all behave exactly as they do sequentially. A failing lane does not abort its siblings.
**Known limits:** convergence cycles stay sequential by design (`review → replan → re-review` has a genuine data dependency), so this speeds up each pass rather than reducing the number of passes; there is no concurrency bound, so every selected lane dispatches at once; and reviewer instances sharing one adapter dispatch concurrently against that single provider, which is the most likely way to hit a limit.
**Reference:** [Configuration](CONFIGURATION.md#parallel-reviewer-lanes-for-gsd-review-3034) · [Enable parallel reviewer lanes](how-to/enable-parallel-reviewer-lanes.md) · [Commands](COMMANDS.md)
---
### 166. Machine-Readable State Contract (`.planning/state.json`)
**Purpose:** External tools that display GSD project state — a workbench, a dashboard, an editor extension — had to parse `STATE.md` and `ROADMAP.md` heuristically. Those are human surfaces: their shape drifts as the templates evolve, and every consumer ends up carrying a brittle second parser that silently reports wrong numbers after an upgrade. GSD now publishes a small, versioned JSON snapshot instead, so the reader binds to a contract rather than to markdown (#3227).
**Behavior:** At every step boundary, GSD writes `.planning/state.json` — `contract`, `flavor`, `milestone`, `phases[]`, `next`, `updated_at`. The boundaries are `state begin-phase` / `planned-phase` / `advance-plan` / `complete-phase` / `milestone-switch`, `phase add` / `add-batch` / `insert` / `remove` / `complete`, and `milestone complete`. The write is best-effort and completely invisible to the command that triggered it: it cannot change an exit code, cannot change stdout, and cannot fail a workflow. Readers prefer the file when it is present and fall back to markdown when it is not.
**Requirements:**
- REQ-SC-01: `contract` is semver, `1.0.0` at introduction. Consumers gate on the MAJOR version; `1.x` changes are additive only. Every key is ALWAYS present — an unknown value is `null`, never an omitted key, because an omitted key is itself an observable a consumer would bind to.
- REQ-SC-02: `phases[]` carries `{number, name, status}` per phase, `status` drawn from exactly `complete | in_progress | pending`. `number` is a string (`"01"` and `"2.1"` are both real ids and neither survives a number cast); `name` is `null` when the roadmap gives a phase no name, never a fabricated placeholder.
- REQ-SC-03: `next` is the same recommended action the `/gsd` front door routes, derived from the smart-entry classifier itself rather than from a second copy of its routing table.
- REQ-SC-04: A missing `ROADMAP.md`, a missing or unreadable `.planning/`, an unwritable target, or any other failure NEVER errors the parent command. A directory that is not a GSD project stays untouched — the publisher will not create `.planning/` in order to publish into it.
- REQ-SC-05: The skills own the file; readers never write it. It is a derived cache — safe to delete, regenerated at the next boundary.
**Composed, never re-derived.** Milestone identity comes from `getMilestoneInfo`; phase rows from `locateProgressTable`, the same `## Progress` locator the progress counters use, so `state.json` can never disagree with the rest of GSD about which phases are complete; the recommended action from `classifyProject`. This module introduces no second answer to any question GSD already answers.
**Why it does not reuse `planning inspect`'s schema.** The two surfaces answer different questions and have opposite shapes. `planning inspect` is a rich, diagnostic-carrying **pull** query a consumer runs; this is a small **push** artifact a consumer watches. Publishing `planning inspect`'s payload at every `phase add` would mean opening every plan, summary and requirements document on a hot path, and freezing a much larger surface as a contract.
**It costs up to three bounded git calls per boundary.** Deriving `next` from the smart-entry classifier means inheriting its git signals — `git status --porcelain`, and `git log @{u}..HEAD`. Each is timeout-bounded and swallows every error, so nothing can hang or fail because of it, but a command like `phase add` did not previously touch git at all. "Invisible to the parent command" is exact about exit code and output; it is not a claim about latency.
**Known limits:** an empty `phases: []` cannot be told apart from "no `ROADMAP.md`" or "roadmap unreadable" — the `1.0` schema carries no diagnostic channel, and `planning inspect` is the surface that does. A roadmap phase marked `Deferred` is reported as `pending`, because the roadmap vocabulary has four values and this contract has three; inventing a fourth wire value would break every existing reader. `phases[]` is not milestone-scoped, so a long-running project lists every phase it has ever had.
**Reference:** [Consume the state contract](how-to/consume-the-state-contract.md) · [Consume the planning snapshot](how-to/consume-the-planning-snapshot.md)
---
### 167. Stated Failing Direction
**Command:** `/gsd-plan-phase` (automatic), `gsd-tools check verify-failure-directions <N>` (#3172)
**Behavior:** A plan's `<automated>` block is the thing that decides whether work is done, and nothing checked that the command inside it could fail. In the motivating case six plans shipped 21 commands that could not run at all — `cargo test -p <pkg> --lib` against a package with no library target. They read as rigour and were not falsifiable, so three separate executors each rediscovered the defect and improvised a substitute at execution time. Every runnable `<automated>` command now needs a `<fails_when>` sibling naming what output constitutes failure:
```xml
<verify>
<automated>npm --prefix apps/api test -- auth.spec.ts</automated>
<fails_when>non-zero exit, or "0 passed" in the summary line</fails_when>
</verify>
```
`gsd-planner` emits it; `gsd-tools check verify-failure-directions <N>` verifies it deterministically; `/gsd-plan-phase` runs the probe before the plan-check pass and hands the JSON to `gsd-plan-checker`, whose check 8f blocks on `severity`.
**Why this shape and not the two obvious alternatives.** *Validating command shape* — teaching the checker Cargo's `--lib`/`--bin` target resolution, then pytest's node-ids, then the next one — always trails the newest toolchain. *Executing each command at plan time* is the strongest signal but means running planner-invented commands, with whatever side effects they carry, during planning. Requiring a stated failing direction needs no toolchain knowledge at all, and it is the only one of the three that catches the dangerous case: the motivating command exited non-zero, so it failed loudly, but the same class of error with a command that exits 0 on a no-op passes green and silently. Naming the failure signal is what makes that visible.
**Presence, not quality — deliberately split.** The probe is deterministic and owns the blockers: a statement is missing, blank, or a whole-value placeholder (`TBD`, `TODO`, `N/A`, `NA`, `none`, `unknown`, `TBA`, `?`, `-`). Whether the statement names the *right* signal is prose judgment, so `gsd-plan-checker` raises a vacuous statement (*"the command fails"*) as a WARNING only. Every BLOCKER stays reproducible; judgment stays advisory.
**It reports, it never prescribes.** The payload names the command with no stated failure mode and stops there. A prescribed statement would be copied verbatim and carry zero information — reproducing the original defect one level up.
**Not findings:** the Nyquist `MISSING — Wave 0 …` sentinel (not runnable, so it has no failure mode to state), an empty `<automated>` body (check 8a owns command presence), and a `<verify>` with no `<automated>` at all.
**Known limits:**
- Presence only. A statement that is present and specific can still name the wrong signal; that is caught, if at all, by judgment rather than by the probe.
- **Breaking for plans authored before this shipped.** A phase planned earlier has no `<fails_when>` anywhere and blocks on re-check until statements are added or the phase is re-planned.
- The adjacent **vacuous pass** — a command that runs successfully and asserts nothing, such as a test-name filter matching zero tests and exiting 0 — is a distinct problem and is explicitly out of scope.
See [State a failing direction](how-to/state-a-failing-direction.md) and [`gsd-tools check verify-failure-directions`](COMMANDS.md#gsd-tools-check-verify-failure-directions).
---
### 168. Runtime Identity
**Purpose:** The predecessor package `get-shit-done-cc` publishes a binary named `gsd-tools`, and so does this one. They answer some of the same verb names with **different semantics**. [#3129](https://github.com/open-gsd/gsd-core/issues/3129) is the worked example: `phases.clear` **archives** here and **deletes** there. Both print success-shaped output, and `.planning/` is gitignored by default, so a user lost 43 phase directories with no error, no warning, and nothing recoverable from git. The failure was silent in both directions — the workflow could not tell it had reached the wrong handler, and the handler could not tell it had been called by a workflow written for a different contract (#3146).
**Behavior:** two independent defenses, one structural and one asserted.
**Structural — the `PATH` branch.** The launcher's `PATH` resolution branch looks for **`gsd_run`** instead of `gsd-tools`. Only this package publishes `gsd_run`; the predecessor publishes `gsd-tools` and `gsd-sdk`. Our `gsd_run` follows its own symlink chain and executes the `gsd-tools.cjs` sitting **beside it**, so resolving it cannot land on a foreign handler.
**Asserted — every other branch.** The path-based branches (a project-local install, a runtime config directory) have no such guarantee: they trust their configured location. So once resolution finishes, and before any verb runs, the preamble probes the tool it picked with `runtime-identity --raw` and matches the answer **anchored** against the compact payload. It exports the result as a two-valued `GSD_IDENTITY_STATUS` (`ok` / `unverified`) and, when it is `unverified`, prints one actionable line naming both plausible causes. The same `gsd-tools runtime-identity` verb remains available by hand, so a human or a support thread can settle "which tool am I actually running?" in one command.
**The match is anchored at both ends, not a substring.** A substring search for `@opengsd/gsd-core` accepts the decoy `{"packageName":"get-shit-done-cc","note":"@opengsd/gsd-core"}`, which any colliding package could publish. The preamble instead requires the payload to *begin* with `{"packageName":"@opengsd/gsd-core"` **and to end with a closing brace**, so a truncated answer fails as well. Closing on `}` costs nothing in future-proofing: a JSON object's own brace is always the last character, whatever type the last value has.
**The status is a value, not prose.** `GSD_IDENTITY_STATUS` exists so the gate can be tested — and read by a later step — without anyone parsing the warning text.
**The byte budget is why the assertion arrived second.** The preamble is inlined into 113 shipped files, several of which sat within **single-digit bytes** of frozen size ceilings — `agents/gsd-verifier.md` had 2 bytes of headroom — and those caps are red lines, not budgets. A first attempt to inline an assertion broke five of them. What made it fit was collapsing the resolver's twenty near-identical `elif [ -f … ]` arms into a single candidate-list helper, which is worth far more bytes than the assertion costs: the preamble is now **1,876 bytes smaller** than the version that carried no assertion at all, so every one of the 113 files moved *away* from its ceiling.
**It fails closed.** If no `gsd_run` is reachable, the resolver falls through its remaining path-based branches and finally errors with an install command. It does not fall back to executing whatever `gsd-tools` happens to be on `PATH` — that fallback *was* the vulnerability.
**A doubly-sourced preamble cannot build a recursive launcher.** `command -v gsd_run` finds the shell *function* on a second source and would return the bare string `gsd_run`, defining the function in terms of itself. `unset -f gsd_run` leads that branch, so the second source resolves exactly as the first did. (An executability guard was tried here instead and removed: it rejected the bare name, fell through every branch, and hit the resolver's `exit 1` — which, in a *sourced* script, kills the caller's shell.)
**Known limits:**
- **The assertion warns; it does not yet stop the run.** The rollout is warn-then-fail. It cannot hard-fail yet because an `@opengsd/gsd-core` older than the `runtime-identity` verb answers exactly as a foreign package does — neither answers — and at rollout the old-version case is the common one. The warning therefore names both causes. A later release turns `unverified` into a refusal.
- An installation old enough to predate `bin/gsd_run` ([#381](https://github.com/open-gsd/gsd-core/issues/381)) is not reachable through the `PATH` branch and must be upgraded or invoked through one of the path-based branches.
- The probe costs one extra process launch per preamble source. It is a pure local read of baked coordinates, deliberately kept off the SDK bridge for that reason.
**Reference:** [`runtime-identity`](COMMANDS.md#runtime-identity) · [Diagnose which gsd-tools is running](how-to/diagnose-a-foreign-gsd-tools.md)
---
_Generated by `scripts/gen-features.cjs` — add a fragment under `docs/features/` and run `--write`._
---
### 3348. Context Drift Gate
**Purpose:** Warns (or optionally blocks) before `/gsd-plan-phase` reuses an existing
`RESEARCH.md`, `PATTERNS.md`, `VALIDATION.md`, or `SPEC.md` that predates a decision added to the
phase's `CONTEXT.md` after that artifact was derived from it. Deterministic — compares git commit
time (falling back to mtime for uncommitted edits), no model call. Sibling to the existing
codebase-drift and schema-drift gates in the `drift` capability. Configure with
`workflow.context_drift_precheck` (on/off) and `workflow.context_drift_action` (`warn`/`block`).
---
### 3884. "Failure Is a Value" — Strict Argv Rejection and the `--pick` Absence Contract
**Purpose:** ADR-3473 §8.4 states the rule directly: absence, emptiness, and
failure are three different things, and a routine that cannot tell them
apart eventually reports the wrong one. Before this change, `gsd-tools` had
two silent instances of exactly that collapse.
**Half one — a stray positional corrupted state, silently (#3358).**
`parseNamedArgs` read only the flags it recognized and dropped everything
else — an unrecognized `--flag` or an extra positional argument (for
example, a stray phase number appended after `state.planned-phase`) was
silently discarded rather than rejected. The caller's own positional read
(`args[2]`, etc.) still worked, so the command ran anyway, on the wrong
phase, and overwrote the previously-current phase block with no error at
all. The fix makes `parseNamedArgs(args, spec)` return the command-routing
hub's own `Result` shape (`{ok:true,data} | {ok:false,kind:'InvalidArgs',...}`)
and requires every call site to declare `positionals: number | 'rest'` — the
count of leading argv slots the caller itself reads directly. An unknown
flag or an unexpected positional past that boundary is now a loud,
non-zero-exit `InvalidArgs` failure instead of a token quietly falling on
the floor. A duplicate flag, a negative-number value (`--plans -1`), and a
documented free-text tail (`init quick <description>`) are deliberately
left alone — none of them are the defect this closes, and forbidding them
would just break working call sites for no gain.
**Half two — `--pick` on an absent field answered `''` at exit 0, exactly
like a present-but-empty one (#3365).** `--pick <field>` extracted one field
from a command's JSON output, but a missing key, an out-of-range array
index, a partially-missing dotted path, or non-JSON output (including a
`--raw` command's output) all rendered the same way: empty stdout, exit
`0`. That is indistinguishable from a field that genuinely holds `null` or
`''` — a real answer. The shell idiom `X=$(… --pick F) || X=default` could
therefore never observe the failure it was written to react to; only a typo
in the verb name would ever make it exit non-zero. `--pick` now exits `1`
with a diagnostic on stderr (`pick_field_absent` naming the field and the
available top-level keys, or `pick_output_not_json` when the output could
not be parsed as JSON at all) whenever the field cannot be resolved. A
present field's value — including `0`, `false`, `null`, and `''` — is
unchanged: those are answers, not failures, and remain exit `0`. See
[CLI-TOOLS.md's `--pick <field>` contract](CLI-TOOLS.md#--pick-field-contract)
for the full outcome table and [json-errors.md](json-errors.md) for the two
new reason codes.
**Why not just default to zero for an absent count.** The sub-issue's own
Done-when checkbox suggested treating an absent field the same as a
zero-valued one. That is rejected on the merits: it demotes *"I could not
answer"* to *"the answer is zero"*, which would make a count-gated shell
guard fire unconditionally on the very projects that could never resolve
the count in the first place — the opposite of what a gate is for.
**Consequence for `scripts/lint-unreachable-guard-drift.cjs`.** That guard's
Detector A existed specifically because the old `--pick` behavior made a
`--pick … || echo <default>` line's fallback arm permanently unreachable.
Once `--pick` exits non-zero on absence, that premise is false and the
shape the detector forbade becomes the *correct* idiom — so Detector A was
retired rather than kept. Detector B (the unrelated `cat`/`ls`-over-a-glob
nullglob hazard) is untouched. See
[Resolve unreachable-guard findings, Shape A](how-to/resolve-unreachable-guard-findings.md#shape-a---pick-with-an--echo-fallback-resolved-upstream-3884).
**Known limits:**
- A value token beginning with `--` still cannot be passed to a declared
value flag (`--summary "--force is now default"` now fails loudly instead
of silently dropping the value) — strictly better, but no `--flag=value`
escape was added.
- `--pick` still cannot distinguish an absent field from a `null` one *on
stdout alone* — the distinction is carried entirely by exit code.
- The ~10 `gsd-tools.cjs` call sites of `parseNamedArgs` get no
compile-time check (that file is hand-written JavaScript, not `.cts`);
enforcement there is the runtime throw on a stale legacy call shape plus
behavioral tests.
- This phase does not sweep every routine in `gsd-core` for `Result`
conformance — it applies the rule to the argument-projection seam and the
`--pick` extractor it names, not the whole codebase.
---
### 3885. No Silent Swallow, No Verdict From Dropped Data
**Purpose:** ADR-3473 §8.5 states the rule directly: a failure or a gap in the
input must not be absorbed into an output that reads as authoritative. A
routine that drops data it could not read or could not resolve, and then
reports a clean result anyway, turns a diagnosable gap into a confidently
wrong answer. This closes four instances of that collapse found across
`gsd-tools`.
**`intel query` no longer crashes past ~12000 levels of nesting (#3427).**
`searchJsonEntries` / `matchesInValue` recursed with no depth bound at all —
an intel JSON file nested deeply enough overflowed the call stack with an
uncaught `RangeError` instead of a diagnosis. The original SDK-era bound
(`MAX_JSON_SEARCH_DEPTH = 48`, lost in the ADR-0174 consolidation) is
restored, paired with a `truncated` result field: a match at or above the
ceiling is not returned, and the result says so rather than reporting a bare
"not found" that is indistinguishable from a genuine miss. A match at depth
48 (inclusive) or shallower is unaffected; the bound is on nesting depth, not
breadth or total node count, so a shallow object with many siblings still
works unchanged.
**`phase-plan-index` no longer blames the author for an edge the tool itself
dropped (#3427).** A `depends_on:` token that resolves to no plan in the
phase (typo, or a stale cross-phase reference) silently dropped that edge,
making the dependent plan a DAG root — its own docstring recorded the intent
as "ignore this edge, never a throw." The tool then compared the resulting
degraded wave against the plan's declared `wave:` and reported the *author's
correct* declaration as a mismatch. The unresolved token is now named in its
own `warnings[]` entry (plan and token together), and the wave-mismatch
warning is suppressed for that plan only — a plan with no dropped edges and a
genuinely wrong `wave:` still warns as before. The token is escaped
(quoted, control characters and embedded newlines backslash-escaped) before
it is embedded in the warning text, so a `depends_on` value crafted to
contain a newline or a quote cannot forge a second, fabricated warning entry.
**A code-review run where every lane failed no longer writes `REVIEWS.md`
from nothing (#3352).** `review.md`'s aggregation step wrote `REVIEWS.md`
regardless of whether any lane actually produced results — a run where every
lane failed still emitted a completed-looking review artifact, and the
per-lane outputs and `.err` files that would have explained the failure were
then destroyed by the run's own cleanup. `REVIEWS.md` is now withheld when
the aggregate has zero lines (every lane failed, not merely skipped under a
lower budget), the run reports the failure instead, and per-lane outputs and
non-empty `.err` files are preserved beside the phase's artifacts before
cleanup runs.
**Unreadable directories are distinguished from absent ones (#3473 B5).**
Four call sites collapsed an `EACCES`/`EIO` on a phase directory into the
same "nothing here" result as a directory that genuinely does not exist —
`countPhasePlansAndSummaries` (`hasContext:false`), `runGapAnalysis`, and two
guarded blocks in `init.cts` (`context_path` absent). Each now distinguishes
"could not read" from "does not exist" and names the discarded path and
error in a dedicated field (`context_read_error` / `phase_dir_read_error`)
rather than silently reading as absent.
**Audited, no defect found:** every retry-set / swallowed-catch call site
this phase's rule covers that had not already been fixed by a prior PR
(`withPlanningLock`, `acquireStateLock`, `atomicRenameWithRetry`,
`renameWithRetry`) was reviewed and found to already fail loudly on a fatal
errno rather than folding it into a retry.
**Known limits:**
- A match deeper than 48 levels is still not surfaced by `intel query` — it
is reported as truncated rather than as absent, but the value itself is
not returned. Raising the ceiling is a separate decision.
- `phase-plan-index`'s `waves` / `wave` fields remain computed from the
degraded DAG when an edge is dropped — this phase stops the tool from
manufacturing a false verdict about it, but does not invent the missing
edge. A consumer that schedules work from `wave` (e.g. `--wave N`
filtering) is still working from the degraded assignment.
- `review.md`'s evidence preservation is bounded by what a lane actually
wrote — a lane that produced no output at all leaves nothing to preserve.
---
### 3897. Runtime Marker Resolution, Derived Codex Sandbox, and In-Phase Short-Form Dependencies
**Purpose:** ADR-3473 §8.3 states the rule directly — one implementation per
invariant, not a hand-maintained copy that quietly drifts from the rule it
stands in for. This closes three instances across `gsd-tools`: a resolver
that never read the install-time signal it was documented to read, a
Codex sandbox map that was fully redundant with the tool contract it stood
in for, and a dependency-resolution tier lost when the SDK lineage was
retired.
**A non-Claude install now resolves its own runtime with no config
needed (#3897).** `resolveRuntime`'s ladder was `GSD_RUNTIME` env var →
project `config.runtime` → `'claude'` — the per-install `.gsd-runtime`
marker the installer has written beside `VERSION` since #2297 was read by
four separate hand-rolled copies (`model-resolver.cts`, and two more inside
`gsd-cursor-subagent-start.js`), but never by `resolveRuntime` itself. A
Codex, Cursor, or other non-Claude install with no `GSD_RUNTIME` set and no
`runtime` key in `.planning/config.json` therefore still resolved `claude`
everywhere `resolveRuntime` is consulted (slash-command style, `query
teams-status`, `validate agents`'s agent-directory selection, and 19 other
call sites). The marker is now the third rung — `GSD_RUNTIME` → `config.runtime`
→ install marker → `'claude'` — so those installs resolve their own runtime
by default. The marker's contents are never trusted verbatim: they are routed
through the same name-normalization the env rung already uses, so a marker
holding an unexpected or hostile value degrades exactly like an unexpected
`GSD_RUNTIME` value would.
**Codex sandbox permissions are derived from each agent's own tool contract,
not a hand-maintained map (#3897).** `generateCodexAgentToml` looked up
`sandbox_mode` in an 11-entry `CODEX_AGENT_SANDBOX` map, falling back to
`read-only` — silently — for every role the map didn't name. Measured against
all 35 shipped roles, the map's 11 entries agree with deriving `sandbox_mode`
from each role's declared `tools:` frontmatter (`workspace-write` when it
declares `Write` or `Edit`, `read-only` otherwise) with **zero disagreements**,
so the map is deleted rather than clamped. The fallback, however, was
under-granting: 16 of the 24 roles that hit it declare `Write`/`Edit` and
would derive `workspace-write`. Pending a decision on whether Codex actually
enforces `sandbox_mode` (a question the derivation can't answer on its own),
those 16 are held at `read-only` by an explicit, self-invalidating hold
list — a hold whose role no longer derives broader, or that names a role
that no longer exists, fails loudly instead of being silently honored.
**Every one of the 35 emitted `.toml` files is byte-identical to before this
change** — the fix is in provenance (an explicit, reviewable rule instead of
a silent default), not in any installed agent's actual permissions today.
**`validate agents` now reports Codex sandbox drift (#3897).** A new
`sandbox_posture` field — report-only, exit 0, same shape as the existing
`codex_posture` — flags any installed Codex `.toml` whose `sandbox_mode`
disagrees with what its role's tool contract derives. Populated only when
the active runtime is `codex`.
**`depends_on` accepts the bare plan number (#3897).** A plan's frontmatter
could already reference a dependency by its full id (`"03-01-auth-hardening"`)
or its canonical phase-plan prefix (`"03-01"`). A third form — the bare
plan number alone (`"01"`) — existed in the retired SDK lineage but was lost
when that lineage was consolidated; a plan written with it silently dropped
the edge entirely, collapsing into wave 1 regardless of its declared
dependency. That form is restored, scoped to the **same phase only**: `"01"`
resolves to the sibling plan whose canonical id ends `-01`. This is an
observable behavior change — a phase whose plans used the bare form and had
silently collapsed into a single wave will now execute in its actual declared
waves. Two plans in the same phase sharing a bare form resolve first-write-wins,
by sorted plan-file order — deterministic, but arbitrary where the collision
happens, matching the retired behavior exactly.
**Known limits:**
- The 17 held Codex roles are pinned at `read-only`, not widened. A faithful
derivation from the tool contract would widen them, because they declare
`Write` or `Edit`; the previous hand-maintained map never listed them and
they fell through a silent `|| 'read-only'` default instead. Deriving *and*
holding keeps emitted TOML byte-identical for all 35 roles today while the
derivation becomes the single owner of the rule. Widening them is a follow-up
once Codex's actual enforcement of `sandbox_mode` is confirmed — until then a
hold is reversible and a widened sandbox is not. The hold list is
self-invalidating: an entry naming a role that no longer derives broader, or
that has no file in the shipped roster, fails rather than rotting into the
subset map this change deletes.
- The bare plan-number form is ambiguous by construction across two plans in
the same phase that share a short form; first-write-wins is deterministic
but not a conflict warning. Prefer the full or canonical id when a phase's
plan numbering risks a short-form collision.
- The install marker never feeds model-tier resolution (`model_profile_overrides`,
`model_policy.runtime_tiers`) — that still reads `config.runtime` alone, and
reporting-only host detection (`agent_runtime`) is a separate, pre-existing
ladder this change does not touch.
---
### 3910. The Raw Terminator Is Banned by Construction
**Purpose:** Make a bare `process.exit(...)` a lint error everywhere it matters, so the
"nothing fails with success" defect class ADR-3889 exists to close cannot silently reopen
through a new call site.
**What changed (ADR-3889 Phase 6, #3910):**
- New rule `local/require-registered-exit` (`eslint-rules/require-registered-exit.cjs`) flags
any `CallExpression` shaped exactly like `process.exit(...)`. It does **not** flag
`process.exitCode = N` — that assignment is the correct drain-then-exit pattern `runMain`
itself uses, and the two are structurally distinct (an assignment target is never a
`CallExpression`).
- Registered on four globs: `src/**/*.cts`, `scripts/**/*.cjs`, `hooks/**/*.js`,
`gsd-core/bin/**/*.cjs` (`eslint.config.mjs:420-426,545-547,574-576,601`). Registering on
`src/**/*.cts` — not only the emitted `gsd-core/bin/lib/*.cjs` mirrors, which are globally
eslint-ignored (ADR-457) — is load-bearing: a rule registered only on the emitted surface is
blind to the real sources, the same way `n/no-process-exit` went invisible (#3496).
- The dead `n/no-process-exit: 'off'` carve-out for `hooks/**` is deleted: Phase 7 (#3911)
migrated every enforcement hook onto `terminateNow`, so it protected nothing.
- Exactly two allowlist entries, repo-wide:
1. The body of `terminateNow` in `src/cli-exit.cts` — detected **structurally** (any
`process.exit()` lexically nested inside a function named `terminateNow`, *and* the file's
basename is `cli-exit.cts`), not by path+line, so it does not rot when the function moves.
2. `gsd-core/bin/gsd-tools.cjs`'s `ensureRuntimeBuild` bootstrap-failure path, via an inline
`// eslint-disable-next-line local/require-registered-exit` with a stated reason — it runs
*before* `./lib/cli-exit.cjs` is even required, so the registered-exit seam does not exist
yet at that point in the process's lifetime.
- The last raw terminators in `src/**/*.cts` were migrated onto the seam, most notably
`src/io.cts`'s `error()`: it changed from an uncatchable `process.exit(1)` to a catchable
`throw new ExitError(1)` (stderr output is byte-identical; `runMain` projects the exit code).
`terminateNow` could not serve this site — ADR-3889 §1 makes exit codes 0 and 1
unallocatable, so `nameForExitCode(1)` throws. That control-flow change required three
interceptor fixes so an `ExitError` reaches `runMain`: `command-routing-hub`'s `dispatch()`
now rethrows it, and the profile-pipeline router's detached `.catch()` no longer calls
`error()` — it writes stderr and sets `exitCode` in place.
- **Known limits (documented and test-pinned, not endorsed):** the rule matches the literal
`process.exit(...)` shape only, with no scope/flow analysis. It does not catch
`process['exit'](0)` (computed member access), `const e = process.exit; e(1)` (aliasing to a
local binding before calling), or `process.exit.call(...)`/`.apply(...)` (indirect invocation).
Catching these needs binding/scope-aware analysis, out of scope for this issue; pinning tests
in `tests/eslint-rules.test.cjs` assert today's non-detection so a future widening is a visible
choice, not a silent one.
See [Resolve a raw-terminator finding](../how-to/resolve-a-raw-terminator-finding.md) for what to
do when this rule fires, and [ADR-3889](../adr/3889-process-exit-contract.md) for the exit-code
registry the seam is layered over.
---
### 3911. Hooks Declare Their Crash Policy
**Purpose:** Give every shipped enforcement hook (`hooks/*.js`, `hooks/*.sh`) a
named, auditable termination vocabulary instead of a bare `process.exit(N)`
scattered per file — and make a hook's fail-open/fail-closed choice a
declaration a reviewer can see, rather than an inference from which literal
integer follows `process.exit(` in its outer `catch`.
**What changed (ADR-3889 Phase 7, #3911):**
- `hooks/lib/hook-exit.js` (hand-written) exposes `allow(payload)` → exit 0,
`deny(payload, stderrPayload?)` → exit 2, and `crash(onCrash, payload)`,
which dispatches to `allow`/`deny` per a `HOOK_ON_CRASH` policy the caller
must supply — `crash()` has no default policy, so a hook cannot fail open
by omission.
- Every one of the 19 enforcement hooks under `hooks/*.js` now declares
`const ON_CRASH = HOOK_ON_CRASH.ALLOW` (or `DENY`) once, with a
hook-specific comment naming why, and calls `crash(ON_CRASH, payload)` from
its outer catch instead of a bare `process.exit(0)` / `process.exit(2)`. No
hook's effective exit code changed — this is a naming-and-declaration
migration, not a behavior change.
- `hooks/lib/cli-exit.js` and `hooks/lib/exit-code-registry.js` are new,
generated, git-tracked copies of the exit-code seam (`src/cli-exit.cts` /
`gsd-core/bin/shared/exit-codes.json`), so a shipped hook can terminate
correctly on a raw, unbuilt clone without depending on `gsd-core/bin/lib/`
tsc output. Generated by `scripts/gen-hooks-cli-exit.cjs` and
`scripts/gen-exit-code-registry.cjs`, both `--check`ed by
`npm run lint:generated-sync`.
- `terminateNow` gained an optional third argument, `stderrPayload`, so a
deny can send a full JSON body to stdout and a distinct plain-text reason
to stderr — needed because `gsd-write-guard.js` (Kimi's native hook bus
reads stderr verbatim back to the model) always sent only the bare reason
string on fd 2. The two streams are now written in independent try/catch
blocks: previously a payload that failed to serialize on fd 1 aborted
before fd 2 ever wrote, producing a deny with an empty stderr reason.
- Two hooks are deliberately **not** migrated to `deny()`:
`gsd-read-injection-scanner.js` (PostToolUse — its harness reads the block
decision from the JSON response body, not the exit code) and
`gsd-cursor-subagent-start.js` (follows Cursor's own `subagentStart`
protocol, which reads `permission: "deny"` from the JSON body at exit 0).
Both still use `allow()`/`crash()` for their no-op and crash paths.
- `gsd-phase-boundary.sh`, `gsd-session-state.sh`, and
`gsd-validate-commit.sh` gained `set -euo pipefail`, and
`gsd-validate-commit.sh`'s three swallow-and-pass sites (the opt-in config
read, JSON command extraction, and the `isGitSubcommand` classifier) now
distinguish a genuine negative from "could not run" — on the latter they
emit a stderr diagnostic and exit 0 instead of silently allowing every
commit (#3838).
See [Declare a hook's crash policy](../how-to/declare-a-hook-crash-policy.md)
for the full how-to, and
[ADR-3889](../adr/3889-process-exit-contract.md) for the exit-code registry
this vocabulary is layered over.
---
### 3912. gsd-tools Declares Outcomes, Pinned at v1
**Purpose:** Give every `gsd-tools` terminating path a declared outcome name, and project that
declaration through the versioned exit contract ([ADR-3889](../adr/3889-process-exit-contract.md)
§4) — without changing a single exit code for a caller that has not opted in.
**Reference — what changed (ADR-3889 Phase 8, #3912):**
- `error(message, reason)` now maps its `reason` argument onto a declared outcome name (`USAGE`,
`NO_INPUT`, `UNAVAILABLE`, `INTERNAL`, `FAIL`) via a fixed table closed over all 25
`ERROR_REASON` members (`src/io.cts`'s `REASON_TO_OUTCOME`). Under the default contract version
`v1`, the declaration is recorded but `error()` still throws `ExitError(1)` unconditionally,
byte-identical to every prior release. Under `v2` (`--exit-contract=v2` /
`GSD_EXIT_CONTRACT=v2`), it throws `ExitError(projectOutcome(outcome, 'v2'))` instead — e.g.
`SDK_MISSING_ARG`/`SDK_UNKNOWN_COMMAND` project to `64` (`USAGE`), `CONFIG_KEY_NOT_FOUND` to `66`
(`NO_INPUT`). All 278 call sites are untouched; 226 pass no reason and default to `UNKNOWN` ->
`FAIL` -> exit `1` under both versions.
- `output()` now declares `DEGRADED` whenever its payload carries a **serializable** `error` value
(any key order — `{found:false, error}` counts the same as `{error, found:false}`). The
discriminator is survives-`JSON.stringify`, not mere key presence: `{ error: undefined }` does
**not** declare `DEGRADED`, because `JSON.stringify` drops an `undefined`-valued property before
the payload reaches the wire.
- A third `globalThis` cell (`src/cli-exit.cts`'s `PENDING_OUTCOME_KEY`) holds the pending declared
outcome between `output()` and `runMain`. Semantics: **last declaration wins, cleared on
consumption** — a later clean `output()` call in the same invocation undoes an earlier degraded
one, and `runMain` clears the cell on every exit so a second `runMain` in the same process never
inherits a stale declaration.
- **Precedence** for the code a void-returning `main()` ends up with, highest first: (1) an explicit
`main()` return, (2) a non-zero `process.exitCode` `main()` already set directly, (3) the pending
declared outcome, (4) otherwise `0`. Projection may only ever **set** a code, never **lower** one
— a review pass wrongly concluded the cell was fail-closed by construction; without rule 2, `state
validate --strict` briefly exited `0` where it must exit `1`.
- **`v1` is byte-identical.** `DEGRADED` projects to `0` under `v1` and to `80`
(`exitCodeFor('DEGRADED')`) under `v2` — that asymmetry is
[ADR-2980](../adr/2980-payload-carried-error-is-a-degraded-result.md)'s compatibility boundary,
deliberately preserved, not a bug to reconcile.
**Explanation — why this is the shape it is:**
ADR-2980 ratified `output({error})`'s exit-0 population on measured blast radius (`output` has 170
direct callers) and Hyrum's Law grounds — a CLI exit code has no `/v2/` of its own, so normalizing
it in place would have broken every caller already treating exit `0` as a soft signal. Its own
"Revisit if" clause named the missing piece: *"a future `gsd-tools` major version provides a
compatibility boundary that a CLI exit code otherwise lacks."* ADR-3889 §4 built exactly that
boundary — a versioned projection selected by flag or env var, defaulting to today's behavior — and
this phase is what wires `error()` and `output()` onto it. Declaring an outcome is unconditional and
immediate; only its *projection* onto an integer is deferred behind the version switch, so the
population ADR-2980 ratified keeps exiting `0` until a caller explicitly asks for something else.
The count matters here too: an AST re-measure for this phase found **64** `output({error})` call
sites across the same nine modules ADR-2980 named — not the 60 that ADR itself recorded, the drift
concentrated in `frontmatter.cts`, `phase.cts`, and `roadmap.cts`. The `v2` projection is asserted
over the enumerated 64, not a restated 60; see ADR-2980's amendment for the module-by-module
breakdown.
See [Adopt the v2 exit contract](../how-to/adopt-the-v2-exit-contract.md) for how to opt in and what
it means for a CI gate, [`docs/json-errors.md`](../json-errors.md#outcome-declaration-and-the-versioned-exit-contract-adr-3889-4-3912)
for the full reference, and
[ADR-2980](../adr/2980-payload-carried-error-is-a-degraded-result.md) /
[ADR-3889](../adr/3889-process-exit-contract.md) for the decisions.
---
### 3951. Reachable Lint Rules and a Non-Destructive Quick-Task Append
**Purpose:** Make two ESLint rules cover the code they were written to govern, and stop
`quick-tasks-append` from overwriting curated `progress.*` values on a body-only write.
**What changed:**
- **`local/no-adhoc-markdown-parsing` reaches its whole registered surface.** The rule short-circuited
unless a file's path matched a flat `src/*.cts` pattern, so it **self-gated on its own filename**.
Two consequences: 28 `.cts` files in `src/` subdirectories sat inside the `src/**/*.cts` glob it was
registered on and were silently skipped, and the rule could not be extended by configuration at all —
widening the glob alone left it inert. Both halves now move together, and a test pins that the gate
and the registration agree in *both* directions.
- **The rule now also covers `tests/**` and `scripts/**`**, which surfaced **80 hand-rolled markdown
parses across 43 test files**. Seventy are routed through the existing `markdown-sectionizer` and
`markdown-table` seams; ten are suppressed with a stated reason (six of those are a shell-pipe
detector whose regex merely resembles a table).
- **`local/no-adhoc-regex-escape` sees property access.** Its unsafe-`new RegExp` arm examined only
bare identifiers, so `new RegExp(obj['key'])` — the shape runtime data actually arrives in — was
invisible. That is why it never fired on a known ReDoS. It now inspects `MemberExpression`, with an
exemption keyed strictly on the property being `source` (18 safe sites), plus provenance exemptions
for `_SOURCE` constants reached through a required module (3 sites).
- **`quick-tasks-append` can write the canonical row.** Optional `--quick-id`, `--slug` and
`--directory` let a caller that has a real quick task emit the same row `/gsd-quick` renders. Omit
them — as `fast.md` does, having neither an id nor a directory — and the row is byte-identical to
before.
- **A body-only append no longer re-derives progress.** The route was the only body-only STATE.md
writer not passing `{ resync: false }`, so appending one row triggered a full re-derive of the
disk-derived `progress.*` frontmatter and replaced curated values. Reproduced: a project with two
real phase directories and a curated `total_phases: 25` collapsed to `2` on append.
**Found by the widening:** `tests/config-field-docs.test.cjs` asserted that
`workflow.subagent_timeout`'s documented default is not `600` — but read the **Type** column instead
of **Default**, so it compared `'number'` against `'600'` and could never fail. The guard against
regressing to the old seconds-based default had been inert. It is now row-scoped and real.
**Known limits:**
- #3426 and #3239 are **not** closed by this. Their hand-rolled scans in
`tests/package-legitimacy-gate.test.cjs` are built from line filters and `split('|')`, not the
regex-literal fingerprints this rule detects — measured at zero violations even with the gate
bypassed. They need new detectors, which is a separate design.
- The 10 suppressions are suppressions, not fixes. Each names why the raw markdown text is the
subject of that assertion.
- The `src/` subdirectory hole was **latent** — zero violations existed there when it was fixed. It is
closed because "no violations today" is not a property that keeps holding, not because it was
hiding anything.
---
### 3970. Per-Task External-Tracker Content-Resolution Seam
**Purpose:** Let a capability declare that an external issue tracker — beads, Linear, Jira,
GitHub Issues — owns a task's *content* (`<action>`/`<verify>`/`<acceptance_criteria>`/
`<read_first>`/`<done>`), not just its status, so `execute-plan.md` can resolve that content
from the tracker at execution time instead of reading it inline out of `PLAN.md`.
**What changed (ADR-3646, #3970):**
- A new optional feature-body manifest field, `taskContentResolver`, declares a `trackerPrefix`
(matched against a task's `<task tracker-id="beads:GSD-42">` attribute — everything before the
first `:`) and a bounded `invoke` (`binary`, `args` carrying the `{{id}}` placeholder,
`timeoutMs`).
- `execute-plan.md`'s per-task loop gains one new, unconditional call before that task's
`read_first` gate: `gsd_run task resolve-content --plan <path> --task-id <tracker-id> --raw`.
A task with no `tracker-id` attribute is unaffected — the call is only made when the attribute
is present, and resolves instantly to a no-op for every project that declares none.
- **The safety property is a real process exit code, not a prose dispatch.** No capability
registered for the tracker, or resolution succeeds with empty content, exits `0` with
`resolved: false` and falls back to inline `PLAN.md` — the one legitimate pre-migration
boundary case. Resolution succeeding with non-empty content exits `0` with `resolved: true` and
its `content` supersedes the task's inline fields for every downstream gate in the execute step.
A resolver that is declared but fails — tracker unreachable, id not found, timeout, malformed
JSON — makes `task resolve-content` itself **exit non-zero**, which `execute-plan.md` treats as
a **hard halt**: stop, surface the tracker-id/prefix/stderr, never fall back to stale
`PLAN.md` content.
- `execute:task` is a new dispatch shape below wave granularity, deliberately **not** one of the
12 existing loop extension points (`discuss:pre` … `ship:post`) and not routed through
`gsd_run loop render-hooks <point>` / `activeHooks`. It exists because the existing
`step`/`gate` prose-dispatch mechanism cannot deliver a hard-halt guarantee while dispatch
reliability at that layer is an open concern (#3647) — see ADR-3646's Context and Rejected
Alternatives for the full reasoning.
See [Develop a task-content resolver capability](../how-to/develop-a-task-content-resolver-capability.md)
for the authoring walkthrough, [Capability manifest → `taskContentResolver`](../reference/capability-manifest.md#taskcontentresolver)
for the field reference, and
[`loop-hook-dispatch.md`](../../gsd-core/references/loop-hook-dispatch.md#the-executetask-point-a-different-shape)
for how `execute:task` differs from the twelve prose-dispatched points.
---
### 4014. Unreadable-Directory Scope Signal
**Purpose:** ADR-3473 §8.4 ("failure is a value") applies to filesystem
listings, not only command argv. `#3885` (B5) gave `roadmap analyze`,
`gap-checker`, and `init`'s JSON bundles a `context_read_error` /
`phase_dir_read_error` string naming an unreadable phase directory — but the
underlying `has_context` / `hasContext` boolean stayed `false` either way,
so a consumer branching on that boolean alone still cannot tell "genuinely
no context file" from "could not read the directory at all." This closes
that gap with a typed signal, reusing ADR-3180's existing frozen `SCOPE`
enum rather than a new vocabulary.
**`findContextMdIn` (`src/planning-workspace.cts`) now reports its own
scope.** Called with a directory path, it returns `{ file, files, scope }`
instead of a bare filename-or-null, and never throws — an unreadable
directory reports `scope: 'unreadable'` (previously it threw, forcing every
caller to hand-roll its own `try`/`catch`); a genuinely absent directory
(`ENOENT`) reports `scope: 'complete'`, the same "real empty" answer as
today. The array-input call form (an already-read listing) is unchanged.
**Five downstream call sites gain an additive `scope` field, none renamed
or removed:** `roadmap analyze`'s `AnalyzePhase.context_scope`,
`gap-checker`'s `phase_dir_scope`, and `init`'s `context_scope` on all three
JSON bundles (`init plan-phase`, `init phase-op`, `init manager`) —
including `cmdInitManager`, whose own read failure previously vanished into
a bare empty `catch {}` with no signal of any kind. `getPhaseFileStats`
(`src/core-utils.cts`) — the shared listing owner behind `roadmap analyze`
and `init`'s `has_context` — no longer lets its own failed read get masked
by an unrelated, already-successful `scanPhasePlans` scope on the same
phase directory.
**Known limits:**
- `context_read_error` / `phase_dir_read_error`'s message text is now a
fixed "Could not read phase directory `<path>`" rather than embedding the
underlying OS errno text — `findContextMdIn`'s directory-string form
reports only the `SCOPE` discriminator, not the raw caught error. The
field's presence and type are unchanged; only its message detail is
coarser than before #4014.
- `init.cts`'s three call sites call `findContextMdIn` for the scope signal
and then still run their own, pre-existing `fs.readdirSync` on the same
path for the rest of their output — an intentional, additive-only choice
to avoid altering already-complex failure control-flow at those sites,
not a performance optimization.
---
### 4836. Graphify CLI Preferred for Planner and Researcher Graph Queries
**Purpose:** `gsd-planner` gets **one** knowledge-graph query per phase and
`gsd-phase-researcher` gets two or three. That single shot decides which modules
the plan treats as related, and therefore how tasks are ordered into waves. It
was spent on the built-in reader, which seeds by case-insensitive **substring**
match over a node's label and description and then expands a hardcoded two hops
— so the phase "User Authentication" seeds on `author`, `authoring`, and
`unauthorized` with exactly the same weight as `authenticate`, and when the
inflated result exceeds `--budget` the trimmer drops edges by confidence tier.
The `graphify` CLI, already a hard dependency of `/gsd-graphify build`, ranks
seeds (IDF weighting, trigram fuzzy matching) and applies context filters before
traversal.
**Both prompts now prefer the CLI and fall back to the built-in reader.** The
branch is `command -v graphify`, the same degradation shape the repo already
uses for Context7 → `ctx7` in `references/research-documentation-lookup.md`. No
new config key: a `graph.json` can only exist if `graphify update .` ran, which
requires the binary, so binary presence is a self-satisfying gate. The fallback
covers edge cases — a CI checkout with a committed graph, a binary since removed
— not the common path. No new tool grant either: both agents already have
`Bash`.
**The planner additionally runs `graphify affected`.** The reference states its
own goal as "which subsystems may be affected by changes in this phase", which is
literally reverse traversal by relation. The built-in reader only approximates it
with undirected two-hop expansion, and has no equivalent verb, so `affected` is
skipped on the fallback path.
**`gsd-tools graphify status` now returns `graph_path`.** The CLI takes the graph
location as `--graph`, and the prompts must not re-derive
`.planning/graphs/graph.json` for it — that would point the CLI at a
non-existent local mirror in exactly the umbrella multi-repo setup
`graphify.graph_path` (#1825) exists to serve. `status` already resolves the
override, so it now reports the absolute path it resolved, on both the
graph-present and the graph-missing branch. For the same reason the presence gate
in both prompts is now the `status` call itself rather than a bare `ls` of the
default location.
**Known limits:**
- **The two paths return different shapes.** `graphify query` emits prose and has
no `--json` flag; `gsd-tools graphify query` emits JSON with per-edge
confidence tiers and `budget_met`/`budget_estimate`. Both are consumed by a
model, and nothing machine-parses this block, but the prompts now say so
explicitly instead of implying a stable shape.
- **`--budget` means different things on the two paths** — rendered output on the
CLI, estimated payload bytes in the built-in reader (#2738). Same flag name,
different unit.
- With `graphify` absent from `PATH` the fallback runs and the injected graph
context is byte-identical to before.
---
_Generated by `scripts/gen-features.cjs` — add a fragment under `docs/features/` and run `--write`._
<!-- FEATURES:END -->
## Related
- [Commands](COMMANDS.md)
- [Configuration](CONFIGURATION.md)
- [docs index](README.md)