Commit Graph

736 Commits

Author SHA1 Message Date
0xdhx
238bee7b03 fix(#4881): trust worktree.baseRef:"head" in harness mode; a WorktreeCreate hook withholds it and only a measured fork base restores it (#4921)
* fix(#4881): trust worktree.baseRef:"head" in harness mode and keep the #4868 observation as what restores it under a WorktreeCreate hook

The pre-dispatch base check still derived its harness-mode verdict from the
retired #48 premise that the harness never reads worktree.baseRef: with
"head" set and HEAD diverged from origin/HEAD it degraded every wave with
baseref-head-ignored-by-harness. #4868 inserted an observation of a clean
prior harness worktree at HEAD ahead of that comparison, but an
execute-phase run never has one at the moment it checks — the base-check
runs before any dispatch, a degraded wave creates no worktrees, and a wave
that did run in worktrees has them removed and HEAD moved before the next
check — so the common case was unchanged (#4881 repro states 1 and 4).

Re-scope of the closed #4752 onto current next, with #4868 kept:

- branch a trusts "head" in both isolation modes, the way the harness is
  measured to behave (#4588: three settings layers, three OSes), and the
  spawn-time exit-42 guard stays the observation-based backstop;
- a Claude Code WorktreeCreate hook in any settings file the check reads,
  or a file that does not parse, withholds that trust — the hook creates
  the worktree without applying the setting — and the inferred comparison
  runs, degrading with baseref-head-bypassed-by-hook;
- the #4868 observation (b2) now sits behind branch a: it is reached only
  when "head" was not trusted outright, and on a hook host it is what
  restores the trust — a hook that forks from HEAD leaves exactly that
  evidence, one that forks elsewhere never does. Its per-HEAD cache is
  unchanged. It is skipped under an explicit --observed-fork-base, which
  outranks an inference from a prior worktree;
- --observed-fork-base <sha> threads a measured fork base through the
  evaluation (strict full-hex, TypeError otherwise).

Tests: the #4868 rows are unchanged and still reachable (they run with the
setting unset); the #4752 rows re-land, with the one exit-128 row updated
to the degrade #4734 pinned since; a new #4881 block pins the start-of-run
state (baseref-head, git never consulted), the hook + observation
composition in both directions, and that an explicit observation skips
the probe.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RPXQQPzHGintWtbhBoQCRS

* docs(#4881): rewrite the eight surfaces that still state the harness ignores worktree.baseRef:"head"

Every prose surface #4868 left untouched still asserted the retired #48
premise as verified fact, starting with the step file the orchestrator
reads. Each now describes the measured behaviour, the WorktreeCreate-hook
exception, the --observed-fork-base input, and the #4868 observation as
what lifts the hook degrade; docs/CLI-TOOLS.md gains the
fork-from-head-observed reason row #4868 did not document.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RPXQQPzHGintWtbhBoQCRS

* chore(#4881): set changeset fragment pr to 4921

* fix(#4881): withhold the #4868 observation under the hook interlock

The hook interlock this PR added withheld the worktree.baseRef:"head"
trust but still let b2's prior-worktree observation restore it, and that
observation cannot be attributed to the hook. The evidence is a clean
agent worktree sitting at the orchestrator HEAD; nothing on disk records
which creator left it there, so one the plain harness created BEFORE a
WorktreeCreate hook was configured — with HEAD unmoved since — reads as
evidence for the hook. It was the one fail-open branch in a mechanism
documented as fail-closed.

Keying the observation cache by hook configuration does not close it.
observeHarnessForkFromHead has two legs: a HEAD-keyed cache and a live
probe over .claude/worktrees/agent-*. A hook-keyed cache simply misses,
and the miss falls through to the probe, which re-finds the same stale
worktree and re-confirms. The probe takes no hook input at all. So the
observation is not consulted under the interlock rather than re-keyed:
on a hook host the only admissible positive signal is an explicit
--observed-fork-base measurement of the dispatch in hand, and absent one
the inferred comparison runs and a mismatch degrades with
baseref-head-bypassed-by-hook, leaving the spawn-time exit-42 guard as
the backstop.

Scoped to the case branch a. declined to trust: "head" set AND a hook
(or an unparseable layer) in the harness's path. With no "head" setting
b2 is #4868's own arm and is unchanged, hook or not — re-scoping that
trust is a separate question this PR does not open, and a test pins the
boundary.

Cost, stated: a hook host with "head" set, on a branch diverged from
origin/HEAD and passing no observation, now runs sequentially. It still
runs parallel when HEAD matches origin/HEAD. No workflow threads an
observation today.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

* refactor(#4881): drop the now-unused forkRef message-builder parameter

buildMsgBaserefHeadIgnored stopped reading forkRef when the message was
made mode-neutral, and the parameter was retained with `void forkRef;`
for symmetry with its two sibling builders, which do read it. Symmetry
is not reason enough to keep a dead parameter on a module-private
function with one caller, so drop it (#4921 review).

Behaviour is unchanged; the message text is pinned by an existing
full-string assertion, which is what covers the only real risk here —
transposing the two remaining arguments at the call site.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

* test(#4881): exercise both FULL_SHA_RE alternatives at their boundaries

The invalid-observation list pinned 39 and 41 hex around the 40-hex
SHA-1 arm but left the 64-hex SHA-256 arm's own +/-1 boundary
unexercised, which the repo's boundary-coverage convention asks for
(#4921 review). Adds 63, 65 and a 64-length non-hex string.

The regex already rejected all three -- this is coverage of correct
behaviour, not a fix -- so it carries no negative control against a
pre-fix base. That it is not vacuous was shown instead by widening the
arm to {63,65}, under which the row fails by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

* docs(#4881): finish the surface sweep the hook interlock owes

Self-found by this round's own pre-push adversarial review, over six passes.
Eight sites, four classes.

FIVE were prose still asserting the stance the interlock overturned, which
reads as live canon to anyone arriving cold: docs/CLI-TOOLS.md,
docs/CONFIGURATION.md and gsd-core/references/planning-config.md each still
said the hook degrade is lifted "unless/until a clean prior harness worktree is
observed"; observeHarnessForkFromHead's own header still said a qualifying
worktree "can only exist if the harness forked from HEAD"; and a test name
still called the observation "required on a hook host" when it is now
inadmissible there. My own sweep had grepped for "restores"/"lifts" and missed
every one -- the ordinary failure of a grep, which returns what you thought to
search for.

The SIXTH is the same class one step worse: gsd-core/workflows/execute-plan.md
still said flatly that Claude Code's isolation="worktree" "forks from
origin/HEAD, not live local HEAD" -- in a paragraph THIS PR already edits, a
few sentences after the clause it corrected. A tombstone makes only its own
line clean; adjoining text asserting the dead stance is the other half of the
same defect. Now qualified on the setting, with the hook exception named.

The SEVENTH is a proof-strength overstatement that predates this PR, with a
driven counterexample: a worktree created from an older base and since `git
checkout --detach`ed onto HEAD is clean, sits at HEAD, and satisfies the probe
identically, so "can only exist" was false. The worktree's own reflog does
retain that original checkout -- the information is not lost, the probe simply
does not consult it.

The EIGHTH is that the header described only one of the function's two legs. A
cache hit returns the prior conclusion without reading any worktree, so "the
probe reads a worktree's present state" was true of the live probe and false of
the cache. The header now separates them, and names the cache's blindness as a
third reason the observation is inadmissible under a hook.

Gaps seven and eight are inherited from #4868 and accepted there for the
no-hook case. Nothing about the mechanism changes here; only what the header
claims for it.

No behavioural change -- comments, prose, and one test's registered name.

Emitted-Drift-Ack-Growth: execute-plan.md — the Pattern A paragraph gained a qualifying
 clause: it stated flatly that Claude Code forks from origin/HEAD, which is the premise
 this PR retires, a few sentences after the clause the PR had already corrected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-23 20:15:28 -04:00
Michel Moreira
1e3e1f7cd8 enhance(#4570): allow disabling planner stall detection (#4585)
* enhance(#4570): allow disabling planner stall detection

* docs(#4570): add changelog fragment

Emitted-Drift-Ack-Growth: plan-phase.md — the explicit opt-out gate covers all five planner and checker spawn classes
Emitted-Drift-Ack-Growth: settings-advanced.md — the toggle prompt and bounded-recovery warning expose the new setting

* chore(#4570): refresh compact-content baseline

* fix(#4570): preserve default-on watchdog fallback

* docs(#4570): qualify the chunked-mode orchestrator rules with the toggle

The two chunked-planning-mode stall-watch imperatives read as
unconditional, with the opt-out stated only in a following bullet. State
the PLANNER_STALL_DETECTION_ENABLED condition inline, matching the three
sites already qualified in plan-phase.md.

* chore(#4570): refresh compact-content baselines against rebased next

* fix(#4570): sync planner stall launcher

Keep the stall-detection helper aligned with the canonical runtime launcher.

* chore(#4570): refresh compact baseline after rebase

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-23 18:46:23 -04:00
Michel Moreira
af822a8024 fix(#4776): resolve the artifact-exists prompt under --auto (#4832)
* fix(#4776): resolve the artifact-exists prompt under --auto

/gsd-ui-phase <phase> --auto stopped at 'UI-SPEC.md already exists for
Phase {N}. What would you like to do?' whenever the file was on disk —
which is most often after an earlier run wrote the contract as a draft
and ended before its checker ran, exactly the state a re-run exists to
verify. Step 4 had no --auto arm; step 9.5 in the same file has had one
since it was added, which is how the drift went unnoticed.

Step 4 now auto-selects Skip: the existing UI-SPEC is left untouched and
the run proceeds to the checker. Skip is the only non-destructive choice
— Update re-runs the researcher, which rewrites the whole contract and
drops answers a person already recorded in it, and View exits without
verifying anything.

spec-phase.md's artifact-exists arm auto-selected 'Update it', which is
the same defect with the opposite sign: an unattended run regenerating a
spec nobody is watching. Per the decision recorded on #4776 — an
unattended run reuses an existing artifact rather than regenerating it —
it now auto-selects Skip and leaves the spec unchanged.

The max-revision-iterations escalation (Force approve / Edit manually /
Abandon) is deliberately untouched and pinned by a test: accepting
blocking findings is a decision a person makes.

Closes #4776

* chore(#4776): add changeset fragment

Emitted-Drift-Ack-Growth: ui-phase.md — --auto arm added to the existing-UI-SPEC branch (#4776)
Emitted-Drift-Ack-Growth: spec-phase.md — --auto arm reworded to reuse the existing SPEC (#4776)

* fix(#4776): extend the reuse-as-is --auto fix to the 3 sibling files

The PR's original scope claim -- that ai-integration-phase.md,
eval-review.md and ui-review.md were "scoped by triage to their own
issues" -- was false; no such issues existed, and it contradicted #4776's
own most recent (2026-09-16) triage comment, which explicitly widened the
fix to require all 5 files under one recommended fix.

Applies the same reuse-as-is pattern: ai-integration-phase.md mirrors
ui-phase.md's 3-way Update/View/Skip shape (auto-selects Skip);
eval-review.md and ui-review.md have only Re-audit/View (auto-selects
View, the only non-regenerating choice). None of the three had any prior
--auto handling at all -- each has exactly one AskUserQuestion call site
total, and it is the one this fix resolves, so an --auto run through any
of them no longer stalls anywhere.

Emitted-Drift-Ack-Growth: ai-integration-phase.md — the --auto arm reusing an existing AI-SPEC is the deliverable (#4776)
Emitted-Drift-Ack-Growth: eval-review.md — the --auto arm reusing an existing EVAL-REVIEW is the deliverable (#4776)
Emitted-Drift-Ack-Growth: ui-review.md — the --auto arm reusing an existing UI-REVIEW is the deliverable (#4776)

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-23 11:13:34 -04:00
0xdhx
b90eef28e8 enhance(#3829): report code review severity counts and record a per-finding disposition (#3861)
* enhance(#3829): report code review severity counts and record a per-finding disposition

`code_review_gate` extracted `status:` from REVIEW.md's frontmatter and discarded
the `critical`/`warning`/`info`/`total` values sitting in the same range, so its
output was byte-identical for a review with one `info` finding and a review with a
Critical. Nothing anywhere recorded what happened to a finding: no file under
`gsd-core/workflows/` branches on `issues_found`, and `gsd-verifier.md` has zero
references to REVIEW.md. A phase therefore reached `phase.complete` with Criticals
standing and no trace they had been seen.

Both halves were approved on the issue; the gate stays advisory.

A — severity surfacing, in `execute-phase.md`. The gate states the breakdown it
already parsed, accepting `blocker:` as the documented tier-equivalent of
`critical:`. The breakdown is shown only when all four counts are numeric
(`REVIEW_COUNTS_OK`); otherwise the countless message stands, because gating on
the total alone still emits `6 findings —  critical` for a review carrying a total
and nothing else.

Frontmatter is extracted by an `awk` that emits only when it saw the CLOSING
delimiter, after stripping CR. A `sed` range re-opens on a body `---` and runs to
EOF: first-match protects a key the frontmatter always carries, but not an
optional one, so a review with no `findings:` block and a body `total:` line would
have reported the body's number. An unterminated block would leak the whole body
the same way.

Every read is guarded and `|| true`-terminated. This step is advisory, and under
`set -e`/`pipefail` a non-matching `grep` exits 1 — an assignment whose command
substitution fails would take the step down with it. A REVIEW.md that is missing,
a directory, or unreadable now leaves the counts empty and execution continues.

B — per-finding disposition, in a new lazily-read step file,
`gsd-core/workflows/execute-phase/steps/code-review-disposition.md`, referenced
from the gate in the established plain read-and-execute form. One row per finding
ID, defaulting to `open`, and:

- `fixed`/`skipped` are reconciled from REVIEW-FIX.md, whose section headings are
  matched WHOLE — a prefix match let `## Fixed Issues Verification` classify every
  finding beneath it as fixed — and only when the fix report names the SAME
  finding. Finding ids are reused across re-reviews, so matching on the id alone
  let a stale fix report declare a brand-new CR-01 already fixed.
- headings inside fenced blocks are ignored; a quoted example is not a finding.
- an id listed under both sections resolves by first occurrence, not row order.
- a recorded disposition is preserved together with the reason in its Source cell,
  escaped pipes included, and a hand-mangled row missing its trailing pipe still
  keeps its decision.
- a decided finding the current review no longer reports is CARRIED and marked;
  `--auto` rewrites REVIEW.md each iteration, so this is routine, and dropping the
  row would erase the record that it was seen. An untriaged `open` row for a
  vanished finding is not carried. A review reporting nothing still reconciles an
  existing ledger rather than freezing it.
- a run that changes no disposition rewrites nothing, so a re-executed phase does
  not produce a docs commit whose only delta is a timestamp.

The record is a sibling artifact, not a section inside REVIEW.md: `--auto`'s
re-review loop rewrites REVIEW.md every iteration, so a ledger kept inside it
would not survive the next pass, and REVIEW.md has a single writer that this step
is not.

B lives in an extracted step file because `execute-phase.md` was 91,493 bytes
against a 98,304 hard cap the size-budget test calls a red line, and because
`scanWiredKinds` caps a call site's dispatch-coverage region at 6000 characters —
an inline version pushed the `kind == "gate"` paragraph out of that window, which
silently drops `gate` from the covered set and fails
`gen-capability-registry --check` while pointing at the capability rather than at
the prose that displaced it. Extraction is what that size test's own message
prescribes, and it leaves the file at 93,854 bytes.

The tests execute the shipped script rather than modelling it. Three adversarial
review rounds each refuted "the mirror is faithful", and mutation testing agreed:
with a hand-written model, deleting the carried-row logic from the shipped file
turned nothing red. The suite now extracts the embedded script — undoing exactly
the four shell double-quote escapes — and runs it, so all ten mutations of its
behaviour are caught.

* chore(#3829): set changeset fragment pr to 3861

* fix(#3829): keep execute-phase.md under both size ceilings and propagate the launcher probe

The first push failed `full test (macos-latest, 24, shard 3/3)`. Two things it
caught that the CI-selected scope for this diff does not run, and that I
therefore did not run either:

1. `execute-phase.md` is governed by TWO ceilings, not one. The XL hard cap in
   `tests/workflow-size-budget.test.cjs` (98304) was satisfied at 95179, but the
   frozen ADR-857 pre-phase-6 ceiling in `tests/claude-orchestration.test.cjs`
   (93600) was not. The whole budget from base is 2107 bytes, which the inline
   reporting half alone did not fit. That half now lives in the extracted step
   file alongside the disposition half, and the parent carries only the paragraph
   that reads and executes it — 91529 bytes, 36 over base.

2. The step file calls `gsd_run`, so it owes the hermes runtime-home probe that
   `tests/runtime-launcher-parity.test.cjs` (E) requires of every workflow file
   that does. Propagated with `node scripts/sync-runtime-launcher.cjs`, the
   remedy that test names.

Verified with the FULL unit suite this time rather than the scoped selection —
14 shards, 0 failures — plus `npm run lint:ci`, and a re-run of the ten mutations
of the shipped disposition script, all still caught.

* fix(#3829): stop the disposition step instructing the agent to execute itself

Blocker 1 and Minor 7 of the round-1 review are one defect. The step file
carried a copy of execute-phase.md's pointer paragraph, so it named its own
path as something to "read and execute" — unbounded self-recursion at runtime
— and that copy is also the duplicated paragraph, sitting immediately above
the full instruction it duplicates.

Removing the copy resolves both. execute-phase.md remains the only surface
that points here, which is what it always intended.

Two structural tests guard it. Both are red against the pre-fix file: no
behavioural test could see either defect, because they execute the node
script through the process seam and so never read the prose that tells the
agent what to load.

* fix(#3829): re-derive the ledger paths in the block that uses them

Blocker 2. The disposition block reads REVIEW_FILE, DISPOSITION_FILE and
PADDED, all derived in the step's FIRST shell block. Each fenced block is
dispatched as its own shell, so all three are empty by the time the second
block runs: the ledger write lands on a bare `-REVIEW-DISPOSITION.md` path
and the review read finds nothing. The step then reports success having
produced no artifact — the feature's central acceptance criterion, silently
unmet, with no error to notice.

The tell was already in the file: the gsd_run shim preamble is re-emitted in
the second block for exactly this reason. These three paths belong beside it,
and now are.

The guard test asserts the general property rather than the instance — every
block derives what it reads, inheriting only the step's declared inputs
(PHASE_DIR, PHASE_NUMBER) — so a third block added later cannot reintroduce
it. Red against the pre-fix file.

* test(#3829): assert the counts mirror against the shipped shell, and execute its guards

Major 4, with Minor 6 and part of Minor 9.

The disposition builder stopped being a mirror three rounds ago, and the
reason given then was that a hand model of a shell-embedded script drifts
while the tests stay green. parseGateCounts kept its mirror anyway. That
argument does not stop applying at the boundary between the step's two shell
blocks, so the mirror now loses its authority: it is asserted against the
shipped awk and greps, run under `set -euo pipefail` in a real shell, across
every fixture it is exercised on.

Negative-controlled in both directions. Dropping `blocker:` from the mirror
alone fails the parity test; replacing the shipped awk with the leaky
`sed -n '/^---$/,/^---$/p'` range fails it on the unterminated-frontmatter
fixture. Divergence in either half is now red, which is what the finding asks
for. Skipped on win32, where there is no bash to compare against.

Minor 6: the zero-count edge is covered — `0` is numeric, so a zero-finding
review reports `0 findings — 0 critical, …` rather than falling back to the
countless form. A guard written against truthiness would have failed here
silently, and now cannot.

Minor 9, partially: running the block makes its advisory guards behavioural,
so the four `src.includes()` assertions that stood in for them are retired —
a missing and an unreadable REVIEW.md are now proven not to abort under
`set -e`, rather than asserted to contain a string. The remaining docs-parity
assertions are kept deliberately; see the PR discussion.

* test(#3829): add the render/re-parse fixed-point property for the ledger

Major 3. RULESET.TESTS.property-based-testing asks for at least one fc
property on a parsing/transformation contract, and the ledger is one with a
fixed point stated in its own prose: re-running the gate preserves every
disposition except `open`, and rewrites nothing when nothing changed.

Two properties, both driving the SHIPPED script rather than a model of it:

  idempotency — a second run reports `unchanged` and leaves the file
                byte-identical. Without it, the timestamp alone dirties the
                tree on every phase re-run.
  round-trip  — a hand-recorded decision AND the reason beside it survive
                render -> re-parse -> render, escaped pipes included. The
                Source cell is where a human writes why something was
                deferred, so losing it loses the only thing that instruction
                asks for.

Negative-controlled per property: disabling the unchanged-check fails the
first and only the first; discarding the carried source cell fails the second
and only the second.

numRuns is 40 rather than the shared 200 because each case spawns the shipped
script twice through the process seam. The seed stays pinned, so a failure
still reproduces; the deviation is stated in the file header rather than made
silently.

* fix(#3829): state a stale fix-report match instead of dropping it silently

Minor 5, plus the finding-id census this round owes.

Exact-title coupling stays — ids are reused across re-reviews, so a stale
REVIEW-FIX.md must not mark a brand-new CR-01 as already fixed. What changes
is the silence. A row that stays `open` because the report named a different
finding under the same id is indistinguishable, to any reader, from a row
that stays open because no report mentioned it. The gate now names the ids it
could not reconcile, on both report paths, and stays advisory throughout.

The census (RV4, self-found — the review did not ask for this). The script
enumerates finding-id prefixes in three places: the heading matcher, the
ledger re-parser, and the severity map's keys. The DOMAIN those enumerate is
owned elsewhere — gsd-code-reviewer.md's body template and its
Label-equivalence paragraph — so it can acquire a member without this script
changing.

Reached: CR, BL, WR, IN — 4 of 4, all present. Not reached: none today. What
follows if that changes is the payload: an unlisted prefix is not mis-tiered,
it is INVISIBLE — the finding never enters the order list and gets no row at
all, so the artifact silently under-reports the review it is meant to record.
Adding a prefix to two of the three copies fails the same way, and additionally
drops carried rows on the next run.

Two guards rather than a rewrite: hoisting the alternation into one constant
means rebuilding three regexes inside a double-quoted shell string, which is
the exact class of edit that produced both of this round's blockers. The
guards make the drift loud instead, and are negative-controlled against each
of the two ways it can happen.

* docs(#3829): keep the feature reference descriptive, not instructional

Minor 8 — a Diataxis mode mix. "Set `deferred` by hand and put the reason in
the Source cell" is a how-to instruction sitting in a reference doc. The
information belongs there (a reader needs to know the field exists and what
preserves it); the imperative does not.

Rewritten to describe the field instead: `deferred` is the one disposition the
gate never writes, and the reason recorded beside it survives re-runs. The
same pass records Minor 5's new behaviour, since the reference described the
title coupling but not what happens when it misses.

The imperative form is kept where it belongs — inside the ledger the gate
renders, which is where a reader meets the field and the only place an
instruction has an audience.

docs/FEATURES.md regenerated from it; `gen-features.cjs --check` is green.

* fix(#3829): close six defects found by reviewing this round's own fixes

None of these came from the maintainer's review. They came from adversarially
reviewing the five commits above before pushing them, and two are worse than
anything the round was opened to fix.

1. A foreign fence marker swapped an example for a finding. The heading scanner
   toggled fenced/not-fenced on ANY fence marker, so a ~~~ line inside a ```
   example closed the fence and the example's real close reopened one. Driven:
   a review quoting ~~~ inside a fenced example produced a ledger recording
   CR-77, the illustration, and omitting CR-01, the actual finding. A
   confidently-written artifact wrong in both directions at once. The open
   marker's character and length are now remembered, and a fence closes only on
   the same character at least as long, per CommonMark.

2. The disposition block had no status gate at all. The prose above it says it
   runs only when the review reports issues — but block 1 computes
   REVIEW_STATUS, emits nothing, and its shell is discarded, so no later block
   could act on that condition even in principle. A prose gate on a value
   nothing downstream can see is not a gate, and a clean re-review would rewrite
   a ledger it was never meant to touch. Re-derived in block 2's own shell.

3. A numeric breakdown could still be internally false. `total: 0` beside
   `critical: 1` is four valid numbers rendering `0 findings — 1 critical, …`.
   Numeric was necessary and not sufficient; an inconsistent breakdown is now
   withheld for the same reason a partial one is.

4. The carried-marker strip ate hand-written prose. It removed a trailing
   `(not in the current review)` unboundedly and unconditionally, so a deferral
   reason that merely ENDED in that phrase lost it — the one field a human
   writes into this artifact. Now bounded to one occurrence, and only on rows
   the marker can legitimately be on. The no-growth property it exists for is
   re-pinned.

5. parseGateCounts diverged from the shipped pipeline in two ways no fixture
   reached. The shipped reads are `cut -d: -f2 | tr -d ' '`: `tr` removes
   INTERNAL spaces (`1 0` -> `10`) where `.trim()` keeps them, and `cut` takes
   only the second colon-field where a tail capture keeps the rest. The mirror
   models the pipeline now, and both counterexamples are fixtures — a parity
   assertion that agrees only on well-formed input asserts very little.

6. The prefix census guards were both partly vacuous. The drift guard read the
   two regex alternations and not the severity map, so a set could agree in both
   regexes while mis-tiering in the map. The domain guard scanned only `### XX-01:`
   headings — and BL appears in no heading at all, only in the Label-equivalence
   prose, so the guard passed purely because BL happened to be hard-coded and
   would have missed the next prose-defined prefix exactly as it missed BL. Both
   widened; the domain the guard now sees is BL, CR, IN, WR.

Each fix fails a named test on reversion and none fires on the ordinary path.
The property generator now deliberately produces the reserved suffix from (4),
which a generator drawn only from innocuous characters could never reach.

Also corrected: the previous commit's account of the empty-path failure. The
script did not write a bare `-REVIEW-DISPOSITION.md`; it threw on reading the
empty review path and the trailing `|| echo` swallowed it as a non-blocking
skip. Same silent outcome, different mechanism, and the comment said the wrong
one.

* fix(#3829): the tests now run what bash runs — and six fixes to the fixes

A second adversarial pass over the previous commit. It found a regression that
commit introduced, and the reason it slipped through is the finding worth
keeping.

THE FIDELITY GAP. Every test here extracts the embedded script as TEXT and
runs it. Bash does not: it expands the double-quoted `node -e "..."` argument
first, so a backtick inside it is COMMAND SUBSTITUTION. The previous commit put
one in a code comment. Bash duly ran it, failed with `+: command not found`,
and handed Node a script two bytes shorter than the one 122 green tests were
exercising. No behavioural test could see this, because none of them ever
asked bash what it would actually pass. One now does, and it is the general
guard: it catches an unescaped backtick, an unescaped $, and any other
expansion the extractor cannot model.

Then, in the shipped step:

- A padded count silently disabled the sum check. `$((08 + …))` fails on base
  inference; it does not abort — the expansion sits in an `if` condition, where
  set -e does not fire — so the check simply never ran and an inconsistent
  breakdown passed with a stray diagnostic as its only trace. `10#` on every
  operand.
- The status guard made the script's own reconciliation unreachable. A clean
  review with an EXISTING ledger must still be reconciled — decided rows
  carried, stale `open` rows dropped — or the ledger freezes showing findings
  as open that the review no longer reports. The guard now skips only when
  there is nothing to reconcile.
- The carried marker is no longer stripped at parse time at all. Bounding the
  strip still ate a carried row's human-written reason. No-growth is a property
  of the RENDER, so it is enforced there: a marker already present is not
  appended again. Nothing is stripped, nothing doubles.
- Fence openers are bounded to three leading spaces, per CommonMark.
- parseGateCounts matched `[ \t]` where the shipped grep uses `[[:space:]]`,
  which covers form feed and vertical tab. Third counterexample of the same
  class, and a fixture.
- The census drift guard checked only one direction, so a tier for a prefix the
  regexes never admit stayed green as dead code that reads as coverage.

TWO OF MY OWN TESTS WERE VACUOUS, and the controls are what said so. The
leading-zero test asserted exit 0 and a consistent verdict — both true before
the fix. The clean-review test drove the node script directly, which never
executes the shell guard at all: it passed unchanged with the guard made
unconditional. Both are rewritten to test the layer the defect lives on, and
both now fail when their fix is reverted.

Every fix in this commit fails a named test on reversion, each mutation
verified to have applied before its verdict was read.

* fix(#3829): the carried marker can no longer outlive the carry

A third adversarial pass. Its most important finding is a defect the SECOND
pass talked me into, which is worth recording as plainly as the fix.

THE MARKER BECAME A LIE. Pass 2 objected that bounding the carried-marker strip
still altered a human-written reason, and proposed storing the cell verbatim
instead. That objection was a preference, not a defect — its own driven output
showed exactly one marker, which is correct — and adopting it created a real
one: once the generated marker is stored it can never leave, so a carried
finding that REAPPEARS in a later review still renders "not in the current
review". The ledger then contradicts its own contents. Driven both runs.

The strip is back, bounded to one occurrence and unconditional. The residual
ambiguity is irreducible — a reason ending in exactly that phrase is
indistinguishable from the marker — and it costs nothing real: on a carried row
the render puts the phrase straight back, and on a current row the phrase was
self-contradictory to begin with. The unbounded quantifier is what had to go,
not the strip. The property now states that contract rather than asserting a
verbatim survival the code deliberately does not provide.

Also:

- An ABSENT REVIEW.md abandoned the ledger it was meant to reconcile. The guard
  proceeds when a ledger exists, then the script read the review unconditionally,
  threw, and the trailing fallback swallowed it — the freeze the reconciliation
  path exists to prevent, reached through the door the guard opened.
- Counts are length-bounded as well as digit-only. Bash integers wrap at 2^64,
  so a 20-digit count arrived at the sum as 0 and an inconsistent breakdown
  passed.
- A closing fence must carry only whitespace after its marker; a line with an
  info string is an opener's shape and ended the fence early.
- parseGateCounts matched [ \t\n\v\f\r] where the shipped grep uses
  [[:space:]], which under this UTF-8 locale matches EM SPACE. `\s` is the
  faithful model. Fourth counterexample of that class, and a fixture.
- The agent-domain scan required [A-Z]{2,}, so a one-letter prefix like `C-01`
  — explicit and parseable, not prose — was invisible to it.

AND THE FIDELITY GUARD PAID FOR ITSELF INSIDE ONE SESSION: writing this round's
first draft I put backticks around a token in a code comment again, in the very
commit whose subject is that mistake. The probe failed, named it, and no test
of behaviour could have. Two of my own tests also had to be rewritten: one
asserted things true before its fix, and one drove the node script directly
where the defect lived in the shell.

383 pass across the touched files and the two size ceilings; ten lint gates
green; every fix fails a named test on reversion, each mutation verified to have
applied before its verdict was read.

* docs(#3829): the Source reason is preserved, but not verbatim — say so

Found by claim-auditing the response comment before posting it, which is the
one place this would have been caught: the doc and the code were written in
different commits and only a reader holding both notices they disagree.

The feature reference said the hand-written reason is "preserved verbatim
across re-runs". It is not, and deliberately so — a reason ending in the
literal phrase "(not in the current review)" loses that trailing phrase,
because it is indistinguishable from the carried marker the gate appends.

The exception is stated rather than dropped, with the reason it is the better
trade: storing the marker instead means it never leaves, and a carried finding
that later reappears goes on claiming it is absent from the very review that
reports it. A ledger wrong about its own contents beats losing a duplicated
phrase, but only if the doc admits which one it chose.

FEATURES.md regenerated; gen-features --check and lint:docs green.

* fix(#3829): the gate now emits the counts it computes (B1a/B1b)

Block 1 computed REVIEW_STATUS and the four counts and printed none of
them, then the prose below asked the agent to display four of them. The
shell exits at the closing fence and the agent sees only stdout, so those
values were unobtainable: REQ-REVIEW-08 was unreachable in every shipped
path and the fence was decorative.

The rule was already stated one block down -- "a prose-only gate on a
value no later block can see is not a gate" -- and applied only to block
2. It now governs the block that is this step's primary deliverable.

Both arms emit, and the status gate is mechanical rather than prose:
a clean/skipped/absent review prints nothing, an inconsistent or partial
breakdown prints the countless form, and the full breakdown prints
otherwise. Driven against the review's own case (critical: 1, warning: 9,
info: 8, total: 18) with no appended emitter:

  Code review: 18 findings - 1 critical, 9 warning, 8 info.
  Consider running: /gsd:code-review 1 --fix

* test(#3829): the counts harness stops manufacturing the output it asserts on (B2)

runShippedGateCounts extracted the shipped fence and then APPENDED its own
printf of the six internal variables before running it. Every counts
assertion was green against a script that existed only inside the test
process: the shipped fence emitted nothing, the tested fence emitted six
lines because the test added them. That is why B1a shipped past a suite
that looks like it covers exactly that surface -- the green was
structurally incapable of turning red for it.

The emitter now lives in the fence, so the harness reads the fence's own
stdout and synthesizes nothing. Parity with the mirror moved up a level
with it: renderGateMessage() renders both arms from the mirror's parsed
counts and the assertion compares the WHOLE emitted message, so a drift
in any parsed value changes the string or the arm it selects. Asserting
on the observable is strictly stronger than asserting on five
intermediates, and it can express what the old probe could not -- an
absent review now reports NOTHING, which is a different fact from
reporting a countless review.

A fifth src.includes() assertion converted with it (round 1 retired
four). It pinned the PROSE stating the countless condition, so it went
red when the emitter moved into the fence while the behaviour it named
was untouched -- the pin arguing for its own conversion.

Negative control: reverting the shipped echo now turns 16 tests red.
Before this commit the same reversion turned zero red, which is the
finding.

* fix(#3829): the disposition column is an enum, not any lowercase token (B3)

ADR-227 requires a trust boundary to validate semantic SHAPE and to coerce
a failure to the contract's safe default. The ledger is a trust boundary by
construction -- the rendered instruction tells a human to hand-edit it --
and the prior-row parser captured column 3 as ([a-z]+), checked against
nothing.

One transposed character was enough. `| CR-01 | critical | opne | - |` is
not the literal 'open', so it beat the default, was excluded from the
`open:` headline count, and was carried forward forever. The ledger then
reported the phase fully triaged off a typo.

The asymmetry is what made this a correctness bug rather than a style
point: a typo OUTSIDE [a-z] ('Deferred') already failed to match, lost the
decision and reset the row to open -- safe. A typo INSIDE [a-z] was unsafe.
The parser failed open in the one direction that matters. A row that fails
the enum now yields no prior entry and the row falls back to 'open', by the
same path the capital-D case already took.

The property test could not have caught this: DECIDED is drawn from the
vocabulary, so no property built on it can present an out-of-vocabulary
token. Added JUNK, the arbitrary for the complement, deliberately
lowercase so it stays inside the old capture's own character set -- the
unsafe half is the token that LOOKS like a decision and is not. The new
property also asserts the headline count agrees with the row it renders,
which is the half the defect actually reported wrongly.

Negative control: the new property fails against the ([a-z]+) capture and
passes against the enum.

* fix(#3829): a finding the heading parser cannot match is surfaced, not dropped (B4)

Two independent parsers produce two numbers one paragraph apart -- the
counts from REVIEW.md's frontmatter, the rows from `### <ID>:` heading
matches against a closed CR|BL|WR|IN alternation -- and nothing reconciled
them. A finding the alternation could not reach contributed no row, no note
and no diagnostic, and the ledger then declared `open: 3 of 3` over a set
strictly smaller than the console line had reported one paragraph earlier.
Two findings recorded nowhere, and neither artifact said so.

The PR's own argument for the closed alternation -- that an unlisted prefix
produces no row rather than a MIS-CLASSIFIED one -- is the wrong trade under
this repo's fail-safe rule. A dropped finding is demoted below every finding
that parsed, and an unparseable finding is precisely the one a human most
needs to see.

Block 2 now derives the frontmatter total (anchored inside the findings:
mapping, digit-and-length-bounded like block 1's) and hands it to the
script, which reconciles it against the CURRENT review's matched findings --
order.length, never rows.length, which also counts carried rows and would
either understate the shortfall or invent one. Surfaced exactly as the
stale fix-report case already is: a non-blocking `unparsed: N` key plus the
console line, both naming the two numbers so the claim is checkable.

  Code review disposition recorded: 3 of 3 finding(s) open (2 finding(s)
  recorded NOWHERE: the review reports 5, but only 3 matched the expected
  heading shape `### <CR|BL|WR|IN>-NN: <title>`)

The key is emitted only when there IS a shortfall, so an ordinary ledger
gains no noise key and the unchanged-run check is unaffected.

Four tests, including three negative controls the round owed itself: a
clean review gains no key, an absent/non-numeric total reconciles nothing
rather than fabricating a shortfall, and a total SMALLER than the row count
cannot render `unparsed: -1`. Reversion control: dropping the key turns the
first red.

* fix(#3829): pass --raw to the commit_docs config-get (#3763)

Not from the review -- from a gate the base range added after it. #3763
lands `tests/config-get-raw-guard.test.cjs`, and this branch was its sole
offender: a config-get command substitution without --raw feeds
JSON.stringify output into a bash string comparison, where it silently
never matches for string values. The consumer here is exactly that:

  if [ "$COMMIT_DOCS" = "true" ]

Every other shipped call site in the tree already passes --raw
(spike.md, fast.md, new-milestone.md, sketch-wrap-up.md, ...), so this is
sibling convention, not a new posture.

Worth recording because the two readings are both correct and they
disagree: round 2's review cleared this exact line under ADR-3409 as "the
safe member of that family", since `query config-get <key>` with no --pick
exits 1 on absence and the fallback arm is reachable. That is still true --
--raw does not change it. The base then moved and added a gate that reads
the same line for a different property.

* fix(#3829): scope the count reads to the findings: mapping, not just the frontmatter (m1)

`^[[:space:]]*total:` matches any indented key anywhere in the block, so a
top-level key later named `total:`, `info:` or `critical:` was picked up
ahead of the nested one. The block's own extensive comment is about scoping
the FRONTMATTER, and the scoping stopped one level short of the mapping the
values actually belong to. `status:` was never exposed -- it is anchored to
column 0 because it IS top-level.

The reads now run over the `findings:` block alone, selected by awk and cut
at the next column-0 key. Block 2's REVIEW_TOTAL derivation (added with B4)
already used that filter; this brings block 1 to it, so the two agree by
construction rather than by coincidence.

The mirror models the same scoping, and two fixtures drive it: a top-level
`total: 999` ahead of a nested `total: 1`, and top-level `critical:`/`info:`
ahead of theirs. Reversion control: unanchoring the shipped reads turns them
red.

* fix(#3829): severity comes from the section heading, not just the id prefix (M3)

gsd-code-reviewer.md emits findings under '## Critical Issues' /
'## Warnings' / '## Info', and that heading is the reviewer's own statement
of a finding's severity. The walker already visits every line -- the
fix-report path tracks '## ' sections -- so the signal was in hand and
discarded in favour of the id prefix alone.

A reviewer who mis-numbers a Critical as WR-04 while filing it under
'## Critical Issues' produced a row reading 'warning'. The ledger's Severity
column is the whole basis for triaging it, and it then disagreed both with
the review it summarizes and with the frontmatter count line block 1 prints
from findings.critical.

Section first, prefix as fallback: a finding under no recognized section --
a review that does not use the documented headings, and every row carried
from an earlier review -- keeps the prefix mapping, BL- included. Sections
are matched WHOLE, exactly as the fix-report sections are, so
'## Critical Issues Verification' does not re-tier what sits under it, and
a heading inside a fenced example does not govern.

Five tests: both mis-numbering directions, the prefix fallback across all
four prefixes, the lookalike heading, and the fenced-example case.
Reversion control: prefix-only turns the first two red.

Sixth src.includes() assertion converted with it -- it pinned the exact
source LINE of the enumeration loop, so it went red when that loop was
reformatted while the property it names was strictly widened. It now
asserts the property: every finding id, in order, once each.

* fix(#3829): an untriaged row is carried too, not silently deleted (M1)

The carry-forward kept a prior row only when its disposition was not
'open', so an untriaged row for a finding the current review no longer
reports was dropped entirely. Combined with the reconciliation gap that
left EVERY row open, a re-review deleted the whole ledger.

The re-review loop rewrites REVIEW.md on every iteration, so REVIEW.md does
not retain it either: run 1 records CR-01 open, the re-review renumbers it
to CR-02, run 2's ledger contains neither. That is #3829's complaint
verbatim -- "no trace of what happened to them" -- reproduced by the
artifact built to prevent it. The old justification, "nothing was decided
about it", is exactly the state #3829 says must leave a trace.

Every prior row is now carried, and the carried marker is what keeps it
honest: the row does not claim the finding is live, it records that it was
seen and never triaged. Two costs, stated rather than discovered: a
renumbered finding shows twice until the old row is triaged, and a carried
untriaged row persists until decided. Both are bounded by the phase's own
findings, both are legible from the marker, and both beat a silent delete.

Five tests updated -- they encoded the dropped-untriaged behaviour as the
contract -- plus one new test for the renumbering case M1 names. Reversion
control: restoring the guard turns six red.

Two self-inflicted defects caught while writing this, both by probes round
1 built:

  - Four unescaped backticks in a comment inside the double-quoted node -e
    argument, which bash ran as command substitution. The extractor-parity
    probe fired ("--auto: command not found"). Third time that trap has
    been sprung in this PR, third time the probe caught it.
  - The reworded ledger footer contained the literal carried-marker phrase,
    and the marker-accumulation assertion counts it across the whole file,
    so a doc line read as a second marker. The assertion was right.

* test(#3829): cover the count-length threshold at limit-1, limit and limit+1 (M2)

The guard is `?????????*` -- nine or more characters -- so the limit is
8 digits accepted, 9 rejected. The only cases were 'x', single digits and a
20-digit value, none of which pins the boundary. RULESET.TESTS
boundary-coverage is a hard rule here and it was unmet.

All three points asserted, with the sum kept consistent at each so the
LENGTH rule is what decides the verdict rather than the sum check
incidentally agreeing.

Reversion control is the off-by-one M2 names: dropping one `?` moves the
limit to 7 digits, which no test could previously notice, and now turns
this one red.

* feat(#3829): wire the disposition ledger into the fix path (B1c/B1d)

REQ-REVIEW-09 was unreachable in every shipped path. execute-phase.md's
code_review_gate invokes review with neither --fix nor --auto, so
<NN>-REVIEW-FIX.md cannot exist when the gate runs and every row it writes
is `open` by construction. The operator then runs /gsd:code-review N --fix
by hand -- the very suggestion the step prints -- which writes REVIEW-FIX.md
and never touched the ledger. A phase with 23 findings, all fixed, ended at
`open: 23 / total: 23`: the artifact that exists to distinguish a triaged
finding from a forgotten one asserted that 23 triaged findings were
forgotten. Worse than recording nothing, because it looks authoritative and
is inverted.

Taking remedy (i), not (ii). Narrowing the docs to say the ledger reflects
the previous phase execution is a legitimate choice, but it ships a feature
whose central artifact is inert and then documents the inertness.

ONE ADAPTATION, because the prescribed site does not exist. The review says
to wire code-review.md's --fix/--auto path. code-review.md is not the writer
(gsd-code-fixer writes the report, code-review-fix.md commits it), and more
decisively it has no point that is AFTER the report exists: it delegates
through code-review/steps/dispatch-fix.md, which calls
Workflow(code-review-fix.md) and then exits the workflow. There is nothing
downstream of that call to wire to.

The site is code-review-fix.md, immediately after commit_fix_report. That
is where the report is on disk and committed, it is the canonical
implementation for all fix logic by dispatch-fix.md's own statement, and it
additionally covers a direct invocation of that workflow -- which a wiring
in code-review.md would have missed.

The same step, not a second copy: it consumes PHASE_DIR and PHASE_NUMBER,
both already parsed from the init JSON, and it is idempotent, so a phase
that reaches the gate and then a fix run ends with one ledger reflecting
both rather than two competing ones.

Driven end to end: the gate writes `open: 2 of 2`, the fix path reconciles
to fixed/skipped and `open: 0`. Two tests -- one pins the wiring and its
ordering relative to commit_fix_report and present_results, one drives the
two call sites in sequence. Reversion control: removing the step turns the
first red; the second covers the reconciliation the wiring makes reachable
rather than the wiring itself.

* fix(#3829): a reflowed fix-report title is the same title (m2)

The stale-fix-report guard compared titles with trim() equality. The strict
instinct is right -- ids are reused across re-reviews, so a stale
REVIEW-FIX.md must not mark a brand-new CR-01 as already fixed -- but
gsd-code-fixer.md writes '### {finding_id}: {title}' under no contract that
the title is copied byte-for-byte from REVIEW.md. A fixer that reflows a
long title produced a spurious mismatch note, left a genuinely-fixed row
'open', and told the reader the report named a different finding. That
false-positive mode was acknowledged nowhere.

Whitespace is normalized, and only whitespace: a wrapped title is the same
title, and it is the one divergence that carries no information. Case
changes and truncation stay strict on purpose -- they are the shapes a
genuinely DIFFERENT finding takes, and widening to them would trade a
visible false positive for the silent false negative the strict match
exists to prevent. The residual is now stated in the step rather than left
to be rediscovered.

The note's wording changed with it. It asserted the report "names a
different finding"; both causes reach that branch and the step cannot tell
them apart, so it now reports the observation -- "titles its finding
differently from the review ... a stale report, or a re-titled one" --
rather than a conclusion it has not earned.

Three tests: the reflow case reconciles cleanly, the re-cased case still
reports, and the stale case still reports with the new wording. Reversion
control: restoring the strict comparison turns the reflow test red.

Seventh src.includes() converted -- it pinned the comparison EXPRESSION, so
it went red when the comparison gained normalization while the property it
names was unchanged.

* docs(#3829): describe the flow that ships, not the one implied (m3)

Both reference pages said "/gsd-code-review <N> --fix records fixed and
skipped, which the gate reconciles from REVIEW-FIX.md" -- true in the
abstract, materially misleading in practice, because no shipped path
performed that reconciliation. With B1c/B1d wired it is now real, and the
pages say WHERE it happens rather than leaving a reader to assume the
in-phase gate does it: the gate runs before any fix report exists and
writes all-open, and --fix is what records what happened.

The round's other behaviour changes land here too, since a reference page
that lags the artifact is worse than none:

  - the disposition column is a closed vocabulary, and a value outside it
    falls back to open rather than being treated as a decision
  - severity comes from the section heading when the review uses one, and
    from the ID prefix otherwise
  - an unparsed shortfall is stated rather than dropped
  - titles are compared ignoring whitespace, so a reflowed title still
    reconciles, and a mismatch is reported as an observation rather than as
    a claim that the report is stale
  - EVERY row is carried now, triaged or not, with the cost of the
    renumbered-finding double-entry stated rather than left to be found

docs/FEATURES.md regenerated from the fragment; lint:generated-sync and
lint:docs both exit 0.

* chore(#3829): migrate the emitted-drift ack from a fragment to commit trailers

ADR-3942 landed on next in #3954: the acknowledgment is a git commit
trailer now, and tests/emitted-drift-acks/ no longer exists.

Worth noting for anyone reading the rebase: this did NOT surface as the
modify/delete conflict the migration guidance predicts. This branch ADDED
its fragment rather than modifying an existing one, and the base deleted
only the files that were already there, so the replay was clean and the
fragment survived silently into a directory that no longer exists. Quieter
than a conflict, and worse -- the gate is what catches it, not git.

Two Growth keys rather than the fragment's one: round 2 wired the ledger
into code-review-fix.md, so that file grew too. Both key on the bare
filename, per the Growth namespace.

Emitted-Drift-Ack-Growth: code-review-fix.md — #3829 review round 2, blocker 1c/1d: REQ-REVIEW-09 was unreachable in every shipped path because the in-phase gate runs before any REVIEW-FIX.md exists, so every ledger row it wrote was open and nothing ever reconciled them. This file gains one step, record_disposition, that reads and executes the same lazily-read step after commit_fix_report. It is the only point in the fix flow that is after the report is on disk: code-review.md delegates here through steps/dispatch-fix.md and exits, so it has no such point at all. Growth is one step of prose, no logic is duplicated, and the step is idempotent so the two call sites converge on one ledger.

* chore(#3829): the changeset describes the round's behaviour, not round 1's

It renders into CHANGELOG, so it carries the same misleading implication
minor 3 was about: "the gate ... reconciling fixed/skipped from
REVIEW-FIX.md" reads as though the in-phase gate does it, when the gate
runs before any fix report exists. Says where it happens, and picks up the
round's other user-visible changes -- carried untriaged rows, section-based
severity, the disposition vocabulary, and the unparsed shortfall.

* fix(#3829): a dotted phase number no longer aborts the step

Found by this round's own adversarial review, in its MISSED section: no
finding asked about it, and it is the most serious thing in the round after
the two blockers.

Both callers explicitly accept a dotted phase -- code-review.md:60 and
code-review-fix.md:36 both validate ^[0-9]+(\.[0-9]+)?$ and name "03.1" in
their own error text -- and both fences reconstructed the path with
`printf "%02d" "${PHASE_NUMBER}"`, which cannot format one. Driven with
PHASE_NUMBER=3.1: bash prints `invalid number` and exits 1, and under
`set -euo pipefail` that aborts the step on its FIRST line. An advisory
gate that promises never to block took the phase's entire review report
down with it, and the newly wired fix-path call site inherited the same
defect.

Pad the integer part and carry the sub-number verbatim, so 3.1 -> 03.1 and
3 -> 03, with both arms falling back to the raw value rather than aborting.
Driven: 3.1 now reads 03.1-REVIEW.md and writes
03.1-REVIEW-DISPOSITION.md; the integer path is unchanged.

Two other findings from the same review, both about claims rather than code:

MINOR 2's TEST WAS MIS-NAMED, and the reviewer was right to refute the
claim. It called itself the "reflowed" case while substituting triple
spaces, which is not a reflow. Driven: a genuinely WRAPPED heading is still
not reconciled, because a `###` heading is one line by definition and the
continuation is a separate paragraph. Not widened -- absorbing whatever
follows a heading into the title would swallow arbitrary prose and make the
stale-report check meaningless, and the kept failure mode is the safe one
(a visible mismatch note, never a wrong "fixed"). The test is renamed to
what it covers and the bound is now pinned by its own test.

THE SHELL-SHARING GUARD DID NOT GUARD. Negative-controlling it -- rather
than reading it -- showed that deleting block 2's real REVIEW_FILE
derivation left it GREEN, on the exact defect it was written for. Block 2
prefixes its `node -e` with `REVIEW_FILE="${REVIEW_FILE}" ...` to put the
values in the child's environment, and the detector counted that
self-referential pass-through as a derivation. Pass-throughs are now
excluded, and the control fires. Pre-existing, not introduced here: the
original column-0 anchor matched that same line.

Also worth recording: my first attempt at that control silently patched
nothing and reported clean. Same lesson this PR already learned once.

* fix(#3829): validate the phase number before formatting it, and make the shell guard executable

Three findings from the round review's continuation pass, all confirmed by
driving them.

1. MY OWN DOTTED-PHASE FIX WAS WRONG on the fallback path. `printf "%02d"
   abc` writes `00` to stdout BEFORE it fails, so
   `$(printf ... || printf %s ...)` CONCATENATES the two: `abc` became
   `00abc`, empty became `00`, and a legitimate `08.1` became `0008.1`
   because bash reads the leading zero as octal. An unset PHASE_NUMBER also
   aborted under `set -u` -- in the step that promises never to abort.

   Validate, then format: never format and fall back on failure. Driven
   across every edge the review named -- 3.1 -> 03.1, 3 -> 03, 08.1 -> 08.1,
   09 -> 09, 1.2.3 -> 01.2.3, and abc / empty / -1 / unset carried verbatim
   with exit 0.

2. THE SHELL-SHARING GUARD STILL DID NOT GUARD. Excluding pass-throughs was
   not enough: a structural predicate recognises assignment TOKENS, never
   assignments that derive a usable value, so `REVIEW_FILE=`,
   `REVIEW_FILE=$REVIEW_FILE` and a commented-out assignment all evaded it.
   No regex closes that class.

   The authority moves to execution -- the third time this PR has learned
   that lesson. The real second fence now runs in a fresh shell with nothing
   but the step's two declared inputs and must write the ledger at the
   correct derived path. All four mutations are caught: empty assignment,
   self-reference, commented-out, and deletion. The textual check stays as a
   cheap fast-fail and is labelled as one.

3. THE TITLE-BOUND CORRECTION HAD NOT REACHED THE DOCS. The step comment and
   both docs pages still said a reflowed title reconciles, contradicting the
   bound pinned one commit earlier. Superseded prose left standing reads as
   current to anyone arriving cold, so all three surfaces are rewritten
   rather than annotated, and FEATURES.md regenerated.

Also hoisted `HAS_BASH` to the file's other top-level constants. `const` is
in the temporal dead zone until its declaration runs, and a
`{ skip: !HAS_BASH }` option object is evaluated eagerly, so a bash-gated
test added above the old mid-file declaration threw a ReferenceError that
aborted its whole describe and CANCELLED its siblings -- while the summary
line still read `fail 0`. It caught three separate additions in this round
before I stopped moving tests and moved the constant.

* fix(#3829): refuse an out-of-shape phase number instead of carrying it into a path

Self-found while writing the prompt for the next review pass, which is the
honest provenance: I asked the reviewer whether a path traversal was
reachable through PHASE_NUMBER, then checked before dispatching.

It was, and I had introduced it. The previous commit's fallback carried an
unusable phase number VERBATIM, and PHASE_NUMBER is interpolated into a file
path:

  PHASE_NUMBER='../../etc/passwd'
  -> REVIEW_FILE=/tmp/phase/../../etc/passwd-REVIEW.md

The `printf "%02d"` it replaced had at least mangled that to `00`. A fix
that makes a path more reachable than the bug it replaced is a regression,
whatever it does for the case it was written for.

Both callers already validate ^[0-9]+(\.[0-9]+)?$ (code-review.md:60,
code-review-fix.md:36), so this is defense in depth rather than a live
exploit -- but the step has two call sites now and should not take either
caller's word for its own inputs. It validates the WHOLE value and, on
failure, builds no path at all: PADDED is empty and each fence refuses by
name rather than coercing. Block 1 declines to report counts read from a
path made out of the bad value; block 2 declines to write, which also keeps
it clear of the bare-name ledger defect round 1 closed.

Driven across the shape boundary: 3.1 / 3 / 08.1 / 09 accepted; abc, empty,
unset, 1.2.3, -1, 3., .1, +1, "3 1" and ../../etc/passwd all refused with
exit 0 and a named diagnostic. Reversion control: restoring carry-verbatim
turns the traversal test red.

* fix(#3829): bound the phase number's length, and make the shell guard prove derivation

Third adversarial pass. Two of its three refutations were already closed by
the previous commit (the ../escape and 1/../../escape traversals, and the
unset-input abort); these two were not.

1. A 54-DIGIT PHASE NUMBER WRAPPED SILENTLY. The validator accepted any
   all-digit value, so `$((10#$_int))` overflowed 64 bits and PADDED became
   `-7908320945662590977`. Length-bounded now at 8 digits, exactly as the
   counts already are and for the identical reason -- and the counts guard
   sitting twenty lines away is why this one is embarrassing rather than
   subtle. Driven at the boundary: 8 digits accepted, 9 rejected.

   The bare `${PHASE_NUMBER}` in the suggestion line is hardened to
   `${PHASE_NUMBER:-}` while here. The empty-PADDED guard makes it
   unreachable today, but it is one refactor away from an unbound-variable
   abort under `set -u`, in the step that promises not to abort.

2. THE EXECUTED SHELL GUARD PROVED THE FENCE WORKS, NOT THAT IT DERIVES.
   A single-phase probe is satisfied by a hardcode, and the review
   demonstrated exactly that: replacing the derivation with
   `case ... in 1) PADDED=01 ;; 7) PADDED=07 ;; *) PADDED=07 ;; esac`
   breaks every real phase and passed the entire suite. It now runs two
   distinct phases, 7 and 3.1 -- a hardcode cannot satisfy both, and the
   dotted one additionally pins the integer-part split.

   The claim "given only the declared inputs" was also overstated: the test
   spreads `...process.env` (it needs PATH and HOME). The DERIVED names are
   now explicitly deleted from that environment, so the claim is true rather
   than merely intended.

Also rewrote a comment that had become false: it pinned a describe to the
end of the file because of the HAS_BASH temporal-dead-zone constraint, which
the hoist removed. Superseded prose left standing reads as current to
anyone arriving cold.

The changeset's "the gate stays advisory and never blocks" is now verified
rather than asserted: both fences exit 0 under an unset PHASE_NUMBER and a
traversal-shaped one.

* fix(#3829): validate both inputs, refuse before building a path, and never write through a symlink

Fourth adversarial pass. Four findings, all confirmed by driving them.

1. PHASE_DIR WAS NOT VALIDATED AT ALL. Unset, both fences died with
   `PHASE_DIR: unbound variable` under `set -u` -- the same class as
   PHASE_NUMBER, which I had just spent two commits fixing while its sibling
   input sat one line away. The step declares two inputs; it now validates
   two.

2. THE LENGTH BOUND WAS ON THE WRONG THING. The nine-character glob applied
   to the WHOLE value rather than the integer part, so it falsely rejected
   `12345678.1` (a legal 8-digit phase) while accepting `1.123456`. Each
   component is bounded on its own now; the sub-number is bounded too, since
   it is likewise interpolated into a filename.

3. REJECTED VALUES STILL HAD PATHS BUILT FROM THEM. The refusal guard sat
   AFTER the assignments, so an unusable input still assembled
   `${PHASE_DIR}/-REVIEW.md` and stat'ed it before refusing. The guard is
   now the first thing after validation, and both fences construct paths
   from validated locals rather than from the raw environment.

4. THE LEDGER WRITE FOLLOWED SYMLINKS. From the review's MISSED section, and
   the sharpest thing in it: `fs.writeFileSync` follows a symlink, so a
   pre-existing symlink at the ledger path replaced the contents of whatever
   it pointed at -- outside the phase directory, with the link left intact
   so nothing looked wrong. Driven, and the target's contents were gone.
   This PR introduces the artifact, so it owns the check: an existing ledger
   that is not a regular file is not a ledger, and the advisory gate says so
   and steps over.

The executed shell guard now draws its phases AT RUN TIME. Fixed fixtures
cannot establish derivation -- the review defeated the one-phase version
with a hardcode, then defeated the two-phase version by adding one more arm
to the same case. Any finite sample loses that race. A phase picked per run
cannot be enumerated in advance; the drawn values print in every assertion
message so a failure stays reproducible. Control: the review's three-value
hardcode now fails on three consecutive runs.

Eighth src.includes() converted -- it pinned the literal `${PHASE_DIR}`
interpolation and went red when construction moved to a validated local,
while "writes a REVIEW-DISPOSITION sibling" was untouched. It now asserts
that property, and that REVIEW.md is not written.

The changeset's "stays advisory and never blocks" is verified rather than
asserted: 8 of 8 hostile-input cases across both fences exit 0 -- both
inputs unset, PHASE_DIR unset, a traversal-shaped phase, and a missing
phase directory.

* fix(#3829): check the ledger path before reading it, and pin the write-safety behaviour

Fifth adversarial pass, and the last one this round. Three fixes, three
disclosed residuals.

FIXED

1. A FIFO AT THE LEDGER PATH BLOCKED FOREVER. readFileSync on a FIFO never
   returns, so the step documented as "advisory, never blocks" blocked
   indefinitely -- the literal counterexample to its own headline claim. The
   non-regular-file check ran after that read.

2. THE UNCHANGED-RUN FAST PATH BYPASSED THE CHECK. A symlink whose target
   already matched the rendered ledger read through the link, reported
   `unchanged`, and never reached the refusal.

   Both fixed by the same move: the check is now the FIRST thing the script
   does, before any read or write of that path. Ordering was the defect, not
   the predicate.

3. THE COMMIT TEST FOLLOWED THE LINK the script had just refused. `[ -f ]`
   resolves symlinks, so the guard and its consumer disagreed about the same
   path and the helper could still be handed one. `[ ! -L ]` added.

   And the behaviour shipped with NO regression control -- I hand-drove it
   last commit and did not pin it, which the review caught by grepping for
   the words. Five tests now: symlink, symlink-with-matching-target, FIFO,
   directory, and an ordinary ledger as the negative control so the refusal
   is not a blanket one. mkfifo goes through the process seam like every
   other spawn here.

DISCLOSED, NOT FIXED -- these are stated in the step rather than carried
silently:

- TOCTOU between the lstat and the write. Node exposes no portable
  O_NOFOLLOW write, and an attacker who can write into the phase directory
  mid-run already has what the check would protect. It narrows a real
  accident; it is not a security boundary and the docs claim none.
- A hard link passes isFile() by construction.
- The REVIEW.md and REVIEW-FIX.md reads still resolve symlinks. They are
  reads of files the operator owns, in their own phase directory.

Also narrowed a comment that overclaimed. The randomized guard's domain is
FINITE -- 88 integer and 792 dotted values -- so a mutation enumerating all
880 passes forever, and Math.random() is unseeded, so "reproducible" means
only that the drawn values are printed on failure. Raising the bar is what
it buys; proving derivation is not, and nothing short of reading the fence
is. The previous comment claimed otherwise and was refuted.

* test(#3829): make the write-safety controls portable to the Windows lane

CI caught what neither the local suite nor five adversarial review passes
could: every one of those ran on Linux.

The FIFO test gated on `mkfifo`'s exit code. On the Windows lane mkfifo
EXISTS and exits 0 while producing something that is not a FIFO, so the
guard passed, the test ran against an ordinary path, the ledger wrote
normally, and the assertion failed for a reason unrelated to the behaviour
under test. It now gates on `lstatSync().isFIFO()` -- what was actually
created, not what the command claimed. Control: with the shipped guard
disabled the test still goes red on Linux, where the FIFO is real.

The two symlink tests are skipped on win32, following this repo's existing
convention for symlink-planting tests (tests/settings-jsonc.test.cjs:389
skips the same class; tests/unreachable-guard-drift.test.cjs:726 records the
reason -- symlink creation requires elevated privileges on Windows CI). The
privilege happened to be available on the lane this round, which is exactly
why the convention is not "try it and see".

* fix(#3829): a bare `|` in a deferral reason is prose, not a parse failure

Review round 3, the one blocker. The Source cell is the one field this ledger asks a
human to hand-edit, and "waiting on team A | team B to align" is an ordinary thing to
type there. The prior-row capture admitted a pipe only when escaped, so a bare one
failed the WHOLE line: prior.get() was undefined, the row fell through to `open` with
an empty Source, and the console line read "1 of 1 finding(s) open" — a Critical a
human explicitly deferred, with a documented reason, rendered indistinguishable from
one never triaged, and the reason gone. The exact ambiguity #3829 exists to remove,
reachable by one missing backslash.

The Source cell is the LAST column, so it is now captured through to the end of the
line, less an optional trailing pipe; a bare `|` inside it is prose. The render
escapes a bare pipe on the next write so the table stays a table, and the escaped
form re-parses to itself, so the second run reports `unchanged` — the fixed point
holds. The ledger's own instruction line says so instead of asking the human to
escape.

Why the property never caught it: SOURCE_CELL only ever appended a PRE-ESCAPED pipe,
so the arbitrary built to stress this cell could not reach the one input that broke
it. It now also emits a bare pipe, and the round-trip expectation is the escaped
form of what the human wrote. A fixed regression case drives the reviewer's exact
input through two runs and asserts the decision, the reason, the headline count and
convergence. Negative-controlled: both new tests fail against the previous capture.

The src.includes() pin on the old capture text is retired for the behavioural case —
it was pinning the defect.

* fix(#3829): escape every bare pipe in one write, whatever precedes it

Round 3, found by the adversarial pass over the round's own fix rather than by the
review. The first escape used /(^|[^\\])\|/g, which CONSUMES the character before the
pipe: adjacent bare pipes were escaped one per run (A||B -> A\||B -> A\|\|B, a third
run to converge, breaking the advertised second-run fixed point), and an escaped
backslash before a pipe (A\\|B) hid the pipe behind the wrong parity and left it bare
in the rendered table. The property generator emits at most one bare pipe, which is
the one case the old form got right, so no property reached either.

Scan as pairs instead: an escaped pair (backslash + anything) is kept verbatim and only
a pipe outside one is escaped. One write, then a fixed point. Regression case drives
`A||B and C\\|D` through two runs; it fails against the previous escape.

* fix(#3829): the script leaves by return, so an explicit exit cannot drop its verdict line

Round 3, from the adversarial pass over the round's own fix. The embedded node script
printed its verdict and then called process.exit(0) -- on the 'unchanged' branch only;
the 'recorded' branch fell off the end. Node's "A note on process I/O" documents
process.stdout writes to pipes and sockets as asynchronous on POSIX, and process.exit()
as forcing exit before pending asynchronous stdout writes complete -- so on a POSIX
lane the caller can see exit 0 with no verdict line. This is a hardening against that
documented hazard, not a reproduced defect: the reviewer's empty-second-run stdout,
which first pointed here, turned out to be its own sandbox -- a bare console.log child
printed nothing there either -- and that attribution is withdrawn.

The script now runs inside main() and leaves by return on all four early-exit paths,
so the event loop drains stdout before the process ends. Same exit status either way,
and the || echo fallback is unaffected. A structural test pins the absence of the call
(comment-stripped; dotted, bracketed and whitespace-split spellings). The empty-review
docs-parity pin that asserted the literal process.exit(0) line is retired -- the
round-2 describe drives that property behaviourally.

Also widens the property generator: SOURCE_CELL now reaches adjacent pipes and a
backslash of either parity before a pipe, the two shapes the first render escape got
wrong while passing every input the generator could then produce -- checked against
an independent parity-walk oracle rather than a copy of the render's own scan.

* fix(#3829): record what an --auto iteration fixed, instead of reporting it open

Round 5's major. `record_disposition` runs once, after the whole capped-at-3
`--auto` loop converges — but this workflow keeps ONE final version of REVIEW.md
and REVIEW-FIX.md rather than per-iteration copies, and deletes the .iterN.md
backups on convergence. A finding fixed in iteration 1 was therefore absent from
the final review (it was fixed, so the re-review stopped reporting it) AND from
the final fix report (overwritten by the last iteration), so the row fell back to
the gate's `open` and rendered `open ... (not in the current review)` — the same
bytes a finding that vanished for an unrelated reason produces. That is the one
distinction #3829 exists to make, undone by the artifact built to make it.

The precise site was the two-arm `applied` construction: for an id the current
review does not report, `sameTitle(undefined, h.title)` is false and
`title.has(id)` is false too, so the entry entered NEITHER `applied` NOR
`staleFix`. It was dropped in silence.

Four changes, one defect:

- A third arm. When the review does not report an id at all there is no title to
  disagree with, so this is not the stale-report case — it is what a finding
  looks like once it has been acted on. Record it. The id-reuse hazard stays
  closed by the arm below it: when the review DOES report the id, a title
  mismatch still goes to `staleFix` and is never applied, so a renumbered
  finding cannot inherit an earlier iteration's `fixed`.
- Rows for decided ids the review no longer reports, carried and marked. A
  decision the ledger cannot render is a decision lost — the same silent drop
  the carry-forward loop already refuses for prior rows, one source over.
- The .iterN.md fix-report backups are read alongside the final report, newest
  first, so the most recent statement about an id wins — the precedence a
  duplicate id already gets within one report.
- The shell guard proceeds on a fix report, not only on an existing ledger. A
  direct `/gsd-code-review N --auto` writes no gate ledger, and a converged loop
  leaves `status: clean`, so a fully successful multi-iteration run recorded
  nothing at all.

And the backups now go in `cleanup_iteration_backups`, after the ledger has read
them. #3190's rule is untouched — spent scratch on convergence, retained on
degradation — only the timing moved; deleting them inside the loop erased every
early fix before anything read it. `CONVERGED` does not survive the loop's shell
and is re-derived from the final review's status, which is exactly how the loop
sets it; anything but a proven-clean review retains.

Seven new regression tests plus an ordering test, all eight reversion-controlled
against pre-fix code — every one fires. One is the negative control that matters:
a reused id whose title differs must stay `open`, never inherit `fixed`.

Residual, stated: an id appearing only in an iteration fix report takes its
severity from the id prefix rather than a section heading, because `sectionSev`
is built from the current review. That is the documented fallback for carried
rows, not a new gap.

* test(#3829): pin the two PADDED derivations against a silent desync

Round 5's minor 1. Each fenced block runs in a fresh shell and must derive what
it reads, so the PADDED derivation — the traversal fence between an
attacker-influenceable phase number and a file path, plus the per-component
length bound — is duplicated verbatim. Both copies were independently tested and
nothing asserted they stay in step, which is the shared-parallel-surface shape
CLAUDE.md requires a parity test for, on security-relevant validation logic
rather than incidental repetition.

Compared line by line rather than through a normalizing rewrite: a normalizer
has to be told what may differ, and whatever it is told to tolerate stops being
asserted. Exactly one line may differ — each block refuses by its own name — and
the test names both forms. It also asserts the slice is substantial, since a
parity test over an empty slice passes vacuously.

Control: dropping one `?` from block 2's length bound, which moves that copy's
limit to 7 digits while block 1 keeps 8, turns it red. That is the exact silent
divergence the finding describes.

One correction to the finding's own statement, since it is worth recording: the
cited lines are :324 and ~:480, which are node-script lines; the derivations are
at :48-83 and :211-246. And they are 35-of-36 identical rather than
byte-identical — the refusal message differs, deliberately.

* docs(#3829): state the PHASE_DIR trust boundary instead of carrying it

Round 5's minor 2 asked that the assumption behind PHASE_DIR's validation be
confirmed rather than silently carried forward at the two new call sites. It is
confirmed, and the comment that stood here was wrong about it: "PHASE_DIR is the
step's other declared input and gets the same treatment" describes something the
code does not do.

Both inputs have the SAME provenance — each caller binds them from
`gsd_run query init.phase-op` (code-review-fix.md:7,17; execute-phase.md the
same) — so neither is raw user input and neither is more trusted. The asymmetry
is not about trust. It is that only one of them has a shape: PHASE_NUMBER
carries a documented contract, `^[0-9]+(\.[0-9]+)?$`, asserted by both callers,
so a value outside it is provably wrong and is refused. PHASE_DIR's contract is
"a filesystem path", which admits `..`, absolute and relative forms and
symlinked parents alike; no predicate separates a legitimate planning directory
from an illegitimate one, so a shape check would reject working setups while
proving nothing.

So the emptiness check is adopted as what it actually is — the guard against
`PHASE_DIR: unbound variable` aborting a step that promises never to block — and
the shape check is declined, with the reason written where the next reader meets
it rather than left to be re-derived.

The residual is restated in place rather than left in a PR comment: PHASE_DIR
may itself be a symlink and the ledger is then written through it, outside the
phase directory, deterministically. Left alone deliberately — the write goes
where the caller pointed. Not a security boundary, and nothing here claims one.

* docs(#3829): record why HAS_BASH is a platform assumption, not a probe

Round 5's minor 3 is DECLINED, and the reason is the repo's own contract rather
than a judgement call — written at the constant so the next reader does not
"fix" it and re-enable what the rule exists to prevent.

The gap is real and confirmed: 22 tests carry `{ skip: !HAS_BASH }`, so block
1's bash severity-reporting path has no Windows-lane coverage. But
`local/no-unguarded-nonportable-exec`
(eslint-rules/no-unguarded-nonportable-exec.cjs, DEFECT.WINDOWS-TEST-PORTABILITY)
REQUIRES this guard around `sh -c` / `bash -c` in tests, and its own remedy text
names `if (process.platform !== 'win32')` as the sanctioned form, because these
constructs fail under Windows Git Bash. So the constant is the repo's answer to
this question, not an oversight in this PR.

Swapping it for a runtime `bash` probe would light 22 tests up on a lane the
rule has already determined they cannot pass — trading a legible, rule-encoded
skip for a red matrix. Reversing that is the rule's decision; a change here
belongs with a change there.

* docs(#3829): describe how --auto's iterations reach the disposition ledger

The reconciliation section described the `--fix` path accurately and said
nothing about `--auto`, which is where round 5's major lived. It now states
that the loop overwrites its fix report each pass, that the re-review drops a
finding once it is fixed, that the gate therefore reads the per-iteration
backups newest-first, and that the backups are removed after the ledger has read
them rather than before. It also states the converged-with-no-ledger case: a fix
report on disk is reason enough to record.

FEATURES.md regenerated (176 features / 21 groups). Changeset extended to name
the shipped behaviour rather than only the `--fix` half.

* fix(#3829): clear lint-workflow-shellcheck, a gate the base range added

Not from the review. The rebase onto `next` brought in `lint-workflow-shellcheck`
(#4109), whose baseline was generated before this PR's new step file existed — so
that file's findings are new by construction and `lint:ci` exited 1 on the
rebased head before this round touched anything. The last green CI run predates
the gate. Caught locally rather than by a red push.

Three fixes and one baseline entry, split by whether the finding is real:

- STRUCTURAL (not ShellCheck, not baselineable): the guard's
  `for _f in "…${PADDED}-REVIEW-FIX.iter"*.md` is the bare `for x in $VAR` shape
  that word-splits differently under bash and zsh. Wrapped in
  `$(printf '%s' "$PADDED")`, the linter's own prescribed remedy.

- SC2097/SC2098, and this one was a genuine latent bug rather than a lint nit:
  `FIX_REPORT_FILE="${_pd}/${PADDED}-REVIEW-FIX.md"` sat in the same env-prefix
  list that sets `PADDED`, so its `${PADDED}` expanded the OUTER variable, not
  the one two entries earlier. Both happen to hold the same value here, which is
  exactly why it would have kept being wrong quietly. Built before the command
  now.

- SC2317 ×3 is baselined, not fixed. It fires on
  `return 0 2>/dev/null || exit 0` — the deliberate idiom that lets a fence
  refuse whether it is sourced or executed — and the verdict is a false
  positive: the `exit 0` is reached precisely in the executed case. Rewriting a
  dual-mode refusal to satisfy a wrong unreachability claim trades a real
  behaviour for a clean report. Baseline 207 -> 210.

`lint:ci` exits 0. 173 tests pass across the two touched files.

* fix(#3829): a reused finding id no longer inherits the old finding's decision

Found by this round's own adversarial review, which drove it rather than
reasoned about it — and it refuted the arm I had named as my strongest
suspicion, so it is recorded as a correction, not a discovery.

Finding ids are reused across re-reviews: the --auto loop renumbers. `row()`
inherited a prior decision on an id MATCH ALONE, with nothing checking it was
the same finding. Driven: a prior `CR-01 fixed` row against a review reporting a
brand-new CR-01 rendered the NEW finding `fixed`. A false decision in the
artifact whose entire purpose is telling triaged from forgotten — the same
failure mode round 4's blocker was, reached by the other door.

I had argued this was closed by the stale-report arm. It is not: that arm guards
the FIX-REPORT path only. The PRIOR-LEDGER path had no title check at all.

- The ledger now records each finding's title, in the FRONTMATTER rather than a
  fifth table column: the Source cell is the field a human hand-edits and the one
  that must escape pipes, and a second free-text column doubles that surface for
  no reader benefit.
- A prior decision is inherited only when the recorded title still matches. An
  ABSENT prior title inherits, deliberately — a ledger written before titles were
  recorded carries none, and refusing there would reset every decision in it,
  which is the loss this guard exists to prevent, caused by the guard.
- A decision whose id has been reused is PRESERVED under a `superseded:` key
  rather than dropped. The review's driven refutation was precisely that the
  mismatch was surfaced while the decision was lost. It cannot keep a row — the
  id is taken, and two rows under one id is an ambiguity, not a record — so it is
  carried in the frontmatter, re-emitted every run, deduped by id+title, and
  named on the console.
- And an iteration-derived decision now cites the report it actually came from.
  The Source cell hard-coded the unsuffixed `<NN>-REVIEW-FIX.md`, so a decision
  read out of an iteration backup cited a file that may not exist. A citation the
  reader cannot follow is worse than none. Also the review's finding.

Five new tests. Four fail against the pre-fix step; the fifth — that a ledger
with no recorded title still inherits — is a BACK-COMPAT guard and passes both
ways by construction. It is not a reversion control and is not counted as one.

* fix(#3829): follow the cleanup move through, and stop miscalling a converged run

Three loose ends the earlier cleanup relocation left, two of them found by the
round's own review and one by the suite.

**#3190's own test still pinned the old placement.** T6 asserted the `.iterN.md`
removal lives inside `auto_iteration_loop` — exactly what moving it broke. Its
SEMANTICS are unchanged and still asserted: removed on convergence, retained on
degradation, creation intact. What it now pins additionally is the ordering that
forced the move — the ledger reads the backups BEFORE they are removed — and that
the loop no longer removes what it just wrote. Rewritten rather than deleted: the
assertion was superseded, the guarantee was not.

**`CONVERGED` had become a decoy.** With the removal gone from the loop, the flag
was set in two places and read in none. Deleted, and the prose that still said
"the loop sets it" rewritten to what is true: the loop breaks on exactly one
condition, a clean re-review, which leaves REVIEW.md at `status: clean` — and that
is what `cleanup_iteration_backups` re-derives from.

**A converged final iteration reported the opposite of what happened.** The
post-loop message keyed on the iteration COUNTER alone, so a run that converged ON
iteration 3 exited with `ITERATION == MAX_ITERATIONS` and printed "Reached maximum
iterations. Remaining issues documented in REVIEW-FIX.md" over a run in which
every finding was fixed. Convergence is re-derived from the review the loop left
behind — the same signal the cleanup step reads, so the two cannot disagree.

* docs(#3829): retract two claims this round made and could not support

Both were caught by the round's own adversarial review, both were driven, and
both would have reached the maintainer. Recording the retraction where the claim
was made, rather than only in a PR comment.

**The env-prefix "latent bug" does not exist.** An earlier commit in this round
claimed that `FIX_REPORT_FILE="${_pd}/${PADDED}-REVIEW-FIX.md"`, sitting in the
same `node -e` env-prefix list that sets `PADDED`, expanded the OUTER variable
rather than the one two entries earlier — reading ShellCheck's SC2097/SC2098 as
a defect report. Driven in bash and in dash: assignments in one prefix list take
effect left to right, and the later entry DOES see the earlier one. The warning
is a false positive here. The split is kept, but for readability only; the
comment no longer describes it as a fix.

**The HAS_BASH decline rested on a rule that does not govern these call sites.**
It cited `local/no-unguarded-nonportable-exec` as REQUIRING the
`process.platform !== 'win32'` guard. Checked, and wrong on both halves: the rule
fires only on a file that also chmods an exec bit with an octal literal, and this
file has none — so it never runs here — while
`eslint-rules/lib/platform-guard.cjs` accepts four guard shapes plus
`os.platform()`, not one. A constraint that exists is not a constraint that
applies, and I did not check which.

The decline stands on narrower and honest grounds: whether these fences PASS on
the Windows lane is UNVERIFIED. What evidence there is points at divergence
rather than absence — the rule's subject line is that `bash -c` constructs "fail
on Windows Git Bash", and this PR already measured `mkfifo` existing on that
runner, exiting 0, and creating no FIFO. So a probe would not be a clean win; it
would light 22 tests on a lane whose shell semantics are known to differ and
unknown in detail. That is a measurement to make deliberately, not a change to
make in passing. The gap is real and is now stated as a gap.

* fix(#3829): close four defects the review drove out of the first title fix

The round's own adversarial review re-ran against the reworked tree and refuted
two more claims. Every item below is its finding, verified before acting.

**An iteration-only decision recorded no title, so the reuse guard leaked.**
`applied` stored `{d, src}` and the row took its title from the current review —
which does not report the finding at all. The row shipped with no title, and the
next review reusing that id hit the title-ABSENT back-compat exception and
inherited the old `fixed`. The exact defect the title machinery exists to close,
surviving through the hole opened for legacy ledgers. `applied` now carries the
title it was decided under.

**A changed decision was dropped in favour of the obsolete one.** The dedupe was
a has()-guard, so re-superseding a finding whose decision had since changed left
the older record standing. It now replaces.

**Re-spaced titles double-recorded.** The dedupe keyed on the raw title while
`sameTitle()` collapses whitespace; the key now agrees with the comparison.

**And the frontmatter was not valid YAML.** `title: Parser: loses data` is
rejected outright by a real reader, and the `superseded:` line format was not
YAML at all. Values are emitted as JSON scalars — YAML 1.2 is a JSON superset —
and superseded records are properly nested. Round-tripped through js-yaml in the
tests.

One more, self-inflicted while fixing the above: the parse registered each
carried superseded record TWICE, once at `- id:` under an empty-title key and
again at `title:`. Records doubled on every run. They are collected during the
walk and registered once, complete.

**T6 was vacuous.** The review flipped `= "clean"` to `!=` in the cleanup and the
rewritten T6 still passed — it greps for `FINAL_STATUS`, `rm` and "retained"
occurring somewhere, never wiring them to a branch. T6b now EXECUTES the fence in
both directions against real files. It fails on that exact mutation.

**And a converged final iteration printed two success messages** — the loop's
break already reported it. This branch now stays silent and exists only to
withhold the degradation warning.

Three CI gates the base range brought in, all tripped by this round's own text:

- `/gsd-code-review` in a comment — runtime workflow artifacts take the colon
  form. Now `/gsd:code-review`.
- The preamble-ordering parity test: my PHASE_DIR comment wrote the literal
  `gsd_run` before the shim preamble. Reworded.
- Prompt-stuffing: the file passed 50K. I trimmed 5.8K of my own commentary
  first; even removing every added comment leaves the added CODE over the line,
  and the file entered this round at 44,523 — 89% of the budget. Added to
  SIZE_ONLY_WORKFLOWS with the same reasoning the two existing entries carry, and
  the same acknowledgement: splitting is the real fix.

* test(#3829): extract the cleanup fence without an ad-hoc markdown regex

T6b's helper used `/```bash\n([\s\S]*?)\n```/`, which trips two of the repo's own
rules: `local/no-adhoc-markdown-parsing` (use the sectionizer, not a hand-rolled
fence regex) and `local/no-crlf-fragile-split` (a bare `\n` against readFileSync
content is wrong under Windows autocrlf).

Line-scanned now, CRLF-normalized first — the same shape `bashFences()` in
tests/code-review-pipeline-regression.test.cjs already uses, which solved this
first. `npm run lint` is clean and T6b still fails on the inverted-branch
mutation it exists to catch.

* fix(#3829): withdraw the superseded-decision store; keep the identity guard

Three adversarial passes over this round each found real defects, and passes 2
and 3 were entirely inside the `superseded:` block added in pass 1 — a second
identity scheme, keyed on (id, title), living beside the row store keyed on id.
Pass 3 refuted it on three separate counts: a legacy title that merely looked
like JSON lost its quotes and fabricated a record; a finding that was deferred,
superseded, then returned and fixed left an active row and an obsolete
superseded record standing together, reporting `unchanged` forever; and my own
test for the replacement path never passed the earlier ledger in, so it guarded
nothing.

The construct had no terminal state. It is withdrawn.

**What survives is the safety property.** The ledger records each finding's
title, and a recorded decision is carried forward only while the id still names
the same finding. That is what stops a renumbered `CR-01` inheriting an earlier
`CR-01`'s `fixed` — a false decision in the artifact whose purpose is telling
triaged from forgotten, and the same class as round 4's blocker.

**What is given up, and it is disclosed rather than hidden.** On a detected
reuse the earlier decision loses its row. The drop is reported on the console
naming the id and what had been decided, the previous ledger is committed so the
row remains in git, and docs/features/code-review-pipeline.md states the
limitation.

Two defects from pass 3 are fixed rather than deleted, because they are in the
guard and not the store:

- **Known-empty and NOT-KNOWN were conflated.** `### CR-01:` yields an empty
  title; that is a title. While it emitted no `title:` key it read back as a
  pre-format ledger and inherited across a reused id — the same leak, three
  passes running. Emitted whenever the title is known, empty included; a carried
  row no source knows stays absent, which is the legacy-compatible read.
  Underneath it was a falsy fallback: `(act && act.t) || priorTitle.get(id)`
  discards `''`. Now a typeof check.
- **JSON.parse ran on legacy values.** A pre-format ledger whose bare title was
  written `"quoted"` was parsed and lost its quotes, so the decision stopped
  matching. The frontmatter now declares `titles: json` and the parse is gated on
  it; a ledger without the marker keeps its scalars.

One defect from pass 3 is NOT mine and is not fixed here: a converged run prints
a success message from the loop break AND another from `present_results`. Both
predate this round. My earlier claim that "the duplicate is gone" was true only
of the pair I introduced; the pre-existing pair stands, and widening this round
into `present_results` is not warranted.

188 tests pass. The three new tests fire against the pre-simplification step.
`lint:ci` exits 0. The step file is 55,590 chars, down from a 62,220 peak.

* docs(#3829): stop the ledger promising a preservation it no longer makes

Fourth review pass. No machinery defects this time — both findings are claims in
text this step SHIPS, which is the class this whole stack exists to prevent.

**The rendered ledger still said "Re-running the gate preserves every row and
every disposition."** That was true until the same round gave the step an
intentional drop for a reused finding id, and then it was false in the artifact's
own user-facing footer. It now states what the step does, including the one
exception, where a reader actually meets it.

**And the console asserted "the previous ledger is in git."** Committing the
ledger is gated on `commit_docs`, and a failed commit is swallowed — so under
`commit_docs=false` the overwritten decision may exist nowhere. The note reports
the drop and stops there; asserting a recovery path that may not be there is the
same overclaim in a smaller font.

Two residuals from the same pass are DECLINED and documented rather than fixed,
because both would need the second identity scheme just withdrawn:

- A pre-titles ledger carries no titles, so its decisions inherit on the id
  alone. Refusing there resets every decision in every existing ledger, which is
  the loss the guard exists to prevent.
- Two genuinely distinct findings sharing both an id and a title are
  indistinguishable to an (id, title) key.

The pass also refuted the `titles: json` marker on a ledger written by
`b86ea6065^`, which emitted JSON titles before the marker existed. Declined:
that revision is an intermediate commit on this unpushed branch and has never
been released. The PR's published head writes no titles at all, so a real ledger
is either pre-titles (unmarked, bare — handled) or written by the shipped version
(marked). The unmarked-JSON state cannot reach a user.

Test pinned, and it fails against the pre-correction step.

* docs(#3829): fix four wrong citations and one false size justification

All four came out of a claim-audit of this round's own response comment — an
audit of the text, not the code, which is where the remaining errors were.

- **The caller citation was wrong.** The in-code note said both inputs bind from
  `gsd_run query init.phase-op`. `execute-phase.md:85` uses `init.execute-phase`;
  only `code-review-fix.md:22` uses `init.phase-op`. The substantive point is
  unchanged — both are orchestrator-derived, neither is raw user input — but the
  citation was not checked.
- **A leftover "the prior row is in git."** Removed from the console note last
  commit, left standing in the comment two lines above it.
- **The docs still carried the promise the ledger had just dropped.** The
  rendered footer was corrected; the same sentence in
  `docs/features/code-review-pipeline.md` was not.
- **The SIZE_ONLY_WORKFLOWS justification was false.** It claimed the added CODE
  alone exceeded the threshold. Removing every round-added comment leaves 47,148
  chars against a 50,000 limit, so the file CAN fit — the claim was wrong, and an
  exemption defended on a wrong premise is worse than no exemption.

So the entry is re-justified on what is actually true, and earned first: another
**10,188 chars** of this round's own commentary are cut (62,220 → 52,032, from a
44,523 baseline that was already 89% of the budget). Fitting under is possible
only by stripping essentially all remaining explanation from logic three review
passes found defects in. That is the wrong trade in a file whose house style is
heavy in-fence documentation, and the entry says so rather than implying the
file had no choice.

One measurement corrected while checking: the Windows-lane skip count is **37**,
not the 22 the review cited nor the 26 I first counted. Twenty-two and 26 count
`{ skip: !HAS_BASH }` CALL SITES; a skip on a `describe` cancels its subtests.
Forced the constant false and counted what actually skips.

236 tests pass. `lint:ci` exits 0.

* fix(#3829): the drop report is conditional, and two published claims were not

A fifth adversarial pass, run against the two commits that went out AFTER the
fourth pass and were never reviewed, refuted three claims this round published.

1. The drop is NOT reported unconditionally. `row()` reports only a RECORDED
   decision (`was.d !== 'open'`); a prior row still at `open` is replaced in
   silence. The behaviour is right — `open` records no decision to lose — but
   the shipped ledger legend and BOTH feature docs asserted the report happens
   every time. Text corrected in all three places, which is the same defect
   class this round already corrected once for the preservation promise.

2. The test guarding that console wording was VACUOUS: it ran with no prior
   ledger, so no reuse occurred and its `is in git` assertion could not have
   failed however the console was worded. Driven through a real drop now, with
   the drop asserted as a precondition. A new test covers the `open` arm and
   fails on the pre-fix legend.

3. The HAS_BASH gap is now MEASURED rather than assumed, on native Windows with
   Git Bash 5.2.37 / MINGW64 first on PATH, node v25.2.1:

       HAS_BASH left alone:  179 tests, 127 pass,  0 fail, 52 skipped
       HAS_BASH forced true: 179 tests, 140 pass, 24 fail, 15 skipped

   So 37 of the skips are this guard's, confirming the count the round
   published — and unskipping is NOT a clean win: 24 fail, clustered on
   `bash -c` quoting and spawn failures, exactly the divergence the eslint
   rule's subject line names. The guard stays; it now documents a measured gap.
   The stale "the count is 22" comment is gone.

4. The size-exemption justification was wrong a second time. The overshoot is
   ~2.3K normalized chars, not "essentially all remaining explanation": the
   round's committed peak was 59,246 chars (not 62,220, which was never
   committed), and it entered at 44,466 chars, not 44,523 — both earlier
   figures mixed bytes into a character measurement. Rewritten to the numbers
   the scanner actually produces.

Also: the shipped comment said both callers validate the phase shape without
naming that they validate PADDED_PHASE, not the raw PHASE_NUMBER this step is
handed.

* fix(#3829): renumber this PR's two REQs, which #3661 took while the branch sat

The rebase onto current `next` surfaced a REQ-number collision, not a text
conflict. #3661 landed `REQ-REVIEW-08` (`workflow.code_review_point`) on
`docs/features/code-review-pipeline.md` while this branch also claimed 08 and
09 for severity surfacing and the per-finding disposition. Two different
requirements under one identifier is the kind of thing that reads as correct in
both diffs and is wrong in the merged tree.

Base numbering wins, because it shipped: `REQ-REVIEW-08` stays #3661's. This
PR's two become **REQ-REVIEW-09** (severity surfacing) and **REQ-REVIEW-10**
(per-finding disposition). Swept the whole tree rather than the conflict hunk —
two references sat in files git merged cleanly and never flagged:

- `gsd-core/workflows/code-review-fix.md:450`, the prose stating why
  `record_disposition` is the step's only reachable call site.
- `tests/code-review-pipeline-regression.test.cjs:1782`, the comment on the
  test that pins that call site.

`docs/FEATURES.md` is regenerated from the fragment rather than hand-edited;
`node scripts/gen-features.cjs --check` is green (178 features, 21 groups) and
`lint:generated-sync` exits 0.

Two things stated rather than quietly carried. The `Emitted-Drift-Ack-Growth`
trailer on the round-2 commit still reads `REQ-REVIEW-09` for what is now
REQ-REVIEW-10 — it is a historical acknowledgment of that commit's growth, and
its purpose is unaffected, so it is left rather than rewritten across 52
replayed commits. And `docs/INVENTORY-MANIFEST.json` appeared stale immediately
after the replay, reporting two missing `cli_modules/` entries; that was the
lane's pre-rebase build output, not manifest drift. Rebuilding in the replayed
lane and re-checking shows it in sync and unmodified. Regenerating before the
build would have committed the deletion of two base-added entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): reach the title round-trip with a generator that can break it

Round 6's only finding. The round-5 title tracking introduced a fresh parser
(the `titles: json` / `  - id:` / `    title:` frontmatter walk) and a fresh
bijective contract (`JSON.stringify(oneLine(t))` out, `/^    title: (.*)$/`
plus `JSON.parse` back in), and `tests/code-review-disposition.property.test.cjs`
was untouched since round 4 with no reference to `title` at all. Every heading
the generator built was `'### <id>: finding number <i>'` — never a colon, a
quote, a backslash, or the empty string.

You were right that this is the round-3 shape again, and I would rather
demonstrate that than assert it. Two mutations to the shipped step, each a
plausible edit rather than a contrived one:

  A. render `titles: raw` instead of `titles: json`, so the re-parser never
     JSON.parses and stores the quoted scalar as the title;
  B. `yv = (t) => oneLine(t)` — the bare scalar, no JSON at all.

    mutation A — new generator: FAIL     old generator: pass (3/3)
    mutation B — new generator: FAIL     old generator: pass (3/3)

Both ship past the pre-round suite. The gap was reachable, not theoretical.

What changed:

- `TITLE`, a new arbitrary drawn from the class the render's own comments say
  the escaping is for — `:` (why `yv()` exists), `"` and `\` (what
  stringify/parse must round-trip), the empty string (the known-empty vs
  not-known distinction the render draws explicitly) — plus scalars that MIMIC
  the ledger's own frontmatter grammar (`findings:`, `titles: json`, a nested
  `    title: ` line, `  - id: CR-99`), unicode, surrounding whitespace, and one
  title long enough to outrun a scanner assuming short scalars.
- `FINDINGS` now carries a title per id, so all four properties run the cycle
  over the title contract instead of over a constant. `IDS` keeps the old
  id-only shape it is built from.
- A fourth property asserting the round trip in the two places it is observable:
  the stored scalar must `JSON.parse` back to the trimmed heading title, and a
  hand-recorded decision must survive the next run.

The second half is the one that matters, and its construction is the point.
The decision is made by EDITING THE RENDERED LEDGER IN PLACE, never by writing
a bare row the way the existing properties do. A bare row carries no
frontmatter, so `priorTitle` is empty, `sameFinding()` returns true through its
`!priorTitle.has(id)` back-compat arm, and the title contract is never
consulted — the property would pass over a completely broken round-trip. Both
mutations above go green against the bare-row form. That collapse is why the
property is written this way, and the comment says so in place.

So the assertion is the consequence, not the JSON: a lossy round-trip does not
corrupt a title, it makes `sameFinding()` false and resets a human's `deferred`
to `open` with the reason gone — this PR's own founding failure mode, reached
through the field the round-5 work added.

BOUND, stated rather than quietly omitted: the generator emits no CR or LF. A
`###` heading is one line by definition, so a newline is not an input the
heading parser can be handed; `oneLine()` guards the value's other producers,
not this one.

Two things found while writing it, both corrected here rather than left:

- `runOnce` now returns stdout. The reuse report is a CONSOLE note, not a
  ledger key, so my first draft's `assert.doesNotMatch(ledger, /^reused:/m)`
  was vacuously true forever — a test that cannot fail.
- `expectedTitle` is a TRIM, not a `\s+` collapse. Collapsing is `sameTitle`'s
  COMPARISON rule; `oneLine()` is the STORAGE rule and preserves internal
  whitespace. The collapse form fails on an internal tab against entirely
  correct code, which is how a test gets weakened instead of believed the first
  time it goes red.

The file header claimed "two properties" while three were running; it now
states four, one line each.

239 tests pass across the four pipeline files, 0 skipped. `lint:ci` exits 0
(`lint-workflow-shellcheck`: 203 baseline findings, 0 new).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): the prefix census guard now says four sites, because round 5 added one

Self-found, from re-deriving the round-1 finding-id census this round rather
than carrying the round-1 verdict forward.

The census guard's comment says the prefix set is "written out three times —
the heading matcher, the ledger re-parser, and (by its keys) the severity map".
That was true when it was written. Round 5's title tracking added a fourth
copy: the frontmatter `- id: ((?:CR|BL|WR|IN)-\d+)` matcher that rebuilds
`priorTitle`.

The guard itself did not fall behind, and the reason is worth keeping visible:
`idAlternations()` scans the extracted script by PATTERN rather than walking a
fixed list of sites, so the new alternation was absorbed with no edit. Verified
by running the extractor at this head — three alternations found, one distinct
set, severity map keys `CR,BL,WR` with `IN` on the documented `info` default,
0 domain members not reached.

Only the prose fell behind. Corrected, with the pattern-scan rationale stated
in place so the next reader does not helpfully convert it into the hand-listed
enumeration it deliberately is not — which would be exactly the defect this
guard exists to catch, in the guard.

Census discharge for this round: re-derived at the rebased head over the
extracted shipped script, 3 enumeration sites reached, 0 not reached; the
domain (the prefixes `gsd-code-reviewer.md` can emit, walked across both its
heading template and its prose Label-equivalence paragraph) is unchanged since
round 1 at 4 of 4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): catch a duplicate REQ id in a fragment, since nothing did

Not from your review — this is the test the round owed itself, and I would
rather say why than let it look like scope creep.

The renumber commit earlier in this round has no reversion control without it.
I reverted that fix to check, and the first attempt LOOKED controlled: reverting
only the fragment turned `gen-features --check` red. That is the generated-sync
gate noticing the projection went stale, not anything noticing the collision.
Reverting CONSISTENTLY — fragment plus a regenerated `docs/FEATURES.md` — is
silent:

    gen-features --check   rc=0
    lint:ci                rc=0
    pipeline suite         rc=0

with two `REQ-REVIEW-08` entries standing in one requirement list. Nothing in
the repo reads REQ ids at all, so there was no second place for it to be caught.

The failure this guards is a MERGE, not an edit, which is why review does not
see it: two PRs open at once each append "the next" REQ number to the same list,
and whichever lands second is rebased onto a list that already used it. git
merges them as different lines of one file and reports nothing. Neither PR's
diff shows a collision — each is correct against the tree it was written on.
That is exactly how #3661 and this PR both ended up claiming REQ-REVIEW-08.

Scope, stated because it is the part that could be wrong: the check is WITHIN a
fragment, never across the corpus. Two different features legitimately both
carry `REQ-REVIEW-01..07` — the cross-AI review feature and the code-review
pipeline — so corpus-wide uniqueness would be false on the committed tree and
would have to be weakened the day it first ran. A requirement list belongs to
its feature; that is the scope of the identifier.

It lives in `describe('the committed docs/features/ corpus')` because it is an
invariant over the committed corpus, which is that block's stated job, and it
pins no count — the file's own header rules out counts as shared mutable cells
that every feature PR would have to edit.

Control: green on the committed tree (no fragment carries a duplicate today);
red on the restored collision, naming the file and the id. 85 tests pass in
this file.

Happy to drop this if you would rather the round stayed inside the review's
four corners — but then the renumber ships uncontrolled, and I would rather put
that choice in front of you than make it quietly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): finish the census comment correction, which stopped one line short

Found by this round's own pre-push adversarial review, which refuted the claim
the previous commit made about itself.

`7f019d985` said the census comment correction was complete. It corrected one
site and left two, both in the helper block twelve lines above the test it
belongs to:

- `severityMapKeys`' header still read "The THIRD copy: the severity map's
  keys". With three alternations the map is the FOURTH copy, and has been since
  round 5.
- `idAlternations`' header said "adding a prefix to only two of them is silent",
  written when there were two alternations and never updated to three.

This is the defect the original correction was ABOUT, committed inside the
correction: a fragment of prose carries no supersession marker, so a reader
landing on line 2810 gets the dead count stated as current fact, and the fixed
comment eighty lines down does not reach them. Fixing one surface and leaving
its neighbour is not a partial fix, it is the same fix not done.

The region is now consistent end to end, and both headers say the thing that
actually matters — the scan is by PATTERN, not a fixed list of sites, which is
why round 5's new matcher needed no edit here and why converting it to an
enumeration would reintroduce exactly the drift it guards.

WHILE HERE, a disclosure that was narrower than the truth. `7a6680e8f` said the
`Emitted-Drift-Ack-Growth` trailer still names REQ-REVIEW-09 for what is now
REQ-REVIEW-10, and left it deliberately rather than rewrite 52 replayed
commits. That is right, but it is not the whole set: the message BODIES of
`c94106568` ("wire the disposition ledger into the fix path") and `06282f668`
("migrate the emitted-drift ack") both state "REQ-REVIEW-09 was unreachable in
every shipped path", meaning the disposition requirement, which is now
REQ-REVIEW-10.

Same decision, stated at its real size: three historical references, not one.
They are commit history rather than living documentation — git is the record of
what was believed when — and rewriting the branch to correct a number in a
message would cost every review round its correspondence to the commits it
reviewed. The TREE carries no stale reference; `docs/`, the workflows and the
tests all read REQ-REVIEW-09 for severity surfacing and REQ-REVIEW-10 for the
disposition.

Regression file: 181 tests pass, 0 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* test(#3829): the third stale count, and a disclosure that over-counted itself

Both found by re-running this round's pre-push review after the last fix. It
refuted the commit that claimed the region was consistent — for the second time
in a row — and it was right again.

**The third site.** `:2917` said "And the third copy, which is not an
alternation" and `:2919` said "Without this, both regexes can gain a prefix".
Written when there were two alternations; there are three, so the map is the
fourth copy and it is three regexes that can drift.

Worth saying how it survived two passes, because the mechanism is the point and
it is the same one this PR keeps re-learning. Both earlier passes VERIFIED with
a grep built from the strings I had just fixed — `THIRD copy`, case-sensitive,
plus a handful of phrasings I expected. `the third copy` in lowercase matched
none of them, and `both regexes` was not a phrasing I thought to look for. A
grep returns what you already thought of; that is not a verification of prose,
it is a re-statement of your own assumption. The region is now checked by
reading it end to end, and all four count statements agree: three alternations
(heading matcher, ledger row re-parser, frontmatter `- id:` matcher), with the
severity map as the fourth copy.

**And the disclosure over-counted.** The previous commit widened the historical
REQ-REVIEW-09 references from one to three. Three is wrong. There are TWO
underlying statements:

  - `c94106568`'s message body, and
  - the `Emitted-Drift-Ack-Growth` trailer on `06282f668`.

I counted `06282f668` twice — once as "the trailer" and once as "a body" — when
its only mention IS that trailer (`git show -s --format=%B 06282f668 |
grep -c REQ-REVIEW-09` outside the trailer line: 0). Over-counting is the safe
direction and it is still a wrong number in a message, which is the thing this
round has been correcting all along.

The decision is unchanged: both are commit history rather than living
documentation, and rewriting the branch to fix a number in a message would cost
every review round its correspondence to the commits it reviewed. The TREE
carries no stale reference — 08 is #3661's `workflow.code_review_point`, 09 is
severity surfacing, 10 is the per-finding disposition.

Comment-only in one test file; no assertion, regex or extracted-script
expectation moved. Regression file: 181 tests pass, 0 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH

* fix(#3829): join the disposition-step dispatch so REQ-LANG-04 inheritance is provable

`lint-response-language-coverage` (#2529, which landed on `next` after this PR was
approved) reported `execute-phase/steps/code-review-disposition.md` as having no
response-language coverage. The step does inherit it: `execute-phase.md` imports
`references/execute-phase-response-language.md` and dispatches the step with
`Read and execute`. The dispatch stub wrapped, leaving the verb at the end of one
line and the path at the start of the next, and `namesFragmentAsEntryPoint` matches
within a single line — so a genuine inheritance was unprovable to the linter.

Rejoining the verb and the path restores it: `namesFragmentAsEntryPoint` goes
false -> true and the lint reports `OK (165 workflows covered)`. Only line breaks
move — the word stream is identical to the previous revision, and the file is
unchanged at 93,390 bytes, so no growth acknowledgment is owed.

This takes the third coverage form the lint documents — inheritance — rather than
the inline directive the CI message names first. Where inheritance is provable the
lint's own comments say a second copy "buys no coverage and adds a sentence that
can drift", and the step file already sits over the prompt-stuffing threshold.

Swept all 76 fragments in the catalog: this is the only one whose parent's previous
line ends with a dispatch verb. The 17 others that are mentioned without a provable
entry point are table-routed or bare prose references carrying no dispatch verb at
all, and correctly hold the pinned inline directive instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JgX6QQmygeZnQqbc3o8RNC

* chore(#3829): regenerate derived artifacts after rebase onto next

The rebase onto current `next` conflicted on the 19 install-tree goldens and
`docs/FEATURES.md`. Those are generated, so the conflicts were resolved
arbitrarily and the generators re-run (`npm run regen:derived`) rather than
hand-merged — a clean textual merge of a generated file attests the merge, never
the content.

Reconciled per artifact against the base's own committed copy rather than against
the pre-regen tree, because the pre-regen tree is the arbitrary resolution:

  - all 19 `tests/fixtures/install-tree/*.json` now differ from
    `upstream/next` by exactly one key,
    `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`;
  - `docs/FEATURES.md` differs by exactly REQ-REVIEW-09/10 and this PR's own
    reference section;
  - `docs/INVENTORY-MANIFEST.json` differs by exactly the same one step file,
    and needed no regeneration to get there.

Nothing the base added was dropped by the arbitrary resolution: the restored
entries (the `gsd-core/agents/` and `gsd-core/commands/gsd/` families, the
compact templates, the `detail/elaboration.md` files, `gsd-secret-read-guard.js`)
are all base-owned and came back through the generator, which is what the
resolve-arbitrarily-then-regenerate discipline is for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* fix(#3829): repair the rebase's conflict resolution in the regression suite

The rebase onto current `next` hit one add/add conflict in this file: #4209's
external-reviewer-evidence describe and this PR's #3829 block were added at the
same insertion point. Resolving it by keeping both sides was correct in
substance and wrong in mechanics — the conflict boundary cuts through two open
blocks that the SHARED trailing `  });\n});` closes, so each side carries +2
unbalanced braces on its own and concatenating them left the file with 683 `{`
against 680 `}`.

`node --check` fails outright, so the whole file deregistered rather than
failing a test — 188 tests silently stopped existing. Rebuilt the region as a
real three-way merge (ancestor a262ad6b6, ours upstream/next, theirs 77ee739c3)
and closed the first side explicitly before the second begins.

Both feature blocks are present exactly once, braces balance 683/683, and the
file runs 188/188 locally. The sibling markdown file resolved the same way is
unaffected and was checked rather than assumed: prose has no block structure to
unbalance, and it differs from the base by 112 added lines with zero removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* test(#3829): inject the unreadable-review failure in a way root cannot bypass

Round 9 finding 1. `runShippedGateCounts({ mode: 0o000 })` does not simulate an
unreadable review under root: root bypasses POSIX read permission bits, so the
fence's `[ -f "$REVIEW_FILE" ] && [ -r "$REVIEW_FILE" ]` guard stays true, the
fixture is read, and the assertion sees a real breakdown where it expects
silence. Reproduced as reported — `node:24-slim`, euid 0:

    not ok 18 - an unreadable REVIEW.md leaves the counts empty and does not abort
    actual: 'Code review: 4 findings — 1 critical, 2 warning, 1 info.\n...'

**The prescribed remedy does not reach this site, so this adapts it rather than
applying it.** Stubbing `fs.readFileSync` to throw EACCES is the right fix where
the read happens in-process; here the read is performed by a spawned `bash`, so
node's `fs` is not on the code path and the stub would change nothing.

What the guard actually has is two legs, and only `-r` is defeated by root:

  - the `-r` leg keeps the mode-bit fixture and declares the lanes it cannot
    bind on (`win32`, `euid 0`), which is exactly what
    tests/plan-review-convergence.test.cjs:2326 does for its own shell-side
    `-r` arm — the repo's existing precedent for this shape;
  - the `-f` leg is new and root-immune: a DIRECTORY at the review path fails
    `-f` for every euid, reaching the same non-reporting arm with the same
    observable. It binds on the bench lane where the first test is skipped.

Skipping the first without adding the second would have traded a false failure
for lost coverage on the only lane that found this.

Reversion control, run as root: reverting this commit fails exactly
`an unreadable REVIEW.md leaves the counts empty and does not abort` and its
enclosing describe `#3861 round 1 — the counts mirror is asserted against the
shipped shell`, with no other change to the failure set. **That answers the
round's open question** — the review flagged the describe as possibly a second
root cause; it is the first one's rollup, and there is no second.

Local (euid 1000): 189/189, both tests run.
Root: 179 pass / 1 skip, the skip naming its reason, the directory test running.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* chore(#3829): refresh the compact-content benchmark baseline for the gate's edit

Self-found in this round; not raised in review. The base range landed #4139's
compact-content benchmark, whose committed baseline records per-workflow token
counts. This PR replaces a 12-line bash block in `execute-phase.md`'s
code_review_gate with a 5-line dispatch paragraph, which moves that workflow's
measured counts by 17 tokens — so the baseline the base just added drifts
against a tree it was measured before.

    DRIFT: split "execute-phase": off 25631 -> 25614 (-17), on 23380 -> 23363 (-17)
    DRIFT: aggregate: off 106923 -> 106906, on 90275 -> 90258

Refreshed with the remedy the script itself names
(`node scripts/benchmark-compact-content.cjs --write`). The regenerated diff
touches only the `execute-phase` entry and the aggregate — every other
workflow's numbers are byte-identical, which is the reconcile this PR's edit
predicts.

Attributed rather than assumed: `tests/benchmark-compact-content.test.cjs` is
27/27 at `upstream/next` with no PR content, and was 26/27 on this head. So the
drift is this PR's, not base noise — and it is invisible to a diff-scoped sweep,
because the PR never touches the baseline file and the base range is what
created it. The sibling `benchmark:compact-content-variants` was checked in the
same pass and reports up to date, so this is the only one of the pair affected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* test(#3829): cover an EMPTY REVIEW.md, a case the body claimed and no test reached

Found by this round's own adversarial audit of the PR body, not by review. The body
has said since round 1 that the suite covers "a REVIEW.md that is missing, empty,
or a directory". Two of those three were true. The empty one was not.

Every `reviewText: ''` call in this file also passes `writeReview: false`, which
makes the file MISSING, not empty — so the arm the body named had no test at all.
They are genuinely different paths through the shipped fence: a missing file never
gets past `[ -f ]`, while an empty one passes both `[ -f ]` and `[ -r ]` and is
actually opened and read.

Probed the shipped fence directly against a real empty file before asserting
anything: exit 0, empty stdout. So the behaviour was already correct and only the
coverage claim was false — which is the same "documented as covered, not covered"
shape this PR exists to make visible in the review gate, found in its own body.

The explanatory comment is deliberately precise about WHY the scan yields nothing,
because the plausible reading is wrong and a later reader would inherit it: it is
not the `NR==1{if($0!="---") exit}` guard. A zero-byte file gives awk no record, so
that action never runs (NR stays 0); the output is empty because `closed` is never
set. The comment also states what the test does not prove on its own — its
observable is identical to the missing-file case, so "the file was read" rests on
the harness and the fence, not on the assertions.

189 -> 190 tests in this file, all passing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV

* chore(#3829): regenerate derived artifacts after rebase onto next

The rebase onto current `next` conflicted in the 19 install-tree goldens,
`docs/INVENTORY-MANIFEST.json` and the compact-content benchmark baseline.
Those are generated, so they were resolved arbitrarily and regenerated
with their own producers rather than hand-merged: every golden now differs
from `next`'s committed copy by exactly the one PR-owned entry
(`gsd-core/workflows/execute-phase/steps/code-review-disposition.md`), the
manifest by the same entry, and the benchmark baseline by the
`execute-phase` split plus the aggregate.

* chore(#3829): regenerate the platform-conformance tier for the added property test

`next` gained the conformance-tier classifier (#4591) and its CI gate after
this branch was cut. The branch adds `tests/code-review-disposition.property.test.cjs`,
so the generated tier list was one file short (546 != 547). Regenerated
with `node scripts/gen-platform-conformance-tier.cjs --write`; the macOS
tier (`--target macos --check`) already matched.

* chore(#3829): regenerate the two conflicted generated artifacts after the rebase

`next` moved 6 commits past the previous base and conflicted in exactly two
files, both generated:

- `tests/fixtures/compact-content-benchmark-baseline.json` — #4208 (`4cc2a466b`)
  and #4619 (`db4d8a9ba`) both moved the measured token counts, and this branch
  moves the `execute-phase` split too.
- `scripts/lib/platform-conformance-tier.generated.cjs` — #4641 (`4d65c248e`)
  made test-conformance the sole Windows selector and narrowed the tier to
  28.5%, and #4568/#4619 re-ran it after.

Both were resolved arbitrarily during the replay and then regenerated with
their own producers rather than hand-merged, per the generated-artifact rule:
`node scripts/benchmark-compact-content.cjs --write` and
`npm run regen:derived` (which runs `gen-platform-conformance-tier.cjs
--write` for both the default and the macOS target).

Reconciled against `next`'s own committed copies rather than the pre-regen
tree:

- the benchmark baseline differs from `next` by exactly the `execute-phase`
  split (`offTokens` 25827 -> 25810, `onTokens` 23576 -> 23559 — the 17-token
  delta this PR's step-file extraction has carried since round 5) plus the
  `aggregate` that sums it;
- the conformance tier differs from `next` by exactly one added entry,
  `tests/code-review-fix-pipeline-regression.test.cjs`. Under the narrowed
  28.5% selector that is the file the classifier now picks from this PR's
  test set. The arbitrary resolution had carried 282 stale lines computed
  under the pre-#4641 selector (`--numstat` on this commit: 3 insertions,
  282 deletions), and regeneration collapsed them.

The full derived sweep was run, not just the two named producers: all 19
install-tree goldens, `docs/INVENTORY-MANIFEST.json`, `docs/FEATURES.md`,
the macOS conformance tier and the exit-code registries regenerated
byte-identical, so nothing else drifted under the new base.

* fix(#3829): accept N-segment phase ids, and bound length per component

The base range added `scanMarkdownSingleSegmentPhaseRegex` (#4568,
`a2331c01f`), which refuses the single-optional-segment phase regex on
phase-carrying markdown lines under three roots — `gsd-core/workflows/`,
`gsd-core/references/` and `agents/` (`lint-phase-id-drift.cjs:301`,
`:373-387`). It flagged two lines in this step file. Chasing the flag
turned up two real defects behind it, so this commit is those rather than
the comment edit the flag literally asked for.

## Defect 1 — the step refused ids both its callers accept

The comments asserted that both callers validate `^[0-9]+(\.[0-9]+)?$`.
#4568 had widened those two call sites to `^[0-9]+(\.[0-9]+)*$`, so the
prose was stale. Correcting only the prose would have shipped a comment
promising N-segment support over code that refused it, because the step
carried a third `case` arm:

    *.*.*)        _ok=0 ;;   # more than one dot: not the documented shape

`23.1.2` took the refusal arm, `PADDED` came back empty, and the step
printed `Code review reporting skipped (unusable phase number ...)` and
wrote **no ledger** — for a phase id both of its callers accept.

It degraded loudly rather than silently; there is a diagnostic on stdout.
The traversal-fence test asserted that refusal as *correct*, listing
`1.2.3` among the values that must be rejected, so an arity bound and a
shape bound sat folded into one `case` arm with a test pinning the pair.

The arity arm is gone. Deleting it alone would have left the step
**wider** than its callers in one direction — `1..2` has an empty
segment, which `^[0-9]+(\.[0-9]+)*$` refuses and the retired arm had been
masking — so a third arm replaces it:

    *..*)         _ok=0 ;;   # EMPTY SEGMENT

## Defect 2 — the length bound was not per-component, though its comment said so

Removing the arity arm made a second defect reachable. The bound read:

    case "$_pn" in *.*) case "${_pn#*.}" in ?????????*) _ok=0 ;; esac ;; esac

`${_pn#*.}` is the whole tail after the first dot — one component only
while an id has at most two. With N-segment ids accepted, that form
rejects `1.1234567.1`, whose every component is a legal 7 digits, purely
because the tail measures 9 characters. The comment directly above it has
read **"LENGTH-BOUND EACH COMPONENT SEPARATELY"** since before this PR,
and had itself named the composite bound as "too strict" — the same
mistake, surviving one level up.

Both fences now walk the segments and bound each:

    _rest="$_pn"
    while [ -n "$_rest" ]; do
      case "$_rest" in
        *.*) _seg="${_rest%%.*}"; _rest="${_rest#*.}" ;;
        *)   _seg="$_rest";       _rest="" ;;
      esac
      case "$_seg" in ?????????*) _ok=0 ;; esac
    done

The `$((10#...))` overflow guard is preserved, per component: bash
integers wrap at 2^64, so an unbounded integer segment would silently
become a negative padded phase.

**The remaining divergence from the callers is a CLASS, not a list:** any
id carrying a component of nine or more characters is caller-accepted and
fence-refused — `123456789`, `1.999999999`, `1.123456789.1`,
`123456789.1`, `1.1.123456789` and so on. That narrowing is deliberate
and is the overflow guard. An earlier draft named two examples as though
they were exhaustive; that wording is withdrawn.

## What is NOT claimed

- The canonical grammar is `PHASE_NUMBER_TOKEN_SOURCE` in
  `src/phase-id.cts:65`, `\d+[A-Z]?(?:\.\d+)*`, added by **#2128**
  (`09be501eb`, 2026-07-10). An earlier draft dated it to #865
  (2026-06-08); that was the first commit to touch the *file*, not the
  one that added the constant, and it is withdrawn.
- #4568 gave the six shell sites **segment-count** parity with that
  grammar, not textual parity: the canonical source permits an optional
  `[A-Z]`, and the shell literals remain digit-only. Driven: this step
  and both callers all refuse `23A.1`, so they agree with each other and
  are jointly narrower than `src/phase-id.cts`. That is a question about
  the six sites rather than about this step, and it is not touched here.
- **#4619 does not produce N-segment ids.** It only transforms an
  already-supplied `{phase_number}` so `$((10#...))` does not abort on
  one. An earlier draft cited it as the producer; that is withdrawn, and
  is stated rather than silently swapped so a reader can see it was
  corrected.
- That the folded `case` arm is *why* nothing caught this is an
  observation about the test's shape, not an established cause.

## Tests

- `an N-SEGMENT phase number reports counts, exactly as its callers
  accept it` — drives `23.1.2` **and** `1.2.3.4`: the retired guard was
  arity-shaped, so a bound merely moved from two dots to three would pass
  a three-segment-only test.
- `the length bound is PER COMPONENT, not over the whole tail after the
  first dot` — drives an 8-char and a 9-char **middle** segment, the
  position the old form got wrong.
- `the fence agrees with its callers across a probed set spanning both
  boundaries` — example-based, and says so: a finite probe cannot prove
  congruence over an infinite language, and one review pass demonstrated
  that by injecting a `2) _ok=0` arm this test still passed. It is a
  regression pin over the values that actually broke.
- The traversal list loses `1.2.3` (legal at this base) and gains `1..2`
  and `1.2.` — the malformed-dot cases the arity guard had masked.

**Negative controls, re-measured against reconstructed fences:**

    fence state              N-seg   probed   per-comp
    fully pre-fix            FAIL    FAIL     FAIL
    shape-fix only           PASS    FAIL     FAIL
    this tree                PASS    PASS     PASS

Two tests, not one, catch Defect 2: the caller-agreement probe includes
`1.1234567.1`, so the whole-tail bound breaks it too. An earlier draft
claimed the per-component test failed alone — that table was written
before `1.1234567.1` was added to the probe and was not re-measured
afterwards. It is corrected here from a fresh run. The N-segment test
correctly does not fire on Defect 2; it predates the bound work and is
insensitive to it.

Found by this round's own adversarial review passes.

* docs(#3829): name this step's two dispatchers correctly, in code as well as in the PR body

`code-review.md` is not a call site of this step. It carries an identical
`^[0-9]+(\.[0-9]+)*$` validator, which is why it kept getting cited as one, but
it never dispatches `code-review-disposition.md`. The two dispatchers are
`execute-phase.md` (`code_review_gate`) and `code-review-fix.md`
(`record_disposition`) -- and only the second validates anything.

This round corrected that in the PR body and simultaneously wrote the old
conflation into the shipped comments, so the file asserted at line 24 what the
body denied in public, and contradicted its own line 48. Four false assertions,
each duplicated because the fenced block is emitted twice:

- "Both callers explicitly accept ... (code-review.md:63, code-review-fix.md:39)"
- "#4568 widened both of this step's callers" -- it widened the one that
  validates; the other has no validator to widen
- "Both callers already validate ... (code-review.md:63, code-review-fix.md:39)"
- "a SHAPE (..., asserted by both callers)" -- asserted by one

The same conflation had propagated into three comments in
`code-review-pipeline-regression.test.cjs`; corrected there too.

Adds the one fact that follows from naming the dispatchers correctly and that
nothing else in the tree records: `execute-phase.md` applies NO shape gate, so
this fence is not mirroring an upstream guarantee -- it IS the guarantee. A
later reader who believes the caller validates will "simplify" it away.

Comments only. The executable shell is byte-identical to 03adf9474 (verified by
stripping comment lines and diffing). 200/200 regression + property, 47/47
prompt-injection security scan, eslint and lint:generated-sync clean.

The wider [A-Z]-axis divergence between these sites and the canonical
`PHASE_NUMBER_TOKEN_SOURCE` is tracked separately as #4660 and deliberately not
restated here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xdw628PpLYveWfJu7kDyvZ

* chore(#3829): regenerate the two conflicted generated artifacts after the rebase onto next

Both conflicted during the replay onto `eb49ff98d` and were resolved arbitrarily, then
regenerated with their own producers (`npm run regen:derived`,
`node scripts/benchmark-compact-content.cjs --write`) rather than hand-merged. Reconciled
against next's committed copies: the conformance tier differs by the one entry this PR
adds, the benchmark baseline by the `execute-phase` split (the same 17-token delta this
PR's step-file extraction has carried since round 5) plus the aggregate that sums it. The
rest of the derived sweep — 19 install-tree goldens, INVENTORY-MANIFEST, FEATURES, the
macOS tier, the exit-code registries — regenerated byte-identical.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* fix(#3829): a carried row keeps the severity the ledger recorded, instead of re-inferring it from the prefix

The ledger always wrote a severity for every row (table cell and frontmatter key) and
nothing read either back: the prior-row regex skipped the cell as `[^|]*`, the frontmatter
walk collected only titles, and a carried row was rebuilt through sev() from the id prefix,
because sectionSev holds only the findings the CURRENT review reports. So a WR-04 the
reviewer filed under `## Critical Issues` was recorded critical, a human deferred it, and the
next run -- the review no longer reporting it -- silently re-recorded it warning. The one
artifact whose purpose is remembering a finding's severity lost it on the second run, in the
unsafe direction (round 11, reproduced by executing the shipped script twice).

Both persisted copies are now read back, enum-validated (ADR-227, as the disposition column
already is): the table cell first, the frontmatter `severity:` as the fallback for a
hand-mangled cell. Severity precedence is the current review's SECTION, then the RECORDED
value, then the id PREFIX, and the recorded value is inherited only while the id still names
the same finding -- the identity rule the disposition already obeys -- so a reused id starts
from its own review. sev() moves below sameFinding() because it now depends on it.

Tests: a new describe drives the reviewer's exact case (WR-04 under `## Critical Issues`,
deferred by hand, dropped by the next review -> stays critical) plus five controls: recorded
outranks prefix under no recognized section; the current section still outranks recorded; a
REUSED id does not inherit; a mangled cell falls back to the frontmatter and a mangled pair
to the prefix; a bare pre-severity row still infers. A fast-check property assigns each
finding a section independent of its prefix, carries every row through an empty review, and
asserts the section severity survives and the third run reports unchanged.

Negative control, measured against the pre-fix step: the two carry tests, the mangled-cell
test and the property fail; the three precedence/back-compat controls pass at both ends, as
they pin behaviour that predates the fix. Every prior carried-row test used CR-01/IN-01,
whose prefix already matched, so the lossy path had returned the right answer by coincidence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* fix(#3829): a malformed REVIEW.md is reported as unparsed, not passed over as clean

A REVIEW.md with three criticals and an unterminated frontmatter yielded REVIEW_STATUS='',
and the counting arm then printed nothing -- byte-identical to a clean review. The guard
that scopes the frontmatter scan was right to yield no values from an unterminated block; the
reporting arm was wrong to treat 'no status' as 'no review'. Block 2 said `status: none`
rather than `clean`, which is why a careful reader could still separate them (round 11,
Minor).

Both fences now record whether the file was actually READ, separately from what it yielded.
A read file with no parseable status -- unterminated frontmatter, no frontmatter, no
`status:` key, a zero-byte file -- prints `Code review status unparsed: ...` with no
breakdown (there is none to trust) and no --fix suggestion (nothing proves there are
findings). Absent, directory and unreadable stay silent: nothing was read, so nothing is
described. Block 2's skip line names the same distinction, `status: unparsed` vs `none`.

The counts mirror follows the shell: a mirror is always handed a text, so its empty-status
arm is the unparsed one, and the existing 'unterminated frontmatter' and 'no frontmatter at
all' parity fixtures now bind the new message on both sides. The EMPTY-file test from round 9
changes its assertion deliberately: its observable is no longer identical to the missing-file
case, which is the point. Five new tests drive the arm, its three shapes, the three shapes
that stay silent, and block 2's wording.

Negative control, against the previous step: the unterminated, no-status, no-frontmatter and
empty-file tests fail, both parity fixtures fail (the mirror moved and the shell had not),
and block 2's `unparsed` assertion fails; the stays-silent controls pass at both ends.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* docs(#3829): name the unlocked read-modify-write and reference #3780 rather than solving it

The ledger is rendered whole from a prior read with nothing serializing two writers, and
this step has two dispatchers plus an invited hand-edit, so the window is real. It is the
shape #3780 reported for WINDOWS.md under parallel executors, which #4681 closed with a
cross-process lock in src/broken-windows.cts. Not taken here, deliberately: the step is a
shell-embedded script with no dependency on the compiled tree, and adopting the lock module
is its own change. Stated at the write site and as a residual in the feature doc; no lost
update has been reproduced (round 11, Minor).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* docs(#3829): cross-reference the two "review disposition" ledgers in both directions

ADR-3806 canonizes a `## Review Dispositions Ledger` section inside PLAN.md for reviews-mode
planning: append-only per round, over REVIEWS.md findings. This PR's
`<NN>-REVIEW-DISPOSITION.md` is a sibling file beside REVIEW.md for the code-review pipeline,
rewritten idempotently with rows carried. Adjacent names, opposite durability rules, and
neither document mentioned the other -- the round-11 review checked the ADR gate against
3806, cleared it, and flagged exactly that mis-read hazard.

An in-place dated amendment section on ADR-3806 (contributor-standards "Amending an accepted
ADR", pattern 1) and a paragraph in the pipeline feature doc, each naming the other and the
axis on which they differ. docs/FEATURES.md regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* chore(#3829): regenerate the compact-content benchmark baseline after the rebase onto next

`next` moved three commits while the round was in flight and the baseline conflicted again;
resolved arbitrarily during the replay and regenerated with its own producer. It differs from
next's copy by the `execute-phase` split this PR has carried since round 5, plus the aggregate.
The rest of the derived sweep regenerated byte-identical.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* fix(#3829): accept letter-variant phase ids, matching the canonical grammar #4744 widened the dispatchers to

#4744 (#4660) landed on `next` while this round was in flight: it widened the six
shell/markdown phase-number mirrors -- `code-review-fix.md:39`, this step's validating
dispatcher, among them -- to the canonical grammar's letter axis (`12A`, `3A`, `23A.1.2`),
and added a `lint-phase-id-drift` ratchet that flags any digit-only mirror left in the
workflow tree. Rebased onto that base, this step was the one it flagged (two fences, two
sites): `12A` was refused by name and wrote no ledger, for a phase id its own dispatcher
now accepts -- the round-10 class ("the step refused phase ids its validating dispatcher
accepts") re-opened by the base. Found by running the base range's modified gates against
the rebased tree, not by the review.

Both fences now admit an uppercase letter in the character class and pin WHERE it may sit --
only as the last character of the integer part, at most once -- so `23a`, `A23`, `2A3`,
`23AB` and `23.1A` stay refused. The per-component length bound is on the DIGITS (the letter
is one character the `$((10#...))` overflow guard has no stake in, so `12345678A` is within
it exactly as `12345678` is), and the letter is carried verbatim after the padded digits,
`3A` -> `03A`, as `src/phase-id.cts` pads it. The two fences stay line-identical except for
their refusal message (the parity test holds), and every comment literal of the old shape
reads the canonical one.

Tests: a new fixture drives `12A`, `3A`, `23A.1.2` and `12345678A` through the shipped
fence to the padded path; the traversal-fence list gains the five wrong placements; the
caller-agreement probe's regex gains the letter axis with both-direction cases, and
`123456789A` joins the deliberate over-bound narrowing. Negative control against the
pre-widening step: the letter-variant test, the caller-agreement probe and the base's
`scanMarkdownLetterlessPhaseMirror` gate all fail; the traversal-fence test passes at both
ends (the five new placements were already refused, by the narrower class).

Also re-anchors the fence and test comments' `code-review-fix.md` / `code-review.md` citations by
content (the validator, not a line number): the line numbers had drifted by one against the rebased
base, and drift again on every rebase.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9FANs6AXuXNa7fCSahTQe

* chore(#3829): regenerate the two conflicted generated artifacts after the rebase onto next

Both files conflicted during the replay and were resolved arbitrarily rather than
hand-merged, then regenerated with their own producers — `npm run regen:derived`
and `node scripts/benchmark-compact-content.cjs --write`.

Reconciled against next's own committed copies rather than against the pre-regen
tree, because a clean textual merge of a pinned-number file attests the merge and
not the numbers:

- `tests/fixtures/compact-content-benchmark-baseline.json` differs from next by
  exactly the `execute-phase` split (offTokens 26264 -> 26200, onTokens 24013 ->
  23949) and the `aggregate` that sums it. That is the step-file extraction this
  PR has carried since round 5, re-measured against the new base; no other entry
  moved.
- `scripts/lib/platform-conformance-tier.generated.cjs` differs from next by
  exactly one added entry, `tests/code-review-fix-pipeline-regression.test.cjs` —
  the file the classifier picks from this PR's test set.

The full derived sweep was run, not just the two named producers: all 19
install-tree goldens, docs/FEATURES.md, docs/INVENTORY-MANIFEST.json, the macOS
tier and the exit-code registries regenerate byte-identical under the new base,
so nothing else drifted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): stop block 2 computing an unparsed shortfall from a self-contradicting findings block

Round 12, Minor. Confirmed, and the premise is slightly stronger than stated: block
2 does not merely skip block 1's `critical + warning + info == total` cross-check —
it never derived the three severity counts at all, so the check's inputs were
absent. It bounded `total` for digits and length only and handed it to the
`unparsed:` reconciliation.

So a REVIEW.md whose `findings:` block disagrees with itself (`total: 10` beside
`critical: 1, warning: 1, info: 1`) made block 1 print the countless form —
breakdown suppressed as untrustworthy — while block 2 still computed a shortfall
from that same untrusted number. Two trust models for one field, one fence apart,
with the weaker one downstream. It fails in the safe direction, which is why the
review did not raise it as a blocker; it is still a real inconsistency.

Block 2 now derives `critical`/`blocker`, `warning` and `info` through the same
`findings:`-anchored filter block 1 uses, with the same digit-and-length bound and
the same `10#` on every operand, and blanks `total` when the three disagree with it.

**Adapted, not applied verbatim — and the divergence is the point.** The finding
says to re-apply block 1's cross-check. Block 1's gate is `REVIEW_COUNTS_OK`, which
demands all four counts be numeric, because block 1 DISPLAYS all four and
`6 findings —  critical` is the half-filled line that rule exists to prevent. Block
2 displays none of them; it uses `total` alone, against the number of headings the
row parser matched. Applied verbatim, the all-four rule blanks a perfectly usable
`total: 5` on a review carrying no severity keys and SILENTLY DROPS an `unparsed:`
shortfall this step reports correctly today — trading a safe-direction over-report
for a silent under-report, which is the wrong way round and is the exact failure
class the `unparsed:` key was added to close. Only the CONTRADICTION ports: absent
counts are not a disagreement, because there is nothing to disagree with.

Driven against the shipped fence, not a mirror:

  consistent 1+1+1=3          -> total 3       CONTRADICTION total:10 -> withheld
  blocker: alternation        -> total 2       counts absent          -> total 5 (kept)
  leading zeros 01+01+01=03   -> total 03      one count absent       -> total 5 (kept)
  no findings block           -> withheld      non-numeric count      -> total 5 (kept)

Six regression tests drive the second markdown fence end to end through `runHook`
under bash, asserting on the rendered ledger. Negative controls fire in OPPOSITE
directions, which is what pins the narrowing rather than only the fix:

- revert the fence fix          -> `a contradicting findings block yields no
                                   shortfall` and the `blocker:` twin go red
- apply the VERBATIM all-four   -> `a total with NO severity keys still reconciles`
  prescription instead             and the partial/non-numeric case go red

Restored tree: 214/214.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* docs(#3829): state the one input that produces no unparsed shortfall

The reference page said the shortfall "is stated" whenever `total:` exceeds the
parsed headings. After the round-12 fix that is conditional, and a doc asserting
the unconditional form describes behaviour the step no longer has.

Names the boundary in both directions, because the narrowing is the part a reader
would otherwise get wrong: a `findings:` block whose three severities are all
present, numeric and do not sum to `total` produces no key — the same input on
which the console line already withholds the breakdown — while counts that are
merely absent, partial or non-numeric are not a disagreement and still reconcile
from `total` alone.

`docs/FEATURES.md` regenerated; `gen-features --check` green (182 features, 21
groups).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): close two holes the round's own adversarial pass found in its first attempt

Neither is from the maintainer's review. Both were found by the pre-push adversarial
pass over this round's own claims, which refuted them by execution.

**1. A malformed severity could suppress a real shortfall.** The sibling frontmatter
reads use `cut -d: -f2 | tr -d ' '`, and `tr -d` deletes INTERNAL spaces, so
`critical: 1 0` arrives as the perfectly numeric `10`. That is long-standing in those
reads — its mirror is pinned as a fixture from round 1 — and it was INERT in block 2
until this round made that block read the severities at all. At that point a repaired
number could satisfy the new sum test and suppress an `unparsed:` shortfall that is
genuinely owed. Driven, pre-fix: `critical: 1 0 / warning: 0 / info: 0 / total: 5`
against three parsed headings emitted no `unparsed:` key where `unparsed: 2` was
correct.

The three severity reads now trim the ends only, so an internal space survives into
the digit check and fails it — `_sum_ok=0`, nothing is suppressed. Fail-safe in the
only direction that matters: when the frontmatter is malformed the step declines to
suppress rather than trusting a repaired number.

**Scope, stated:** only the SUPPRESSION inputs are strict. `REVIEW_TOTAL`'s own read
still uses `tr -d ' '`, unchanged and identical to block 1's — narrowing it would
change the `unparsed:` computation itself, which is pre-existing behaviour and wider
than this round. So `total: 1 0` is still read as `10` by both blocks, as before.

**2. The repointed #4748 gate could not see a later rebinding.** Its derivation slices
stop AT the first anchored assignment, so inserting the canonical lookup and then
overriding it with `REVIEW_FILE="${_pd}/WRONG-REVIEW.md"` left every assertion green —
the slice pins a line, not the path the fence actually consumes. A new test pins the
whole file instead: the only `REVIEW_FILE=` bindings permitted are the canonical
lookup (exactly twice, once per fence, each being a fresh shell) and the identity
pass-through that hands it to the embedded node script as an env prefix.

Negative controls, both the adversarial pass's own mutations, against the restored
tree at 215/215 and 164/164:
- restore `tr -d ' '` on the severity reads -> the internal-space test reds
- insert the WRONG-REVIEW override after the lookup -> the rebinding test reds

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): one parser for every count block 2 reads, and widen the rebinding guard

A second adversarial pass, run against the first pass's own fixes, refuted three of
them by execution. Fixes to review findings are the class most likely to carry a new
defect, which is why that pass exists; all three were real.

**1. `cut -d: -f2` takes the SECOND FIELD, not the scalar.** So `critical: 1: junk`
arrived as the perfectly numeric `1`, and `1+0+0 != 5` was read as a contradiction
that SUPPRESSED a shortfall genuinely owed. The previous fix trimmed the ends but
still cut at the wrong place, so it closed the internal-space shape and left this one
open. `-f2-` keeps everything after the first colon; the malformed scalar stays
malformed and `total: 5` still reconciles.

**2. The deliberate asymmetry was wrong, and it FABRICATED.** The previous fix parsed
the severities strictly and left `total` lenient, on the reasoning that narrowing
`total` was out of scope. Driven: `critical: 5 0` with `total: 1 0` repaired only the
total to `10`, rejected the severity, skipped the contradiction check, and invented
`unparsed: 7` against three parsed headings. Both uniform policies behave sanely —
strict rejects the malformed total, lenient detects `50 != 10`. A field is either
trustworthy or it is not; parsing one leniently and its sibling strictly is the shape
that fabricates. Every count this block reads now goes through one parser.

**Scope, restated because it moved:** the previous commit said `REVIEW_TOTAL`'s read
was deliberately unchanged. That is no longer true and the reasoning behind it did not
survive contact — the asymmetry it protected is what produced the fabrication. Block
1's reads are still untouched; its own all-four gate runs over consistently-parsed
values, so it has no equivalent split.

**3. The rebinding guard missed an indented or exported assignment.** `^REVIEW_FILE=`
let both `  REVIEW_FILE=...` and `export REVIEW_FILE=...` through, and each executes
exactly like a bare one. The predicate now absorbs leading whitespace and an optional
`export` before the accept-list decides.

Driven after the fix, against three parsed headings:

  critical: '1: junk'  total: 5      -> total 5 kept, unparsed: 2 reported
  critical: '5 0'      total: '1 0'  -> total rejected, no unparsed key
  critical: '1 0'      total: 5      -> total 5 kept (unchanged)
  consistent / contradiction / blocker / leading-zero / absent — all unchanged

Negative controls, each the adversarial pass's own mutation, against 381/381:
- `-f2-` back to `-f2`            -> the second-colon test reds
- `total` back to lenient `tr -d` -> the fabricated-shortfall test reds
- an INDENTED rebinding           -> the rebinding guard reds
- an `export` rebinding           -> the rebinding guard reds

Residual, disclosed: a duplicate `critical:`/`blocker:` key is still resolved by
`grep -m1` taking the first match. Duplicate keys are invalid YAML and the same
first-match rule is long-standing in the sibling reads; detecting them is a wider
change than this round.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): one parser for the whole step, and prove a contradiction from a partial sum

A third adversarial pass, run against the second pass's fixes. Three findings, plus
one this round's own negative control caught afterwards.

**1. An absent severity still bounds the sum from below.** Counts are non-negative,
so a missing one can only ADD: when the severities that ARE present already sum to
MORE than `total`, the block disagrees with itself whatever the absent value is.
Requiring all three before comparing missed that — driven: `critical: 4`,
`warning: 4`, no `info:`, `total: 5` reconciled against a total the present counts had
already refuted. The comparison is two-armed now: EQUALITY when all three are known, a
LOWER BOUND when they are not. An UNDERshoot stays reconcilable, because that is
exactly what the absent count explains.

**2. Block 1 now uses the same parser, so the console and the ledger cannot
contradict each other.** Tightening block 2 first left the two fences disagreeing about
the same bytes. Driven: `critical: 1 0` with `total: 1 0` repairs to 10 and 10, which
SUM — so block 1 reported `10 findings — 10 critical, 0 warning, 0 info.` from a
`findings:` block containing no such numbers, while the ledger recorded three rows and
no shortfall. Block 1's reads move to `cut -d: -f2-` plus an end-trim; both fences now
take the countless arm on that input. The counts mirror moves with them — its whole job
is modelling the shipped pipeline, and it modelled the retired one.

This is wider than the review's finding and I want that visible: the finding was about
block 2 alone. But a disclosed divergence between a console line and a ledger is the
confusion this PR exists to remove, so it is fixed rather than documented.

**3. `REVIEW_FILE+=-wrong` executes and was missed.** The rebinding guard matched only
`=`; `+=` appends (driven: `REVIEW_FILE=good; REVIEW_FILE+=-wrong` prints `good-wrong`).

**4. My first negative control for (2) was VACUOUS, and that is the reason for the new
`BLOCK 1 withholds a breakdown built from REPAIRED counts` test.** Reverting block 1's
parser left the suite green: on every fixture that existed both parsers landed on the
same arm, so parity could not see the difference. A SELF-CONSISTENT repaired breakdown
separates them, and the test pins it directly rather than through parity.

**Correction to the previous commit's claim.** It said moving `total` to a strict read
"changes nothing for a well-formed review". That is false: `total:\t5\t` is valid YAML
(`yaml.parse` returns 5) which the old `tr -d ' '` rejected and the new trim accepts.
The change is an improvement, not a no-op, and the claim was the wrong shape.

Also from the third pass's MISSED: the fixtures exercised malformed `critical` and
`total` only, so they did not pin the four-field symmetry the fix claims. Every field
now gets every malformed shape.

Negative controls, against the restored tree at 383/383 (220 in the pipeline file):
- revert block 1's parser        -> the repaired-counts test reds (was vacuous; now fires)
- revert the overshoot arm       -> the absent-severity contradiction test reds
- a `+=` rebinding               -> the rebinding guard reds

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): finish the mirror update, and give status the same parser as the counts

A fourth adversarial pass. The important finding is that the PREVIOUS commit's mirror
update was HALF APPLIED, and the suite could not see it.

**1. The counts mirror was still on the retired parser.** That file carries TWO
helpers: `first`, used for `status:`, and `firstIn`, used for ALL FOUR counts. The
previous commit updated `first` and the comment above it, and left `firstIn` on
`split(':')[1].replace(/ /g,'')` — so the shipped block had moved to `-f2-` + end-trim
and the mirror had not, while the parity assertion stayed green.

It stayed green for the same reason this round's earlier negative control was vacuous:
on every fixture that existed, both parsers reach the COUNTLESS arm, so the rendered
message is identical and parity cannot see the divergence. Two fixtures now separate
them — `a self-consistent repaired breakdown` (10 == 10+0+0, so the retired parser
renders a full breakdown from a `findings:` block containing no such numbers) and
`tab-separated counts` (valid YAML the retired `tr -d ' '` made non-numeric). Reverting
`firstIn` reds both.

This is the same shape this PR's round-3 reply already recorded about itself: a fix
verified with a grep built from the strings just fixed. The region is checked by
reading it end to end now.

**2. `status:` kept the retired parser after the counts moved off it, and it is the
read where truncation costs most.** `cut -d: -f2` turned the valid YAML scalar
`status: clean:junk` into the bare `clean`, so an unusable status took the CLEAN arm
and suppressed BOTH the console report and the ledger. Driven, both parsers side by
side. The whole scalar matches no arm now, so the step reports. One parser for every
scalar this step reads, in both fences.

**Two claim corrections, no code change:**

- The previous commit implied block 1's console output was preserved for every
  well-formed review. It is not: `critical:\t1` is valid YAML that the retired
  `tr -d ' '` left non-numeric (countless form) and the trim now reads (full
  breakdown). That is an improvement, and the claim was the wrong shape. The
  `tab-separated counts` fixture pins it.
- "Block 1 and block 2 can no longer contradict each other about counts" was
  overstated. It is true of the PARSER, which is what changed. They can still differ
  when the body carries MORE findings than `total:` declares: the reconciliation
  reports a shortfall only, and the excess direction is deliberately clamped so a
  review under-declaring its own total cannot render `unparsed: -1` — pinned by the
  round-2 test `a total SMALLER than the rows is not reported as a negative shortfall`.
  That is pre-existing and out of this round's scope; stating it rather than widening
  scope again.

Negative controls, against the restored tree at 223/223 (387 across both files):
- revert the mirror's `firstIn`  -> both new parity fixtures red
- revert the `status:` parser    -> the clean-arm suppression test reds

`npm run lint:ci` exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): pin the reads to LC_ALL=C, so the parser cannot depend on the machine

A fifth adversarial pass. It refuted the claim that the shipped reads and their JS
mirror are equivalent, and the counterexample is a locale.

**The POSIX character classes are locale-defined, and glibc's C.UTF-8 disagrees with
both C and en_US.UTF-8.** Driven, same sed, same input, three locales:

    LC_ALL=C          clean<U+2003>  ->  clean<U+2003>   (kept)
    LC_ALL=C.UTF-8    clean<U+2003>  ->  clean           (trimmed)
    LC_ALL=en_US.UTF-8 clean<U+2003> ->  clean<U+2003>   (kept)

C.UTF-8 classifies U+2003 — and U+1680, U+2000-U+200A, U+205F, U+3000 — as BOTH
[[:space:]] and [[:blank:]]. So `status: clean<U+2003>` trimmed to the bare `clean`,
took the CLEAN arm, and silently suppressed both the console report and the ledger —
but only on machines whose locale said so. That is the same suppression the previous
commit fixed for `clean:junk`, with a machine-dependent trigger instead of a parse one.

`[[:blank:]]` is NOT the fix — it is locale-defined too, and C.UTF-8 puts U+2003 in it
as well. Nor is `[ \t]`: POSIX bracket expressions provide no escape at all, so `\t` there is a
backslash and a `t`. GNU sed's default reading of it as TAB is an extension, and the
same GNU sed asked for conformance shows the other reading on this host:

    sed -E          's/[ \t]+$//'  draft  ->  draft
    sed --posix -E  's/[ \t]+$//'  draft  ->  draf

A POSIX-conforming sed is therefore expected to truncate `status: draft` to `draf`.
That expectation is derived from POSIX plus the `--posix` demonstration above; it was
NOT driven against a BSD/macOS sed, because this host has none. The portable
fix is to pin the locale: under C the class is exactly {space, tab, NL, VT, FF, CR},
which is precisely what the mirror already spells out literally. The two now agree by
construction rather than by coincidence of the machine.

All 20 read sites (10 `grep`, 10 `sed`, both fences) are pinned. The mirror's four
anchors move from JS `\s` to the same literal class, closing the divergence in the
other direction — `\s` matches a U+2003 indent that the pinned `grep` does not.

**A second gap, found by this round's own control rather than by the reviewer.** The
mirror has two helpers, and the previous commit proved `firstIn` (the counts) was
pinned by a fixture. `first` (the status) was NOT: reverting it left the suite green.
Every pre-existing status fixture left both parsers on the SAME arm — `issues:found`
truncates to `issues`, which is no more `clean` than `issues:found` is — so the status
mirror could drift unseen, exactly as `firstIn` had. The fixture that separates them
is one where truncation FLIPS the arm: `status: clean:junk`.

That is the third time this round a mirror edit was invisible to the fixtures that
existed, and the question that finds it every time is: what input actually separates
the two versions?

**Two prose corrections in the step**, which had gone stale rather than wrong-headed:
the block-2 comment still said "the sibling reads use `tr -d`" after they had all been
moved off it, and the mirror's class comment claimed an equivalence it did not yet have.

Negative controls, each driven against the committed tree:
- drop LC_ALL=C from the shipped seds  -> 5 red, incl. both locale-invariance tests
- revert the `first` status mirror     -> `a status whose truncation would flip the arm` reds
- revert the mirror anchors to `\s`    -> `locale-invariant on a unicode-space indented count key` reds

The locale-invariance tests are the durable guard: the parity fixtures only run under
whatever locale the suite inherits, so they can catch this only on a machine that
already has the bug. These drive the same input under both locales and assert the
shipped fence does not care.

238/238 across the two pipeline files. `npm run lint:ci` exits 0, with
lint-workflow-shellcheck reporting 212 pre-existing findings and 0 new.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* fix(#3829): pin the awk selectors too, and assert the invariant instead of claiming it

A sixth adversarial pass, and it refuted the previous commit's central claim. That commit
pinned all 10 `grep` and all 10 `sed` reads to LC_ALL=C and then said the parser no longer
depends on the machine. It does: the `findings:` MAPPING SELECTOR is an `awk`, and both
copies of it were left unpinned.

`awk '/^findings:[[:space:]]*$/{f=1; next} f&&/^[^[:space:]]/{exit} f'` resolves its two
character classes through the ambient locale exactly as grep's and sed's did. Driven, on
`findings:<U+2003>`:

    fence 1, LC_ALL=C        Code review found issues.
    fence 1, LC_ALL=C.UTF-8  Code review: 1 findings — 1 critical, 0 warning, 0 info.
    fence 2, LC_ALL=C        TOTAL=''
    fence 2, LC_ALL=C.UTF-8  TOTAL=1

Under C.UTF-8 the opener matched and the mapping opened; under C it did not. The same
review rendered a breakdown on one machine and the countless message on another, with
every grep and sed already pinned.

**Why it was missed is the more useful part.** The previous commit's census counted
`grep` and `sed` sites and reported zero unpinned — because it SEARCHED FOR THE TOOLS IT
HAD JUST EDITED rather than for the tools that were there. That is the same shape as this
round's other three misses: a check built from the thing just changed cannot see what the
change forgot. So this commit does not just add the fourth and fifth pins; it replaces the
claim with an assertion the next edit cannot fool:

  `every locale-sensitive tool in the step is pinned to LC_ALL=C` walks the step file and
  fails on ANY unpinned `grep`/`sed`/`awk`, naming line and call. `cut -d: -f2-` and
  `tr -d '\r'` stay exempt, and the exemption is principled rather than residual: neither
  resolves a character class or a collation — one splits on a single ASCII byte, the other
  deletes one literal byte.

The mirror's block boundary moves to the same literal classes, for the same reason the
anchors did last commit — `/^findings:\s*$/` and `/^\S/` model neither pinned side.

**A correction to the previous commit's message, made in place.** It asserted that BSD sed
reads `[ \t]` as a literal backslash and `t`, stated as driven fact. It was not driven —
this host has no BSD sed. The claim is now stated as what it is: POSIX bracket expressions
provide no escape, GNU's TAB reading is an extension, and GNU sed asked for conformance
demonstrates the other reading here (`sed --posix -E 's/[ \t]+$//'` turns `draft` into
`draf`). The conclusion is unchanged; the evidence class was overstated.

Also measured while establishing that LC_ALL=C is safe for non-ASCII, and worth recording
because it makes the pin a strict improvement rather than a wash: on a REVIEW.md carrying a
single invalid UTF-8 byte, GNU grep under C.UTF-8 reports `binary file matches` and emits
nothing, blanking EVERY read; under C the value parses and is rejected on its merits. Valid
UTF-8 is untouched either way — the C space class is entirely bytes < 0x80, which no UTF-8
multibyte sequence contains, so the trim cannot split a character.

Negative controls, each driven and restored:
- unpin the four `awk` selectors  -> 3 red, incl. the new invariant test naming both lines
- revert the mirror block boundary to `\s`/`\S` -> both `findings:` opener tests red

241/241 across the two pipeline files. `npm run lint:ci` exits 0 — after it caught a real
defect in the new test itself: `split('\n')` on readFileSync content is banned here
(DEFECT.WINDOWS-CRLF-TEST-PORTABILITY), and it now uses `splitLines()`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): make the pin invariant see calls that are not piped

The invariant added in the previous commit keyed on `| grep|sed|awk`, which is true of
every call in the step today and is exactly the wrong thing to rely on. A guard written
around the shapes that happen to exist cannot see the shape a later edit introduces —
`awk '...' < "$f"` or `$(grep ...)` would have walked straight past it, which is the same
property that let the two awk selectors sit unpinned through a commit claiming the parser
was locale-independent.

It now blanks the PINNED calls and treats anything still naming one of the three tools on
a non-comment line as an offender, so the check is "every call is pinned" rather than
"every piped call is pinned". Comment lines stay exempt: the step's prose names unpinned
forms while explaining why they were retired.

Driven both ways against the invariant alone:
- inject a NON-PIPED unpinned `awk '...' < "$REVIEW_FILE"` -> reds (the old form did not)
- unpin the four piped `awk` selectors                     -> still reds (no regression)
- unmodified tree                                          -> green

241/241 across the two pipeline files, `npm run lint:ci` exits 0. No shipped behaviour
changes in this commit; it only widens what the test can see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): close the invariant's path-qualified hole, and state the limit it keeps

A seventh adversarial pass. It confirmed the shipped fix — census clean at 24 pinned calls,
whole-fence output identical under C and C.UTF-8 on every input built to separate them, and
both fences byte-identical on a well-formed review carrying `café 東京` and `naïve résumé` —
and then refuted the claim I made about the GUARD, not the feature.

`/usr/bin/awk '/[[:space:]]/{exit}' < "$REVIEW_FILE"` passed the invariant. The preceding-
character class shielded any match preceded by `/` or `.`, so a path-qualified call was
invisible. A path-qualified call is still a call; the class no longer shields either. The
substrings that motivated the exclusion are unaffected — `parsed`, `passed` and `awkward`
have a word character on one side or the other, so the boundary still rejects them.

**The rest of that finding is disclosed rather than fixed, deliberately.** `$AWK "$f"` and a
command name computed inside the embedded `node -e` block also evade the check, and they are
not closable by this mechanism: it scans text, not shell or JavaScript command structure. The
same pass that found them showed that widening the regex further only trades those false
negatives for false positives on quoted strings and awk program text. So the test now STATES
its boundary instead of implying it has none — the previous comment claimed a census "a future
edit cannot fool", which was exactly the kind of overclaim this round has been correcting.
What it catches is enumerated there, driven, along with which direction its false answers go.

Driven against the invariant alone, each mutation reporting its own substitution count:
- `/usr/bin/awk ... < "$f"` (the exact evasion) -> reds; before this commit it did NOT
- `env awk ... < "$f"`                          -> reds
- `LC_ALL=C.UTF-8 awk ... < "$f"` (wrong pin)   -> reds
- unmodified tree                               -> green

241/241 across the two pipeline files, `npm run lint:ci` exits 0. No shipped behaviour changes
in this commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): stop the guard's comment promising a closed list of what it misses

An eighth adversarial pass. It confirmed the delta was test-only and that both fences are
byte-identical across it — exit, stdout, stderr and rendered ledger — and then refuted the
comment again, with two more evasions: a command name fragmented in shell (`a''wk`), and an
executable command substitution on a physical line starting with `#` inside a multiline
quoted argument, which the comment exemption skips.

Both are real. Neither is the point. Three passes running have each found one more evasion of
a TEXT scan, which is the actual finding: **the list cannot be closed.** A comment that
enumerates residuals is false the moment someone is cleverer than the enumeration, and fixing
it by appending the newest example just resets the clock.

So the comment no longer claims an inventory. It says what this is — a regression guard against
the accident that has now happened twice in this round, a read added or edited without its pin
in a file where every other read has one — and what it is not: a proof. The examples are marked
as illustrations. The operative instruction is the one that survives any future evasion: treat
anything it reports as real, and never treat its silence as proof a new read is pinned.

No logic changed; the guard catches exactly what it caught before. 241/241 across the two
pipeline files.

`npm run lint:ci` exits 0 — after it caught this commit twice over, which is worth recording
because both were in prose I had just written to be careful:
- the previous message's example path `docs/grep.md` read as a genuine docs reference from this
  file, and lint-docs-guard-registration demanded a baseline entry for a path that does not
  exist. The illustration is now `bin/grep-wrapper`, and the reason is stated inline.
- an earlier commit's `split('\n')` on readFileSync content tripped the CRLF-portability rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): the guard's two directions are not symmetric, so stop saying they are

A ninth adversarial pass. It confirmed the delta before it was comment-only with the guard's
logic byte-identical, and that all six advertised forms are still caught — and then found the
one absolute the rewrite left behind.

The comment said "treat anything this guard reports as real" three lines above admitting the
guard produces loud false positives. Driven: `echo ok # grep is discussed` is reported, and
labelled `(unpinned)`, though it is prose and no unpinned read exists. Both sentences were
mine, in the same comment, written in the same edit that was supposed to remove overclaiming.

The instruction is now the accurate one, which is that the two directions are NOT symmetric:
a REPORT is cheap to adjudicate — read the line, a trailing comment or a path is obvious — while
SILENCE proves nothing, because the known evasions are silent and so is any evasion nobody has
thought of yet. Investigate every report; never read silence as proof a new read is pinned.

Comment-only. Guard logic untouched, `gsd-core/` byte-identical to cac648aab — three consecutive
passes have now confirmed the shipped behaviour unchanged. 241/241 across the two pipeline files,
`npm run lint:ci` exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): a report needs its context to adjudicate, not just its own line

A tenth adversarial pass, and the last one this round runs. It confirmed `gsd-core/` is the
SAME TREE OBJECT as at cac648aab (105b5ed9b) with the guard's caught and silent sets unchanged,
and refuted one more sentence of the same comment.

"A report is cheap to adjudicate (read the line)" is false, driven: the identical reported
physical line `  grep` is a COMMAND after `:` and an ARGUMENT after `printf '%s\n' \`. The guard
reports both, and the reported line alone does not distinguish them — the preceding line is what
settles it. The comment now says so, with that counterexample in it.

This is the fourth consecutive pass to find a defect in this one comment and none in the shipped
code, which is itself the result worth recording: the shipped fix has been frozen since cac648aab
and confirmed byte-identical by four passes, while the prose describing a best-effort text scanner
took four attempts to stop overclaiming. Writing an accurate description of what a heuristic does
NOT do turns out to be harder than the heuristic.

Comment-only; guard logic untouched. 241/241 across the two pipeline files, `npm run lint:ci`
exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* test(#3829): re-home the lookup's phase-id gate beside the step, not in a block upstream removed

#4781 (#4628) removed #4748's letter-axis work from `tests/nsegment-phase-grammar.test.cjs`,
including the `#4748 — the REVIEW.md lookup` describe block. This PR had four tests living in
that block, because that is where the gate was when #3829 moved the lookup out of
`execute-phase.md` and into the lazily-read step file.

Those four tests assert properties of THIS PR's step file, not of #4748's sites. Rebasing onto
the removal would have deleted them silently — the branch would still be green, with its own
coverage quietly gone. They move here instead, unchanged in substance, beside the step they
guard: an unrelated upstream revert can no longer take this PR's coverage with it.

One assertion did NOT come along. The old block also checked that `execute-phase.md`'s init
parse list names `padded_phase`; #4781 removed that field from the list, and the assertion is a
property of #4748's site rather than of this step. Carrying it here would only have pinned
someone else's revert to this PR.

The step never depended on that field in the first place — it computes PADDED itself, validating
PHASE_NUMBER for shape and traversal and padding the digit run through `10#` while carrying an
optional letter verbatim. That self-containment is why the removal costs this PR nothing but the
tests' address.

5 pass, 0 fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* chore(#3829): regenerate the derived artifacts the round-12 rebase invalidated

The base moved from 092d9256b to 003d982c8 (six merges) while this round was in flight, so the
branch was rebased and the derived state had to be re-derived rather than hand-merged.

`platform-conformance-tier.generated.cjs` regains `tests/configured-entrypoint-validation.test.cjs`,
added upstream by #4249. Resolving the conflict hunk-by-hunk in favour of this branch had dropped
that entry; `lint:generated-sync` caught it, which is what that check is for.

`compact-content-benchmark-baseline.json` carries the measured values at the new base rather than
this branch's stale pair: split "execute-phase" off 26200 -> 26242, on 23949 -> 23952
(8.59% -> 8.73%); aggregate off 108243 -> 108285, on 91595 -> 91598 (15.38% -> 15.41%). The
benchmark reports drift and exits 0 either way, so a stale baseline does not announce itself here
— it announces itself in CI.

**The growth acknowledgment is back, and the reason is worth stating.** Against the previous base
this PR left `execute-phase.md` a net -103 bytes: the extraction removed more than the dispatch
paragraph added, so the file ended up smaller than the base's copy and the
`Emitted-Drift-Ack-Growth` trailer became false and was dropped. #4781 then rewrote that file
upstream, and against the new base the same extraction nets +36 bytes (93421 -> 93457) — which is
the figure this PR originally reported at round 2. The size delta was never a property of this
change alone; it is a property of this change against whichever base it sits on, and it has now
been both signs in one round. The trailer was restored then. Round 14 rebased onto `029acd915`, where #4830's
re-land moved `execute-phase.md` again and the same extraction is -103 once more
(93564 -> 93461), so the trailer is false a second time and this commit no longer
carries it. Third sign flip, same reason each time.

`lint:generated-sync` and `benchmark-compact-content --check` both clean afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B9qJPShuor4FqtZSX25jks

* chore(#3829): regenerate derived artifacts after rebase onto next

The rebase onto `c9a5cc3e1` conflicted in the 19 install-tree goldens.
Those are generated, so they were resolved arbitrarily and regenerated
with their own producers (`npm run regen:derived`, then
`benchmark-compact-content.cjs --write`) rather than hand-merged — a clean
textual merge of a generated file attests the merge, never the content.

Reconciled per artifact against the base's own committed copy rather than
against the pre-regen tree, because the pre-regen tree is the arbitrary
resolution:

- every install-tree golden now differs from `c9a5cc3e1`'s copy by exactly
  one entry, `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`,
  which is this PR's own new step file;
- the compact-content benchmark baseline by the `execute-phase` entry
  (offTokens 26184 -> 26167) and the aggregate that sums it.

`docs/INVENTORY-MANIFEST.json` was in the at-risk set but regenerated
byte-identical, so it carries no change here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RJPv9LNaPV3Cd3NCpfFKGb

* chore(#3829): regenerate derived artifacts after rebase onto 029acd915

The rebase onto current `next` conflicted in two generated files — the
compact-content benchmark baseline and the platform conformance tier. Both
were resolved arbitrarily and regenerated with their own producers
(`npm run regen:derived`, then `benchmark-compact-content.cjs --write`)
rather than hand-merged: a clean textual merge of a generated file attests
the merge, never the content.

Reconciled per artifact against the base's own committed copy rather than
against the pre-regen tree, because the pre-regen tree is the arbitrary
resolution. Every differing key belongs to a file this PR actually touches:

- `tests/fixtures/install-tree/claude.json` differs from `029acd915`'s copy
  by exactly one entry,
  `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`, this
  PR's own new step file;
- `scripts/lib/platform-conformance-tier.generated.cjs` by exactly one
  entry, `tests/code-review-fix-pipeline-regression.test.cjs`, a test this
  PR adds;
- `tests/fixtures/compact-content-benchmark-baseline.json` by the
  `execute-phase` entry (offTokens 26344 -> 26280) and the aggregate that
  sums it, which this PR moves by editing `execute-phase.md`.

The other 18 install-tree goldens, `docs/FEATURES.md` and
`docs/INVENTORY-MANIFEST.json` were regenerated too and came back
byte-identical, so they carry no change here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* test(#3829): follow #4748's REVIEW.md-lookup gate to the step that now owns the lookup

Self-found while rebasing onto `029acd915`, not raised in review.

#4830 (`8a5166598c`) re-landed #4768's letter-suffix work on `next`, restoring the
`#4748 — execute-phase.md resolves the REVIEW.md path from init's padded_phase`
block in `tests/nsegment-phase-grammar.test.cjs`. That block had been removed by
#4781, which is why round 13 re-homed this PR's own four tests out of it. The
restored block anchors on a line this PR deletes:

    expected 1 line(s) containing "REVIEW_FILE=\"${PHASE_DIR}/${PADDED}-REVIEW.md\"", found 0

It fails at the describe level, so all four of its tests go with it. It was green
before this rebase only because the base did not carry the block yet.

Putting the line back is not available. #3829 moved the lookup into the lazily-read
step file because `execute-phase.md` did not fit under ADR-857's frozen pre-phase-6
ceiling (93600); the parent is at 93461, and the three lines this gate anchors on
cost 183 (measured, not computed: `git show 029acd915:... | sed -n '1168,1170p' | wc -c`). The block would also be dead code — the step performs the lookup.

So the gate follows the lookup. Two of its four assertions are properties of the
lookup and are re-pointed at the step file: that no fence hands PHASE_NUMBER to
`printf "%02d"`, and that both lookups are preceded by a PADDED binding that pads
the digit run through `10#` and carries the letter verbatim. The regression control
on the lookup line itself comes along, now over both fences. The init-parse-list
assertion stays on `execute-phase.md`, which still names `padded_phase`.

Two assertions do NOT come along, and they are the two that were properties of the
INLINE site rather than of the lookup: the `PADDED="{padded_phase}"` literal binding
(the step derives PADDED itself, validating PHASE_NUMBER for shape and traversal
first), and the composition run over the three live lines. The step's executable
coverage — a composition run plus a padding-agrees-with-the-canonical-normalizer
matrix over letter ids — already exists in `tests/code-review-pipeline-regression.test.cjs`
under "#3829 — the step's REVIEW.md lookup resolves a letter-suffixed phase without
a shell re-pad". Mirroring it here would be a second implementation of one grammar.

That coverage is NOT equivalent, and the difference is worth stating rather than glossing. At its
original site #4748's gate was a DATAFLOW pin: `execute-phase.md` bound init's own
`{padded_phase}`, so the lookup could not disagree with the canonical normalizer because it never
computed anything. The step reconstructs the value in shell, so that pin is not available at this
address and AGREEMENT with the normalizer is what replaces it. A follow-on commit adds a
fast-check property asserting that agreement over generated ids, because the existing 13-shape
matrix samples 2 of 26 letters and cannot see a divergence outside its own points.

One consequence is disclosed rather than absorbed: `padded_phase` is now parsed but unused in
`execute-phase.md` (`:95`). Removing it from that parse list is #4830's call on its own site, not
this PR's, so the assertion that it is still named stays.

#4748's property is unchanged: a letter-suffixed phase resolves its own REVIEW.md,
and an already-padded `08` does not read as octal.

Negative-controlled rather than asserted. Against the shipped step: 162 pass, 0 fail.
Dropping `$_let` from both PADDED bindings reds "every lookup is preceded by a PADDED
binding that carries the letter run"; dropping `10#` reds it too. The step file was
restored byte-identical after each control.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* test(#3829): assert the step's padding against the canonical normalizer by property, not by 13 points

From this round's own pre-push adversarial review, not from the maintainer's.

The prior commit re-points #4748's REVIEW.md-lookup gate at the step that now owns the lookup. The
review's finding was that this is not coverage-equivalent, and it is right: at the original site the
gate was a DATAFLOW pin — `execute-phase.md` bound init's own `{padded_phase}`, so the lookup could
not disagree with the canonical normalizer because it never computed anything. The step reconstructs
the value in shell, so agreement with the normalizer is what has to replace the pin.

That agreement was already asserted, but over a 13-shape matrix. Its words: "future canonical
grammar changes could therefore diverge without this gate detecting them." Correct — the matrix
samples 2 of 26 letters and a bounded set of segment shapes, and this PR has twice been told that a
generator which cannot reach the interesting input is a fixture with extra steps (round 3's
SOURCE_CELL, round 5's title generator). Same defect, third address.

So the agreement is now a property over generated ids: a digit run inside the step's own 8-digit
bound, an optional single A-Z, and up to two dot segments, asserted equal to
`normalizePhaseName(id)` through the shipped shell derivation of BOTH fences. Milestone `N-N` forms
are outside the step's accepted domain and are asserted nowhere here rather than silently passed.

Negative-controlled on two mutations, and the second is the one that justifies the property rather
than the matrix:

  %02d -> %03d                      matrix RED,   property RED
  drop the letter when outside {A,B} matrix GREEN, property RED

The second is the added coverage, demonstrated rather than argued: the matrix is structurally
unable to reach a letter it does not enumerate. The step file was restored byte-identical after
each control.

numRuns is 25 — each case spawns bash twice through the process seam, once per fence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* fix(#3829): pad the phase id as a string, so the step agrees with the canonical normalizer

Found by the property the previous commit added, on its second adversarial pass. This is a real
divergence in shipped behaviour, not a test-only correction.

`normalizePhaseName` (src/init.cts, via gsd-core/bin/lib/phase-id.cjs) left-pads a phase id's digit
run to a MINIMUM of two and otherwise PRESERVES it — `padStart(2, '0')`. The step re-derived the same
value arithmetically, `printf "%02d" "$((10#$_dig))"`, which does not preserve: it collapses every
leading-zero run longer than two.

  id          normalizePhaseName   the step (before)
  8           08                   08
  08          08                   08
  008         008                  08     <-- diverges
  0008        0008                 08     <-- diverges
  00000008    00000008             08     <-- diverges
  0008A       0008A                08A    <-- diverges

Consequence: for such an id the gate resolves `08-REVIEW.md` while init emits `008-REVIEW.md`, so it
finds no review and says so — advisory, and therefore silent. That is the failure class #4748 exists
to close, reached by a different road: not a letter this time, but a leading-zero run.

The fix is a string pad that implements `padStart(2, '0')` exactly, at both fences:

  case "${#_dig}" in 1) PADDED="0${_dig}${_let}${_sub}" ;; *) PADDED="${_dig}${_let}${_sub}" ;; esac

It is strictly less machinery than what it replaces. `10#` existed only to stop bash reading a
leading zero as octal inside `$(( ))`; with no arithmetic there is no octal hazard to guard, so the
remedy is retired rather than kept. `scripts/lint-phase-id-drift.cjs` — which exists to flag
unsanctioned `printf "%02d"` re-pads of phase-carrying variables — is green, and now has one less
re-pad to tolerate.

Two dependent assertions move with it, and both are now stated as the PROPERTY rather than as one
spelling of the remedy: the round-14 gate in `tests/nsegment-phase-grammar.test.cjs` asserts the
binding does no arithmetic and carries both `${_dig}` and `${_let}`, and the regression pin in
`tests/code-review-pipeline-regression.test.cjs` follows the new form.

The property's generator is widened in the same commit, because its first cut could not have found
this: it built the digit run with `String(fc.integer(...))`, which can never produce a leading zero,
so it had silently LOST the `08`/`09` coverage the 13-shape matrix beside it already had. The run is
now generated as a digit string, and segment depth goes to four (the repo exercises `1.2.3.4`). The
step's grammar is unbounded in depth; four is a stated bound, and it is this property's residual.

Negative-controlled, and the matrix is the control's control — it stays GREEN on both:

  revert to the arithmetic pad          matrix GREEN, property RED
  drop the letter when outside {A,B}    matrix GREEN, property RED

The step file was restored byte-identical after each. 407 pass / 0 fail across both test files;
lint:ci, gen-features --check, gen-platform-conformance-tier --check, benchmark-compact-content
--check and gen-install-tree-fixtures all clean, with no regenerated artifact moving.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* test(#3829): hold #4748's gate by execution, and retire the comments the string pad made false

Third adversarial pass on this round. Two findings, both fair, neither a correctness defect.

FIRST — the gate pinned a SPELLING, not the property. Its objection was concrete: an equivalent
multi-line string pad would have failed a regex that matches one `case ... esac` line. That is a
false-positive generator, and a guard that false-fires is a guard that gets deleted.

The split is now honest about what each layer can hold. A STATIC gate can hold the DEFECT SHAPE —
no arithmetic in the binding above each lookup — and that is all it asserts. Correctness is held by
EXECUTION: the base gate regains a composition test that runs the shipped derivation slice of both
fences against `normalizePhaseName`, over `3A 8 9 08 008 0008A 23A.1.2`.

That composition test is the one I removed two commits ago, and removing it was the weaker call. At
#4748's ORIGINAL site the gate could be static because the property was a literal binding of init's
own `{padded_phase}`; nothing could disagree, because nothing computed. At this site the step
derives the value, so the property is behavioural and only execution holds it. `008` is in the list
because it is the case the arithmetic pad got wrong and no prior fixture covered.

Controlled three ways, and the middle one is the finding being answered:

  arithmetic pad restored            4 fail   caught
  EQUIVALENT multi-line string pad   0 fail   no false positive
  drop the letter outside {A,B}      1 fail   semantic drift caught

SECOND — the step carried comments the fix had made false, in four places. Two explained `10#` as
part of the live phase derivation; two justified the eight-digit bound by bash integer overflow.
Neither described the code any more. They are rewritten to current truth rather than annotated,
because a fragment carries no supersession marker and overturned prose reads as canon:

  - the octal rationale now says the pad performs no arithmetic and needs no `10#`, and notes that
    `10#` survives in this step only on the severity COUNTS, which really are numbers being added;
  - the length bound now states that overflow is unreachable since the pad stopped converting, and
    that the bound stays for the reason it always also had — every component is interpolated into a
    filename, and filesystem components are finite.

408 pass / 0 fail across both files. lint:ci, lint-phase-id-drift, gen-features --check,
gen-platform-conformance-tier --check, benchmark-compact-content --check and the
prompt-injection-scan security suite are all clean, and no regenerated artifact moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* docs(#3829): correct four comments, including one this round's own rewrite got wrong

Fourth adversarial pass. Comment-only; no code, no test logic, no regenerated artifact moves.

ONE OF THESE IS MY OWN ERROR, introduced two commits ago. Rewriting the length-bound rationale, I
replaced the dead integer-overflow justification with "the bound stays because every component is
interpolated into a filename and filesystem components are finite". That is false, and it was driven
false: the bound is PER SEGMENT, every segment is joined into ONE filename component, and depth is
unbounded. Thirty 8-digit segments yield a 278-character PADDED and a 294-character name against a
NAME_MAX of 255. Replacing a dead rationale with a wrong one is worse than leaving the dead one, so
the comment now states what the bound actually does and names the composite-length gap as a residual
of this validator that predates the pad change. It is not fixed here; it is stated.

The other three are stale rather than wrong:

- `overflow guard` named the length check in two places. Nothing overflows any more -- the pad does
  no arithmetic -- so it is the digit bound, and is called that.
- The `#4748` block header in `tests/nsegment-phase-grammar.test.cjs` still said the step pads
  through `10#`, still said the composition run did not survive the move, and still said the
  executable coverage was "cited rather than copied" -- while the composition test sat twenty lines
  below it. All three were true when written and none survived this round. The header now records
  why a STATIC assertion cannot hold a BEHAVIOURAL property, and that the PR's fast-check property
  is a different instrument over the same contract rather than the same test twice.
- The severity-count comment said a base-inference failure "takes the whole advisory step down under
  `set -e`". It does not: the arithmetic sits inside an `if` condition, a TESTED context, where
  `set -e` is inert, and the consistency check is SKIPPED instead -- which the regression suite
  already records. `10#` stays; only the account of what it prevents is corrected. This one predates
  the round and is corrected because it is adjacent and factually wrong, not because it blocked
  anything.

408 pass / 0 fail across both files; lint:ci, lint-phase-id-drift, gen-features --check,
benchmark-compact-content --check and the prompt-injection-scan security suite all clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nAJFVZcCZHtSgYmWM5WNP

* chore(#3829): regenerate the compact-content benchmark baseline after the rebase onto fac0e9de8

The rebase onto current `next` conflicted on this generated fixture, as it has in every
recent round: the base regenerates it for its own token deltas and this branch regenerates
it for `execute-phase.md`'s, so both sides rewrite the same keys. Resolved arbitrarily
during the replay and regenerated with its own producer
(`scripts/benchmark-compact-content.cjs --write`), never hand-merged.

Key-level drift against the base's committed copy is exactly two entries and both are this
PR's own: `splits.execute-phase` (the workflow this PR edits) and `aggregate`, which is the
sum over the splits and therefore moves whenever any split does. Zero foreign keys moved.

The full generator sweep was re-run after the replay -- gen-features, gen-inventory-manifest,
gen-platform-conformance-tier (both targets), gen-install-tree-fixtures and the benchmark --
and this fixture is the only artifact that moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V5xnzzipRHPHq8Z3doQvn3

* chore(#3829): regenerate the compact-content benchmark baseline after the rebase onto b956bb7c6

`next` moved again while this round was running -- #4902 landed at 06:27Z and touches this same
generated fixture -- so the branch went back to CONFLICTING within minutes of the previous push.
This is the second rebase of the round, not a correction of the first.

Resolved arbitrarily during the replay and regenerated with its own producer
(`scripts/benchmark-compact-content.cjs --write`), never hand-merged. Key-level drift against the
new base's committed copy is again exactly `splits.execute-phase` and `aggregate` (the sum over
splits) -- zero foreign keys.

The full generator sweep was re-run after this replay as well; this fixture is the only artifact
that moved. Build inputs were untouched by the base range, so the lane's existing build stands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V5xnzzipRHPHq8Z3doQvn3

* fix(#3829): refresh the launcher preamble in place, without hoisting it above the guard

Found by this round's own post-rebase validity gate, not by review: a reader flip-test over the 85
tests that read the two at-risk workflow files was clean before the replay and failed after it, on
`runtime-launcher-parity (#373)` invariant (B).

The cause is a real interaction. #4902 landed on `next` mid-round and rewrote the canonical launcher
preamble; invariant (B) counts occurrences of that exact snippet, so this step's older copy matched
zero times even though it sat in the right place. The remedy the invariant names --
`node scripts/sync-runtime-launcher.cjs` -- fixes the count but also HOISTS the preamble to the top
of the block, and that is wrong here: it moved the shim ahead of the status guard, and six of this
PR's own tests exist to pin that ordering (`runDispositionGuard` asserts the block opens with its
guard, then the shim). Running the tool verbatim turned one red into seven.

So the preamble text is refreshed to the current canonical snippet IN PLACE, at the offset it
already occupied. Both constraints hold at once, verified by execution rather than by reading:
invariant (B) sees exactly one canonical occurrence and it precedes the first `gsd_run` call, while
the guard still opens the block (shim at offset 19897 of the second fence, and the test wants > 0).
`runtime-launcher-parity` + `code-review-pipeline-regression` together: 282 tests, 281 pass.

Not fixed here, and not ours: `(K2) end-to-end: the resolved local tool honors
git.allow_default_branch_commits (#4834)` -- the one remaining failure -- fails identically on a
detached worktree at pristine `b956bb7c6` carrying none of this PR's content (37 tests, 36 pass,
same single failure). Reported rather than chased.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V5xnzzipRHPHq8Z3doQvn3

* fix(#3829): record the shortfall when NO finding in a review parses

Found by this round's own pre-push adversarial review, not by the maintainer -- and it is the
same defect class round 2 raised as Blocker 4, surviving in the one corner that round's fix did
not reach.

The unparsed reconciliation exists so a finding the CR|BL|WR|IN heading parser cannot match is
SURFACED rather than dropped. It reported faithfully whenever SOME findings parsed. It reported
nothing at all when NONE did, on a phase with no prior ledger and no fix report: a REVIEW.md
declaring two Criticals, both written under a prefix the alternation does not carry, produced no
ledger, no console line and no diagnostic. That is precisely the silent drop this reconciliation
was added to close, reachable exactly where the evidence is weakest -- the run in which not one
finding was understood.

Two exits discarded it, and the first one is the one that actually fired. The shortfall was
derived beside the render, while the exit that stands down for "nothing to record" keys on
order.length and sits ~200 lines earlier; it returned before the value existed. The later
rows.length exit had the same hole but was unreachable for this input. So the derivation moves
above the earlier exit -- order is final from the heading walk and never grows again, so the value
is unchanged -- and both exits now decline to fire while a shortfall is outstanding. The result is
a zero-row ledger carrying an unparsed key: an honest record that the review declared findings and
none of them were understood, which is strictly better than the file not existing.

Scoped, not removed. A genuinely clean review is untouched: a declared total of 0 is not greater
than order.length, so unparsed is 0 and both returns still fire exactly as before. The new test's
companion pins that, and it is why the exit was relaxed conditionally rather than deleted.

Driven at every step rather than reasoned about. Before the fix, three cases through the shipped
script: two unmatched findings on a first run wrote NOTHING; two unmatched plus one matched
reported unparsed: 1; two unmatched against an existing ledger reported unparsed: 2. Only the
first was silent, which is why the mechanism read as covered. After the fix the first renders a
ledger with unparsed: 2 and names it on the console, and the other two are byte-unchanged.

The new regression test was negative-controlled against the pre-fix step file and goes red there
(1 pass / 1 fail over the pair); post-fix both pass. Its companion clean-review control is green
on both sides, so the pair is not passing by accident.

The PR's six test files: 554 tests, 554 pass, 0 fail, 0 skipped. lint:ci exits 0. The generator
sweep still produces no drift.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): pin the SECOND exit that discarded the shortfall, and prove the pair kills it

Self-found by the round review of the previous commit, not by the maintainer: that commit added
the !unparsed conjunct to BOTH exits but tested only one of them. Deleting the later one left
every new test green -- a surviving mutant, which is coverage in name only.

The two exits are reached by different inputs, which is why one fixture cannot pin both. The
earlier exit stands down the moment a fix report exists, so an input carrying one sails past it
and lands on the later rows.length return. The fixture therefore needs a fix report that
contributes NO row: an id the alternation CAN match becomes a carried row, rows.length is 1, and
the later guard never decides. The first draft of this test used CR-99 and was vacuous for exactly
that reason -- it passed with the guard deleted. It names SEC-03 now.

Mutation-controlled in all three directions, since a test that kills no mutant pins nothing:

  earlier exit loses !unparsed   -> test 1 RED,  test 2 green, test 3 green
  later exit loses !unparsed     -> test 1 RED,  test 2 RED,   test 3 green
  both exits neutered (over-fire)-> test 1 green,test 2 green, test 3 RED
  unmutated                      -> all three green

Every test kills at least one mutant and no mutant survives all three, so the pair covers both
roads to the drop and the clean-review control covers the over-fire the relaxation could have
introduced.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): close the fourth cell — the later exit's own over-fire

Self-found again by the round review, which drove the mutant set rather than trusting the matrix
the previous commit asserted: neutering ONLY the later exit survived all three tests. The previous
commit's matrix was accurate and incomplete, which is the more dangerous shape -- it reads as a
closed argument.

The guards form a 2x2 and only three cells were pinned. T1 pins the earlier exit's under-fire, T2
the later exit's under-fire, T3 the earlier exit's over-fire. Nothing pinned the LATER exit's own
over-fire, and it is reachable: a fix report -- even one naming no matchable id -- makes the
earlier exit stand down, so control arrives at the later exit carrying a genuinely clean review and
no shortfall. Neuter that exit and a phase with nothing to report grows a zero-row ledger reading
"0 of 0 finding(s) open", with all three earlier tests green through it. T4 is that case.

Five mutants driven over the four tests:

  earlier exit loses !unparsed        -> T1 RED
  later exit loses !unparsed          -> T1 RED, T2 RED
  both exits -> if(false)             -> T3 RED, T4 RED
  ONLY later exit -> if(false)        -> T4 RED          (the survivor this commit kills)
  ONLY earlier exit -> if(false)      -> all four green

The last one is reported as an EQUIVALENT mutant rather than an open cell, and the distinction is
the point of stating it: the earlier exit's over-fire condition is a strict subset of the later
exit's -- it adds only fixReports.length === 0 -- so removing it is subsumed and changes no
observable behaviour. A test cannot kill a mutant that does not alter output, and pretending
otherwise would mean writing one that asserts on internals.

The PR's six test files: 556 tests, 556 pass, 0 fail, 0 skipped. lint:ci exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* fix(#3829): keep the embedded record-builder inside a Windows command line

Block 2 runs the disposition record-builder as `node -e "<script>"`, so the entire
script is a single argv entry. Windows caps a command line at 32767 characters
(CreateProcess) and Node surfaces the overflow as ENAMETOOLONG from spawn -- the
process never starts. Linux's ~2 MB ARG_MAX cannot see that cliff at all.

The previous fix (bd83bc88e, hoisting the shortfall derivation above the earlier
exit) grew the extracted script from 31941 to 33353 characters. There were 826
characters of headroom; it spent 1412. On the next CI run ubuntu and macOS stayed
green and `conformance test (windows-latest, 24, shard 1/3)` went red with 83
failures -- 79 reporting `spawn_failed` out of runShippedDisposition, the other 4
asserting on a ledger that was never written. That same shard was SUCCESS at
596968aa0, the head before that commit.

Measured on native Windows (node v25.2.1), bisected: the largest `-e` argument that
still spawns is 32728 characters; 32729 fails. Both payloads driven directly:

    pre-fix   33353 chars  ->  SPAWN-FAIL ENAMETOOLONG
    post-fix  15916 chars  ->  SPAWN-OK

Note the shape of it: the fix that makes this gate report a silently-dropped finding
was itself silently dropped on Windows, because the whole script stopped launching.

What changes here is placement -- not content, not behaviour. 17 long rationale
comment blocks move out of the quoted payload into a new "Design notes for the
embedded record-builder" section in this file's prose, each anchored to the code line
it preceded so the pairing survives the move. Comment runs shorter than five lines
stay inline, where adjacency is cheap. Nothing is deleted.

    extracted script  33353 -> 15916 chars (16812 under the measured limit)
    executable code   byte-identical at 10853 bytes, verified by diffing the payload
                      with all comment lines stripped from both sides
    this file         78542 -> 79538 bytes -- the prose moved, it did not grow

Verified green: this PR's own test file at 250 tests across all 36 suites, 0 fail;
changeset-lint, docs-lint, default-flip-documentation, lint:ci, gen-emitted-baseline,
workflow-size-budget (133 tests), and lint-workflow-shellcheck (204 pre-existing
findings, 0 new).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): pin the embedded script under the Windows command-line budget

Nothing guarded the size of the `node -e` payload, so the regression the previous
commit fixes was invisible to every Linux gate and surfaced only as 83 Windows
failures that said `spawn_failed` and never mentioned length. Without a guard the
next addition to the script re-breaks Windows exactly the same way, and finds out
the same expensive way.

The test asserts the extracted script stays under a 24 KiB budget -- the 32767
CreateProcess cap less roughly 8 KiB of deliberate headroom, so the script has
somewhere to grow before this fires.

It is bounded from BELOW as well, and that half is the point: a pure length
assertion passes when the extractor returns '', which is exactly what a moved fence
or a renamed delimiter would produce. A guard that reports a comfortable 0 bytes is
the vacuous-oracle shape. The lower bound makes a broken extractor fail loudly here
instead of reporting success.

Negative-controlled rather than assumed. Against the PRE-fix step file the test goes
red on the real payload (33353 > 24576); against the fixed file it passes at 15916.
A new test that has only ever been run against fixed code can be green because it
hit a branch the bug never lived on.

This PR's own test file: 250 tests, 250 pass, 0 fail, 0 skipped, across all 36
suites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

* test(#3829): say what the command-line guard does not prove

The round review's MISSED, adopted. The guard counts the extracted JavaScript, but the
32767 cap applies to the whole serialized command line -- executable path, quoting and
backslash escaping included -- so a quote-heavy payload expands on the way out, and the
32728 figure it cites is one host's measured threshold rather than the CI runner's.

The budget is unchanged and still correct; only the claim around it moves. 24576 leaves
roughly 8 KiB for both effects, which is a practical margin, not a proof that everything
the guard admits will spawn. Stating that in the test is cheaper than having a future
reader infer a guarantee the assertion cannot make.

Comment-only. No assertion, no budget and no behaviour changes.

This PR's own test file: 221 tests, 0 fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfaJkgoq87837ZqWtTxqqn

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-22 21:04:09 -04:00
Michel Moreira
76ef60ba25 enhance(#4836): prefer the graphify CLI for planner and researcher graph queries (#4874)
* enhance(#4836): prefer the graphify CLI for planner and researcher graph queries

The planner gets one knowledge-graph query per phase and the researcher two
or three, and that single shot decides which modules the plan treats as
related — and therefore how tasks are ordered into waves. It was spent on
the built-in reader, which seeds by case-insensitive substring match over a
node's label and description and then expands a hardcoded two hops. The
phase "User Authentication" seeds on `author`, `authoring` and
`unauthorized` with the same weight as `authenticate`, and when the
inflated payload exceeds `--budget` the trimmer drops edges by confidence
tier — so the highest-confidence tier can be discarded to fit a payload
that bad seeding inflated in the first place.

The graphify CLI is already a hard dependency of /gsd-graphify build, and
it ranks seeds (IDF weighting, trigram fuzzy matching) and applies context
filters before traversal. Both prompts now prefer it and fall back to the
built-in reader, branching on `command -v graphify` — the same degradation
shape the repo already uses for Context7 to ctx7. Binary presence is a
self-satisfying gate: a graph can only exist if the binary built it, so the
fallback covers edge cases (a CI checkout with a committed graph, a binary
since removed), not the common path. No new config key and no new tool
grant — both agents already have Bash.

The planner additionally runs `graphify affected`. The reference states its
own goal as "which subsystems may be affected by changes in this phase",
which is literally reverse traversal by relation; the built-in reader only
approximates it with undirected two-hop expansion and has no equivalent
verb, so `affected` is skipped on the fallback path.

`graphify status` now reports `graph_path`, the resolved absolute graph
location, on both the present and the missing branch. The CLI takes the
graph location as `--graph`, and the prompts must not re-derive
`.planning/graphs/graph.json` for it: that would point the CLI at a
non-existent local mirror in exactly the umbrella multi-repo setup
`graphify.graph_path` (#1825) exists to serve. For the same reason the
presence gate in both prompts is now the `status` call itself rather than a
bare `ls` of the default location, which was already blind to the override.

Known limit, stated in both prompts rather than implied: the two paths
return different shapes. `graphify query` emits prose and has no `--json`
flag; the built-in emits JSON with per-edge confidence tiers and
budget_met/budget_estimate. `--budget` also counts rendered output on one
and estimated payload bytes on the other (#2738) — same flag name,
different unit. Both are read by a model and nothing machine-parses the
injected block. With graphify absent from PATH the injected context is
byte-identical to before.

Closes #4836

Emitted-Drift-Ack-Growth: gsd-phase-researcher.md — the CLI-first branch, the reason it is preferred, and the output-shape warning are the deliverable; a pointer to a part would not be read at the decision point.
Emitted-Drift-Ack-Growth: gsd-planner.md — one sentence in the load_graph_context step pointer, so it stops naming the default graph path the reference no longer assumes.

* docs(#4836): record the CLI-first graph query in the planner and researcher entries

* chore(#4836): add changeset fragment

* enhance(#4836): name the full domain word in the planner's query-term examples

The reference's own example — phase "User Authentication" → term "auth" — is
the exact collision the CLI-first path exists to avoid, and it stays a
collision whenever the fallback path runs, since that path matches the term as
a substring of label and description.

* fix(#4836): surface graph_path on the unparseable-graph status branch

graphifyStatus() returned graph_path on the exists:true and exists:false
outcomes but not on the third, error, outcome (graph.json present but
unparseable). The planner/researcher prompts gate CLI-first dispatch on
exists, not on this outcome, so a corrupt graph file made them fall
through to the CLI-first branch with the literal <graph> placeholder and
no real path to substitute.

* docs(#4836): note graph_path's trust boundary at the --graph interpolation

graph_path is reflected verbatim into a double-quoted --graph argument the
agent executes via Bash. It comes from graphify.graph_path, a config
surface already trusted elsewhere, so this isn't a new trust boundary --
but it is a new injection site (no --graph flag existed on this call
before). One-line caution for anyone hardening this later.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-22 19:59:45 -04:00
Tom Boucher
b956bb7c67 fix(#4834): gate the launcher PATH arm on runtime identity and prefer config-home installs (#4902)
* test(#4834): failing-first launcher hijack regressions

* fix(#4834): gate the launcher PATH arm on runtime identity and prefer config-home installs

A gsd_run on PATH that cannot prove it is @opengsd/gsd-core (a foreign package, or a
release older than the runtime-identity verb) is no longer accepted by the launcher
snippet's PATH arm; resolution falls through to the hard error when no path-based
candidate matches. The runtime-config-home arm now precedes the PATH arm, restoring
the documented prefer-local-over-PATH order, so an installer-managed install wins
even against a genuine global. The 16-home probe list is factored into _gsd_homes()
and the identity gate into _gsd_id_ok(), keeping the per-copy delta at +141 bytes.

The files whose frozen ceilings had no headroom (gsd-executor, gsd-plan-checker,
gsd-verifier, gsd-planner, execute-phase, execute-plan) now load the resolver by
@-include from gsd-core/references/gsd-run-resolver.md (the onboard.md pattern)
instead of carrying an inline copy. Propagated to all other inlined workflow/agent
copies via scripts/sync-runtime-launcher.cjs; the resolver reference re-copied
byte-equal (parity B2); the hard-error text, docs/how-to/diagnose-a-foreign-gsd-tools.md,
and the CONTEXT.md launcher predicate updated to match (#4834); the quick-batch row-48
guard gains the canonical-preamble sweep carve-out (#4834, per its own #3730/#2529
precedents); the compact-content benchmark baseline regenerated.

Emitted-Drift-Ack-Growth: add-backlog.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: add-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: add-tests.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: add-todo.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ai-integration-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: audit-fix.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: audit-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: audit-uat.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: autonomous.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: check-todos.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: cleanup.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: code-review-fix.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: code-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: complete-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: debug.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: diagnose-issues.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: discuss-phase-assumptions.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: discuss-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: do.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: docs-update.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: edit-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: eval-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: explore.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: extract-learnings.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: fast.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: forensics.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: graduation.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-code-fixer.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-debugger.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-eval-auditor.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-eval-auditor.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-intel-updater.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-intel-updater.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-phase-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-project-researcher.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-project-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-research-synthesizer.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-research-synthesizer.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-ui-researcher.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-ui-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: health.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: import.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: inbox.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ingest-docs.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: insert-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: list-seeds.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: list-workspaces.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: manager.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: map-codebase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: milestone-summary.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: mvp-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: new-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: new-project.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: new-workspace.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: next.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: pause-work.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: plan-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: plan-review-convergence.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: plant-seed.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: pr-branch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: profile-user.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: progress.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: quick-batch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: quick.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: remove-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: remove-workspace.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: resume-project.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: scan.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: secure-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: settings-advanced.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: settings-integrations.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: settings.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ship.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: sketch-wrap-up.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: sketch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: smart-entry.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: spec-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: spike-wrap-up.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: spike.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: stats.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: sync-skills.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: thread.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: transition.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ui-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ui-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ultraplan-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: undo.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: validate-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: verify-work.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)

* docs(#4834): backfill the changeset PR number

* test(#4834): regenerate the compact-content benchmark baseline after the rebase

---------

Co-authored-by: sim <sim@local>
2026-09-21 02:27:36 -04:00
Tom Boucher
eea9247c93 enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select (#4912)
* enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select

Auto-mode used to auto-select a checkpoint:decision's first <option>
unconditionally, making a decision checkpoint's safety depend on option
presentation order rather than an authored choice. Add an optional
auto_select="<option-id>" attribute on the <task> tag: absent, auto-mode
now escalates to a human exactly like gate="blocking-human" does; present,
it names the option auto-mode selects; naming an id with no matching
<option id> is a hard structural-validation error at plan-parse time
rather than a silent fallback to the first option. gate="blocking-human"
continues to win over everything, unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4095): anchor auto_select/id attribute regexes past hyphenated decoys

An isolated adversarial review of the auto_select work found that both new
attribute regexes used \b as their left anchor, which is a word boundary,
not a "start of attribute name" boundary. A decoy attribute ending in the
same word (e.g. data-id="...") sitting before the real id="..." on the
same <option> tag matched first, silently corrupting the extracted option
id. Anchor on (?:^|\s) instead so only the real attribute name can match.
Adds a regression test reproducing the exact decoy-attribute shape, plus a
Unicode option-id test from the same review pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4095): register auto-select-attribute.test.cjs in the docs-guard lane

lint-docs-guard-registration failed: the new test reads docs/reference/
plan-md.md but was not registered, so a future edit to that doc could
silently desync from the test without the guard catching it on the PR
that changed the doc. Registered alongside its direct precedents
(precondition-element.test.cjs, reversibility-tagging.test.cjs), which
read the same file for the same reason.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4095): fit the decision bullet under execute-phase.md's frozen byte ceiling

The remote gsd-test run caught what local checks missed: execute-phase.md
carries a frozen ADR-857 Phase-6 byte ceiling (93600) with only 36 bytes
of headroom before this change, and the original checkpoint:decision
wording pushed it to 93772 (over the ceiling). Cascaded into failures in
phase6-capstone-conformance, execute-phase-completion-reconciliation,
claude-orchestration, and the compact-content drift-report test.

Also caught: tests/package-legitimacy-gate.test.cjs anchors a
"decision is conditional, not unconditional" safety check on the literal
phrase "first option" in the decision bullet — which #4095 deliberately
removes, since there is no more unconditional first-option pick. The test
was asserting an assumption this change intentionally makes obsolete;
re-anchored on tokens that still identify the bullet ('decision',
'auto-spawn') without weakening what the test actually verifies (the
bullet must still carry a blocking-human carve-out).

Also fixed a word-order mismatch between my own new test's regex and the
actual doc text it was asserting against (tests/auto-select-attribute.test.cjs).

Regenerated the compact-content benchmark baseline
(tests/fixtures/compact-content-benchmark-baseline.json) to match the new
byte counts.

Emitted-Drift-Ack-Growth: gsd-executor.md — +7 bytes (49138 -> 49145), from the auto_select carve-out added to the checkpoint:decision auto-mode bullet; already trimmed once to fit the 49152 hard cap.
Emitted-Drift-Ack-Growth: execute-phase.md — +12 bytes (93564 -> 93576), from the same carve-out in the orchestrator's decision bullet; kept 24 bytes under the frozen 93600 ADR-857 ceiling after two rounds of trimming for clarity vs. margin.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4095): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 23:47:01 -04:00
0xdhx
88b5775dc8 enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI (#4477)
* enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI

gsd-ui-auditor is chartered to audit interaction and handed a capture
driver with no interaction verb: `npx playwright screenshot` cannot
click, fill, hover, press or snapshot, so a hover state, an open menu,
a focus ring or a form's validation state never appears in its
evidence and every Experience Design finding degrades to code reading.

Implements the shape approved at triage, not a new capability:

- capabilities/ui/capability.json declares `workflow.ui_interaction_capture`
  (boolean, default false) on the capability that already owns the
  auditor (ADR-894 one-owner invariant); capability-registry.cjs regenerated.
- gsd-core/workflows/ui-review.md reads the key through gsd_run and hands
  it to the auditor as `interaction_capture:` in the spawn <config> block —
  the auditor carries no gsd_run resolver, so the key travels by value.
- agents/gsd-ui-auditor.md gains an anchored interaction-capture section
  AFTER the static block. With the key on and a Chrome binary resolved it
  starts the `chrome-devtools` CLI (chrome-devtools-mcp, floor ^1.8.0) on
  an --isolated profile, opens the dev URL the static block reached,
  takes the a11y snapshot for element uids, captures the baseline and a
  Tab focus-ring state, drives the UI-SPEC's interactive components, saves
  console output, and stops the daemon unconditionally. Key off, no dev
  server, or no Chrome: one status line, and the Playwright-only static
  path runs exactly as before — the static fence is untouched.

Needs only Bash: no MCP server, no tools: change. Chromium-only by
nature; Firefox/WebKit stay on Playwright. `wait_for` is MCP-only, so
readiness is polled through evaluate_script.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* test(#4223): bind the interaction-capture shape and containment

- manifest, generated registry, config schema and config-set/loadConfig
  all know workflow.ui_interaction_capture as a default-off boolean, and
  hand-written non-booleans fall to the slice default
- the orchestrator reads the key and hands it down; the auditor never
  grows a gsd_run dependency
- the static fence stays Playwright-only and the interaction fence
  chrome-devtools-only, so key-off is today's path
- the interaction fence runs under bash with a stub driver on PATH: key
  off / absent / no dev server / no Chrome invoke nothing; the happy path
  starts first and stops last on the [selected] pageId with the documented
  flags; a failed capture is removed and not counted; new_page and start
  failures still honour the stop-only-if-started rule; CHROME_BIN and
  CHROME_DEVTOOLS_MCP_VERSION overrides flow through
- docs/CONFIGURATION.md row shape; registered in the docs-guard lane

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* docs(#4223): document workflow.ui_interaction_capture and its how-to

- docs/CONFIGURATION.md: one row in the workflow.* table, default-off
- docs/AGENTS.md: the gsd-ui-auditor entry names the key and what the
  interaction-capture section adds, skips and never claims
- docs/how-to/enable-ui-interaction-capture.md: turn it on, read the
  `**Interaction captures:**` outcomes, what it does not do, turn it off
- docs/README.md: index the how-to beside live-DOM verification

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* chore(#4223): add changeset

Added-type fragment; pr: carries the issue number until the PR exists.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): use the /gsd:ui-review namespace form in the auditor's prose

Claude-facing source (agents/, workflows/) uses the /gsd:<cmd> namespace;
the hyphen form is retired there and the slash-command-namespace guard
rejects it. docs/ keep the hyphen form by convention.

Emitted-Drift-Ack-Growth: gsd-ui-auditor.md — #4223: the anchored default-off interaction-capture section (prose + one bash fence) appended after the static Playwright block inside <screenshot_approach>, plus one `**Interaction captures:**` line in each of the two report templates, one completion-checklist line and one Step-3 sentence. The static fence is byte-identical to next; nothing was removed or reordered.
Emitted-Drift-Ack-Growth: ui-review.md — #4223: a two-line config-get read + true/false normalisation in step 0 and one `interaction_capture:` line in the spawn <config> block with a three-line note on why the value travels by prompt. No step, gate, or dispatch shape changed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): per-run daemon session, bounded navigation, and step failures that count

Three findings from the pre-file adversarial review of the interaction
fence, folded in:

- `--sessionId <epoch>-<pid>` on every driver call. `start` restarts
  whatever daemon shares its session and --isolated isolates only the
  browser profile, so two concurrent audits — or an audit beside the
  operator's own CLI daemon — would otherwise stop each other. The CLI
  accepts hex and dashes only; the id is validated by the test stub.
- `new_page --timeout 30000`: the one verb that takes a bound, placed
  before every verb that does not, so a hung page is caught first.
- a failed take_snapshot or press_key now increments the failure count
  and is named on stdout; two clean screenshots can no longer read as
  `0 failed` after the step that gives the interactions their uids failed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): subshell-unique session id, CRLF-safe page-id parse, stale-snapshot removal

Second review round, both reviewers:

- session id is `<epoch>-<BASHPID>-<RANDOM>`: `$$` is inherited by a
  subshell, so two audits forked from one parent in the same second
  shared an id and could stop each other's daemon (driven by the reviewer)
- `tr -d '\r'` before the `[selected]` parse so a CRLF-emitting driver
  under Git Bash still matches the `$` anchor, and `|| true` on the
  assignment so a failed new_page cannot abort the block under
  `set -e -o pipefail` before the unconditional stop
- a failed take_snapshot removes any snapshot.txt it left or inherited
  from a reused directory, so stale uids never drive the interactions
- `<config>` placeholder is `{interaction_capture}`, lowercase like its
  `{phase_dir}` / `{padded_phase}` siblings — the block is a prompt
  template, not a bash heredoc
- how-to: the `not captured` row no longer claims the daemon started

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): check new_page's exit status before parsing its output; regression cases for the edges

Third review round:

- a new_page that prints a page line and then exits non-zero is a failed
  navigation, not a page id: the exit status is checked in an `if` before
  the output is parsed (driven by the reviewer against the previous
  `|| true`, which masked exactly that)
- regression cases for what the last two rounds added: CRLF driver
  output, a stale snapshot removed on failure, partial-output new_page
  failure, and the whole fence under `set -e -o pipefail` (both the
  failed-navigation path and the happy path)
- the harness whitelist gains `date`; the session-id assertion now
  requires all three parts, so a silently empty epoch cannot hide again

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): keep gsd-ui-auditor under the DEFAULT-tier size cap; changeset pr placeholder

- the three review folds pushed agents/gsd-ui-auditor.md to 25179 bytes,
  over the 24576-byte hard cap tests/agent-size-budget.test.cjs enforces;
  the interaction section's comments are tightened to the same content
  in fewer bytes (23559 now). No bash changed — the fence's own tests and
  the real-browser run are unchanged.
- .changeset/vivid-yaks-fly.md carries the policy placeholder `pr: 0`, which
  the post-create backfill rewrites to the PR's own number.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* chore(#4223): set changeset fragment pr to 4477

* test(#4223): compare the fence's status path with the separator the fence uses

On the windows-latest lane the happy-path case failed on `\interaction` vs
`/interaction` alone: the fence joins "$SCREENSHOT_DIR/interaction" with a
literal slash, and the assertion built its expectation with path.join. Every
other case in the file passed on that lane, including the CRLF and
errexit/pipefail ones.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* test(#4223): drop the inert file-header allow-test-rule marker

Review round 1 on #4477: the `source-text-is-the-product` marker sat at
line 2, outside no-source-grep's 8-line lookahead of every readFileSync
site (the first is ~60 lines down), so it suppressed nothing. It was also
unnecessary: every read in this file is a .md/.json path, which the rule
does not trigger on. Deleted rather than relocated — there is no site to
relocate it to. Negative control: `eslint` on the file is clean without it.

* chore(#4223): regenerate the platform-conformance-tier lists for the new test

Review round 3 on #4477. `next` gained chore(#4591)'s platform-conformance-tier
gate after this branch opened; its two committed lists must name every file
under tests/, and this PR's tests/ui-interaction-capture.test.cjs had never
been in them. Once the branch was updated against next the lists were stale
and three jobs went red on head 575667dd: lint-tests (gen-platform-conformance-tier
--check), conformance test (macos-latest) at 546 !== 547, and shard 1/3's
fragment-single-edit-propagation, which sees the same staleness as regen:derived
touching files beyond the fragment edit under test.

Regenerated with the repo's own generators, no hand-editing. The general tier
goes 546 -> 547 and the macOS tier 196 -> 197, each by exactly this one entry;
both --check arms are clean. Verified the red is this PR's own file and not
base drift: at upstream/next both generators report "list matches" (546 / 196),
and our committed copies were byte-identical to next's before this commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CBQTeGX1JYHF5DRWp4wvZ

* fix(#4223): bound, confine and trap the chrome-devtools driver fence

Round 4 — three findings in one fence, interleaved on the same lines, so one
commit:

- Every driver call is time-bounded. `cdt <ceiling> <verb>` runs the client as a
  background job in its own process group (`set -m`) under a watchdog that kills
  the whole group at the ceiling — TERM, then KILL two seconds later. One pid is
  not enough: npm forwards SIGTERM only to its direct child, so killing `npx`
  alone leaves the client holding the fence's stdout and a `$(cdt … new_page)`
  capture blocked past the ceiling (driven against a real npx tree by the
  round's adversarial review; the pid-only first cut of this commit had exactly
  that hole). The watchdog is an exec'd bash (`"$BASH" -c`), never a `( … )`
  subshell: a subshell inherits bash's saved copies of the caller's stdio (the
  fds ≥10 a function-level `>/dev/null` redirect leaves behind) and holds them
  open, so a runner waiting for EOF waited out the whole 60 s ceiling whenever a
  watchdog outlived its kill — measured as the intermittent 30 s test run the
  review flagged; 0/60 after. It polls the job's process GROUP (`kill -0 --
  -pgid`, every 0.1 s) and stands down by itself once the group is empty;
  nothing ever signals it. The group, not the leader pid: a child can outlive
  the leader while holding the `$(cdt … new_page)` pipe, and a leader-pid poll
  stood down at once and left the substitution open-ended (driven by the
  round's adversarial review at 6× the ceiling; a pgid cannot be reused while
  any member lives, which a bare pid can). The daemon `start` launches is
  spawned detached (its own session) and never in that group. Two platforms
  forced the never-signalled shape. Under bash 3.2.57 the earlier `trap … TERM;
  sleep & wait $!` form ignored its TERM in 3 of 300 fast calls and slept out
  the whole ceiling — CI's macos job hanging 30 s right after `start`. On Git
  Bash a signal to a watchdog still starting up hung the fence's `wait` for it:
  18 of 20 fence tests at the harness's 30 s cap in 3 of 3 full-file runs,
  while a fence slowed by xtrace, or three tests run alone, never hit it (a
  startup race; the mechanism is not pinned further). Polling: 0/300 slow calls
  and 0 orphaned sleeps under 3.2.57 and 5.2, the fence suite 20/20 in 3 of 3
  full-file runs on Git Bash 5.2.37 (fractional `sleep 0.1`: driven on GNU,
  msys and busybox sleep; BSD sleep documents it). A clock that cannot launch
  (`sleep … || exit 0`) stands the watchdog down rather than firing at once
  and killing a healthy call — by design that leaves a hung call unbounded,
  the pre-round-4 behaviour, instead of failing a healthy one. A hung call
  returns once its group is gone: at the ceiling, plus up to the 2 s
  TERM-to-KILL grace. The KILL after the grace is sent only to a group that
  is still alive: a pgid freed during the grace can be reused, and an
  unconditional KILL could hit an unrelated group (the round's review).
  `start` (npx fetch + Chrome launch) gets CHROME_DEVTOOLS_START_TIMEOUT
  (180 s), every verb CHROME_DEVTOOLS_STEP_TIMEOUT (60 s). timeout(1) is absent
  on macOS and this agent carries no gsd-tools resolver, hence a bash watchdog
  rather than either.
- --allowUnrestrictedPaths -> --workspace "$INTERACTION_DIR": the driver may
  write under the run's interaction/ directory and nowhere else. Relative, like
  every --filePath (unchanged from rounds 1-3): the daemon resolves both against
  one cwd (chrome-devtools-mcp 1.9.0 spawns it with cwd: process.cwd() and
  path.resolve()s both), and a relative path needs no dialect translation — an
  absolute `pwd -P` path is an msys path on Git Bash, which a Windows-native
  daemon cannot resolve (CI's windows conformance shard caught the first cut).
  --workspace is a 1.9.0 flag (absent from 1.8.0's `start --help`, verified),
  so the documented floor moves from ^1.8.0 to ^1.9.0, where
  --allowUnrestrictedPaths is deprecated.
- `stop` is owed by an EXIT trap after a successful `start`, not by position
  (it replaces any earlier EXIT trap — none exists in this file); the explicit
  call keeps it in order, a flag makes the trap a no-op afterwards, and only the
  shell that installed the trap may act: a subshell copy of the fence state
  carries CDT_STARTED=1 and, under a timing race CI's ubuntu job hit (reproduced
  locally at 3/40 under load: the second `stop` came from a subshell pid, never
  main), issued a second `stop`. The identity is `$(exec /bin/sh -c 'echo
  "$PPID"')`, not $BASHPID — macOS ships bash 3.2, where BASHPID does not exist
  and CI's macos conformance job showed the guard comparing empty to empty. The
  fence was driven under bash 3.2.57 for the injected-subshell, errexit
  failed-new_page, errexit failed-resize, hung-start, hung-new_page and happy
  paths. A failed resize_page is a counted failed step now, not the one bare
  command an errexit runner could abort on.

Prose in the section is tightened to pay for the mechanism: 23559 -> 24517
bytes against the 24576 DEFAULT-tier cap.

Tests: the stub driver hangs as a real child tree (sh waiting on a child that
holds stdout — never an exec), so a pid-only kill fails the new
aHungNewPageWhoseChildHoldsStdoutIsStillCutOffAtTheCeiling test (negative-
controlled: it blocks for the harness's whole cap on the old wrapper). A hung
start and a hung capture are cut off within ceiling + grace + slack and still
reach stop; an injected bare failure under errexit reaches stop through the
trap, exactly once; an injected subshell call of cdt_stop issues nothing; the
happy path issues exactly one stop; every driver call site names a ceiling and
the only bare $CDT is the wrapper's own spawn; the start line carries
--workspace with the capture directory, every --filePath lies under it, and no
code line carries --allowUnrestrictedPaths. A driver whose leader exits at
once while a child keeps holding the capture pipe is still cut off at the
ceiling (negative-controlled: a leader-pid poll blocks for the harness's whole
cap). A watchdog whose clock cannot launch leaves a 300 ms driver call alone
(negative-controlled: the trap form kills `start` in under 20 ms). The harness
EXPORTS its stub-only PATH — unexported, the exec'd
watchdog fell through to bash's compiled-in default PATH and never saw the stub
dir — and ships `sleep` there as an exec-wrapper script (portable to Git Bash,
pid-preserving).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj

* fix(#4223): gitignore gate covers the capture directory, and upgrades an existing file

Round 4 Blocker. The gate enumerated image extensions, so snapshot.txt (the
accessibility tree, with entered form values) and console.txt (which can carry
tokens) were committable by `git add .`. The gate now ignores `interaction/` as
a directory — the next artifact type is covered by construction — and it appends
whatever an existing .gitignore lacks instead of writing once. The write-once
form was the same defect one step later: every project that had already run an
audit would never have received the new pattern at all.

Tests run the gate fence under bash: a fresh file carries every pattern; an
image-only file from an earlier audit gains interaction/ and keeps its own
header without duplicating present lines; a second run appends nothing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj

* test(#4223): declare the interaction-capture anchor as a comment marker

The #4324 colon-token gate (slash-command-namespace) landed on next after this
branch was opened and reads `<!-- gsd:ui-interaction-capture -->` as an
unconvertible /gsd: command token. It is a section anchor of the same family as
gsd:live-dom-families and gsd:write-continue, so it is declared in
COMMENT_MARKER_TOKENS rather than renamed. Found by running the base-added
gates against the merged tree; CI at ca8d2508 predates the gate.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-09-20 05:45:15 -04:00
Tom Boucher
0977d0a475 fix(#4797): run-with-timeout's .cmd mediation folds onto projectSpawnInvocation — spaces in the shim path no longer break it (#4892)
* test(#4797): failing-first on Windows — a .cmd/.bat shim under a spaced path must run

* fix(#4797): run-with-timeout's .cmd mediation folds onto projectSpawnInvocation — spaces in the shim path no longer break it

The private /d /s /c argv-array copy quoted the shim path (any space in the
path) and cmd.exe /s stripped the first-and-last quote of the whole /c string,
so the pre-space fragment became the program name: exit 1, empty stdout, on
every repo whose path contains a space. The declared seam wraps the whole
command line in one extra quote pair with windowsVerbatimArguments — the
reporter verified the shape against the repro on Windows.

* chore(#4797): backfill changeset PR number (4892)

---------

Co-authored-by: sim <sim@local>
2026-09-20 05:08:22 -04:00
0xdhx
8a5166598c fix(#4830): re-land #4768's letter-suffix phase-id fix and its lint-phase-id-drift ratchets on current next (#4873)
* fix(#4830): re-land #4768's letter-suffix phase-id fix and its lint-phase-id-drift ratchets on current next

Commit 740ba0d8a (#4781) removed every change #4768 had merged for #4748:
the first-non-digit split at execute-phase.md's two arithmetic sites, the
init-emitted `padded_phase` the REVIEW.md lookup binds instead of
`printf "%02d"`, the canonical-grammar extractions in autonomous.md and
plan-review-convergence.md, the `.changeset/zesty-wolves-tumble.md`
fragment, and the three lint-phase-id-drift ratchets with their tests.
The guard and the code it guarded left together, so nothing went red.

This is a cherry-pick of 092d9256b onto current `next`, resolved against
the #4683 threat-id fields on execute-phase.md's Parse-JSON line, with the
changeset `pr:` reset to the placeholder and the compact-content benchmark
baseline regenerated against the current base.

(cherry picked from commit 092d9256b8)

Emitted-Drift-Ack-Growth: autonomous.md — restores #4768's canonical-grammar extraction and its explanatory comment for --from/--to/--only
Emitted-Drift-Ack-Growth: execute-phase.md — restores #4768's first-non-digit split at two arithmetic sites and the padded_phase binding for the REVIEW.md lookup
Emitted-Drift-Ack-Growth: plan-review-convergence.md — restores #4768's canonical-grammar phase extraction and its comment
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AX7LXxc3uAkGki6iaYiAMP

* chore(#4830): set changeset fragment pr to 4873

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-20 00:01:21 -04:00
Tom Boucher
6dcc0428dd fix(#4784): Xcode gates pass -project and derive the destination from available simulators (#4879)
* test(#4784): failing-first — Xcode gates must pass -project and derive the destination from available simulators

* fix(#4784): Xcode gates pass -project '$XCODEPROJ' and resolve the destination from available simulators, skipping loudly when none exist

Also surfaces the -collect-test-diagnostics never / workflow.test_gate_timeout
guidance on a 124 timeout (a real iOS suite measured ~600s of sysdiagnose
collection after the tests passed).

* test(#4784): reachability-honest test shapes — apostrophe-confused prose (the path the false positive actually reaches past the splitter) and the file's single-quote convention

* fix(#4784): review fold-ins — unanchor the simulator extraction (real simctl lines end in a state suffix), escape apostrophes for the bash -c layer, assign the command variables on the skip path, behavioral extraction test

The isolated reviewer's Finding 1 was critical: the first cut's $-anchored sed
matched ZERO real simctl lines (every device line ends in (Shutdown)/(Booted)),
so the gates would have skipped on every machine. The 46,454-test green could
not see it (content assertions do not execute the pipeline); the new behavioral
test runs the gate's own extraction against a real-format line.

* chore(#4784): backfill changeset PR number (4879)

---------

Co-authored-by: sim <sim@local>
2026-09-19 19:18:30 -04:00
Tom Boucher
11b3091df0 fix(#4738): record opencode's staged skills in the install manifest (#4847)
* test(#4738): opencode manifest must record its staged skills (failing first)

* fix(#4738): record opencode's staged skills in the install manifest

* test(#4738): use the centralized temp-dir helper in the manifest tests

* fix(#4738): retire the dead hostBehaviors vocabulary entry, tighten detector asserts, temper changeset

* chore(#4738): backfill changeset PR number (4847)

---------

Co-authored-by: sim <sim@local>
2026-09-18 04:52:32 -04:00
Tom Boucher
c9a5cc3e12 fix(#4683): detect cross-plan threat-ID duplicates before execution (#4828)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; two orthogonal agent reviews ran (isolated adversarial REQUEST-CHANGES with all six findings dispositioned, plus a bypass/consumer-lens APPROVE) and the sha-pinned bench passed 46362/0 on the merged head.
2026-09-17 19:06:55 -04:00
Tom Boucher
be1b76dddd fix(#4700): queue the headless mempalace mine on the palace lock (#4821)
* test(#4699): add failing-first coverage for skipping complete phases in next_phase

* fix(#4699): skip already-complete phases in the next_phase cascade

Both next-phase scans selected the numerically lowest phase above N
without consulting completion state, so completing a reopened phase
persisted an already-[x] phase as STATE.md current_phase while
roadmap.analyze correctly named the outstanding one (issue repro:
completing 2 with phases 1 and 3 already [x] returned next_phase 03).

The cascade collects the complete phase numbers from the roadmap
checkboxes (milestone-scoped, comparePhaseNum-deduped) and skips them in
both the disk scan and the roadmap scan; a [x] checkbox row and its
heading sibling both name a phase that is never next. Heading-only and
checkbox-less roadmaps behave exactly as before.

* test(#4699): align the negative-control expectation with the disk spelling

* test(#4699): pin the STATE.md persistence and the all-later-complete tail corner

Review findings: the regression never asserted STATE.md current_phase
(the issue's actual harm), and the all-later-phases-[x] corner
(is_last_phase true, next_phase null) was unpinned. A changeset fragment
is included.

* docs(#4699): backfill changeset PR number

* test(#4700): add failing-first coverage for the queued headless mine

* fix(#4700): queue the headless mempalace mine and surface skipped captures

The capture's mine ran in the foreground with no lock handling: MemPalace
wraps every mine in a per-palace lock, so any concurrent writer (two
phases finishing a stage at once, a git-hook refresh mining the same
palace) made it exit 1 (MineAlreadyRunning) and the onError: skip step
silently dropped the capture — unlost for CONTEXT/PLAN/SUMMARY files that
can be re-filed, unrecoverable for execute:wave:post problem-fix pairs.

The mine now queues via --daemon --background (MemPalace #2029: the daemon
holds a job refused the lock and runs it when the holder exits), and the
report step gains the queued and skipped outcomes per #4700's requirement
that a skipped capture never stay silent. Option 2 (write_routing.cli) is
unreleased at MemPalace 3.9.0; option 3 (retry) re-enters the same lock
race — both declined in the PR body.

* fix(#4700): queue the wave:post problems fragment's headless mine too

The issue names the execute:wave:post problem-fix pair as the
unrecoverable loss (no source file to re-file later); the
capture-problems fragment's headless mine ran foreground like the capture
capability's did. Same fix: --daemon --background, with the lock-deferral
rationale inline.

* docs(#4700): backfill changeset PR number

* fix(#4682): register the stale-reverification part in the capability registry

The new steps/ part is a shipped workflow file; gen-capability-registry
--check requires it in the committed registry.

---------

Co-authored-by: sim <sim@local>
2026-09-17 06:30:36 -04:00
Tom Boucher
2bfff17ff8 fix(#4682): route stale verification to the verifier regeneration path (#4818)
* test(#4682): add failing-first coverage for stale verification routing

* fix(#4682): route stale verification to the verifier regeneration path

The stale routing entry sent users to /gsd-verify-work — but verify-work
never rewrites VERIFICATION.md (its only write is the human_needed
canonicalization), so following the advice re-ran UAT, reached the same
stale check, and looped. init's projector and execute-phase's generic
next_command presentation both mirror this entry, so the dead end appeared
on three surfaces.

The stale entry now routes to execute-phase, and execute-phase's
all-plans-complete resume tree gains a stale arm (as a steps/ part, keeping
the spine under its frozen ADR-857 ceiling) mirroring the missing route:
skip cross_ai_delegation/execute_waves/checkpoint_handling, continue at
aggregate_results, and let verify_phase_goal re-dispatch the gsd-verifier —
regenerating VERIFICATION.md and its digest, marked phase or not. The
non-stale fall-through. Staleness detection, the digest format (#4623),
every other routing entry, and the #3684 resume arms are untouched.

Emitted-Drift-Ack-Growth: verify-work.md — stale stop rewritten to dispatch the verifier and re-check (#4682)
Emitted-Drift-Ack-Growth: execute-phase.md — VERIFY_STATUS == stale resume arm added to condition 3 (#4682)

* test(#4682): register the stale-reverification part and align projected commands

The new steps/ part must be registered in the inventory manifest and the
per-runtime golden install trees (regen:derived); the projected stale
next_command is /gsd-execute-phase <phase> (formatGsdSlash prefixes the
runtime surface), the human_needed bare-report probe keeps routing to
verify-work (unchanged semantics), and init-manager's recommended action
follows the new command.

* test(#4682): prefix the remaining stale routing assertions with the runtime surface

Nine stale next_command assertions and the human_needed bare-report probe
still carried the unprefixed or flipped forms from the earlier line-number
edit; all now assert the shipped /gsd-execute-phase <phase> projection,
with the human_needed probe reverted to its unchanged verify-work routing.

* test(#4682): align the last stale projection assertions with the execute-phase route

* docs(#4682): backfill changeset PR number

* test(#4682): refresh the compact-content baseline after the rebase

The rebase onto the #4670 squash brought verify-work.md's bounded
reconciliation text into this branch; the committed compact-content
baseline now reflects the post-rebase split sizes. Local --check is
clean; the previous bench drift (+243) was the baseline, not the diff.

* fix(#4682): carry the response_language directive in the stale-reverification part

The new steps/ part is its own coverage unit for lint-response-language-coverage;
it takes the shared canonical directive line like its sibling execute-phase
parts.

---------

Co-authored-by: sim <sim@local>
2026-09-17 02:23:03 -04:00
Tom Boucher
651511d1e3 fix(#4670): bound the commit-claim window to the plan's own history (#4813)
* test(#4670): add failing-first coverage for the bounded commit-claim window

* fix(#4670): bound the commit-claim window to the plan's own history

The reconciliation measured plan_head_before..HEAD — a window that grows
with every later plan's task and SUMMARY commits plus execute-phase's own
phase-completion commit — so an honest plan flagged commit_claim_mismatch
as soon as anything landed after it (real project: claims 3/5/2 measured
20/10/5).

The executor now also records plan_head_after (HEAD at its measurement
moment, after the last task commit, before the SUMMARY commit), and
verify-work reconciles exactly against plan_head_before..plan_head_after
with a merge-base ancestry check; SUMMARYs without the anchor fall back to
the legacy warning path instead of an unsound BLOCKER. Both #3968 failure
modes (claimed commits never made; task commits lost) still block, driven
by the issue's own fixture scenarios.

Emitted-Drift-Ack-Growth: verify-work.md — reconciliation gains the bounded window and legacy fallback (#4670)
Emitted-Drift-Ack-Growth: gsd-executor.md — plan_head_after anchor documented in the measurement protocol (#4670)

* fix(#4670): name the history-rewrite case and keep the executor under its cap

Review findings: the BLOCKER enumeration named only the two #3968 causes,
so an honest plan whose recorded window was rewritten afterwards (rebase,
amend, cherry-pick) got a mislabeled diagnosis — the clause now names that
case with the manual-recount remedy. The executor's growth crossed the
LARGE-tier hard cap (49152), so the plan_head_after documentation is
compressed to the minimal capture + frontmatter write (verify-work.md
carries the semantics), the #2751 PROSE_ALLOWLIST entry is re-pointed at
the shifted line (#4670 moved it from 823 to 825), the compact-content
baseline is regenerated, and the changeset records the two un-established
edges (shared-base waves, subrepo ledgers).

Emitted-Drift-Ack-Growth: gsd-executor.md — plan_head_after anchor documented in the measurement protocol (#4670)

* docs(#4670): backfill changeset PR number

* docs(#4670): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 22:06:03 -04:00
Tom Boucher
652796e903 fix(#4665): route --fix past the empty-scope exit (#4810)
* test(#4665): add failing-first contract coverage for the --fix empty-scope recovery

* fix(#4665): route --fix past the empty-scope exit

check_empty_scope exited the entire workflow whenever REVIEW_FILES was
empty — before dispatch-fix — so with #3661's incremental scoping, a phase
whose only post-review changes were planning artifacts could never run
--fix against its standing REVIEW.md findings, and the skip output did not
even mention the flag.

The skip is now a self-contained guarded fence (explicit REVIEW_FILES
emptiness check): it fires only when --fix is absent OR the phase's
REVIEW.md does not exist. Otherwise the workflow proceeds directly to
dispatch-fix, which delegates to code-review-fix.md — the canonical fix
implementation that already documents handling an existing REVIEW.md —
while the fresh-review steps (structural pre-pass, reviewer lanes,
spawn_reviewer, commit_review) are skipped: nothing new to review, nothing
to commit. dispatch-fix.md's route docstring is synced.

Emitted-Drift-Ack-Growth: code-review.md — check_empty_scope gains the --fix recovery branch (#4665)

* docs(#4665): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 19:09:19 -04:00
Tom Boucher
85545a77a5 fix(#4663): gate the canonicalization on the uat-passed predicate (#4809)
* test(#4663): add failing-first contract coverage for the blocked-uat canonicalization gate

verify-work.md's complete_session step flips VERIFICATION.md to passed on
'zero issues' alone, so a session whose every UAT row is blocked (a session
that observed nothing) canonicalizes the report. Pins the deployed contract
the fix must satisfy: the flip runs the unflagged phase uat-passed predicate
inside the human_needed branch, frontmatter.set sits inside a passed==true
guard, a refusal message carries the blocker count and keeps
human_needed, and an indeterminate pre-check fails closed. All four new
assertions are RED until the workflow grows the guard.

* fix(#4663): gate the canonicalization on the uat-passed predicate

complete_session flipped VERIFICATION.md to passed whenever the session
recorded zero issues and the status was human_needed — but blocked rows are
not issues by this workflow's own rule, so a 0-passed / 0-issues / N-blocked
session (one that observed nothing) rewrote the canonical report to passed.
Every later reader (transition.md's preliminary check, resume paths,
validate-phase, verification.status) then inherited the unearned pass while
the phase-close predicate correctly refused it.

The flip now runs the phase-close predicate in a new --uat-only form before
canonicalizing: UAT rows evaluated (at least one pass, no
pending/blocked/failed/unexplained-skip row), VERIFICATION-status blockers
skipped — they must be, because the report still reads human_needed at
pre-check time and that status is itself a blocking verification entry, so
the full predicate could never pass there and the flip would deadlock
(found by isolated review, probed). The --require-verification call stays
the transition gate; the refusal branch reports the blocker count and keeps
human_needed; an indeterminate pre-check fails closed.

Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663)

* test(#4663): align the canonicalize pre-check needles with the shipped line

The workflow line carries a 2>/dev/null redirect the needles did not
include, so both pre-check assertions fail against the committed fix
(fixed-string grep verified). Reviewer-found; needle and message aligned.

* fix(#4663): reword the canonicalize prose and refresh its size baseline

The rationale paragraph mentioned the flagged transition-gate call by its
flag, putting a --require-verification literal before the first
phase uat-passed occurrence and breaking the existing ordering pin; the
prose now describes it without the literal. verify-work.md's growth also
drifted the committed compact-content baseline; regenerated via
benchmark-compact-content.cjs --write (derived artifact, report-not-gate
contract).

Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663)

* docs(#4663): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 17:26:24 -04:00
Tom Boucher
5e729445d3 fix(#4657): give the ui consideration probe a text_en language channel (#4804)
* test(#4657): add failing-first coverage for the ui probe's text_en channel

Mirrors the #3717/#4156 test shape onto the UI adapter: a failing-first
proposeConsiderations regression (Danish text + English text_en must classify
as its English equivalent, not land in the #1110 unclassified sentinel),
proposeElements/analyzeCoverage/CLI end-to-end pairs, fail-closed text_en
validation cases (empty/whitespace/non-string, unconditional under an
elements override), a ui-phase.md Step 9.5 workflow-prose contract test, a
reference-doc Inputs parity test, and a fast-check property proving any
cue-matching prose classifies identically under a cue-free Danish rendering
plus text_en. All new assertions are RED until src/ui-consideration-probe.cts
and the workflow/reference docs are updated.

* fix(#4657): give the ui consideration probe a text_en language channel

Element gains an optional text_en; classifyElement's own signature stays
untouched (a locked, directly-tested export) and the text_en ?? text
selection is pushed to the two classification call sites (proposeConsiderations,
proposeElements) instead. text_en is validated fail-closed: an empty or
whitespace-only value throws rather than silently winning the ?? fallback and
degrading classification to zero kinds.

Mirrors #3717/#4156 onto the UI adapter: ui-phase.md Step 9.5 gains the
Non-English projects section (mirroring spec-phase Step 5.5) and the
ELEMENTS_JSON shape comment documents the field with both zero-applicable
guard arms named; the reference doc's Inputs section, the PROBE.ui CONTEXT
predicate (with both derived indexes regenerated), and the nav-override
test expectation stay in sync. The ui-phase contract test carries the
site-scoped allow-test-rule marker and its cluster is registered in the
test-file-count allowlist ratchet.

Emitted-Drift-Ack-Growth: ui-phase.md — Non-English text_en section, ELEMENTS_JSON shape comment, and two-arm guard wording (#4657)

* docs(#4657): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 13:46:40 -04:00
Tom Boucher
3014775a3f fix(#4656): expose coverage.unclassified and widen the zero-applicable guards (#4800)
* fix(#4656): expose coverage.unclassified and widen the zero-applicable guards

* fix(#4656): regenerate golden coverage fixtures and update the rollup pin

Emitted-Drift-Ack-Growth: spec-phase.md — #4656: guard widened to the all-unclassified case, doc claim corrected
Emitted-Drift-Ack-Growth: ui-phase.md — #4656: guard widened identically

* fix(#4656): sync edge-probe doc blocks and coverage pins with the new field

* fix(#4656): key the mandatory confirmation on the widened guard

* docs(#4656): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 10:33:32 -04:00
Tom Boucher
caecaec62e fix(#4648): delegate explore seeds to the plant-seed workflow (#4798)
* fix(#4648): delegate explore seeds to the plant-seed workflow

* fix(#4648): delegate explore seeds to the plant-seed workflow

Emitted-Drift-Ack-Growth: explore.md — #4648 consumer wiring: the seed output now delegates to /gsd:capture --seed (plant-seed) instead of hand-writing a divergent, reader-invisible shape

* fix(#4648): plant-seed extracts an idea-stated trigger into trigger_when

Emitted-Drift-Ack-Growth: plant-seed.md — #4648: write-seed sets trigger_when from the idea text when /gsd-explore passes the conversation trigger inline

* docs(#4648): backfill changeset PR number

* docs(#4648): correct changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 08:24:44 -04:00
0xdhx
003d982c83 fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint (#4749)
* fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint

Two defects in the covered-input fingerprint (#4155), one issue.

1. `computeCoveredDigest` hashed the whole bytes of every declared path
   uniformly, so `.planning/ROADMAP.md` and `.planning/REQUIREMENTS.md` —
   which every phase rewrites as ordinary bookkeeping, and which the closing
   phase's own `phase.complete` / `requirements mark-complete` rewrite AFTER
   the verifier ran — flipped every phase that declared them to `stale` on
   zero implementation change, and from there `isPhaseComplete` →
   `init.manager` → `complete-milestone`'s `ALL_PHASES_VERIFIED` gate.
   Fingerprint v2 leaves any direct child of a planning root out of the
   hash: `.planning/` itself, plus the phase's own planning root (the parent
   of its `phases/`, so `planningDir`'s `<project>/` and `workstreams/<ws>/`
   layouts are covered without the digest knowing what a workstream is —
   `sharedPlanningRoots` / `isSharedPlanningDoc`, defined by position rather
   than a name list so the set cannot drift; a root is accepted only when the
   phase dir sits under a `phases/` directory inside `.planning/`). Such a path is still validated
   exactly as every other covered path (confined, present, a regular file —
   the fail-closed contract is unchanged); only its bytes are ignored, and a
   declaration made only of shared documents fails closed like an empty one.
   A stored digest names its version, and `readVerificationStatus` now
   recomputes under THAT version (`parseFingerprintVersion`,
   `KNOWN_FINGERPRINT_VERSIONS`): a legacy v1 report keeps v1 semantics
   until it is re-fingerprinted, so the upgrade alone stales nothing; a
   version this build cannot recompute fails closed.

2. `verification.fingerprint` received a raw positional slice, so
   `--files a`, `--files "a,b"` and `--files a --files b` all put the literal
   token into the covered set and failed closed as "a covered file is
   missing, unreadable, or escapes the project root" — the message that
   convinced the reporting project the digest was permanently
   unrecomputable. `parseFingerprintFileArgs` accepts every form (plus
   `--files=a,b`, freely mixed with bare positionals), treats any other
   `--flag` and an empty `--files` value as usage errors that say so, and
   the phase-dir argument must now be an existing directory: omitting it
   used to take the first covered file as the phase dir and print a
   plausible digest over the rest at exit 0.

Regression tests (tests/verification-status.test.cjs, #4623 block): the
cross-phase case from the report, the same-phase `requirements
mark-complete` / `phase.complete` cases from the thread, a workstream-scoped
root, v1-preserved / unknown-version-stale, the fail-closed cases (missing,
directory, escaping symlink, all-shared), every `--files` form against the
bare form, the unknown-flag / empty-value / omitted-phase-dir errors, and
AC5's zero-file error. Verified failing against the pre-fix source: 29 of 34
fail, the 7 that pass pin behaviour the fix must leave unchanged.

Docs: CONTEXT.md Verification Module, agents/gsd-verifier.md's
covered_files instruction (rewritten in place — the file sits 21 bytes under
its LARGE hard cap), gsd-core/templates/verification-report.md.

Fixes #4623

Emitted-Drift-Ack-Growth: gsd-verifier.md — the #4155 covered_files instruction now states that planning-root docs are digest-inert (#4623); +18 bytes, under the LARGE cap
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DCMY8P8s6dp4g3Rxu3nNAi

* chore(#4623): set changeset fragment pr to 4749

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 05:26:44 -04:00
Michel Moreira
49f313d611 fix(#4461): make code-review summary extraction shell-safe (#4533)
* fix(#4461): make code-review summary extraction shell-safe

Emitted-Drift-Ack-Growth: code-review.md — use a literal heredoc for shell-safe SUMMARY parsing

* chore: add changeset for #4533

* fix(#4461): keep heredoc outside command substitution

* chore: rerun CI after Windows timeout

* test(#4461): execute the summary heredoc adversarially

* test(#4461): normalize adversarial paths for Git Bash

* test(#4461): keep adversarial fixture valid on Windows

* test(#4461): scope heredoc regression claim

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 03:23:11 -04:00
Tom Boucher
740ba0d8a3 fix(#4628): expose DAG-ready plans and restrict dispatch to them (#4781)
Emitted-Drift-Ack-Growth: execute-phase.md — #4628 consumer wiring: ready_plans parse pointer, not-ready named skip, and waiting condition 2b reference to the ready-wave-gate step file

Co-authored-by: sim <sim@local>
2026-09-16 02:38:43 -04:00
0xdhx
092d9256b8 fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it (#4768)
* test(#4748): pin the letter-axis defect at the seven shell sites outside #4660's six

Extends tests/nsegment-phase-grammar.test.cjs one class over: for each of the
seven sites the live shell lines are read off disk by anchor and executed in
bash against a letter-suffixed fixture. The four `$((10#$PHASE_INT))` split
sites must yield PHASE_N without a shell error for `03A` / `12A` / `3A` /
`03A.1.2` and the commit-scope ERE they build must match both `feat(3A-01):`
and `feat(03A-1):`; the review-file lookup must bind init's `padded_phase`
rather than re-pad in shell; the `--from`/`--to`/`--only` and
plan-review-convergence extractions must return `12A` / `23A.1.2` (and
`23.1.2`) whole; the legacy normalizer must pad `3A` to `03A` and must not
mangle an already-padded `08`. Every pre-existing shape (`06`, `08.5`,
`23.1.2`, `36.14`) is a regression control.

tests/init.test.cjs asserts `init execute-phase` emits `padded_phase` for a
directory-backed `03A`, a ROADMAP-only `4B` (→ `04B`), the existing ROADMAP
fallback `1` (→ `01`), and `null` when the phase is not found.

Negative control against the unfixed tree: 41 failures in the grammar file,
exactly the "(fails before the fix)" cases and the three derived from them
(scope ERE, three-flag extraction, the `08` octal trap); 2 in init.test.cjs,
both the new assertions. Every regression control already green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it

The canonical phase-number grammar (src/phase-id.cts) is digits, an optional
uppercase letter, then dotted segments — `12A`, `3A`, `23A.1.2` are documented
shapes that `init`, `phase-id.cts` and `phase remove` renumbering already
round-trip. Seven shell sites in shipped workflows and references still
assumed digits-and-dots. Four classes, one fix each:

Class 1 — `PHASE_INT=${PHASE_NUMBER%%.*}; $((10#$PHASE_INT))` (execute-phase.md
×2, completion-reconciliation.md, tdd.md). The post-#4619 split stops at the
first DOT, so on `03A` the "integer" is `03A` and bash aborts with `value too
great for base`. Split at the first NON-DIGIT instead (`%%[!0-9]*`): the
integer half is a pure digit run, and the letter rides along in the rest the
way the dotted fraction already did — `03A.1.2` → PHASE_N `3A\.1\.2`, so the
#4003 zero-pad-tolerant scope ERE matches both `feat(3A-01):` and
`feat(03A-1):`. Byte-identical output for every id that worked before.

Class 2 — `PADDED=$(printf "%02d" "${PHASE_NUMBER}")` before the REVIEW.md
lookup (execute-phase.md). `printf` cannot pad a letter id (prints `03`,
exits 1) — and cannot even re-pad an already-padded `08`, which bash reads as
an invalid octal and prints as `00`, so the lookup resolved phases 08 and 09
to `00-REVIEW.md` today. The disk path hands the workflow the directory's
padded number but the ROADMAP fallback hands it the heading's bare one, which
is why the re-pad existed. `cmdInitExecutePhase` now emits `padded_phase`
through `normalizePhaseName`, exactly as the plan-phase and code-review inits
do, and the workflow binds `{padded_phase}` instead of re-deriving.

Class 3 — `grep -oE '[0-9]+\.?[0-9]*'` (autonomous.md `--from`/`--to`/`--only`,
plan-review-convergence.md). Stops at the letter, so `--from 12A` ran from
phase 12 with no error. Now the canonical ERE `[0-9]+[A-Z]?(\.[0-9]+)*`, which
also closes the single-segment dot-axis gap the same shape carried (`23.1.2`
→ `23.1`, #4568's class in a spelling neither lint saw).

Class 4 — the legacy manual normalizer (phase-argument-parsing.md, reached
from mvp-phase.md). Its two branches (`^[0-9]+$`, `^[0-9]+\.[0-9]+$`) left
`12A` unpadded and never padded `3A` to the `03A` a directory carries; its
integer branch also hit the same `printf` octal trap on `08`. One branch for
the whole canonical token now, padding the digit run via `$((10#…))`.
Whether this legacy surface should instead be retired in favour of `init`'s
normalization is the maintainer call the issue names; extending it keeps the
documented contract true either way.

Driven end to end: `init execute-phase 3A` on a fixture with a
`03A-letter-variant/` directory emits `phase_number: "03A"` and now
`padded_phase: "03A"`; on a ROADMAP-only `### Phase 4B:` it emits `"4B"` /
`"04B"`. The issue's own evidence line claimed `padded_phase` was already in
the execute-phase init output — it was not; that key is emitted by the
code-review / plan-phase inits, which is where the claim was read from.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): extend lint-phase-id-drift with three ratchets for letter-hostile phase-id consumers

The rules that landed with #4619, #4568 and #4660 police grammar MIRRORS —
regexes that describe a phase id. The #4748 sites are CONSUMERS of one, and
every existing rule reported clean on them: the shell-arithmetic rule's
`_INT` escape trusts a NAME the dot-only split did not earn on `03A`; the
`[0-9]+\.?[0-9]*` shape is neither the bounded form the single-segment rule
bans nor the unbounded form the letterless rule inspects; and nothing looked
at `printf "%02d"` at all. Three narrow additions, one per shape:

- findDotOnlyIntegerSplitDrift — `X_INT=${<phase-var>%%.*}`; the safe split
  is `%%[!0-9]*`. Keys on the SOURCE variable being phase-carrying.
- findLooseDottedPhaseRegexDrift — `[0-9]+\.?[0-9]*` / `\d+\.?\d*` on a
  phase-carrying line; the canonical form is `[0-9]+[A-Z]?(\.[0-9]+)*`.
  Disjoint from the two sibling regex rules by construction.
- findShellPhasePrintfPadDrift — `printf "%0Nd" …` whose arguments name a
  phase-carrying, non-`_INT` variable; a pad of an `_INT` via `$((10#…))`
  and a `{padded_phase}` binding are the sanctioned shapes.

Same `<!-- phase-id-owner: … -->` sanction, same scan roots as their nearest
sibling (shell idioms over workflows + references, the regex shape over
workflows + references + agents), same documented limit of a per-line
textual scan. The post-#4619 comment that described the `_INT` convention
as proven by `%%.*` is corrected to name the digit-run split. Confirmed
against the base commit: each rule fires on exactly its own unfixed sites
(2+1+1, 3+1, 1+1) and zero violations remain on the fixed tree.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* docs(#4748): add Fixed changeset

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline and acknowledge emitted growth

The three top-level workflow files below grew by the letter-aware split, the
canonical extraction ERE, the `{padded_phase}` binding, and the comment lines
that name the grammar each site now honours. The committed compact-content
benchmark moved with them; refreshed with `benchmark-compact-content.cjs
--write` (aggregate reduction 15.47% -> 15.45%).

Emitted-Drift-Ack-Growth: execute-phase.md — #4748: first-non-digit PHASE_INT split at the plan-selection and TDD-gate sites, `{padded_phase}` binding at the REVIEW.md lookup, and the comments naming why (482 bytes)
Emitted-Drift-Ack-Growth: autonomous.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the --from/--to/--only extractions plus one comment naming the grammar (249 bytes)
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the phase extraction plus one comment naming the grammar (160 bytes)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): name padded_phase in execute-phase.md's init parse list

A `{field}` token inside a workflow bash block is substituted from the init
JSON only for fields the workflow tells the model to parse. `phase_number`
is on that list; `padded_phase` was not, so the `PADDED="{padded_phase}"`
binding at the review lookup would have been a literal — for every phase,
not only letter ones. Found by the pre-file adversarial review (claim 2, the
author's own named suspicion); the test now asserts the parse list carries
the field beside `phase_number`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on its source and widen the printf rule to any %d form

Two false negatives from the pre-file adversarial review of the three #4748
ratchets: `PHASE_PREFIX=${PHASE_NUMBER%%.*}` escaped the split rule because
the destination did not end in `_INT` (the defect is the split, not the
name it lands in), and `printf '%02d'` / `printf "%2d"` escaped the printf
rule because it required double quotes and the zero flag (`%d` cannot parse
a letter id under any width). Both rules now key on the phase-carrying
SOURCE alone; base-site firing counts are unchanged (2+1+1, 1+1) and the
fixed tree stays at zero. The `[[:digit:]]` spelling and the `/phase/i`
heuristic remain the sibling rules' documented limits.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline after the parse-list edit

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): compose init's emitted padded_phase through the live REVIEW.md lookup

The Class 2 site is a `{padded_phase}` template token, which no test can
execute as written. This substitutes the value init emits
(`normalizePhaseName`) into the three live lookup lines and runs them
against a fixture, so the emitted value, the binding, the path construction
and the status extraction are exercised together — `03A-REVIEW.md` and
`08-REVIEW.md` each resolve to their own status. Suggested by the resumed
adversarial review pass (claim C).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): move the #4619 and #4003 source-parity pins to the letter-safe split

tests/execute-phase-decimal-arithmetic.test.cjs and
tests/safe-resume-gate-anchoring.test.cjs pin the four Class 1 sites'
snippet byte-for-byte, so the first-non-digit split reddened both in the
whole-suite run (scripts/ci-test-scope.cjs does not select either file for
a workflow edit — the scoped run was green). The pinned snippet is now the
shipped one, and the behavioural half of the #4619 file gains the letter
case (`03A` → `3A`, `23A.1.2` → `23A\.1\.2`) beside its decimal cases.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on the _INT destination again, tolerating the quoted spelling

Keying on the source alone (the previous commit's widening, from a review
probe) flags `PARENT_PHASE="${PHASE_NUMBER%%.*}"` in
gap-closure-artifacts.md — a correct derivation that wants everything
before the first dot, letter included. The defect this rule polices is a
dot split INTO the name the shell-arithmetic rule trusts as an integer, so
`_INT` is the discriminator on purpose; the quoted spelling that site uses
is now tolerated so the same shape into an `_INT` cannot hide behind it.
Base-site firing unchanged (2+1+1), zero on the fixed tree, and the
parent-phase line is pinned as a silent case.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): use t.after() for the composition test's fixture cleanup

CONTRIBUTING forbids try/finally inside a test body; the per-test cleanup
form is `t.after(() => cleanup(dir))`. Flagged by the filing driver's
test-ruleset gate before the PR was created.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): set changeset fragment pr to 4768

* chore(#4748): refresh the compact-content benchmark baseline after rebasing onto next

Regenerated with `node scripts/benchmark-compact-content.cjs --write` on the
rebased tree (base 0d6bf19bf); `--check` confirms it matches the live recompute.
Only the execute-phase split and the aggregate totals differ from next's copy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 23:31:36 -04:00
Tom Boucher
b5fc8061b3 fix(#4624): persist orchestrator-worktree worker lifecycle records (#4778)
* fix(#4624): persist orchestrator-worktree worker lifecycle records

* fix(#4624): address review findings on the worker lifecycle protocol

* fix(#4624): require the summary path and surface torn records on status --path

* fix(#4624): distinguish no-record from torn-record, tolerate older shims in the sweep

* docs(#4624): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-15 20:25:24 -04:00
Tom Boucher
72be41d96d fix(#4601): reap the whole process tree on run-with-timeout's Windows force stage (#4775)
* test(#4601): pin the run-with-timeout tree reap on win32

* docs(#4601): add changeset for the win32 tree reap

* fix(#4601): reap the wrapped command's whole tree on Windows timeouts

* docs(#4601): backfill changeset PR number

* fix(#4601): keep the mediated shim alive without depending on stdin

---------

Co-authored-by: sim <sim@local>
2026-09-15 18:31:05 -04:00
Tom Boucher
0d6bf19bf1 fix(#4600): explicit --converge overrides the convergence feature gate (#4771)
* test(#4600): pin explicit-flag-overrides-gate precedence

* fix(#4600): explicit --converge overrides the convergence feature gate

PLAN_STRATEGY=converge is set only by an explicit --converge/--cross-ai,
so gating it on workflow.plan_review_convergence made an explicit
operator flag lose to a config default and stop the run with a
question-shaped success. The config remains the default for non-flag
invocation; the flag now wins, and the step says so instead of
gate-and-exit.

* test(#4600): pin the precedence contract sentences as written

* fix(#4600): keep the precedence sentence on one line

The pinned contract phrase wrapped across a line break, so the
writer-contract assertion could not match it.

* fix(#4600): document the flag-overrides-gate precedence on user surfaces

commands/gsd/autonomous.md and docs/COMMANDS.md still said --converge
requires workflow.plan_review_convergence=true; both now state the
override. Changeset typed Changed with the docs update alongside.

* docs(#4600): backfill changeset PR number

* fix(#4600): restore the convergence gate mention in user surfaces

* test(#4600): pin the dispatched convergence run against the config gate

* fix(#4600): override the convergence gate on the dispatched run

Emitted-Drift-Ack-Growth: autonomous.md — #4600: the converge dispatch appends --override-gate inside the PLAN_STRATEGY conditional, and the precedence sentences replace the stale fail-fast instruction
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #4600: config gate 1.5 honors an explicit --override-gate dispatch (token-anchored) while the standalone veto and config-get default are preserved

---------

Co-authored-by: sim <sim@local>
2026-09-15 15:53:42 -04:00
Tom Boucher
779f67cb11 fix(#4378): mint collision-free SEED-YYMMDD-xxx seed ids instead of a shared count (#4754)
* test(#4378): regression tests for collision-free seed ids

* fix(#4378): mint collision-free SEED-YYMMDD-xxx ids, not a shared count

plant-seed derived the next seed id from 'ls .planning/seeds/SEED-*.md | wc -l'.
.planning/seeds/ is shared but each worktree only sees what has merged, so two
workstreams planting before either merges computed the same id and git merged
both files silently.

The id is now the local date plus a 3-char random base36 suffix -- the shape
.planning/quick/ already uses -- computed from knowledge one worktree has alone,
with a same-day regen guard. deriveSeedIdentity learns the new canonical grammar
alongside legacy SEED-NNN (whose parsing never changes), the --enrich parser and
the filename-prefix fallback keep the full new-format id, and the docs that
state the filename shape move to it.

The prefix fallback previously truncated any non-pure-numeric id at
'SEED-<digits>' -- the same one-id-two-answers ambiguity the issue reports,
reproduced one level down.

* fix(#4378): harden seed id generation per adversarial review

- parse-idea: anchor the --enrich extractor to the flag and capture the
  complete id, uppercase-tolerant; a leftmost 'SEED-[0-9]+' truncated an
  uppercase or malformed suffix to its date and enriched an arbitrary
  same-day seed via head -1. Ambiguous and unmatched targets now fail
  closed instead.
- generate-seed-id: tolerate the expected SIGPIPE under pipefail, abort
  loudly when the suffix cannot be drawn (an empty suffix would collapse
  every seed's id to the bare date), and run the same-day regen guard as
  a find existence test (the 'ls <glob>' shape trips the #3409 drift
  guard and degenerates under a stray nullglob).
- deriveSeedIdentity: document the theoretical legacy/new grammar
  ambiguity (6-digit counter + 3-char base36 slug, no frontmatter).
- changeset: state the residual same-day collision bound instead of
  implying zero.

Emitted-Drift-Ack-Growth: plant-seed.md — the counting step became hardened date+random generation with explicit failure modes; growth is the failure handling, not duplicated logic

* fix(#4378): address standards and spec review findings

- tests: move the allow-test-rule marker to its suppression site (the
  file-header placement was inert per CONTRIBUTING site-scoping); add
  width-boundary coverage (5/7-digit dates, 2/4-char suffixes pin the
  documented branch behavior); add a writer-to-reader parity property
  that parses the mint widths out of the shipped workflow so the two
  grammar owners cannot drift; cover uppercase ids end-to-end in the
  reader.
- plant-seed.md: draw/retry restructured as one loop with a loud
  terminal failure; SEED_SUFX renamed SEED_SUFFIX; regen guard drops
  the redundant head -1; the ambiguity error no longer advises an
  impossible 'complete id' for duplicate legacy ids.
- commands.cts: refresh the cmdListSeeds comment still describing
  SEED-NNN as the only canonical form.
- changeset: drop the audit claim the spec axis showed to be an
  overstatement (audit's id display is filename-derived, pre-existing).
- remove a stray untracked artifact file swept into the tree.

* test(#4378): correct boundary expectations to the module's real branch behavior

The first matrix run on the boundary tests caught my hand-trace of the
regex branches, not a module defect: the slug regex's alternation
backtracks to the legacy branch whenever the canonical branch cannot
complete (so the slug is the remainder after the legacy numeric
prefix), and the 7-digit case fails the canonical branch at its 7th
digit before the dash. Pin the verified values.

* docs(#4378): backfill changeset PR number

* fix(#4378): audit seed identity uses the canonical grammar

Review of this PR found the audit surface publishing a fused filename
stem (SEED-081-region for SEED-081-region.md) where list-seeds reports
the canonical id -- one id, two answers across surfaces, the same
ambiguity class the issue files. scanSeeds now derives identity through
the SAME deriveSeedIdentity the list-seeds gate uses (frontmatter id,
then filename id-prefix, then stem), and audit-open acknowledge
resolves --seed-id by scanning for the derived identity, falling back
to the literal stem so callers scripted against pre-canonical output
keep working. Roll-in per the fix-inline rule: found during this PR's
review, same seed-identity seam.

RED probe: pre-fix audit published seed_id SEED-081-region-becomes /
slug 081-region-becomes for a legacy seeded file; post-fix SEED-081 /
region-becomes, matching list-seeds.

* test(#4378): probe timeout uses the class norm after windows-lane timeout

The windows conformance shard failed its bounded sh -c probes at the
local 5000ms bound (cold sh.exe spawn under shard load) while the
identical code passed this PR's two earlier windows waves. The probe
now uses PROBE_TIMEOUT_MS from the class-norm module instead of a local
override, per the helpers/timeouts.cjs convention.

---------

Co-authored-by: sim <sim@local>
2026-09-15 11:41:29 -04:00
Tom Boucher
cbbde6786a fix(#4546): deferred UAT follow-ups no longer block completion and promote to the backlog (#4769)
* test(#4546): failing-first tests for deferred uat follow-ups

* chore(#4546): regenerate derived lists for the deferred-promotion suite

The new verify-work-deferred-promotion suite changes the tests/ tree the
macOS conformance-tier classifier tracks and is a novel file under the
verify prefix in the test-file-count ratchet; both derived lists are
regenerated/registered per their own guards' instructions.

* fix(#4546): deferred uat follow-ups no longer block, and get promoted

Two halves of one disconnect (#1921's deferral design vs the completion
predicate):

- uat-predicate: the item parser now captures the block's reason: line
  alongside result:. A skipped item whose reason carries the
  verify-work writer's 'Deferred follow-up:' template is a deliberate
  deferral -- non-blocking, flagged deferred in the report. Quote-
  tolerant (the writer wraps the value) and case-insensitive. A
  reasonless skip, a non-deferral reason, pending/blocked/issue/
  failed/missing all still block, exactly as before.
- verify-work complete_session: when the Deferred Follow-Ups section is
  non-empty, offer to promote the items to a ROADMAP.md 999.x backlog
  entry reusing next.md's prior_phase_completeness entry shape, with a
  --files-scoped commit. Offer, not auto-mutation -- matches the
  workflow's interactive convention and next.md's own prompt style.

* chore(#4546): refresh compact-content benchmark baseline

verify-work.md grew (the #4546 deferred-follow-up promotion offer in
complete_session); the registered split's token counts moved with it.
Baseline recomputed with the script's own --write.

Emitted-Drift-Ack-Growth: verify-work.md — complete_session gained the deferred-follow-up promotion offer (detection, [P]/[K] choice, the next.md-shaped 999.x entry template, and the --files-scoped ROADMAP.md commit); the growth is the new contract text, not duplication

* fix(#4546): gate/audit agreement and review fixes for deferred follow-ups

- src/uat.cts categorizeItem: a skipped item carrying the deferred
  follow-up template reason now categorizes as 'deferred' (the category
  already existed for deferred-items.md entries) instead of being
  misfiled into the blocked families by keyword match -- the gate/audit
  agreement #3078-CR expects, restored in the permissive direction the
  #1921 design intends. Checked BEFORE the keyword families so '...
  on the release build next version' is not build_needed.
- verify-work.md promotion step: numbering scans for the smallest free
  999.n (count races + non-contiguous history), one backlog entry per
  deferred follow-up, ROADMAP.md-absent behavior specified, idea text
  newline-flattened, Deferred at placeholder harmonized with next.md.
- DEFERRED_REASON_RE: trust assumption documented (authoring contract,
  not a security boundary; non-matching spellings block fail-closed).
- tests: the property now drives evaluateUatPassed and derives
  expectations from the input spec (never restates the matcher),
  includes the no-result-line branch, and pins its seed; the parity
  test drops try/finally for the approved pattern, uses createTempDir,
  sites its allow-test-rule marker at the suppression site, and asserts
  the literal [P]/[K] choices.

* fix(#4546): close promotion-test docstring, drop fc replay-path misuse, refresh baseline

The final matrix run caught three defects in my own review-fix commit:
the parity test file's JSDoc was left unterminated (the whole file
parsed as one comment -- zero tests registered, hence the file-level
'test failed' the runner reported); fast-check's replay-path parameter
was misused as a label (invalid path at replay); and the workflow-text
ambiguity fixes re-drifted the compact-content benchmark baseline.

* docs(#4546): add Fixed changeset for deferred follow-up coverage

* docs(#4546): backfill changeset PR number

* fix(#4546): use the pattern seam escapeRegex for shape-marker matching

The hand-rolled metacharacter escape in the shape-marker assertion
tripped local/no-adhoc-regex-escape, whose named remedy this adopts.

---------

Co-authored-by: sim <sim@local>
2026-09-15 07:30:38 -04:00
BeeHiggs
0967358b8b enhance(#3638): render bracket phase IDs on progress, stats, manager and statusline surfaces (epic #612 PR-5) (#4111)
* enhance(#3638): render bracket IDs on display surfaces

Gate progress, stats, manager, and statusline projections on the bracket convention; validate phase_id_convention and single-source the convention card.

Forward note: the uat.cts bracket co-change remains deliberately deferred to its owning slice.

* chore(#3638): point the changeset at PR #4111

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3638): close bracket display review gaps

* docs(#3638): register phase display modules

* chore(#3638): re-trigger CI after macOS shard SIGTERM

`full test (macos-latest, 24, shard 3/3)` failed on 20ce98cd1 in
`tests/lint-compiled-artifact-sync.test.cjs` — the spawned
`scripts/lint-compiled-artifact-sync.cjs` was killed at 60024ms
(`exited null (signal SIGTERM)`, stdout and stderr both empty), 24ms past
the test's own `TSC_COMPILE_TIMEOUT_MS`. That is the failure mode the
constant's comment already documents ("under CI shard load that compile
can exceed the budget, dying to a SIGTERM with empty piped stdout").

No content change; this empty commit exists only to re-run the matrix,
since re-running a job needs write access on the upstream repository.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 04:47:47 -04:00
0xdhx
aeac47b95a fix(#4465): bound /gsd:undo commit selection to the phase directory and HEAD (#4472)
* fix(#4465): bound /gsd:undo commit selection to the phase directory and HEAD

`--phase` documented a primary path reading `.planning/.phase-manifest.json`,
but nothing in the repository writes that file, so the documented fallback was
the only reachable path:

    git log --oneline --no-merges --all | grep -E "\(0*${TARGET_PHASE}(-[0-9]+)?\):" | head -50

That selection has no milestone bound and no reachability bound, and it feeds
`git revert --no-commit`. On a project that reuses a phase number it stages
deletion of a previous milestone's files, under a confirmation gate that
displays only `{hash} — {message}` — the one field that carries no milestone
discriminator.

Port the #3995 anchor already live in code-review.md: resolve the phase's own
directory via `find-phase` (which resolves through planningDir, so it is
workstream-correct), take PHASE_START as the first commit adding anything under
it, and select over `PHASE_START^..HEAD`. Drop `--all`. Fail closed when no
anchor resolves rather than widening to a repository-wide search, and report
truncation instead of silently capping at 50.

Also resolve the dead manifest read rather than leaving documented-but-
unreachable behaviour, and hoist dependency_check onto the workstream-resolved
planning root — it read a hardcoded `.planning/ROADMAP.md` and
`.planning/phases/`, which are the wrong tree under an active workstream.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv

* fix(#4465): close three defects the pre-create adversarial review found

Root-commit off-by-one: `${PHASE_START}..HEAD` EXCLUDES PHASE_START, so when the
phase's first commit is the repository root the selection silently dropped it
and the undo refused legitimate work. The root branch now selects over `HEAD`.

Empty-selection exit status: `grep` exits 1 on no match, and the removed
`| head -50` had been masking that rc. Both selection pipelines now end in
`|| true` so an empty selection reaches the workflow's own Empty check instead
of aborting the block.

Truncation stop was documented for MODE=phase only; MODE=plan could still cap
silently. Both modes now carry it.

Also documents two residuals the review surfaced rather than leaving them
implicit: a revision range is ancestry and not chronology, so a pre-phase side
branch merged in after PHASE_START stays selectable; and the `--diff-filter=A`
anchor does not follow renames, so an archived phase directory under-selects
(reverts too little or refuses, never too much).

Tests: 12 assertions, negative-controlled. A-H and K-L are RED against the
pre-fix workflow; J is RED against the intermediate revision that carried the
off-by-one; I is green in both directions by design.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv

* chore(#4465): backfill the changeset fragment's pr field

The fragment could not carry `pr:` before the PR existed; both
`scripts/changeset/lint.cjs` and `scripts/lint-docs-required.cjs` require it
and reported `missing_pr` until now. Both are green with it filled in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv

* chore(#4465): acknowledge the deliberate growth of undo.md

The emitted-attribution gate flags undo.md growing 11881 -> 17435 bytes. The
growth is the fix: a one-line `git log --all | grep` selection is replaced by
an anchored, HEAD-bounded selection for BOTH modes, each with its own
fail-closed branch and truncation stop, plus three residuals documented next to
the code they qualify rather than left implicit. A workflow document is the
executable contract, so the residuals belong in it.

Emitted-Drift-Ack-Growth: undo.md — replaces a one-line unbounded commit-subject grep with an anchored HEAD-bounded selection in both --phase and --plan, each with a fail-closed branch and a truncation stop, plus three residuals documented in-workflow (#4465)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv

* test(#4465): execute undo.md's selection fences against a git fixture

The shipped regression test matched substrings in undo.md's bash fences. A
substring cannot tell a live invocation from a dead one, and this is the
boundary/reachability defect class this repo's own conventions want driven
with limit-1/limit/limit+1 execution proof. So the fences the runtime runs —
sliced out of undo.md by content anchor, never by position — are now replayed
with `bash -c` inside createTempGitProject fixtures against the real
`gsd-tools.cjs`, the same fence-execution shape #2308 and #2352 use.

Ten executed cases: the fixture's own negative control (the retired `--all`
grep over-selects the archived milestone and a dead branch), `--phase` and
`--plan` selecting only the current milestone's HEAD-reachable instance, the
single-milestone selection unchanged, limit-1 (a matching pre-phase commit
excluded, PHASE_START itself included), the root-commit branch selecting over
`HEAD`, fail-closed on an absent phase in both --phase and --plan (empty PHASE_DIR and UNDO_RANGE), the
active workstream's phase directory winning over the root's, and
dependency_check's `planning inspect --pick generated_from.planning_root --raw`
resolving the workstream root, the project root, and the `.planning` fallback.

Skipped on win32 with the #2352 precedent's reason: the fences are POSIX bash
driven through `bash -c`; the shape half still runs there.

Negative-controlled against the pre-fix undo.md (upstream/next): 11 of 12
shape assertions red, and the executed block fails at fence extraction.

Shape test L now pins the >50 refusal message in both modes, not only its
heading — the cross-AI round review removed the paragraphs under intact
headings and L stayed green; it now fails on that mutation.

* fix(#4465): refuse an archived-milestone PHASE_DIR in both undo modes

`find-phase` searches the live `phases/` directory first, then every
`milestones/v<X.Y>-phases/` directory in ascending version order, and its
ambiguity check is scoped to one directory: the `matches.length > 1` test sits
inside `cmdFindPhase`'s per-`searchDir` loop (`src/phase.cts`). A phase number
that is not live therefore resolves silently to the OLDEST archived milestone
carrying one, with no warning.

Anchoring there is wrong in both directions at once. The oldest commit adding
that path is the archival move, so the phase's real work predates the window
and falls outside it, while the window runs forward from that archival through
every later milestone -- where the subject grep matches THEIR same-numbered
phase. Driven on a two-archived-milestone fixture, `--phase 03` selected v2.0's
`feat(03-01): add search index` and `docs(03-01): v2.0 phase plan` and excluded
v1.0's own `feat(03-01): implement auth endpoint`. On `git revert --no-commit`
that is the cross-milestone contamination this PR exists to close, recurring.

Both modes now blank PHASE_DIR on an archived resolution and refuse with their
own message, so the existing fail-closed rule stays load-bearing even if the
refusal's prose is not honored.

The "archived or renamed phase directory under-selects" residual claimed this
failure could revert "too little or refuse, never too much". That was wrong in
the direction that matters; it is rewritten to cover renames only, and the
archival case is recorded as refused rather than disclosed.

Tests: a two-archived-milestone fixture, a negative control asserting the
unguarded window selects v2.0's two commits and none of v1.0's, refusals in
both --phase and --plan, and an inertness check on a live resolution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* docs(#4465): close the three Minor review items on undo.md and its test

Purpose line: it still advertised rolling back "using the phase manifest",
the exact `.planning/.phase-manifest.json` mechanism this PR removes. Test B
already reads the whole file, so scope is not why it missed this -- B greps the
hyphenated `phase-manifest` token, the filename, while the purpose line named
the same dead mechanism in prose. Test N pins the <purpose> block itself, which
is spelling-independent; widening B's pattern to /manifest/i instead would fire
on any future sentence that merely mentions one.

Merge-commit anchors: `git log --diff-filter=A -- "${PHASE_DIR}"` does not walk
merge diffs by default. A directory added on a side branch is still found -- the
side-branch commit that added it is in history -- so the uncovered case is
narrower than "introduced via a merge": it is a directory first appearing in the
merge RESOLUTION. The information is not absent from history, only unrequested:
`git log -m` prints that add once per parent. Recorded as a residual rather than
fixed -- the failure is a refusal, and the evil-merge fixture costs more than a
safe-direction branch is worth.

allow-test-rule marker: suppression is site-scoped, and CONTRIBUTING.md pins the
window at MAX_MARKER_LOOKAHEAD_LINES = 8 with only blanks and comments between.
The file-header marker sat 43 lines above the first `readFileSync` with requires
and a function definition in between, so it was inert for both read sites. Moved
to each site. The markers are belt-and-braces today, and NOT because the rule
ignores `RegExp.test` -- it handles `regex.test(tracked)` explicitly
(no-source-grep.cjs:239, :597-605). Neither read is tracked at all:
`looksLikeSourcePath` (:378-390) admits only .cjs/.cts/.js/.mjs/.mts/.ts and
UNDO_PATH is a .md, and the second site's reader is `readFileNormalized`, which
the rule does not recognise as `readFileSync`. They become load-bearing if
either scope widens.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4465): record the archived-milestone refusal in the changeset

The fragment described the PHASE_START bound and the fail-closed rule but not
the archived-resolution refusal added this round, which is a user-visible
behaviour change: `--phase N` on a number that is no longer live now stops
rather than anchoring on an archived directory.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* docs(#4465): disclose the same-number-and-slug residual in undo.md

The round's own cross-model body audit found this stated in the PR description
and nowhere in the deployed workflow -- which is the half that survives merge,
and the half this PR's whole argument says residuals belong in.

The anchor is the CURRENT path and `--diff-filter=A` does not follow renames, so
a later milestone that re-creates the same literal directory (`03-auth` again,
not merely phase `03` again) makes the oldest add at that path the previous
occupant's. The archived-milestone refusal added this round structurally cannot
reach it: `find-phase` returns the LIVE directory, so nothing is under
`milestones/` to refuse.

Driven -- v1.0 and v2.0 both using `.planning/phases/03-auth`, v1.0 archived in
between: PHASE_DIR resolves live, the guard correctly does not fire, the anchor
is `docs(03-01): v1 plan`, and the selection returns all four v1+v2 phase-03
commits. `code-review.md` carries the same residual on the same anchor, where it
is read-only; on `git revert --no-commit` it is not, so the note points at
`/gsd:undo --last N`.

Documented rather than fixed: following renames across a re-created path needs a
phase identity that a directory name does not carry, which is the same wall
residual 1 hits.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4465): refuse a LIVE phase directory whose name is also archived

Self-found this round, from the cross-model audit of the previous commit: that
commit disclosed same-number-and-slug reuse as a residual and asserted closing it
needed a phase identity a directory name does not carry. The audit refuted the
second half by producing a fix, and it is cheap and safe-direction, so the case
is now refused rather than documented.

`--diff-filter=A` does not follow renames, so a later milestone that re-creates
the same literal directory (`03-auth` again, not merely phase `03` again) anchors
on the EARLIER occupant's add commit and the window opens there. The archived
refusal added earlier in this round structurally cannot reach it: `find-phase`
returns the LIVE directory, so nothing is under `milestones/` to refuse. Driven
before the guard -- v1.0 and v2.0 both at `.planning/phases/03-auth`, v1.0
archived in between: PHASE_DIR resolves live, the archive guard correctly stays
inert, the anchor is `docs(03-01): v1 plan`, and all four v1+v2 phase-03 commits
are selected. That is a previous milestone's work staged for `git revert
--no-commit`.

The guard needs no identity reconstruction: the same basename present under an
archived `milestones/v*-phases/` means the path has been used before, so the
anchor is untrustworthy and both modes refuse with their own message. It fails
toward refusing a legitimate undo of the reusing milestone, which `--last N`
covers; the alternative is reverting the earlier one's commits.

Tests: a reused-slug fixture, a negative control asserting the unguarded window
reaches back into v1.0 (all four commits), refusals in both modes, and an
inertness check on a distinct slug. All three new checks red against the pre-fix
fence.

Also narrows the changeset, which still claimed selection "no longer" reaches a
previous milestone -- true of the archived route, not of this one until now --
and corrects the concurrent-workstreams residual, which claimed the window
removed the previous-milestone class "entirely" while this case remained open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4465): harden the reused-path refusal against its own false positives

The cross-model audit of the previous commit refuted four of its claims. All
four were right; this closes them.

A REAL BUG: the refusal message interpolated `${PHASE_DIR}` after the fence had
already blanked it, so it would have rendered "Phase 03 resolves to , but that
directory name is also archived at ...". The live path is now preserved in
`PHASE_DIR_LIVE` before blanking, and a test pins that it survives.

THREE FALSE-REFUSAL ROUTES, each of which could block a legitimate undo:

- `[ -e ]` accepted a regular FILE where an archived phase directory would sit.
  The evidence the refusal claims is "an earlier milestone used this path", and
  only a directory is that. Now `[ -d ]`.
- The glob `v*-phases` accepted milestone directory names `cmdFindPhase` itself
  rejects -- its filter is /^v\d+.*-phases$/, so `vnondigit-phases` is not a
  milestone it would ever resolve. Refusing over one is refusing on evidence the
  producer discards. Now `v[0-9]*-phases`, in the collision check and in the
  archived-resolution guard alike.
- `${PHASE_DIR%/phases/*}` silently left the path unchanged when it carried no
  `/phases/` segment, so the scan ran against the wrong root and read as clean.
  The strip must now have fired. Outside today's producer contract either way --
  cmdFindPhase's live output always carries `/phases/` -- so this hardens a claim
  rather than fixing a reachable defect, and is stated as such.

The commit message and the changeset both asserted the guard proves the name is
"archived"; before this commit it proved only that some entry existed. Both are
narrowed to what the fence now actually establishes, and the changeset headline
no longer claims more than the anchor plus the two refusals deliver -- residual 2
(a range is ancestry, not chronology) is untouched by either.

Tests: a file-not-directory twin, a malformed `vnondigit-phases` twin, and the
preserved live path. All three red against the pre-hardening fence, along with
shape test M.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* docs(#4465): disclose what the reused-path scan does not survive

A fifth audit pass refuted three claims in the previous commit. Two are wording,
one is a real fail-open; none is fixed by more shell, so all three are stated.

The wording: `v[0-9]*-phases` was described as MIRRORING cmdFindPhase's
/^v\d+.*-phases$/. It is not equivalent -- the glob's `*` matches a newline where
the regex's `.` does not, so a directory named `v6<newline>-phases` is accepted
here and rejected there. It tracks the filter closely enough to reject the
malformed siblings that motivated it; it does not mirror it, and the comment no
longer says so. The changeset likewise said the twin is "an archived phase
directory", where the check establishes only that a directory of that name exists
under an archived milestone -- narrowed to that.

The fail-open: this collision check is the only fence in undo.md that relies on
pathname expansion (every other one uses `case`, which `set -f` does not affect).
Under a runtime with globbing disabled the scan is skipped silently and a genuine
collision passes; under `shopt -s failglob` a NON-match aborts the fence. Both sit
outside the shell state this workflow assumes throughout, and defending only this
fence while the rest of the file assumes defaults would be inconsistent -- so it
is a documented residual rather than a hardened one.

Also disclosed: the check is conservative at two edges. `[ -d ]` follows symlinks,
and an empty directory of the right name counts, so either can refuse an undo a
stricter ownership test would allow. That direction is the intended one -- refusing
too often costs a `--last N`, refusing too rarely reverts another milestone's work.

No test is added. Asserting behaviour under `set -f` would pin a shell state the
workflow does not otherwise support, and the two conservative edges are the
documented intent rather than defects.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4465): register the new regression test in the conformance-tier list

Round 3's only blocker. The platform-conformance-tier classifier, its generated
registry and the test that asserts the registry is fresh all landed together on
`next` in `bcd99696d` (#4591 / #4598) on 2026-09-10 08:36 -0400 -- a day after
this branch's last push, so the gate did not exist when the branch was last
green. This PR adds a unit-suite test file, the classifier's content scan selects
it, and a list regenerated on `next` without it is therefore stale the moment the
two meet.

(The review cites `0b928fe28c` for that landing. That commit is #4592, four hours
later the same day, and `git diff-tree` shows it touched only
`scripts/ci-test-scope.cjs`, `scripts/gen-platform-conformance-tier.cjs` and
those two files' tests -- not the generated registry at all. The date and the
diagnosis hold either way.)

Regenerated with `node scripts/gen-platform-conformance-tier.cjs --write` on the
rebased tree: one line, 267 -> 268 entries. The `--target macos` list is
unaffected -- it writes a different file, `macos-conformance-tier.generated.cjs`,
and its classifier does not select this test (198, already matching) -- so
`lint:generated-sync`'s second conformance-tier link needed nothing.

**Two different baselines, stated so the counts are not read as one.** CI's
failure on the prior head reads `546 !== 547`, and this commit's diff reads
267 -> 268. Both are correct and they are not the same tree: at the merge commit
`54b0b197f` the committed list held 546 entries, so the classifier wanted 547.
`4d65c248e` (#4641, "narrow the tier to 28.5%") then landed on `next` at
2026-09-11 17:00 -0400 -- after that CI run started at 13:26Z -- and collapsed
the list to 266, with `a2331c01f` taking it to 267. Hence 268 here. The two
547-entry lists are not the same file set: upstream's includes
`tests/execute-phase-decimal-arithmetic.test.cjs`.

Order matters and is worth stating: the registry is derived from the `tests/`
tree, and the base range removed 282 entries from it and added 3 (`git diff
--numstat` reports `3 282`; the familiar 279 is the net shrinkage, not the count
of entries changed). Regenerating before the rebase would have produced a
547-entry list against a base carrying 267, and a three-way `git merge-file`
control over that pair does conflict. Rebase first, regenerate second.

This clears both reds on the prior head, not one. `lint-tests` is the one the
review named; `test (ubuntu-latest, 24, shard 3/3)` is the same staleness seen
through `tests/platform-conformance-tier.test.cjs:278`, which asserts the
committed list is fresh. It was the only failure in 1267 tests on that shard, and
it completed at 13:38:33Z -- five minutes after the review was submitted at
13:33:54Z, which is why the review recorded the ubuntu matrix as green.

Control: removing the added line reds `real tests/ tree classification matches
the committed list` (1 of 55, `267 !== 268`); restoring it gives 55/55.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R4EJcuDjaaF1QeeT8AxYbj

* test(#4465): pin what find-phase returns for the workstream archive layout

Review round 4 found the archived-phase handling blind to the second archive
layout, milestones/ws-<name>-<date>/phases/, that `workstream complete` writes.
The suite had no fixture for it and its comments described one archive reader
where there are two.

- Names both readers in the resolver-contract comment: find-phase routes to
  cmdFindPhase, which admits only /^v\d+.*-phases$/ under milestones/, while
  the phase locator (listArchiveVersionDirs) also enumerates ws-*.
- Adds a ws-* fixture built the way `workstream complete` builds it (the whole
  workstream dir moved into milestones/ws-feat-<date>/).
- Negative control: a re-created workstream with the same phase slug, run
  without the collision guard, selects the archived generation's commits.
- Pins the actual find-phase contract: a phase living only in a ws-* archive
  resolves to nothing, and nothing is selected.
- Pins that an ordinary deleted plan file inside a live phase is not treated
  as a previous occupant.

runFences gains an explicit env override so a test can activate a workstream
after the leak scrub. Every test here passes on the pre-fix workflow; the
assertions that need the fix land with it in the next commit.

* fix(#4465): refuse a re-created phase path by its history, not an archive-layout glob

Review round 4: the archived-phase handling encoded the archive layout as a
literal `v[0-9]*-phases` glob, while the phase locator owns two layouts. The
second, milestones/ws-<name>-<date>/phases/, is what `workstream complete`
writes.

Driven on the real fences: a workstream `feat` completed into ws-feat-<date>/
and then re-created with the same 03-auth slug resolves to the live path,
the collision glob finds no v*-phases twin, the anchor opens on the archived
generation's first commit, and --phase 03 selects that generation's commits.

The collision check now asks git whether this exact path went EMPTY somewhere
in HEAD's history and came back: `git log -m --no-renames --diff-filter=D`
over PHASE_DIR, then `git ls-tree -d` at each deleting commit to tell a
vacated directory from an ordinary deleted plan file. That answers for every
way a path can be vacated -- the flat archive, the ws-* archive, and a phase
removed and re-added under the same slug, which no layout glob could see --
without a second reader of the layout to keep in sync. The refusal message
now names the commit that vacated the path.

The archived-resolution refusal names both layouts. find-phase (cmdFindPhase)
searches only the flat archives today, so a ws-*-only phase resolves to
nothing and fails closed on the not-found rule; the ws-* arm keeps the refusal
correct if find-phase is ever taught the locator's second layout, and a
stubbed-resolver test pins it for both modes.

The retired glob's shell-state residual (pathname expansion under set -f /
failglob) is gone with it. Its replacement residual is documented in-workflow:
the history check misses a single commit that both moves the directory away
and re-creates it, and over-refuses when a side branch emptied it and the
merge kept it.

* fix(#4465): run undo's path-scoped git calls from the project root find-phase answers against

Found by this round's pre-push review. find-phase prints PHASE_DIR relative
to the PROJECT ROOT -- gsd-tools resolves the root before it dispatches --
while the workflow's shell stays wherever the user invoked /gsd:undo. A git
pathspec is read relative to git's cwd, so from a subdirectory every
PHASE_DIR-scoped call looked in the wrong place:

- the anchor (`git log --diff-filter=A`) came back empty, so a legitimate
  --phase or --plan undo was refused. Fail-closed, but a refusal of valid
  work, and it dates from round 1;
- the collision check's `git log --diff-filter=D` came back empty, so a
  re-created path was never recognised. The anchor failing too is the only
  reason this did not over-select.

Both modes now take PROJECT_ROOT from `planning inspect --pick
generated_from.cwd` -- the directory find-phase itself resolved against --
and run the three path-scoped calls as `git -C "${PROJECT_ROOT:-.}"`. The
project root, not `git rev-parse --show-toplevel`: a project need not sit at
the top of its repository, and a control test pins that choice. Selection,
revert and rev-parse calls are SHA-only and unchanged.

Tests: every behavioural test ran from the fixture root, where the two
readings coincide. New ones run the fences from sub/dir -- normal selection
in both modes, refusal of a re-created path for both archive layouts in both
modes, and a dropped plan file NOT refused, which is how a mis-rooted
ls-tree would fail (it lists nothing, and nothing reads as "vacated"). Shape
test O pins every PHASE_DIR-scoped git call to the root. The two tests
renamed in the previous commit are relabelled as false-positive controls:
they pass with the collision check deleted, and say so.

Reversion controls, all driven: the pre-change undo.md fails 7 tests; one
mode's ls-tree mis-rooted fails 3 in either mode; PROJECT_ROOT from
--show-toplevel fails 2.

* fix(#4465): refuse a phase planned in another repository, and never widen a foreign anchor

Found by this round's second pre-push review, against the previous commit.
Running the anchor from the project root is right while the project root
and the caller share a repository. In a `sub_repos` project they do not:
.planning/ lives in a parent repository and the code in child ones, and
from a child findProjectRoot returns the parent. The anchor then came from
the PARENT's history -- a commit the child does not hold. `git rev-parse
"${PHASE_START}^"` failed in the child, the root-commit arm read that as
"no parent", set UNDO_RANGE=HEAD, and selection ran over the child's whole
history. Driven: both child commits selected, including one older than the
phase. Before the previous commit that case resolved no anchor and failed
closed; the previous commit made it destructive.

Two layers, both modes:

- a same-repository refusal in the guard fence: the caller's git directory
  and PROJECT_ROOT's, each physically resolved, must be the same one
  (per-worktree, so a linked worktree compares correctly). A different one
  sets PHASE_DIR_FOREIGN, blanks PHASE_DIR, and the workflow stops with a
  --last N pointer;
- in the anchor: a PHASE_START that is not a commit this repository holds is
  blanked before the root-commit arm. That ambiguity -- "no parent" read as
  "root commit" when it can also mean "not a commit here" -- dates from
  round 1; the previous commit made it reachable.

Shape test O now also pins that PROJECT_ROOT is assigned only from planning
inspect: a rooted call is only as good as the root it is handed, and a later
reassignment passed every spelling check (review, claim 8). Shape test P
pins both layers in both modes.

Correction to the previous commit's message: it said an empty anchor was
the only reason a missed collision could not over-select before that
commit. The review drove a counterexample -- a tracked `.planning` path
coinciding with the caller's subdirectory resolves a wrong, non-empty
anchor -- so that sentence overstated it.

Reversion controls, driven: the previous commit's undo.md fails 3 (P, the
sub_repos refusal, defense in depth); the refusal made inert fails 2 (P and
the refusal; the anchor check alone still keeps the range empty); the
anchor check made inert fails 2 (P and defense in depth; the refusal alone
still refuses). A verbatim negative control shows the pre-hardening
root-commit arm widening the parent's anchor to all of HEAD.

* fix(#4465): compare the repository, not the worktree, and anchor only inside HEAD's history

Found by this round's third pre-push review, against the previous commit.
Its repository gate compared per-worktree git directories. gsd-tools maps
a linked worktree with no .planning/ of its own to the MAIN worktree
(resolveMainWorktreeCwd), so from such a worktree PROJECT_ROOT is the main
checkout: one repository, one object database, two per-worktree git dirs
-- and the gate refused a legitimate undo. It now compares the COMMON git
directory, physically resolved, which is the repository's identity; a
sub_repos child is still a different one and is still refused.

Driving that case showed the anchor check was too weak for it. The anchor
comes from the main worktree's branch history, and nothing guaranteed the
linked HEAD contains it: a linked branch that left main before the phase
started would get a window bounded by a commit outside its own history.
The check is now "PHASE_START is an ancestor of HEAD" (`git merge-base
--is-ancestor`), which subsumes the previous "is a commit here" test -- a
missing object is not an ancestor either -- and changes nothing in the
ordinary case, where the anchor is read from HEAD's own log.

Tests: a linked-worktree fixture (the linked checkout has no .planning/,
which is what makes gsd-tools map it to main; the test asserts the mapping
happened) is allowed in both modes and selects the phase; the same shape
branched before the phase resolves no range. Shape test P pins the
common-dir comparison, forbids --absolute-git-dir, and pins the ancestry
check.

Reversion controls, driven: the previous commit's undo.md fails 3 (P and
both linked-worktree tests); per-worktree git dirs restored fails the same
3; the ancestry check replaced by the previous cat-file test fails 2 (P
and the branched-before test, which then selects a linked commit whose
history never held the phase); no anchor check at all fails 3 (P, defense
in depth, the branched-before test).

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 01:41:35 -04:00
Tom Boucher
e2bfc06558 fix(#4709): a retired runtime id must not resolve to Claude Code (#4756)
* fix(#4709): a retired runtime id must not resolve to Claude Code

AC#1 of epic #4709 — the last unmet acceptance criterion. Every other phase
(#4711, #4716, #4732, #4743, #4753) is merged; the epic does not close until
this lands.

THE DEFECT, MEASURED

Five runtime-resolution accessors resolved a RETIRED id to a plausible-looking
value, indistinguishable from the same call with a canonical id. Measured on
5d4c98cde7 by executing the built modules:

  getRuntimeLabel('gemini')              -> 'Claude Code'
  getProjectInstructionFile('gemini')    -> 'AGENTS.md'
  getGlobalConfigHomeFragment('gemini')  -> "'.claude'"
  getGlobalConfigDir('gemini')           -> ~/.claude   (byte-identical to 'claude')
  getDirName('gemini')                   -> '.claude'

So asking for a runtime Google sunset on 2026-06-18 wrote into Claude Code's
global config home and labelled the install "Claude Code". Nothing errored and
nothing warned.

AC#1 names four accessors. getDirName is the fifth, found by a reviewer: same
module, same silent-wrong-answer class, and it feeds capability-state's
runtimeConfigDir. Fixing only the four the criterion happened to list would
have left the defect reachable, so it is guarded too.

WHY THE CHECK CANNOT LIVE IN CANONICALIZATION

canonicalizeRuntimeName returns null for 'gemini', 'gemini-cli', 'Gemini',
'GEMINI' AND for ''. After canonicalization a retired id, an unknown id and an
empty string are the same value, so anything keyed off the canonical form
cannot tell them apart — it would have to treat all three alike, which is the
behaviour being fixed. The check therefore runs on the RAW input.

WHAT THIS DELIBERATELY DOES NOT DO

The criterion reads "reject a non-canonical runtime id". Taken literally that
overturns three recorded decisions, so the narrower reading was put to the
maintainer as a blocking question and this implements the answer: RETIRED ids
throw, unknown and future ids keep falling back.

Preserved:

  - The #1529 contract, written into getProjectInstructionFile's own docblock
    as a mapping table ending "unknown / future runtimes -> AGENTS.md (safe
    cross-agent default)". That default exists so a runtime GSD has never heard
    of still gets a working instruction file.
  - ADR-1239 Phase B / #1679, which preserved GLOBAL_CONFIG_HOME_FRAGMENTS
    BYTE-FOR-BYTE when it collapsed a 14-branch chain, with golden install
    parity asserting generated hook output is unchanged across every runtime.
  - The explicit `if (!runtime) return <default>` branch. Empty string is a
    supported input, not a non-canonical id.

The distinction the code encodes: ABSENCE OF KNOWLEDGE IS NOT THE SAME AS
RECORDED RETIREMENT. Unknown means "no information, degrade safely". Retired
means "we know it is gone and we know what replaced it" — and silently
substituting a different product for it is the defect.

ONE INACCURACY IN THE CRITERION, RECORDED RATHER THAN REPEATED

AC#1 says the accessors return "a Claude Code value". True for getRuntimeLabel,
getGlobalConfigHomeFragment, getGlobalConfigDir and getDirName — but
getProjectInstructionFile returns 'AGENTS.md', which is not a Claude value at
all. The defect it points at is real for all of them, so the fix covers all of
them, but the wording is wrong for one.

MATCHING

RETIRED_RUNTIME_DETAILS is a Map keyed by canonical retired id, and
RETIRED_RUNTIME_SPELLINGS maps every spelling to that id. Both are Maps, not
object literals: a literal indexed by a computed key resolves INHERITED
properties, so '__proto__' and 'constructor' were truthy and threw with every
field `undefined`, while isRetiredRuntimeId — which already went through a Set
— correctly answered false for the same input. Two guards disagreeing about one
id is worse than either answer. A Map has no prototype keys, so that hazard is
structural rather than patched. The predicate and the assertion now share one
normaliser and one table and cannot diverge.

Candidates are normalised NFKC + lowercase + strip non-alphanumerics. Folding
the separators makes 'gemini-cli', 'gemini_cli', 'gemini.cli' and 'geminicli'
one key instead of four near-misses found one at a time, and NFKC folds the
full-width 'gemini' a CJK keyboard produces. It stays MEMBERSHIP matching,
never prefix or substring: 'gemini-2.5-pro' folds to 'gemini25pro' and
'gemini-3.1-pro-preview' to 'gemini31propreview', neither a member, so Google's
live model ids — part of Antigravity's real on-disk contract — are untouched.

Homoglyph folding is deliberately not attempted, and a Cyrillic 'і' would slip
through. These values arrive from argv and env, trusted inputs here, and a
mapping broad enough to catch deliberate homoglyphs would start catching
legitimate ids. Stated rather than left for the next reader to discover.

This over-broad-match trap is the recurring shape of the whole epic: an
exclusion or match written wider than its subject. Four occurrences, each
cited: #4716's `gemini-[0-9]` sweep exclusion hid a stale review.models.gemini
row whose value was "gemini-2.5-pro" on the same line; #4753's first
model-display escape was a blanket /^ \d/ that laundered "Gemini 2.5 CLI as a
supported runtime."; its dialect rule then used a +/-24-character window that
let one legitimate reference license a live claim 21 characters away; and its
model rule treated the ABSENCE of a runtime word as a grant, passing five
unqualified live-runtime claims. Earlier drafts of this message and its
artifacts said "five" in one place and "three" in another with nothing cited;
it is four, listed here, and the artifacts now agree.

THE THROW

RetiredRuntimeError carries `code: 'GSD_RETIRED_RUNTIME'` so a caller can
handle this case without string-matching a message that may be reworded, and
the message names the id, the successor and the retiring issue.
assertNotRetiredRuntime runs as the FIRST statement of each accessor, including
before getGlobalConfigDir's explicitDir branch, so an explicit directory cannot
mask a runtime that is gone.

`gsd-tools query project-instruction-file --runtime gemini` answered the new
throw with a raw stack trace — a user-facing regression this change introduced.
Its sibling routeSkillsRoot already emitted a clean single-line error for an
unknown runtime, so that route now maps GSD_RETIRED_RUNTIME through the same
`error()` helper, and a test asserts the contract directly: non-zero exit,
stderr naming Antigravity and #1928, and no stack frame. It was the only
unwrapped call site in that CLI; I checked the rest rather than assuming.

getRuntimeNewProjectCommand is deliberately NOT guarded: its value does not
vary by runtime in a way that makes a retired id a wrong answer, so throwing
would cost callers a crash without correcting anything. Verified by observing
it return the same value across claude, codex, opencode, kimi, antigravity,
copilot and an unknown id.

RECONCILING THE TESTS THAT PINNED THE DEFECT

The full remote matrix went red with 14 failures, and every one was a
pre-existing test asserting the fallback this criterion calls a defect. One had
already been caught locally by review; the matrix found the other thirteen
across four files. They were reconciled by intent, not blanket-inverted:

  - Tests whose SUBJECT is the retired runtime — "gemini falls back on label /
    config-fragment / new-project surfaces", "gemini no longer maps to
    GEMINI.md (defaults to AGENTS.md)", "gemini is no longer a known runtime —
    falls back to AGENTS.md" — had pinned the defect, titles and all. Their
    assertions are INVERTED rather than deleted, so the history of what the
    behaviour used to be stays attached to the test that pinned it.
  - Tests whose SUBJECT is "an unregistered id falls back generically", with
    gemini merely the SAMPLE, still assert a TRUE property that this change
    deliberately preserved. Those keep their assertion and switch the sample to
    a genuinely unknown id, with a retired-id refusal pinned alongside so both
    halves of the distinction sit together.
  - The project-instruction-file parity loop dropped gemini from its
    parametrised runtimes — both sides now refuse, so there is no value to
    agree on — and gained a dedicated refusal-parity test.

A FIFTEENTH was then found by executing the touched suites locally, in process,
one file at a time — `tests/runtime-name-policy.test.cjs:135` asserted
`getProjectInstructionFile('gemini-cli') === 'AGENTS.md'`, and its own comment
read "gemini-cli was an alias for gemini", which is exactly why that spelling
is now a retired one rather than a merely-unrecognised one. Inverted like the
rest.

Two remote runs on this change were avoidable: the first by reconciling the
tests that pinned the old behaviour before shipping, the second by executing
the touched suites locally first. The matrix is the authority; it is not the
discovery mechanism. Local per-file execution is bounded and cheap and is not
the banned `node --test` fan-out.

All five touched suites now pass in process: runtime-name-policy 47/47,
gemini-runtime-removed 32/32, project-instruction-file-parity 12/12,
runtime-homes-legacy-ids-drift-guard 2/2, install 452/452.

COVERAGE

Failing-first, one per accessor as the criterion demands, each proven RED
against 5d4c98cde7 before the fix existed — the table at the top of this
message IS that baseline, and the exports the tests import did not exist yet
either.

Asserting only the throw would pass if every id threw, which would break every
install, so each property is paired with its opposite: every canonical id still
resolves on all five accessors with byte-identical values; '' keeps its
documented branch; a genuinely unknown id keeps 'Claude Code' / 'AGENTS.md' /
'.claude' / ~/.claude. That last one is the load-bearing negative — it is the
decision the maintainer chose to preserve, so a later patch that "tightens" the
guard to reject all non-canonical ids turns it red with the reason attached.

Boundary coverage maps limit-1/limit/limit+1 onto set membership: 'gemin',
'geminix', 'gemini-2.5-pro' and 'gemini-3.1-pro-preview' must NOT throw, the
retired id and its folded spellings must. '__proto__', 'constructor' and
'  CONSTRUCTOR  ' are pinned as must-not-throw, and predicate/assertion
agreement is asserted directly. Several assert.throws calls initially passed a
string as the second argument, which node treats as the MESSAGE rather than a
matcher, so they asserted nothing about the error; they now use a real
predicate checking the code.

The tests live in the owning modules' suites rather than a new issue-named
file: lint-regression-test-names rejects new bug-NNNN/fix-NNNN/issue-NNNN test
files outright and directs the regression to the owning module's suite.
scripts/lib/macos-conformance-tier.generated.cjs regenerated through its own
--write path, since the tracked test-file count moved.

Fixes #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4709): backfill changeset PR number (#4756)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 22:22:16 -04:00
0xdhx
48271de43f fix(#4660): widen the 6 shell/markdown phase-id mirrors to the canonical grammar's letter axis (#4744)
* test(#4660): pin the letter-axis parity defect across all 6 shell/markdown phase-id sites

Extends tests/nsegment-phase-grammar.test.cjs (#4568) one axis over: for each
of the six sites, reads the live regex off disk and asserts it agrees with
src/phase-id.cts's PHASE_NUMBER_TOKEN_SOURCE on the letter axis in BOTH
directions — accepts `12A` / `3A` / `03A` / `23A.1.2`, still rejects `3a`,
`3AB`, `A3` and the other canonical-invalid shapes — and that the two
extracting sites return the full letter-suffixed token rather than its digit
prefix (or nothing).

Negative control against the unfixed tree: 22 failures, exactly the
"(fails before the fix)" cases; every reject-parity case already green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* fix(#4660): widen the 6 shell/markdown phase-id mirrors to the canonical grammar's letter axis

Adds `[A-Z]?` after the leading digit run at all six sites #4568 widened —
the ERE translation of src/phase-id.cts's `\d+[A-Z]?(?:\.\d+)*` — so a
documented, canonical-valid id like `12A` or `23A.1.2` is no longer refused
by the four validating sites (code-review.md, code-review-fix.md,
gsd-code-fixer.md, gsd-code-fixer.compact.md) or truncated to its digit
prefix by the two extracting sites (execute-plan.md's plan-filename grep,
plan-phase.md's --research-phase capture). Behaviour is byte-identical for
every id that matched before; the adjacent comment and error-message text
now names the grammar it mirrors.

Driven: `init code-review 3A` on a fixture with a `03A-slug/` directory and
a `### Phase 3A:` heading emits `padded_phase: "03A"`, which the old regex
rejects and the widened one accepts — nothing upstream of the validator
mangles the id.

At execute-plan.md the trailing `-[0-9]+` is the PLAN number and stays
digit-only; plan and milestone dimensions are out of scope per the brief.
`CASE_FLEXIBLE_PHASE_NUMBER_TOKEN_SOURCE` derives from the canonical source
by a literal `.replaceAll('A-Z', 'A-Za-z')`, so src/phase-id.cts is
deliberately untouched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* chore(#4634): extend lint-phase-id-drift to ban a letter-less phase-id mirror in workflows/ and agents/

Adds findLetterlessPhaseMirrorDrift — the letter-axis twin of the #4568
single-segment rule — flagging the unbounded-segment shape
`[0-9]+(\.[0-9]+)*` (and its \d / doubled-backslash near-variants) whose
digit run is NOT followed by the `[A-Z]?` class, on any phase-carrying line
across gsd-core/workflows/**/*.md, gsd-core/references/**/*.md and
agents/**/*.md. Sanctioned the same way (`<!-- phase-id-owner: ... -->`),
tolerates the case-flexible `[A-Za-z]?` directory-scanning variant so it
cannot force that separate axis to narrow, and is wired into scanAll.
Confirmed zero violations against the real tree post-#4660 fix, and one
violation when a single site is reverted.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* docs(#4660): add Fixed changeset

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* chore: regenerate conformance-tier manifests for the extended grammar test

tests/nsegment-phase-grammar.test.cjs now requires the compiled
gsd-core/bin/lib/phase-id.cjs (to assert the canonical grammar agrees with
each site's live regex), which moves it to a different platform-conformance
tier; `gen-platform-conformance-tier.cjs --check` in lint:ci flagged the
macOS manifest as stale.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* test(#4660): reword a comment that tripped lint-docs-guard-registration

The comment mentioned `docs/CONFIGURATION.md` between two backticked
tokens, which the lint's template-literal detector read as a docs/ path
expression. The test reads no docs/ file.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* chore(#4660): refresh the compact-content benchmark baseline and acknowledge emitted growth

plan-phase.md grew by 4 bytes (`[A-Z]?`), which moves the committed
compact-content benchmark; refreshed with `benchmark-compact-content.cjs
--write`. The six shipped files below grew by the widened regex literal plus
the comment and error-message text that now names the canonical grammar.

Emitted-Drift-Ack-Growth: code-review.md — #4660: `[A-Z]?` at the PADDED_PHASE validator plus a comment/error message naming the canonical grammar and the `12A` example
Emitted-Drift-Ack-Growth: code-review-fix.md — #4660: `[A-Z]?` at the PADDED_PHASE validator plus a comment/error message naming the canonical grammar and the `12A` example
Emitted-Drift-Ack-Growth: gsd-code-fixer.md — #4660: `[A-Z]?` at the padded_phase sink validator plus the defense-in-depth comment and error message updated to the canonical grammar
Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — #4660: `[A-Z]?` at the padded_phase sink validator plus the comment and error message updated to the canonical grammar
Emitted-Drift-Ack-Growth: execute-plan.md — #4660: `[A-Z]?` in the plan-filename phase extraction (6 bytes)
Emitted-Drift-Ack-Growth: plan-phase.md — #4660: `[A-Z]?` in the --research-phase capture (6 bytes)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* chore(#4660): set changeset fragment pr to 4744

* chore: re-trigger Validate Branch Name

The required check-branch context was cancelled on this head by the
workflow's cancel-in-progress group when the changeset pr-field backfill
push landed three seconds after the PR opened; no completed run exists for
the current head, and a fork contributor cannot re-run it. Empty commit to
re-run it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-14 19:32:12 -04:00
Tom Boucher
06845717fe feat(#4740): make the Loop Host Contract role partition normative and enforced (#4742)
* test(#4740): pin the per-step role-family partition

Failing-first coverage for the Loop Host Contract role partition. At this
commit crossCheckRoleFamilies does not exist, so the rows throw
"crossCheckRoleFamilies is not a function" -- the RED proof they bind to
behavior rather than restating it.

ADR-894 section 3 assigns roles per step but parenthesises the assignment as
"(illustrative roles)", and nothing enforced it. The only thing standing in the
way was a single deepEqual in this same file, which is editable prose.

Rows cover: each step's own family accepted; a strict subset accepted; a
foreign role rejected at every step; an unknown role rejected; an unknown step
failing CLOSED; capitalization not silently matched; every offending role
reported rather than only the first; and purity, because buildContract puts the
same array into the generated contract.

Two rows exist because an earlier cut of this suite was vacuous. The purity
fixture is deliberately UNSORTED -- an alphabetically-sorted fixture cannot
fail an in-place sort(), and the mutant was being killed by three unrelated
rows instead. A parity row asserts ROLE_FAMILY and ROLE_TO_AGENT cover the
exact same role-name domain, both directions: they are parallel constants over
one domain, so divergence is the generative-fix class CLAUDE.md names.

Every negative row asserts the offending ROLE NAME and the STEP NAME appear in
the message. A count-only assertion survives a mutant that reports the wrong
role, which the 80% Stryker gate would surface only after a full CI round-trip.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#4740): reject a cross-family agent-role declaration

Orchestration and execution are distinct functions of the loop and must not
drift into one another. That partition was real but unenforced: ADR-894
section 3 calls its own role assignment "illustrative", and the generator
accepted anything. Adding orchestrator to execute-phase.md's agent-roles line
compiled, --check passed once regenerated, and capability-validator.cjs then
began accepting into:"orchestrator" at every execute point.

ROLE_FAMILY maps every role to one of orchestration, planning or execution.
EXPECTED_FAMILY_BY_STEP gives each of the five steps exactly one family.
crossCheckRoleFamilies rejects a cross-family role, a role outside the
vocabulary, and an unknown step. It reports every offender, not the first.

It fails CLOSED on an unknown step, deliberately diverging from
assertPointsCoverage's "unknown step -- caught elsewhere". For points that is
true: the canonical-set and duplicate checks catch it. For roles there is no
second net, so failing open would leave an unknown step as the one input that
bypasses the gate.

crossCheckRoles' orchestrator exemption is untouched. ROLE_TO_AGENT maps roles
to agent FILES and the orchestrator is the host, owning none -- admissibility
and agent-file presence are separate concerns with separate checks.

Additive to section 3's existing rule that contribution.into must be a member
of the step's agentRoles, which is unchanged. That governs what a CAPABILITY
may target; this governs what a WORKFLOW may declare. No capability is
affected, and all five workflows already declare single-family sets, so the
gate is green on the commit that introduces it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4740): make the ADR-894 role assignment normative

Section 3 parenthesises its per-step role assignment as "(illustrative roles)".
That word was accurate about the list's PURPOSE -- it illustrated the shape of
a generated contract entry -- and wrong about its STATUS, because the
assignment was load-bearing from the moment the generator consumed it. Read
literally it makes the partition an example rather than a rule.

Appended as a dated in-place section per docs/contributor-standards.md, which
records that an accepted ADR is never rewritten and names this the default
pattern. Section 3's original body is untouched.

The amendment states the three disjoint families, the one family each step
admits, that a step may declare a strict subset but never outside it, and why
this is a clarification rather than a new decision: the contract is generated
from the workflow markers "so it cannot drift into a lie", and all five
workflows have always declared single-family sets. What was absent was any
statement that it is required, and any check that it holds.

It also pins the distinction that is easy to re-merge: contribution.into being
a member of agentRoles governs what a CAPABILITY may target and is unchanged;
the family rule governs what a WORKFLOW may declare. The CONTEXT.md glossary
entry for the Loop Host Contract records the same, beside the agent-reference
drift guard it already documented.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4740): add changeset fragment

pr:0 placeholder is backfilled with the real number once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4740): backfill changeset pr number

Replaces the pr:0 placeholder with 4742 now that the PR exists. Verified with
GITHUB_BASE_REF=next, the way CI runs them: changeset lint and lint:docs both
go from invalid_pr(0) to ok. Without that env both report success without
evaluating the branch at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4740): stop injecting the orchestrator procedure into executors

claude-orchestration declared a contribution at execute:wave:pre with
into:"executor". loop-hook-dispatch.md defines a contribution as "inject
fragment.inline verbatim into the context for the role named in into", so its
267 lines were injected into EXECUTOR prompts whenever the capability was
enabled. Those lines are orchestration end to end -- construct a wave manifest,
resolve the dispatch backend, invoke the Workflow tool to spawn executors,
bridge per-agent results into the merge chain. An executor can act on none of
it.

Retargeting to into:"orchestrator" would not have been a fix. ROLE_TO_AGENT
carries no orchestrator entry by design: the orchestrator IS the host, and the
host's procedure lives in execute-phase.md. A step's agentRoles enumerates
agents a capability may inject context INTO, so adding orchestrator there would
model the host as an injectable agent -- the same category error pointed the
other way, and it would need an exception carved into the partition the same
issue just made normative.

So the defect is the mechanism, not the label. A contribution injects into an
agent's context; "replace step 3's inline dispatch loop" is a change to what
the HOST does. The contribution channel was serving as a host-behaviour
directive because it was the only channel available at an execute point.

The entry is removed. plan:post into:"planner" is correct and untouched. The
procedure is preserved verbatim at docs/workflow-backend-dispatch.md inside the
capability -- it is the only copy in the repo -- and is no longer injected
anywhere.

Consequence, not softened: the Workflow backend now has no loop wiring.
Detection, emission and config remain and the design is intact, but nothing
dispatches it. Under the separation ADR-1143 itself asserts it never had a
legitimate channel; ADR-1143's own audit already records the end-to-end path
has never been exercised. Wiring it properly needs a host-level mechanism that
does not exist today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4740): invert the stale execute:wave:pre registry assertions

Removing the contribution left four surfaces asserting or describing the old
state. Caught by an isolated review before a verification run was spent, which
is the point of reviewing first: the first of these was a guaranteed CI red.

execute-wave-post-gate-pipeline-e2e asserted against the REAL generated
registry that byLoopPoint['execute:wave:pre'] held exactly one contribution
with capId claude-orchestration. It now holds zero. Inverted to assert exactly
0 -- not a vague >= 0 -- and the #2285 comment above it now explains the
current state rather than the one it was written for.

CONTEXT.md's Claude Orchestration entry claimed two contributions at wired
points. It is now one, and the entry's execute:wave:post label was already
wrong before this change: the manifest said execute:wave:pre. Rewritten to one
plan:post contribution, why the execute-point one was removed, and where the
procedure now lives.

One assertion in claude-orchestration.test.cjs could not fail. It tested for
the prose "(into the executor)" while the doc says "(`into: executor`)", so no
plausible wording matched it and the paired plan:post assertion was carrying
the row. Replaced with a check on the structural claim, and proved RED by
restoring the two-contribution wording before reverting.

The moved procedure keeps section headings that speak as a live contribution --
"When this contribution is active", "Why execute:wave:pre". Preserving the body
verbatim was deliberate, so the headings stay and an editor's note under the
header explains why they read that way.

A sweep of all 17 files referencing byLoopPoint found no further siblings: the
remaining hits are a synthetic capability fixture and an empty-points test that
already expected no active hooks, both correct before and after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 18:30:47 -04:00
Tom Boucher
eb49ff98df fix(#4728): stop presenting the retired Gemini CLI as a supported runtime (#4743)
* fix(#4728): stop presenting the retired Gemini CLI as a supported runtime

#1928 removed the Gemini CLI runtime after Google sunset it on 2026-06-18, and
updated the ENGLISH docs. The locale mirrors and the runtime-loaded workflow
prose were not updated in the same change, and no gate asserts the ABSENCE of a
retired runtime, so both drifted quietly for a year.

The finding that shaped this change: English is already correct. docs/
ARCHITECTURE.md, CONFIGURATION.md, USER-GUIDE.md, how-to/install-on-your-runtime.md
and CLI-TOOLS.md carry zero runtime-axis Gemini references; the only English hits
anywhere are a Gemini 2.5 Pro MODEL line, the GEMINI_API_KEY row, and prose that
correctly documents the retirement. So the docs half of this is translation lag,
not a content decision, and every locale edit here is parity with an existing
English line rather than new wording:

  - install-on-your-runtime.md  English has NO `### Gemini CLI` section  -> deleted
  - USER-GUIDE.md :843          "…, Antigravity CLI, Kilo)"              -> substituted
  - ARCHITECTURE.md             English has NO Gemini CLI table row      -> row deleted
  - ARCHITECTURE.md :24         English holds `Kimi CLI` in that slot    -> Kimi CLI
  - context-monitor.md :3       "`AfterTool` for Antigravity CLI"        -> substituted
  - spike-and-sketch.md :93     "(Codex, Antigravity CLI, etc.)"         -> substituted
  - configure-model-profiles    "Codex, OpenCode, Antigravity CLI, or Kilo" -> substituted
  - COMMANDS.md                 English keeps only hyphen + Codex bullets -> colon bullet deleted
  - FEATURES.md                 source docs/features/multi-runtime-support.md:10
                                lists no Gemini CLI                       -> name removed

ARCHITECTURE.md:24 is the clearest case for reading English rather than
substituting blind: Antigravity ALREADY appears later in that list, so replacing
Gemini CLI with Antigravity would have named it twice. English holds Kimi CLI
there, so that is what the locales get.

The largest single class was hand-duplicated boilerplate. A "Text mode" paragraph
repeated across 34 runtime-loaded workflow files ends "…required for non-Claude
runtimes (OpenAI Codex, Gemini CLI, etc.)". No lint enforces that sentence and no
script syncs it, so every copy was edited. These files are read by the agent at
runtime, so they steer behavior rather than only informing a reader — which is why
this class matters more than its word count suggests.

The slash-command-form section is restructured in all four languages to match
English, which had already dropped its colon-form bullet. That bullet claimed the
colon form is "Gemini CLI only", which was false on its own terms independent of
the retirement: `/gsd:…` is GSD's canonical AUTHORING token, rewritten per runtime
at install time, and NO runtime registers it — VALID_COMMAND_STYLES is
{slash-hyphen, shell-var} and 18 of 19 runtimes declare slash-hyphen. Substituting
the runtime name would have left the claim false with Antigravity's name in it, so
the claim is gone, matching English.

Two anchor regressions were caught and fixed while doing that. zh-CN lost its
explicit {#slash-command-forms-hyphen-vs-colon} anchor while its TOC still linked
it; the anchor is restored. ko-KR and pt-BR never had an explicit anchor and rely
on the slug generated from the heading text, so shortening the heading broke their
own TOC links; those links now point at the new slugs. English's heading lost its
anchor while its TOC still links the old one — that latent English bug is
deliberately NOT copied.

Preserved, because `gemini` is not one thing here and a blanket sweep breaks the
product: ~/.gemini/antigravity{,-ide,-cli} and ~/.gemini as their parent;
~/.gemini/config (#3738); GEMINI.md; hookEvents "gemini"; GEMINI_API_KEY in all
four locales; every gemini-* model id and the Gemini 2.5 Pro references in
ko-KR/pt-BR/zh-CN (ja-JP genuinely lacks that line — the locales have diverged, so
a uniform patch would be wrong); the hook-event dialect notes, which are
RE-ATTRIBUTED rather than deleted because Antigravity inherits that dialect;
reapply-patches.md:93's legacy-install note; host-integration-capability-matrix.md
:27 and :342, which correctly record the sunset and Antigravity's contract;
whats-new-1.7.0.md and FEATURES.md:3506, which document the retirement itself; and
the generated launcher preamble, which belongs to epic #4632 — zero
_GSD_SHIM_NAME lines appear in this diff.

Coverage: a #4728 block in tests/gemini-runtime-removed.test.cjs asserts the
retired name is gone from STRUCTURAL POSITIONS (a level-3 heading, a table row's
first cell, a runtime-example parenthetical) rather than asserting the string is
absent, which would be wrong. It pairs those with positive PRESERVE assertions
over the same files — Antigravity's heading, ~/.gemini/antigravity, GEMINI_API_KEY,
AfterTool — so a patch that deletes too much fails as loudly as one that deletes
too little. The model-axis test pins both the presence in three locales and the
absence in ja-JP, so a later uniform patch that "helpfully" adds it back fails.
The new docs/ reads tripped lint-docs-guard-registration for the first time in
this file, so the test is registered in scripts/docs-guard-registry.cjs.

Not covered here, by design: nothing above would catch a Gemini-as-runtime
reference appearing in a NEW file tomorrow. That is the repo-wide drift guard,
#4729, which must land last — written now it would red on the very references this
change removes.

Fixes #4728

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4728): fix four review blockers, including a vacuous test and my own duplicate

A full matrix run on 31f12d7943 FAILED with 3 real failures, and an isolated
adversarial review returned BLOCK on four blockers. All of it was correct.

1. I committed the exact error I claimed to have avoided. The commit message
   boasted that ARCHITECTURE.md:24 proved the value of reading English rather
   than substituting blind, because Antigravity already appeared later in that
   list. Five hundred lines further down the SAME four files, my
   `Gemini:` -> `Antigravity:` substitution produced TWO consecutive
   `- Antigravity:` bullets, because an Antigravity bullet was already there.
   English (ARCHITECTURE.md:827) merges them into one. Now merged in all four
   locales, reusing each locale's existing words.

2. `--gemini` survived in the runtime-detection CLI flag list in all four
   locale ARCHITECTURE.md files. English:817 holds `--kimi` in that slot and
   already lists `--antigravity` later, so this is another place where
   substituting Antigravity would have duplicated it. Now `--kimi`.

3. Two runtime-loaded workflow files still enumerated Gemini one line ABOVE the
   line I had already corrected -- the "Adaptive (Recommended)" option in
   settings.md:192 and new-project/steps/auto-mode-config.md:95.

4. THE NEW TEST WAS VACUOUS for two of its five files. It matched only
   `non-Claude runtimes (` and `(e.g. `, and neither regex could reach the two
   lines the change actually fixed: health.md:52 reads `non-Claude (Codex, ...)`
   without the word "runtimes", and execute-phase.md:1028 has no parenthetical
   at all. The reviewer proved it by re-introducing Gemini at both lines and
   watching the assertion stay GREEN. That same blind spot is what hid finding 3.

   Replaced with a case-sensitive `/\bGemini\b/` walk over every
   `gsd-core/workflows/**/*.md`, which works because every LEGITIMATE gemini
   reference in that tree is spelled differently and cannot match: Antigravity's
   paths are lowercase with a slash (`~/.gemini/antigravity`), Google's model ids
   are lowercase and hyphenated (`gemini-3.1-pro-preview`), and the env vars are
   uppercase (`GEMINI_CONFIG_DIR`, `GEMINI_SESSION_ID`). A bare capitalised
   `Gemini` there means the retired RUNTIME is being named. The walk asserts it
   found at least 50 files so an empty walk cannot pass vacuously, and it now
   covers the nested `new-project/steps/` directory where finding 3 lived.

   Two allowlist entries, both by line CONTENT and both justified:
   reapply-patches.md's `Legacy: ... pre-#1928` note, and settings-advanced.md's
   `Known provider` menu. The second was escalated by the agent rather than
   decided: Section 8 of that file says model policy is defined "independently"
   of the runtime, so `(Claude / OpenAI / Gemini / Qwen)` is the PROVIDER axis --
   the same axis as the lowercase model ids -- and must keep working.

   Proven to fail, not just asserted: the predicate reports 0 offenders on the
   real tree and exactly 2 on a /tmp copy with Gemini re-injected at
   health.md:52 and execute-phase.md:1028.

Also from the review: a `| Gemini |` COLUMN survived in the locale FEATURES.md
comparison tables (English has none) -- removed from all three, with header,
separator and every body row kept aligned; two ENGLISH runtime-axis sites were
missed by my own parity standard (how-to/execute-a-phase.md:88 and
how-to/verify-and-ship.md:89, the latter doubly stale since #4716 retired the
Gemini reviewer lane); docs/USER-GUIDE.md:12 linked a dead anchor, which I had
found and deliberately left -- record-and-proceed on a known defect is exactly
what the rules forbid, so it is fixed; docs/COMMANDS.md:12 and all four mirrors
still claimed "the hyphen and colon forms are runtime-specific spellings" with
no colon form documented anywhere, so that false sentence is deleted; and ko-KR
had the installer rather than the user doing the targeting.

The other two matrix failures were the compact-content benchmark baseline, which
drifted because this PR changes byte counts, refreshed via the script's own
`--write` path rather than by hand; and this commit's emitted-drift-ack trailers.

Method note on the acks: the failing run measured growth against
origin/next@1110c3b4ee, which is the STALE LOCAL `next` ref -- gsd-test merges
into the local base branch, and this machine's `next` is seven commits behind
origin/next, which is checked out in the main worktree and so cannot be
fast-forwarded from here. The 32 trailers below are computed against the REAL
base (origin/next @ ca8d9d4459) by comparing each tracked file's blob size, which
is one more file than that run reported -- the extra is settings.md, grown again
by fix 3. docs-update.md and map-codebase.md are deliberately NOT acked: they
SHRANK, since there the fix deleted ", Gemini CLI" rather than substituting, and
acking a file no delta consumed is itself an error.

Refs #4728

Emitted-Drift-Ack-Growth: add-tests.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: add-todo.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ai-integration-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: check-todos.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: cleanup.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: complete-milestone.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: do.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: eval-review.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: execute-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: execute-plan.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: health.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: import.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: inbox.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: manager.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: new-milestone.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: new-workspace.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: note.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: onboard.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: plant-seed.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: profile-user.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: quick.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: remove-workspace.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: secure-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: settings.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ship.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: smart-entry.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ui-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ui-review.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: undo.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: update.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: validate-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: verify-work.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4728): add the changeset fragment

The PR body claimed one was present and it was not — caught by
scripts/changeset/lint.cjs reporting fail_missing_fragment, not by the
checklist, which is exactly why the lint exists.

Type Fixed: the diff is prose, and a docs-only fix uses Fixed since there is
no Documentation type.

Refs #4728

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 16:49:52 -04:00
Tom Boucher
b54c1c5848 fix(#4709): retire the Gemini CLI reviewer lane (#4716)
* fix(#4709): retire the Gemini CLI reviewer lane

Google stopped serving Gemini CLI for the free/Pro/Ultra tiers on 2026-06-18 —
the same sunset that removed the gemini RUNTIME in #1928 (shipped 1.8.0). GSD
targets solo developers, so those tiers ARE the user path: the lane spawned
`gemini {{model}} -p -`, a binary that no longer answers for the majority of
users, and five locales documented it as a supported choice.

The lane was re-created after #1928 by the reviewer-lane-as-manifest-data work
(6a9babda69, #2798/#2837, ADR-2782). Per the maintainer that re-creation was an
error in that buildout rather than a considered decision, so this corrects a
mistake and needs no ADR-2782 amendment.

Reviewer roster: 12 lanes / 13 flags -> 11 lanes / 12 flags.

TWO sources of truth had to be removed, not one. Deleting
capabilities/gemini/capability.json left the capability registry at 11 lanes
while src/review-lane-descriptor.cts's hand-maintained REVIEWER_LANES array
still carried its own complete gemini entry at 12 — precisely the disagreement
checkReviewerLaneParity exists to catch. Both are gone; both parity checkers
now run clean against the real tree (lane parity ok/0 violations, docs parity
0 violations).

Surfaces stripped of the dead flag:
- capabilities/gemini/ deleted; registry and capability-matrix regenerated
- src/review-lane-descriptor.cts: REVIEWER_LANES entry, docblock count, and the
  three doc comments that used --gemini as a live example
- commands/gsd/{review,plan-review-convergence,autonomous,progress}.md and the
  four matching skills/*/SKILL.md: argument-hint frontmatter and flag bullets
- gsd-core/workflows/help/modes/{full,full.compact}.md: /gsd-help signatures,
  the detected-CLI list, and the reviewer-title list
- gsd-core/workflows/settings-integrations.md: the integrations wizard no longer
  offers "Gemini" as a model option, and the settable-keys list drops it
- gsd-core/workflows/review.md: the `command -v gemini` probe, the --gemini
  flag, the roster frontmatter, the install pointer to the sunset repo, and the
  jq-less / precedence / self-skip lane lists
- gsd-core/workflows/sync-skills.md: "two runtimes (grok, gemini) resolve to
  ANOTHER runtime's skills root" is now one runtime; gemini never aliased
  anything, it fell through canonicalizeRuntimeName to a fail-closed default
- docs/{CONFIGURATION,COMMANDS,CLI-TOOLS}.md, docs/reference/capability-matrix.md,
  docs/how-to/set-up-cross-ai-review.md — including its `npm install -g
  @google/gemini-cli` instruction and the two rows recommending --gemini
- docs/features/{cross-ai-peer-review,opt-in-parallel-reviewer-lanes}.md as the
  generator inputs behind docs/FEATURES.md, plus the three locale FEATURES.md
  signature lines the docs-parity gate covers (the #2781 class: a flag change
  that never reaches the mirrors)

Counts reconciled against measurement rather than arithmetic: 8 timeout keys of
11 lanes, 11 budget keys, 9 model keys, and four hardcoded literals in
tests/reviewer-lane-declarations.test.cjs (NEW_LANE_ONLY_IDS 5->4, LITERAL_ROSTER
12->11, two roster counts 12->11).

BEHAVIOR CHANGE, accepted deliberately: `gsd config-set review.models.gemini`
now errors with "Unknown config key". An existing key already in
.planning/config.json still parses and is simply never read, so no project fails
to load. This is the repo's own documented policy for exactly this case
(docs/CONFIGURATION.md:327 — "a key left over from a removed reviewer validated
silently and was never read. Such a key is now rejected by config-set"), so no
installer migration ships. Note my first measurement of this was WRONG: I tested
config-get, which reads undeclared keys fine, and generalised. Read and write are
different surfaces and gave different answers.

Antigravity is untouched throughout — its --antigravity/--agy flags,
review.models.agy, ~/.gemini/antigravity configHome, ~/.gemini/config global
skills root (#3738), hookEvents "gemini", GEMINI.md instruction file, and every
gemini-* model id it actually runs on.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4709): changeset for the reviewer-lane retirement

Type Removed: the --gemini flag and its three config keys are user-visible
surface that no longer exists.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4709): close the 24 test failures and the locale-doc gap the gates found

An adversarial review and a full matrix run between them found substantially
more fallout than inspection had. All of it is this PR's own, and all of it is
fixed rather than waved off.

THE MATRIX RUN FOUND 24 FAILURES ACROSS 6 FILES. Inspection had predicted two.
The dominant class was a test helper that looks up a lane by slug and throws
`no declared lane 'gemini'`:

- tests/feat-2483-review-claude-mds-guard.test.cjs (6) — used gemini as the
  "other declared first-party lane" to contrast against claude's env
  suppression. Now qwen, verified from source as a lane that declares no `env`
  (only claude does), so the contrast still holds.
- tests/review-lane-descriptor.test.cjs (6) — the duplicate-flag and
  duplicate-section fixtures deliberately COLLIDED with a real declared lane to
  prove the parity checker reports a duplicate. `--gemini`/`Gemini` no longer
  collide with anything, so the checker reported
  `descriptor_lane_not_in_registry:acme` instead and the tests proved nothing.
  Now collide with `--codex`/`Codex`, reproduced against the real checker.
- tests/review-reviewer-selection.test.cjs (3) — these distinguish KNOWN-but-
  undetected from UNKNOWN. gemini flipped categories, inverting what they
  proved. The known case now uses qwen; `__nope__` stays the unknown fixture.
- tests/review-default-reviewers-resolution.test.cjs (2), and
  tests/settings-integrations.test.cjs (3) — the wizard now offers three
  reviewer CLIs, not four, so the test and its name say three.
- Two count assertions the earlier sweep missed outright:
  reviewer-lane-declarations.test.cjs:359 (`length, 12`) and
  reviewer-docs-parity.test.cjs:681 (`>= 12`).

THE LOCALE-DOC GAP, and why the parity gate stayed green over it. All four
locale mirrors still documented `--gemini` as a live reviewer flag. The
docs-parity checker asserts the PRESENCE of every current flag and never the
ABSENCE of a retired one, so "0 violations" was never evidence those files were
clean — my earlier reading of it as such was wrong. This is the #2781
locale-drift class in the opposite direction. Fixed across 12 locale files:
COMMANDS.md flag lists and table rows, CONFIGURATION.md `review.models.gemini`
rows and reviewer prose, CLI-TOOLS.md config examples, and
set-up-cross-ai-review.md including its install block and its
which-reviewer-to-choose row, which now recommends Antigravity.

ALSO FOUND, and instructive about my own method: docs/CONFIGURATION.md:297 still
carried a `review.models.gemini` row. My sweep had missed it because my grep
excluded lines matching `gemini-[0-9]` to spare Google's model ids — and that
row's example value is `"gemini-2.5-pro"` on the same line. The exclusion built
to avoid false positives created a false negative.

Remaining comment/example sites: src/review-reviewer-selection.cts:309 and
src/config.cts:598 named the dead flag and key as examples;
gsd-core/references/planning-config.md:269 likewise; and
review-reviewer-selection.cts:22 claimed in the PRESENT tense that gemini is a
lane-only reviewer capability. Line 38 of that same docblock says "Before this
phase the five non-runtime reviewers (gemini, ...)" and is left exactly as is —
that is past-tense history, and rewriting it would falsify the record.

Deliberately still deferred to Phase 4, because it is the RUNTIME axis rather
than the reviewer lane: the locale install-on-your-runtime.md `--gemini --global`
instructions, the USER-GUIDE colon-form notes, and the ARCHITECTURE
runtime-detection flag lists.

Both parity checkers green against the real tree; lint:ci exit 0.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4709): backfill the changeset PR number

pr: 0 -> 4716, now that the PR exists. Never guessed ahead of the number.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 03:03:44 -04:00
Tom Boucher
6f99e493e7 fix(#4395): make the debug session manager's own gsd-debugger spawn blocking (#4718)
* test(#4395): prove the manager spawns its debugger without blocking

Failing-first regression coverage for #4395.

debug.md:209 mandates the orchestrator to session-manager spawn carry
run_in_background: false, and says why outright: "Claude Code backgrounds
subagents by default, and only that flag makes the spawn return the
compact session summary directly" (#2196).

The session-manager to debugger spawn, one level down, carries no flag.
Measured: run_in_background appears nowhere under agents/ -- only in
gsd-core/workflows/. So by the rule #2196 itself states, that spawn is
backgrounded, Step 3 ("Handle Agent Return") has no return to inspect,
the manager emits CONTINUE_REQUIRED, the orchestrator auto-resumes per
#2257/#3448, and a second detached debugger races the first on
.planning/debug/<slug>.md.

Row 4 is the load-bearing one: it closes the CLASS by requiring every
subagent spawn under agents/ to declare run_in_background explicitly, so
the next agent that spawns one has to decide rather than inherit a silent
host default. It is scoped to agents/ precisely so it cannot misfire on
the workflows that deliberately use true for parallel fan-out.

Rows 5-7 are pins, not fixes: the #2196 mandate one level up, Step 2 as
the single spawn-format source that the eight continuation sites delegate
to, and the survival of CONTINUE_REQUIRED (which has a legitimate trigger
unrelated to this defect).

Red round: 4 of 7 rows fail. Row 3 needed hardening first -- asserting
only that the two variants AGREE passed vacuously, because two missing
flags are also equal; it now asserts each is present before comparing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4395): make the manager's own debugger spawn blocking

The orchestrator-to-manager hop already requires a blocking spawn and says
why (#2196, debug.md:209): Claude Code backgrounds subagents by default,
and only run_in_background=false makes the spawn return its summary. The
manager-to-debugger hop, one level down, carried no flag -- measured,
run_in_background appeared nowhere under agents/ at all.

So that spawn was backgrounded. Step 3 ("Handle Agent Return") opens
"Inspect the return output for the structured return header" -- with
nothing to inspect, the manager correctly declined to fabricate a terminal
summary and returned CONTINUE_REQUIRED; the orchestrator correctly
auto-resumed (#2257/#3448); the resumed manager reached Step 2 and spawned
a SECOND detached debugger. Both then raced on .planning/debug/<slug>.md.

Every observable in the report follows with no further assumption,
including the count: the reporter saw exactly three collisions in one
invocation, and debug.md:251 caps auto-resumes at three per slug -- one
collision per cycle.

Fixed at the cause, in both shipped variants, kept byte-consistent. The
eight continuation sites say "see Step 2 format", so they inherit it.

The issue offered two remedies. The second -- have the auto-resume path
reconcile a still-running debugger before spawning another -- is not taken:
it treats the symptom, and needs machinery that does not exist (no portable
way to enumerate or stop another runtime's live agents, plus an in-flight
sentinel with staleness and recovery rules, or an orphaned marker deadlocks
the session permanently). With the spawn blocking, the manager cannot reach
Step 4 while a debugger is live, so such a guard would also be unreachable.

#2257, #3448, the anti-loop heuristic, the cap of three, and the
CONTINUE_REQUIRED shape are all correct and untouched. CONTINUE_REQUIRED
keeps its legitimate trigger: the manager genuinely exhausting its own turn
budget mid-investigation.

Also corrects the red-round test to the canonical CALL form. debug.md
writes run_in_background=false inside Agent(...) and run_in_background:
false in prose; the first draft asserted the prose form, which the shipped
call would never have matched.

Emitted-Drift-Ack-Growth: gsd-debug-session-manager.md — the blocking spawn flag plus the note recording why an unstated flag produced colliding debuggers
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.compact.md — same change as its full sibling, kept byte-consistent with it

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4395): add changeset fragment

pr:0 placeholder is backfilled with the real number once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4395): refresh the variant benchmark baseline

Two entries move.

gsd-debug-session-manager.md 4766/4477 -> 4938/4649 is this change: the
blocking-spawn flag plus its explanatory note, added to BOTH variants to
keep them byte-consistent, so the compact sibling grows by the same amount
and the pair's reduction ratio dips 6.06 -> 5.85. The compact file remains
strictly smaller than its canonical sibling, which is what the variant
guard's size check actually requires.

gsd-code-fixer.md 10741 -> 10740 is NOT from this branch -- the file is
untouched here. It has scored 10740 since f334f277dd (#4324) reworded a
line without refreshing this fixture, so the stale number is sitting on
next. Fixed here rather than deferred; the #4350 branch carries the
identical one-token correction, so whichever lands first makes the other a
no-op.

Refreshed with the variant script's own --write.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4395): state the blocking rule agent-wide and correct two overclaims

Review round. Three substantive corrections, one of them to a claim I made
in the previous commit message.

1. The blocking rule is now stated AGENT-WIDE, not per-call. My earlier
   claim that "the eight continuation sites say 'see Step 2 format', so
   they inherit it" was false: exactly ONE of them names Step 2 (the
   compact variant says "Step 2 format" without the "see", which is why
   the first draft of the test matched it zero times there). The other
   sites inherit only because Step 2 holds the sole Agent() spawn literal
   in each file. Both variants now say so outright, and the test pins the
   sole-literal invariant in BOTH variants rather than the prose wording
   in one.

2. debug.md's two auto-resume buckets now restate run_in_background=false
   for the re-spawn. "The same session_params" does not carry it --
   session_params is prompt content, not the spawn flag.

3. The test extractor now also matches single-line Agent(...) calls. The
   class guard was blind to exactly the shape a future offender is most
   likely to take.

Also corrects the diagnosis: remedy 2 is NOT unreachable once the spawn
blocks. This agent's own retained CONTINUE_REQUIRED trigger is "turn
budget exhausted WHILE THE DEBUGGER IS STILL INVESTIGATING", so a harness
turn cutoff mid-wait still double-spawns; debug.md:209 names the same
class from the other side. What this fix removes is the SYSTEMATIC case --
every invocation, because the spawn was always backgrounded. The residual
turn-cutoff window survives, bounded by the existing three-resume cap, and
is stated in the PR rather than denied.

Emitted-Drift-Ack-Growth: gsd-debug-session-manager.md — blocking spawn flag, the note recording why an unstated flag produced colliding debuggers, and the agent-wide restatement the per-site inheritance actually depends on
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.compact.md — same change as its full sibling, kept byte-consistent with it
Emitted-Drift-Ack-Growth: debug.md — both auto-resume buckets restate the spawn flag, since session_params does not carry it
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4395): backfill the changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 02:12:46 -04:00
Tom Boucher
ed819aa4d6 fix(#4379): make the TDD RED-commit pathspec language-agnostic (#4715)
* fix(#4379): make the TDD RED pathspec language-agnostic

The pathspec IS this gate's definition of "a test file", and it listed
only JS/TS conventions. Go's *_test.go matches none of them, so a commit
adding a failing Go test was invisible, RED_COMMIT came back empty, and
every behaviour-adding task halted with TDD GATE TRIPPED. references/tdd.md
already advertises `go test ./...` and `cargo test` as supported, so the
gate was refusing to see tests the docs promised to support.

Two corrections, both measured against a seeded repo rather than reasoned:

- cover the conventions tdd.md advertises: *_test.go, test_*.py,
  *_test.py, *_test.exs, *_spec.rb, *_test.rb.
- drop the `**/` prefix. It does NOT match a path with no directory
  component, so a root-level foo.test.js was invisible even in the
  language the gate did support -- a second defect the report did not
  mention. A bare glob matches at every depth.

Deliberately not widened to ordinary source: a pathspec matching
implementation files would make the gate pass on any in-scope commit,
which is worse than tripping wrongly. Rust is a known gap for that exact
reason and is now documented rather than silently broken.

Driving the shipped pathspec against a seeded repo: before, 0 of 7
language/root conventions matched; after, 7 of 7, with src/impl.go and
src/lib.rs correctly unmatched in both.

Emitted-Drift-Ack-Growth: execute-phase.md — the widened pathspec plus the comment recording why it must not cover ordinary source
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4379): drive the shipped RED pathspec against a real repo

The existing row pinned the pathspec as a literal string, which the fix
makes stale. Re-point it, and add behavioural coverage that EXTRACTS the
pathspec from the shipped workflow and runs git log with it against a
seeded repo -- re-typing the pattern into the test would only assert that
two copies of a string agree.

Rows: every advertised convention is visible; a root-level test file is
not invisible (the half the report missed); existing JS/TS still matches;
implementation files never match, so the gate can still trip; and Rust
inline #[test] stays out of reach, asserted rather than left silent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4379): add changeset fragment

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4379): be honest about the widened pathspec's cost

Adversarial review: the rationale comment claimed the change was safe
without naming what it gives up. `*.spec.*` can match a non-test file
carrying the word (api.spec.json, openapi.spec.yaml), which lets the gate
pass on a commit touching only that. Not new -- `**/*.spec.*` already
matched those at any nested path, so dropping `**/` extends the same
class to the root -- but the comment should say so rather than imply
the widening is free.

Also: the tdd.md list named Ruby and Elixir as recognised while the
detection step above it enumerates only Node/Python/Go/Rust. Say
explicitly that the gate's pathspec is wider than the detected project
types, and why that is deliberate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4379): drop a vacuous row, fold its point into a real one

Adversarial review: the "rust inline #[test] remains out of reach" row
asserted src/lib.rs never matches -- the identical assertion to the
"implementation files never match" row directly below it. It exercised
nothing about #[test] semantics and would have passed against almost any
fix, so it was coverage theatre.

Delete it and move its rationale into the row that already carries the
assertion, where it explains WHY the Rust gap follows from that row
holding: the two cannot both be satisfied by a path-based gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4379): move the pathspec rationale out of a size-capped file

execute-phase.md sits under a FROZEN byte ceiling (ADR-857 Phase 6,
#1168: < 93600). The 20-line rationale comment I added pushed it to
93933 and tripped seven tests, all the same ceiling. Base was 92371, so
the budget was 1229 bytes and the comment spent 1481.

Keep six lines at the call site -- what the pathspec is, why it is not
wider, where to read more -- and move the trade-off detail to
references/tdd.md, which has no ceiling. That is the right home anyway:
the workflow is loaded into context on every run, the reference is read
on demand.

92882 bytes, 718 under. Pathspec line byte-identical; re-proved
behaviour after the trim: 0/7 conventions before, 7/7 after, no
implementation files matched either way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4379): refresh the compact-content benchmark baseline

execute-phase.md changed size, so the committed baseline drifted. The
script's own contract makes it a report that exits 0, but the test
asserts the committed baseline is up to date -- refresh via --write,
which is what the drift message instructs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4379): backfill the changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 00:57:08 -04:00
Tom Boucher
ec22377a9d fix(#4351): run-scope preserved review evidence (#4713)
* fix(#4351): run-scope preserved review evidence

The preserve block copied lane output into a flat .review-diagnostics/
using each file's source basename. A lane slug is stable across runs, so
the destination was stable across runs too -- and cp over an existing
file is a success, so a second review of the same phase destroyed the
first run's evidence with no error and no warning, in the one directory
that exists to outlive the rm -rf beside it.

Copy into one subdirectory per run instead. $RUN_DIR is mktemp -d, so
its basename is already unique per run by construction; the UTC stamp in
front is only a sort key and is omitted if date fails. Destination-only:
nothing inside $RUN_DIR is renamed, because both
prepare_trimmed_prompt_for_reviewer and the lane invocation resolver
depend on those exact basenames.

Verified by extracting the real fence and running it twice against one
phase dir: before, one report survived and it was run 2's; after, both.

Emitted-Drift-Ack-Growth: review.md — per-run diagnostics subdirectory plus the comment explaining why uniqueness cannot come from the clock
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4351): cover repeated runs, resolve through the run subdir

Adds the two-run rows the issue asks for: both runs' reports recoverable,
and a lane failing identically twice leaving two stubs. Both assert on
CONTENT, not a file count -- a clobber producing the same number of
files would pass a count-only check, and the defect is that run 1's
bytes were replaced.

runWriteReviewsFlow gains an optional phaseDir so a caller can run the
flow twice against one phase directory, which is the only arrangement
that can observe the overwrite. Omitted, it mints a fresh one as before.

The existing rows asserted a flat readdir of the diagnostics root, which
the fix makes stale. They now resolve through preservedPath/preservedNames
so they keep asserting WHICH files were preserved rather than silently
becoming assertions about the layout; the layout is pinned once,
explicitly, by oneRunSubdirectoryPerRun_4351.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4351): add changeset fragment

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4351): name the run, not the clock, in the changeset

Adversarial review: the fragment said each run gets its own "timestamped
subdirectory", which reads as though the timestamp provides the
separation. It does not -- collision-safety is mktemp's random basename,
and the stamp is a sort key that is dropped entirely when date fails.
The code comment already said so; the user-facing text now agrees.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4351): backfill the changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:58:38 -04:00
Tom Boucher
1110c3b4ee fix(#4709): stop minting the retired gemini runtime id in shipped surfaces (#4711)
* test(#4709): assert no shipped surface mints a retired runtime id

Extends the #1928 removal guard to the surfaces it structurally could not
reach. Its own docblock scopes it to the installer CLI contract and the
runtime-name-policy exports; it spawns the installer and inspects module
exports, and never reads gsd-core/workflows/**, commands/** or skills/**.

Four structural assertions, all RED on next:
- every RUNTIME= assignment must name a canonical runtime
- the runtime->model-tier table must name only model-catalog runtimes
- runtime selection menus must offer only canonical runtimes
- config-set runtime / model_profile_overrides examples must be canonical

Structural, not textual: each asserts the literal is canonical or the runtime
exists as a catalog key, never that the string "gemini" is absent. That string
is load-bearing across Antigravity's real on-disk contract, so a fifth test
pins that contract from the descriptor (not from a resolved path, which would
read $ANTIGRAVITY_CONFIG_DIR and the real $HOME -- the #4312 defect class).
An over-broad gemini -> antigravity replacement fails there rather than ships.

The menu assertion is scoped by the nearest preceding `question:` matching
/runtime/i, because the same file carries a provider menu (anthropic, openai)
and a budget menu (high, medium, low) whose labels are single lowercase tokens
too and name neither a runtime nor anything the policy should judge.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4709): stop minting the retired gemini runtime id in workflow text

#1928 removed the gemini runtime after Google sunset Gemini CLI on 2026-06-18,
but the removal stopped at the installer boundary. Runtime-loaded workflow text
kept assigning the id, and the name policy's unknown-id fallbacks then applied a
default designed for a never-known FUTURE runtime to an id GSD itself retired:
getRuntimeLabel('gemini') is 'Claude Code', getProjectInstructionFile('gemini')
is 'AGENTS.md', getGlobalConfigDir('gemini') is ~/.claude. A stale id produced a
plausible wrong answer instead of an error.

Those fallbacks are DELIBERATE and are left untouched here -- four docblocks
document them, src/runtime-name-policy.cts:220-222 calls the label default "the
always-safe default, fail-closed", and an existing test in this very suite pins
getProjectInstructionFile('gemini') === 'AGENTS.md'. This commit removes the
REACHABILITY of the retired id instead:

- new-project.md, ingest-docs.md: the runtime-detection cascade mapped
  /.gemini/ and $GEMINI_CONFIG_DIR to RUNTIME=gemini. Both now map
  /.gemini/antigravity{,-ide,-cli}/ and $ANTIGRAVITY_CONFIG_DIR to
  RUNTIME=antigravity, the documented successor. ingest-docs.md was not in the
  original report; the new structural test found it.
- settings-advanced.md: dropped the `gemini` row from the runtime->model-tier
  table. The model catalog has no gemini runtime (runtimeTierDefaults has 18
  keys, none of them gemini), so the row advertised built-in defaults for a
  runtime whose config key is ignored. Its three model IDs were copied from the
  `google` PROVIDER preset -- a provider axis rendered as a runtime axis.
- settings-advanced.md: removed the `gemini` / "Gemini CLI." runtime menu
  option and its group listing, so no menu offers a runtime GSD cannot install.
- settings-advanced.md: repointed the config examples from `runtime gemini` to
  `runtime antigravity`, which ships no built-in tier defaults and is therefore
  the case those overrides actually exist for.
- reapply-patches.md: $GEMINI_CONFIG_DIR -> $ANTIGRAVITY_CONFIG_DIR,
  ~/.gemini/gsd-local-patches -> ~/.gemini/antigravity/gsd-local-patches, and
  the local scan's bare .gemini -> .agents (Antigravity's localConfigDir). This
  file is hand-written, so `npm run sync:launcher` never reached it.
- update.md: bare ~/.gemini and ./.gemini as GSD config dirs -> the real
  ~/.gemini/antigravity and ./.agents.

Antigravity's own Gemini-family surfaces are untouched by design: ~/.gemini as
its configHome parent, ~/.gemini/config for global skills/agents (#3738),
hookEvents "gemini", GEMINI.md as its projectInstructionFile, the
~/.gemini/antigravity{,-ide,-cli} ambiguity probes (#1441), and every gemini-*
model ID. The launcher's own GEMINI_CONFIG_DIR arm is left to #4632, which
absorbed #4347 for it.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4709): record the Gemini -> Antigravity migration research

Primary-source research note behind #4709: the sense taxonomy that separates a
runtime-axis `gemini` (stale) from Antigravity's on-disk contract, Google's
model IDs, and release history (all load-bearing); the PRESERVE table; the
guard-gap analysis; and the per-file inventory with file:line citations.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4709): keep the legacy patches probe, and stop tripping two lint gates

Three review findings, fixed inline.

1. Dropping the global ~/.gemini/gsd-local-patches probe was a regression: a
   pre-#1928 Gemini CLI install put patches there, and a stranded patches dir is
   still the user's work. Restored as an explicitly-labelled legacy arm probed
   AFTER Antigravity, so a live install always wins. This is a directory probe,
   not a runtime home -- it assigns no runtime id, so it does not reintroduce
   the defect this PR closes. The $GEMINI_CONFIG_DIR env probe is deliberately
   NOT restored: that names a runtime config home, which
   tests/declarative-reference-antigravity.test.cjs:307 pins as ignored.

2. The comment added in (1) originally contained the literal string that the new
   structural test matches, so the test flagged its own fix's comment as a mint.
   Reworded. The test was right; a comment in shipped workflow text is as
   readable to a matcher as code is.

3. docs/research/gemini-to-antigravity-migration.md used the colon slash-form
   inside a quoted manifest description. lint-docs-command-form rejects it:
   docs are never passed through the install-time converters, so the colon form
   names a command no runtime registers. Normalised to the hyphen form.

Also verified, rather than assumed: the local scan's .agents entry is
unambiguous. Antigravity is the ONLY runtime declaring localConfigDir '.agents'
across all 19 capability manifests; grok and codex use ~/.agents as a GLOBAL
home, and this scan is local (./$dir).

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4709): retire Gemini CLI from the PR templates, and close two review gaps

Adversarial review findings, all fixed inline.

1. All three .github/PULL_REQUEST_TEMPLATE/*.md still offered "Gemini CLI" under
   "Runtimes tested", and none offered Antigravity. #1928's follow-up dropped
   Gemini CLI from .github/ISSUE_TEMPLATE/*.yml but missed the PR templates, so
   every contributor opening a fix/feature/enhancement PR has been asked for two
   releases which runtime they tested and offered a retired one. Now Antigravity.
   Guarded by a new assertion: runtime checklist labels in the PR templates must
   appear in the runtime label table. Proven non-vacuous by reverting one
   template line and watching the probe report the offender.

2. The #4709 scanning corpus excluded agents/, which also ships runtime-loaded
   markdown including .compact.md variants. Widened: 318 -> 382 files (+64), zero
   new offenders, so the gap was coverage rather than a live defect.

3. gsd-core/workflows/sync-skills.md said "grok and gemini have no dedicated
   installer flag — they alias the codex and claude skills roots respectively."
   The gemini half is wrong twice over: the runtime is retired, and it never
   aliased claude -- canonicalizeRuntimeName returns null for it and the caller's
   fail-closed default merely happens to be claude. Describing that as designed
   aliasing is exactly the confusion this issue is about. Reduced to grok, which
   genuinely does alias the codex skills root.

4. The changeset said the runtime was removed in 1.11. It shipped in 1.8.0
   (CHANGELOG.md:1023 is the enclosing release heading for the #1928 entry at
   :1124). Corrected.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4709): follow the corrected sync-skills prose, and refresh the compact baseline

Three GREEN-run failures, all caused by this PR's own edits.

1. tests/sync-skills-cross-runtime-refuse.test.cjs pinned the literal phrase
   "grok and gemini have no dedicated installer flag" — a test REQUIRING shipped
   text to name a runtime retired in 1.8.0, which is the exact class #4709
   exists to remove. The assertion and its rationale comment now track the
   corrected prose ("grok has no dedicated installer flag"), and the docblock's
   runtime list drops gemini. The remaining assertions in that file — the
   guard's exit, the installer pointer, the $DEST reference, guard-before-copy
   ordering — are untouched, so #3025's contract is otherwise intact.

2. tests/fixtures/compact-content-benchmark-baseline.json drifted because the
   new-project.md edits changed its compacted size (split "new-project": off
   14279 -> 14308, on 12335 -> 12364; aggregate off 107411 -> 107440).
   Refreshed with `node scripts/benchmark-compact-content.cjs --write`, which
   is that script's own documented remedy.

3. emitted-attribution reported four grown workflow files with no
   acknowledgment. Acked below as commit trailers per ADR-3942, which moved the
   acknowledgment out of tests/emitted-drift-acks/*.json fragments and into the
   PR's own commit range (read with three-dot base...head). Exactly the four
   files the gate named are acked — settings-advanced.md and sync-skills.md
   shrank and are deliberately absent, since a trailer no delta consumed is a
   staleAcks error.

Refs #4709

Emitted-Drift-Ack-Growth: ingest-docs.md — the runtime-detection cascade now names Antigravity's three real directories (/.gemini/antigravity{,-ide,-cli}/) and $ANTIGRAVITY_CONFIG_DIR in place of the single retired /.gemini/ arm and $GEMINI_CONFIG_DIR; three correct paths cost more bytes than the one wrong path they replace.
Emitted-Drift-Ack-Growth: new-project.md — same runtime-detection correction as ingest-docs.md, plus dropping "gemini/" from the two GEMINI.md instruction-file sentences so the prose stops contradicting getProjectInstructionFile, which returns AGENTS.md for that retired id.
Emitted-Drift-Ack-Growth: reapply-patches.md — restores the legacy ~/.gemini/gsd-local-patches probe as an explicitly-labelled arm after an adversarial-review finding that dropping it stranded a pre-#1928 user's patches, and repoints the env/global probes at Antigravity; the four-line comment is load-bearing, since a bare retired-runtime path with no explanation is exactly what the next reader would delete.
Emitted-Drift-Ack-Growth: update.md — bare ~/.gemini and ./.gemini as GSD config dirs are replaced by the real ~/.gemini/antigravity and ./.agents, which are longer strings; no content was added beyond the corrected paths.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4709): backfill the changeset PR number

pr: 0 -> 4711, now that the PR exists. Never guessed ahead of the number.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:08:20 -04:00
Tom Boucher
f334f277dd fix(#4324): stop the retired /gsd: prefix reaching users (#4712)
* test(#4324): prove colon tokens the installer cannot convert leak

Failing-first regression coverage for #4324. The install rewrite
(transformContentToHyphen) is gated on an exact match against the
commands/gsd stem list, so any /gsd:<token> whose token is not a
registered stem survives the install and reaches the user as the
deprecated colon form.

The gate is load-bearing -- it is the only thing protecting the
workflow DSL marker family (gsd:section, gsd:protected, gsd:loop-host,
gsd:guard, gsd:dispatch, gsd:plan-revision-conflicts), which
workflow-fragments parses as a literal. So this suite asserts the
shipped text is convertible rather than asserting the transform is
broad, and pins the marker family as explicit negative space.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4324): stop unconvertible colon tokens reaching the user

The install rewrite is gated on an exact match against the commands/gsd
stem list, so a /gsd:<token> whose token is not a registered stem
survives the install and reaches the user as the deprecated colon form.
That gate is load-bearing -- it protects the gsd:section /
gsd:protected / gsd:loop-host marker family -- so the fix is in the
shipped text, and the source stays colon per CONTEXT.md's two-tier rule.

- quick-batch command + skill description: close the command token at a
  boundary so `/gsd:quick`-shaped converts instead of being skipped.
- gsd-code-fixer (both variants): execute-plan and diagnose-issues are
  workflows, not commands, so they never converted and rendered beside
  two hyphenated siblings on the same line. Name them as workflows.
- help topic-mode: the extraction rule hard-coded a colon prefix that
  the converted full.md never ships, so --brief could never match a
  signature line and silently fell back on every topic. Describe the
  signature line without a literal prefix.
- update.md: drop the prefix from prose describing a stale command.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4324): add changeset fragment

pr:0 placeholder is backfilled with the real number once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4324): locate the help summary per reference variant

Adversarial review finding. Restoring the signature-line match (the
#4324 fix) activated a latent defect in the clause next to it: compact
scope emitted "the single non-blank line immediately after" the
signature, and that clause is only correct for full.md.

full.compact.md puts the summary on the signature line itself, after an
em-dash, and its next non-blank line is an unrelated "Usage:" line. Both
variants ship and both are served, so before this commit the compact
variant would have emitted the wrong line as the summary. It was masked
until now only because the stale colon prefix meant no signature line
ever matched at all.

Name the two placements and pick per line, and say explicitly that a
Usage: line is never a summary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4324): de-vacuum the help parity check, narrow the marker waiver

Two adversarial review findings against the #4324 coverage.

The help-parity assertion went vacuous the moment the fix landed: once
topic.md stops spelling a literal prefix, the matched set is empty and
the assertion holds for any rewording, correct or not. It now also
asserts across BOTH served reference variants that each ships signature
lines under the hyphen prefix, that the two genuinely disagree about
where the summary sits, and that topic.md still names both placements
and the Usage: guard.

The marker waiver keyed on "sits inside an HTML comment", which waves
through a real broken reference that happens to be commented out --
`<!-- see /gsd:typo-cmd -->` scored clean. Enumerate the six marker
families instead. Verified the narrowed rule catches that probe and
still passes over the tree; it also surfaced a seventh family,
write-continue, that the broad rule was hiding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4324): normalize the namespace in skill descriptions

Both hyphen-namespace skill converters ran the hyphen transform over the
body but rebuilt the frontmatter description from the raw field, so a
/gsd:<cmd> mention in a command description survived into the installed
SKILL.md -- the exact field the host's skill picker renders, which is
the surface this issue was filed about.

The local flat-command path was already correct because it rewrites the
whole file; only the skills path, used by a global install, was
affected. Confirmed by installing into a fake HOME before and after.

Fixed in both copies: bin/install.js and the src/ source of truth that
compiles into gsd-core/bin/lib.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4324): assert descriptions through the real converters

The previous version of this check called transformContentToHyphen on
the description line itself and passed, while a real install still
shipped the colon form -- the converter never calls that transform on
the description. It asserted a proxy for the behaviour instead of the
behaviour.

Drive convertClaudeCommandToClaudeSkill and
convertClaudeCommandToClineSkill over every registered command and
assert on the emitted description. Verified it fails against the
pre-fix converters and passes against the fixed ones.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4324): regenerate skills after the description change

skills/<name>/SKILL.md is generated by gen-plugin-skills, not
hand-maintained, and lint:generated-sync caught the hand edit. The
regenerated file emits the hyphen form, which also corrects the
assumption behind the scan comment in the namespace test: skills/ is
runtime-emitter output, not colon source.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4324): re-sanction normalizeKimiSkillName's real end line

The description-normalisation fix inserted five lines above
normalizeKimiSkillName in src/runtime-artifact-conversion.cts, moving its
closing brace from 635 to 640. MAJOR-1 pins that line deliberately, so the
planted violation landed INSIDE the exempted body and went unflagged --
0 !== 1.

Re-sanction the value rather than derive it: the array is named
sanctionedRealEndLines, and a pinned line that fails loudly on drift is
the design. Deriving it would remove the human check the name asks for.

Verified by executing all four MAJOR-1 rows against the real tree: each
planted violation is flagged at realEndLine+1 and each unmodified file
stays exempt.

Emitted-Drift-Ack-Growth: gsd-code-fixer.md — names execute-plan and diagnose-issues as workflows rather than as slash commands that do not exist
Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — same rewording as its full sibling, kept byte-consistent with it
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4324): backfill the changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 22:17:48 -04:00
github-actions[bot]
3697b94655 chore: sync next package version to 1.14.0 2026-09-14 00:01:49 +00:00
Tom Boucher
c0b2a05d2f fix(#4594): one canonical dispatch-identity owner — the emitted format and the parser that reads it back (#4693)
* fix(#4594): give dispatch identity one owner for the emitted format and its parser

The isolation guards decided whether a run-scoped sentinel applied to a
dispatch by regex-scraping model-authored prose. The scrape returned values in
a different namespace from the ones the sentinel records, so the comparison
could never succeed:

  sentinel  { phase: "03", plan: "03-02-hardening" }   <- $PHASE_NUMBER / $plan_id
  prose     "Execute plan 02 of phase 03-auth."
  scraped   { phase: "03-auth.", plan: "02" }          <- greedy (\S+), both wrong

#4594 reports only the phase half. Measured against a real phase-plan-index
run, plans[].id is phase-prefixed, plan-numbered AND slugged, while the prose
carries a bare in-phase plan number — so the plan field mismatches too, and the
Claude path is dead rather than latent. A fresh sentinel was therefore
discarded on every executor dispatch and every legitimate ISOLATION=none
degrade was denied, leaving the work unrun.

hooks/lib/dispatch-identity.js is now the single owner of both halves. The two
prompt-body producers emit a canonical marker carrying the same shell values
the sentinel records, so producer and consumer agree by construction. The prose
frame stays as a fallback, bounded by the phase-token grammar ADR-2121 owns and
deliberately reporting no plan — an absent identifier means "cannot compare"
and is safe; a wrong one is a false mismatch and is not.

The prose sentence itself is byte-identical: the executor agent reads it too,
so the marker is purely additive (Hyrum's Law).

An inapplicable sentinel is now named in the guards' deny reason instead of
being dropped silently — the silence is why this survived three producers and
two consumers unnoticed. Interpolated values come from a sentinel file and from
prompt text, so both are length-bounded and stripped of control characters.

ADR-4630 locks the seam and maps the epic's three phases.

Refs #4630
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4594): resolve eight review findings across the dispatch-identity seam

Three orthogonal review engines ran on 43418af144 — the code-review skill's
Standards and Spec axes, and an isolated adversarial security pass — plus a
self-review of the committed diff. Every finding is fixed here; none deferred.

F1 (major, reproduced). A keyless or unknown-key-only marker — the literal
"[gsd:dispatch]" or "[gsd:dispatch run=..]" — matched the marker grammar and
returned source:'marker' with both fields null, suppressing the prose fallback
entirely. Any prompt text containing that literal silently disabled identity
narrowing, so a fresh sentinel applied to a dispatch it was never scoped to,
defeating #3045 SECURITY F2. Prompt text is attacker-influenceable. A marker
that yields neither recognized key is no longer a marker: the scan continues to
later markers, then later texts, then prose. Forward-compatible tolerance of
unknown keys is unchanged.

F2/F3 (major). The first cut duplicated sanitizeForReason,
describeSentinelDiscard and REASON_INTERPOLATION_MAX_LEN byte-for-byte across
both guards — the exact defect class this epic exists to delete, and with no
cold-load justification, since both hooks already require hooks/lib/. They now
live in hooks/lib/isolation-deny-reason.js, and buildSentinelDiscard lives in
isolation-sentinel.js beside the comparison it mirrors, returning the nested
{sentinel:{phase,plan}, dispatch:{phase,plan}} shape instead of a bespoke
four-field bag that renamed the pairs already flowing through the seam.

F4 (hard violation). The visibility test asserted on the deny reason's prose.
CONTRIBUTING.md prohibits raw text matching on hook output, which is why every
deny carries a stable reason_code. The discard is now a structured
sentinel_discarded field on each hook's stdout JSON, and the test asserts that;
the sentence stays for the operator but is no longer the contract.

F5 (hard violation). The 64-character truncation limit had no boundary
coverage. 63/64/65 are now exercised against the single consolidated helper.

F6 (minor). sanitizeForReason stripped C0/C1 controls but not U+2028/U+2029 or
the bidi overrides, so a crafted value could still reflow or reverse the
message. Both classes are stripped, with a test each.

F7 (major). The producer/template parity test was vacuous — it rendered a
marker and re-parsed its own output, and would have passed with both templates
deleted. It now reads the two workflow templates, extracts each marker line,
substitutes the measured values and asserts the owner's parser returns them.
Proven red by deleting one template's marker line before being proven green.

F8 (doc). ADR-4630 and the design notes claimed the marker is guaranteed on the
orchestrator-worktree path because that prompt is built in shell. It is not:
executor-isolation-dispatch.md:131 says plainly that those are template
placeholders, not shell variables, so {plan_id} is model-substituted there too.
A false guarantee in a design lock is worse than a stated limit. Both documents
now say the marker is model-substituted on both paths and that the prose
fallback is the real floor everywhere. The "3 workflow templates" count was
also wrong — 3 prose sites across 2 files, 2 of which carry the marker.

Refs #4630
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4594): refresh the compact-content baseline and acknowledge execute-phase.md growth

Refs #4630.

The dispatch-identity marker and its substitution note grew
gsd-core/workflows/execute-phase.md by 525 bytes (91846 -> 92371), which drifts
two real-tree guards that lint:ci does not run:

- tests/benchmark-compact-content.test.cjs asserts the committed baseline is
  "up to date"; the split for execute-phase.md moved off 25827 -> 25952 and on
  23576 -> 23701, taking its compaction reduction 8.72% -> 8.67%. Baseline
  regenerated with scripts/benchmark-compact-content.cjs --write.
- tests/emitted-attribution.test.cjs requires a growth acknowledgment trailer
  for any emitted file that grows, keyed on the bare filename. Added below.

The growth is two additions and no rewrites: the [gsd:dispatch ...] marker line
inside the Agent() prompt's <objective>, and the note telling the orchestrator
to substitute {plan_id} with the plan's id verbatim. Both are load-bearing --
the marker is what lets a guard hook match a dispatch to the sentinel the
per-plan gate wrote, and without the note the orchestrator has no instruction
telling it the value must not be paraphrased.

Emitted-Drift-Ack-Growth: execute-phase.md — adds the canonical [gsd:dispatch] identity marker and its {plan_id} substitution note, which the isolation guards compare verbatim against the run-scoped sentinel (#4594)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4594): set changeset fragment pr to 4693

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 15:36:08 -04:00
Tom Boucher
bbdf7e8e84 chore(#4654): add local/no-unconfined-path-join and drain it to zero — Phase 4 of #4636 (#4674)
* chore(#4654): add local/no-unconfined-path-join and drain it to zero

Phase 4 of epic #4636 — the ratchet, and the phase that makes the epic hold.

THE MEASUREMENT THAT RESHAPED THE PHASE. An AST census (the repo's own parser,
not grep) found what the epic never enumerated: ADR-4650 named seven containment
implementations; `src/` alone held roughly 24 more hand-rolled gates across ~13
files, several guarding a write or an `fs.rmSync`. Two verified by reading rather
than pattern-matching — `research-store.cts` comments its own as "ensure the
resolved file path stays inside the store dir" immediately before a write, and
`capability-lifecycle.cts` gates `fs.rmSync` with one.

So the epic's Done-when "one containment predicate, used at every site" was FALSE
when Phase 3 reported it satisfied. It is true now: the rule is clean across
src/, scripts/, gsd-core/bin/ and hooks/ with an EMPTY allowlist.

WHY NOT THE RULE THE ISSUE PROPOSED. #4654 proposed flagging `path.join` whose
first argument is a managed root and whose later arguments derive from argv. That
is a taint analysis over 2046 call sites, in ESLint, without type information;
"derives from argv" is not locally decidable. Any approximation either floods or
is trivially evaded, and a rule that fires on hundreds of correct sites earns an
allowlist of hundreds — the opposite of a ratchet. What is actually duplicated is
the COMPARISON, not the join, and that has one recognizable shape.

  Arm 1  X.startsWith(Y + sep)            the hand-rolled containment idiom
  Arm 2  a containment predicate called as a bare statement, answer discarded

Arm 2 is the issue's "asserts the result was narrowed, not merely that a helper
was called". Its example `validatePath(x, root).resolved` is already
structurally impossible — Phase 3 un-exported `validatePath` — so the remaining
expressible failure is ignoring the answer, which is the defect that recurred
five times in this epic. The census found exactly one live instance
(`milestone.cts:1643`); it now returns the proven `ContainedPath` so consumers
stop re-deriving the path the comment above it was extracted to stop them
re-deriving.

The rule deliberately does NOT try to catch validate-one-path-use-another where
the answer is used but a different variable flows onward. That needs flow
analysis; the branded `ContainedPath` from Phase 3 is the defense there, and the
two are complementary.

PER-SITE FAMILY CHOICE, NOT A DEFAULT. Phase 3's lesson binds: collapsing a
lexical site onto the realpath family broke four tests and was caught only by the
matrix. Every migrated site was triaged individually. The six
installer-migrations tree-walks and the six capability-lifecycle gates take the
LEXICAL family because their operands are already realpath-resolved and they
deliberately treat the final component as a link; boundary sites take realpath.

TWO SITES WITH AN INVERTED CONTRACT, which a mechanical swap would have broken.
`installer-migrations.cts:127` and `runtime-artifact-install-plan.cts:144` REJECT
`target === root` by contract, while the canonical comparison ACCEPTS it. Swapped
naively, a migration could `rmdir` the user's config root and a third-party
descriptor could write at configHome itself. Both keep `=== root` as an explicit
additional arm alongside the predicate call — the predicate decides containment,
the call site keeps its own extra condition (ADR-4650 decision 6).

ONE DUPLICATE DELETED OUTRIGHT: `planning-inspect.cts`'s `isWithinRoot` was
byte-identical to `isContainedIn` and said so in its own docstring.
`isContainedIn` is now exported for callers that have already resolved both
operands and need only the comparison, with a doc note that a caller which has
NOT resolved them must use a full predicate instead.

THE MARKER, AND WHY IT IS NOT THE ALLOWLIST. Nine sites are justified holdouts and
carry `// allow-handrolled-containment: <reason>` with a mandatory, reviewable
reason. Two justifications: (a) not a containment decision — an ancestor-walk loop
condition, sub-repo grouping, worktree identity matching, declared-path coverage;
(b) it IS containment but the canonical predicate is unreachable —
`capability-validator.cjs` is a committed pre-build `.cjs` and the compiled
`security.cjs` is untracked build output, so requiring it would break a fresh
clone. `scripts/lib/drift-scan.cjs` runs under `lint:ci` with the same exposure.
The marker was renamed from `allow-lexical-prefix-match` mid-phase because that
name asserted only (a) and would have stated something false at the (b) sites.

A marker suppresses BEFORE the violation counter increments, so a file whose
every occurrence is marked still reports `staleAllowlistEntry` — otherwise a
drained entry lingers and silently re-permits the site later.

DEMONSTRATED RED, per #4654: a hand-rolled copy reintroduced into a real `src/`
file made `npm run lint` fail with the rule's full guidance message; removing it
returned the tree to clean. Both halves recorded — red alone proves nothing,
since a rule red for an unrelated reason looks identical.

DISCLOSED: `defaultRequireFromInstallRoot` (gsd-tools.cjs) previously carried two
distinct rejection messages and two manual realpath calls; routing it through
`tryWithinRoot` collapses them to one message, and a missing module now surfaces
as MODULE_NOT_FOUND rather than ENOENT. No test asserts either message. The
security property is preserved and slightly strengthened — the candidate is
realpathed and containment re-checked, and the dangling-symlink oracle closure
comes along with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4654): record the containment ratchet in CONTEXT.md and the security model

Both entries previously described the seam without the thing that keeps it a
seam. They now state what the rule bans, and — more usefully for whoever reads
this next — what it deliberately does NOT attempt: deciding per path.join call
whether an argument came from user input. That question is not locally
decidable, and an approximation across ~2000 join sites would earn an exemption
list of hundreds, which is the opposite of a ratchet.

Also records the marker's two legitimate justifications and that its reason is
mandatory, so the escape stays reviewable rather than becoming a mute button.

Glossary gate 270 refs exit 0; install-tree goldens and CONTEXT-INDEX.json
regenerated and confirmed byte-identical rather than assumed — which also
confirms eslint-rules/ is not a shipped path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4654): close review findings and the two matrix failures

MATRIX FAILURE 1 — a collapsed message broke a negative-proof test, and my
evidence for collapsing it was wrong. I searched tests/ for the literal string
"resolves outside its install root", found nothing, and reported that no test
asserted it. The test matches a REGEX SUBSTRING, /outside its install root/, so
the literal search missed it. What broke was "NEGATIVE PROOF: a symlinked module
pointing OUTSIDE the install root is not loaded" — the test guarding the exact
property I claimed was preserved. defaultRequireFromInstallRoot now does both
checks again with both messages byte-identical, each routed through the
canonical predicate, which is better than the original since that hand-rolled
both comparisons.

MATRIX FAILURE 2 — shipped migrations are checksum-locked, and a marker cannot
serve there. migrationChecksum hashes plan.toString(), which INCLUDES comments,
so a suppression marker inside a plan body drifts the baseline exactly as an
edit does. Measured: with markers in place, two of the four still differed from
their committed checksums. The four shipped bodies are now byte-identical to
next, and the rule's config excludes those four paths BY NAME rather than by a
directory wildcard, so a NEW migration is still covered. Six containment
comparisons stay un-ratcheted there; that gap is recorded in the rule's Known
gaps, in CONTEXT.md and in the security model rather than left implicit.
Justification (c) is removed from the marker's documented reasons, because a
marker was proven unable to express it.

ADVERSARIAL REVIEW — the sharpest finding was that the rule banned the CORRECT
shape while permitting the incorrect one: startsWith(root) with no separator is
the genuinely unsafe form, since it accepts a sibling such as root-evil, and my
own test blessed it as valid. Flagging every bare startsWith would swamp the
rule, so that stays a STATED gap rather than a silent one. Closed for real: the
template-literal spelling, which the census never saw because it only inspected
plus-concatenation — that surfaced TWELVE more sites, now triaged and migrated.
A separator reached through a const alias is now resolved via scope analysis.
And isContainedIn, exported in Phase 3, was missing from the discarded-result
set, so a bare no-op call went unflagged on the one function the epic funnels
through.

SECURITY REVIEW — the marker could over-suppress two ways: a block comment
worked identically to a line comment, and one marker silently covered every
violation sharing its line. It now requires a Line comment positioned after the
flagged node ends, so it anchors to the node it trails. Four sites had dropped
an unreachable-but-deliberate equality rejection against the root; each is
restored as the call site's own arm. eslint.config.mjs still documented the OLD
marker token, which my rename missed — it would have sent the next author in
circles.

A FALSE GREEN, recorded because it nearly stuck: lint:ci reported exit 0 from a
stale eslint cache while twelve real violations existed. Every lint check here
now clears the cache first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4654): anchor a suppression marker to the violation it actually trails

The matrix caught this; my own test caught it, on its first execution. The case
"two violations on one line: trailing marker suppresses only the one it trails"
expected 1 error and got 0 — both were suppressed.

ROOT CAUSE: the anchoring accepted any Line comment on the node's line whose
range started at or after the node's end. A trailing marker at the END of a line
sits after EVERY node on that line, so that condition held for all of them.
"After the node" does not identify WHICH node the marker trails. The fix reads
as correct and is not.

FIX: deferred reporting. Violations accumulate during traversal instead of being
reported immediately; at Program:exit each marker claims exactly ONE pending
violation — the one on its line whose end is nearest before the marker begins —
and every unclaimed violation is then counted and reported. One marker, one
suppression. An earlier violation sharing the line is still reported, which is
the property the security review asked for and the previous attempt only
appeared to deliver.

The counter now increments at flush time rather than during traversal, so a
suppressed occurrence still does not keep an allowlist entry alive.

AND A TOOL THAT SHOULD HAVE EXISTED BEFORE THE FIRST MATRIX RUN. `node --test`
is hard-blocked here, so this rule's test file could only ever be executed on
the remote matrix — which is why a broken anchoring shipped into a run. ESLint's
programmatic Linter API is not a test runner, and exercising the rule through it
verifies every case locally in seconds. All 24 now pass locally, including the
two-on-one-line case that failed remotely. That loop should have been built
before the rule was first sent to the matrix rather than after it failed twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4654): backfill PR 4674 into the changeset and complete 70-docs.json

The phase gate requires enablementSequence and the Diataxis quadrants; 70-docs
now carries both, with the how-to quadrant skipped for a stated reason rather
than an empty field. The audience for this deliverable is a contributor who
trips the rule, and the task-oriented guidance reaches them in the ESLint
message itself — which names the correct predicate, says how to choose between
the realpath and lexical families, cites the Phase 3 regression caused by
choosing wrong, and gives the marker syntax. A docs/how-to page would be a
second, driftable copy read by nobody at the moment of failure.

enablementSequence is recorded as what it actually is: a VERIFICATION sequence,
not an enablement one. The rule is never off, so there is no off-to-on
transition to describe.

scripts/lint-docs-required.cjs now passes (ok_docs_updated) — it could not
evaluate against the mandated pr:0 placeholder.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 22:17:46 -04:00
Twisted Fate
6edd506cc7 fix(#4558): report a byte-identical restore destination as already_present (#4599)
* fix(#4558): report a byte-identical restore destination as already_present

restore-custom-files treated a destination that is byte-identical to its
backup exactly like a missing one: plan reported it as `eligible`, --apply
re-copied the same bytes and reported `restored`, and both counters
included it. Because update.md drives its restore question off
eligible_count, the workflow re-offered the same no-op restore on every
update and accepting it never settled anything.

Emit a distinct `already_present` outcome for that case. It is excluded
from eligible_count and restored_count, --apply writes nothing for it,
and the backup is left intact. A differing destination is still
skipped_destination_exists and a missing one still restores normally.

Regression tests cover the identical-destination plan/apply paths, the
idempotence-after-success cycle (missing -> restored -> silent plan), and
a mixed backup. Docs for the outcome enum are updated to match.

* fix(#4558): tighten already_present wording after review

update.md's RESTORE_ELIGIBLE == 0 branch now also names the already-present
case, CLI-TOOLS.md no longer calls the follow-up plan run "silent" (the
entry is still reported, just never offered), and a test message reads
correctly.

* chore(#4558): add changeset fragment for #4599

* chore(#4558): acknowledge update.md growth from the restore-outcome guidance

Emitted-Drift-Ack-Growth: update.md — added already_present restore-outcome guidance for #4558

---------

Co-authored-by: TwistedRiCen <16397953+TwistedRiCen@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-13 02:16:40 +00:00
sim
bbc3f131be refactor(#4653): make containment ONE decision, resolved two ways
Satisfies #4653 DW1 and DW9, which were the phase's outstanding acceptance
criteria: every other implementation must be deleted or route its containment
DECISION through the canonical predicate, and no surviving wrapper may decide
WHETHER a path is contained.

Three implementations were being retained with their own comparisons, on the
argument that each needs LEXICAL resolution — a realpath-based predicate is the
wrong tool wherever a symlink must be preserved rather than resolved. That
argument is correct about RESOLUTION and was being used to justify owning the
DECISION too. Those are separable, and separating them is what closes the
criteria honestly rather than by reinterpretation.

  isContainedIn(resolvedTarget, resolvedRoot, pathImpl?)   module-internal

is now the single place this repo decides containment. It is separator-aware, so
a sibling merely sharing a prefix (`<root>-evil` against `<root>`) is still
rejected. Two exported families sit on it and differ ONLY in how a candidate is
resolved before the decision:

  assertWithinRoot / tryWithinRoot                realpath-resolving
  assertWithinRootLexical / tryWithinRootLexical  path.resolve only, no I/O

The lexical pair carries `opts.pathImpl`, so win32 separator semantics stay
testable off Windows — that seam already existed in isPathConfined and would
have been lost by a naive collapse.

The three call sites now take their decision from the predicate and keep only
what is genuinely theirs:

  external-descriptor-trust isPathConfined   delegates outright; pathImpl forwarded
  installer-migrations ensureInsideConfig    delegates; keeps its own message and
                                             its LEXICAL fullPath, which callers
                                             consume for existsSync and journal rows
  gsd-tools.cjs isInsideDir                  delegates; keeps its own `target !==
                                             root` condition, and the separate
                                             symlink refusal above it stands

DW5 is not weakened by this. That criterion binds the symlink oracle and the
ancestor canonicalization; both are untouched. The only change inside
validatePath is three comparison lines becoming one call, and the rejection
string `Path escapes allowed directory: <resolved> is outside <base>` stays
byte-identical because it is an observable CLI contract.

What this does NOT do, stated plainly: the lexical family still cannot see a
symlink. That is a property of lexical resolution, not a gap in the seam, and
the three callers that need it are the three that must pair it with their own
symlink refusal — which is exactly what the fix earlier in this phase added at
the install sites. The doc comment says so at the definition, and CONTEXT.md and
docs/explanation/security-model.md are corrected: they previously described
these three as deliberately NOT routed through the predicate, which is no longer
true.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 18:33:01 -04:00
sim
cd58aaabf4 refactor(#4653): drain the containment duplicates and record the two rulings
Phase 3 of epic #4636, stage 3c. ADR-4650 decision 6: a wrapper may decide HOW
to degrade, never WHETHER a path is contained. Four implementations are drained
on that rule; two are retained, with the reasons recorded rather than assumed.

DRAINED — the containment decision now comes from the canonical predicate:

  scripts/check-glossary-refs.cjs   local isWithinRoot deleted outright.
  src/installer-migrations.cts      ensureInsideConfig keeps its throw and its
                                    lexical fullPath; only the decision moves.
  src/planning-inspect.cts          isPathContained keeps must-exist as its own
                                    condition; only the decision moves.

Two of those are wrappers rather than deletions, and each is a wrapper for a
reason that would have been a silent behavior change if collapsed naively:

- `isPathContained` returns FALSE for a path that does not exist, because
  fs.realpathSync throws ENOENT and its catch swallows it. The canonical
  predicate does the opposite: for a missing target it walks up to the nearest
  existing ancestor and ACCEPTS a not-yet-created path under the root. Its
  callers at planning-inspect.cts:747 and :839 guard a phaseDir immediately
  before readdirSync, so under a naive swap a missing phaseDir would stop
  reporting scope UNREADABLE and start throwing ENOENT out of readdirSync.
  Existence is therefore kept as an explicit local requirement.

- `ensureInsideConfig` returns a LEXICAL fullPath that both callers consume for
  existsSync and for journal entries. The canonical predicate realpath-resolves,
  so if configDir is itself a symlink the two differ. The decision is canonical;
  the returned value stays lexical. Its message is likewise preserved verbatim,
  which is why this uses tryWithinRoot plus an explicit throw rather than
  assertWithinRoot.

`isWithinRoot` in planning-inspect is left in place and documented: it is a pure
comparison over paths the CALLER has already resolved, which readDocument does
inline specifically to keep a third degradation shape (exists-but-unreadable vs
absent) that neither isPathContained nor the canonical predicate expresses. It
is the comparison step of one implementation, not a second implementation.

RETAINED, DELIBERATELY — gsd-core/bin/gsd-tools.cjs. My own design document said
"collapse" and that was wrong. The file carries an explicit comment forbidding
it, and the comment is correct: its three checks reject symlinks OUTRIGHT, which
is strictly stricter than the canonical predicate, not a reimplementation of it.
The canonical predicate accepts a link whose target lands inside the root — for
a restore that is still wrong, because writing through the link overwrites
whatever it points at instead of materializing a regular file. Collapsing would
have reintroduced that hole. The comment is updated to name the current exported
predicate, to record that this was reviewed under this phase and deliberately
not collapsed, and to note that isInsideDir treats target === root as NOT
contained — the one implementation in the repo that does.

THE configHome RULING — retained lexical, and a false safety claim corrected.
isPathConfined stays lexical because two of its callers must validate a
destSubpath BEFORE the mkdirSync that creates it (install-engine.cts:1608,
install-profiles.cts:880), where realpath cannot resolve and a realpath-based
predicate would reject every legitimate install.

Its docstring's justification, however, did not survive being checked. It cited
capability-source.cts:491,577,675 as the upstream symlink rejection that made
the lexical form safe. Read directly: :491 is a blank line before assertSafeId's
JSDoc and :577 is an entry-count budget check. Neither is a symlink check. The
real guards are :585-586 and :671-674. Worse than stale line numbers, the claim
that this "keeps every caller of this function's callers symlink-safe" is false:
that rejection lives in capability-source's staging path and covers only the
capability-loader route to assertDescriptorConfined. Three other callers do not
reach it, and only retired-artifact-cleanup.cts:69 carries its own defense
(its lstatSync check at :77). The docstring now states what is actually true and
cites the lines that actually exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 13:29:54 -04:00
Tom Boucher
a2331c01f1 fix(#4568): widen the phase-number regex to accept N-segment ids at 6 shell/markdown sites (#4646)
* test(#4568): pin the N-segment phase-grammar defect across all 6 shell/markdown sites

Manually traced against the current tree: the validating regex at
code-review.md rejects a 3-segment id (23.1.2), and execute-plan.md's
extraction truncates a 23.1.2-01-PLAN.md filename down to 1.2-01.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4568): widen the phase-number regex to accept N-segment ids at all 6 shell/markdown sites

Widens `?` to `*` on the dotted-segment group at all 6 sites (byte-identical
behavior for 1- and 2-segment ids, character class unchanged): code-review.md,
code-review-fix.md, gsd-code-fixer.md, gsd-code-fixer.compact.md (validating
sites, plus their comment/error-message text), execute-plan.md's plan-filename
extraction, and plan-phase.md's --research-phase flag capture.

Also disambiguates the nsegment-phase-grammar test's plan-phase.md anchor,
which was matching an unrelated earlier `--research-phase` occurrence (line
77's generic-value capture) instead of the targeted site (line 131).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): extend lint-phase-id-drift to ban the single-segment phase regex in workflows/ and agents/

Adds findSingleSegmentPhaseRegexDrift, banning the bounded
`[0-9]+(\.[0-9]+)?` shape (and its \d/doubled-backslash near-variants) on any
phase-carrying line across gsd-core/workflows/**/*.md,
gsd-core/references/**/*.md, and the newly-scanned agents/**/*.md, sanctioned
the same way as the existing shell-arith rule. Wired into scanAll; confirmed
zero violations against the real tree post-#4568 fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4568): add Fixed changeset

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: regenerate conformance-tier manifests for the new test file

The emitted-attribution gate also flags 4 files growing: code-review-fix.md
(+21 bytes), code-review.md (+21 bytes), gsd-code-fixer.compact.md (+9
bytes), gsd-code-fixer.md (+6 bytes). The growth is the fix itself: each
site's validation regex widened from a bounded single-optional-dotted-segment
shape to the unbounded form, and the accompanying comment/error-message text
grew by a few characters to mention the new 3-segment example.

Emitted-Drift-Ack-Growth: code-review-fix.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Emitted-Drift-Ack-Growth: code-review.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the error text (#4568)
Emitted-Drift-Ack-Growth: gsd-code-fixer.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4568): backfill changeset pr number to 4646

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 17:04:28 -04:00
Tom Boucher
db4d8a9bae fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic (#4644)
* fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic

$((10#${PHASE_NUMBER})) is a hard bash/zsh syntax error when PHASE_NUMBER is
decimal (01.1, from an inserted phase) or N-segment (23.1.2) — neither is
valid shell-arithmetic syntax at all, and the failed expansion aborts the
rest of the snippet in a non-interactive shell. safe_resume_gate runs
unconditionally before trusting STATE.md or dispatching any executor, so
execute-phase failed at its own gate before the first executor on any
decimal phase, regardless of workflow.tdd_mode. Regression from #4194.

Fixes all 4 sites: safe_resume_gate and the TDD gate in
workflows/execute-phase.md, the completion-signal spot-check fallback in
workflows/execute-phase/steps/completion-reconciliation.md, and the
executor gate validation example in references/tdd.md. Each now zero-strips
only the leading integer segment into a *_INT variable (via %%.* / #
parameter expansion — always valid shell syntax regardless of what follows)
and keeps the remainder as an escaped-dot string for the anchored commit-
scope regex, exactly as issue #4619 verified in both bash and zsh. A plain
integer phase (12, 01) computes byte-identically to before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4619): pin the decimal/N-segment fix and characterize the pre-fix bug

Behavioral coverage via real bash execution: the old $((10#01.1)) form
throws (characterizes the bug, matching the issue's own reproduction); the
new form resolves 01.1 -> 1\.1 and 23.1.2 -> 23\.1\.2, unchanged for plain
integers (12 -> 12, 01 -> 1); the resulting anchored ERE matches
feat(01.1-03):/test(1.1-3): and correctly rejects feat(01-03):,
feat(01.2-03):, feat(011-03):, feat(12-03): for a decimal phase — mirroring
issue #4619's own verified table exactly. Updates
safe-resume-gate-anchoring.test.cjs's 4 existing source-text assertions
(one per site) to the new fixed text.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): refine the shell-arith drift detector to distinguish safe from unsafe arithmetic

With #4619's fix in place, the guard's original "ban $((10#... outright,
match any occurrence" was too blunt: it flagged a comment merely mentioning
the pattern in prose, the now-safe $((10#$PHASE_INT)) arithmetic on an
already-%%.*-stripped integer, and the always-safe plan-id arithmetic
(plan ids are plain integers, never decimal). Refines the detector to skip
full-line comments and to only flag a captured variable/placeholder name
that contains "phase" and does NOT end in _INT/_int — the naming convention
the #4619 fix establishes at all four sites for "already reduced to a safe
integer." A plan-id variable was never phase-number arithmetic in the first
place and is excluded on the same basis.

This closes epic #4634's D6 ("lint-phase-id-drift... passes with no new
exemptions") and D7 ("a decimal and N-segment phase id survive an
end-to-end execute-phase selection without error") for real — the guard now
reports zero violations across all five .cts/.md rules.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: regenerate conformance-tier manifests for the new test file

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4619): cover the plain-padded-integer near-miss matrix too

Review found the anchored-ERE near-miss coverage only exercised the
decimal case (PHASE_NUMBER=01.1); issue #4619's own worked table also
verifies the plain padded-integer case (01 -> PHASE_N=1) against its own
near-miss set (matches 01-03, rejects 01.1-03/011-03/12-03). Adds the
missing assertion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4619): add Fixed changeset

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4619): correct JS backslash-escaping in safe-resume-gate anchoring test

The test's string-literal assertions for the PHASE_FRAC//./\\.} pattern wrote
only 2 backslash characters in JS source, which single-quoted-string parsing
collapses to 1 real backslash at runtime -- but the workflow/reference files
actually contain 2 raw backslash bytes at that position (needed so bash's
${var//pattern/replacement} produces the correct single-backslash output).
Write 4 backslash characters in the JS source at all 4 occurrences so the
runtime string matches the files' real bytes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4619): refresh the committed compact-content benchmark baseline

The new PHASE_INT/PHASE_FRAC arithmetic lines added to
gsd-core/workflows/execute-phase.md shifted its committed compaction-ratio
baseline. Regenerate via `node scripts/benchmark-compact-content.cjs --write`.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4619): note the safe_resume_gate arithmetic growth in the test header

The emitted-attribution gate flags execute-phase.md growing 91253 -> 91846
bytes (593 bytes). The growth is the fix: the safe_resume_gate and TDD RED
block now derive PHASE_INT/PHASE_FRAC before computing PHASE_N, so a
decimal/N-segment phase number (e.g. 01.1, 2.3.1) zero-strips its leading
integer segment via base-10 arithmetic instead of forcing the whole value
through $((10#...)) and hitting a hard shell syntax error on the first dot.

A blank line previously separated the Emitted-Drift-Ack-Growth trailer from
the Co-Authored-By trailer below it, which splits git's trailer-block
detection: only the last contiguous non-blank run of Key: Value lines at the
end of a commit message is recognized as trailers, so the growth ack was
silently read as ordinary body text and the differential-attribution gate
failed with the growth unacknowledged. Joining the two trailers into one
contiguous block fixes it.

Emitted-Drift-Ack-Growth: execute-phase.md — adds PHASE_INT/PHASE_FRAC derivation to the safe_resume_gate and TDD RED commit-scope grep so a decimal/N-segment phase number zero-strips its leading integer segment via base-10 arithmetic instead of failing on a non-numeric value (#4619)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4208): replace chmod-based restore-failure injection with a root-proof git shim

`tests/commit-files-deletion.test.cjs`'s two restore-failure tests simulated
an unwritable index via a `post-index-change` hook running `chmod a-w` on
the git dir. That relies on the OS enforcing the *owner's own* permission
bits against itself, which uid 0 (a routine identity inside this repo's
Docker-based gsd-test benches) does not: every DAC check short-circuits true
for root, so the write the chmod meant to block silently succeeds, the
restore comes back clean, and the disclosure/rollback behavior under test
never actually gets exercised.

This is CLAUDE.md's own named anti-pattern for I/O-failure injection
("Cross-platform test IO-failure injection" — chmod tricks fail under root
Docker/CI). It is confirmed as the actual root cause here, not a production
defect: `src/commands.cts`'s `restoreRemovedEntries`/rollback-disclosure
logic (added by #4253, merged just before this run) was hand-traced and
manually reproduced end to end on an unprivileged workstation against a
freshly built `gsd-core/bin/lib/commands.cjs`, and it already produces
exactly the `staging_failed` + "could not be restored" / "could NOT be
restored during rollback" results both tests assert. The other
`post-index-change`-based tests in this file (a `sleep` to force a timeout;
a real `update-index` to flip a restored entry's mode) are unaffected
because neither depends on a permission check — consistent with only the
two chmod-based tests failing on the real remote run.

Replaces the chmod fixture with a fake `git` placed ahead of the real one on
PATH that fails only `update-index --add --cacheinfo` — the one call the
restore makes — unconditionally, regardless of privilege level. Every other
git invocation execs straight through to the real binary, so the rest of
each scenario (`rm --cached`, the restore's own `ls-files` verification,
etc.) is exercised exactly as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4619): backfill changeset pr number to 4644

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4619): feed the bash fixture script via stdin, not argv, to fix Windows CI

Passing the script as a `-c "<script>"` argv element made it subject to
Windows' CreateProcess command-line argument encoding, which silently
dropped the escaped-dot backslashes before bash ever saw them (observed on
PR #4644's windows-latest CI shard: `1\.1` came back as `1.1`). Feeding the
same script via stdin instead removes argv entirely from the transport, so
there is nothing for Windows to re-encode. POSIX behavior is unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 15:47:14 -04:00