Commit Graph

2958 Commits

Author SHA1 Message Date
0xdhx
900504f985 fix(#3776): decide nothing-to-commit from the staged diff, not from staging success (#3859)
* fix(#3776): decide nothing-to-commit from the staged diff, not from staging success

`cmdCommit`'s empty-diff guard tested `stagedPaths.length === 0`, but
`stagedPaths` records paths whose `git add` exited 0 — "did staging
succeed", not "is there anything to commit". Staging an already-committed,
unmodified file succeeds while contributing no diff, so the guard was
reachable only when every named path was missing from disk.

The ordinary empty-diff case therefore fell through to `git commit`, where
the only thing converting the failure back to `nothing_to_commit` was a
string match on git's output. Git runs the pre-commit hook before it
decides there is nothing to commit, so a rejecting hook pre-empted that
match and the caller was handed `commit_failed` carrying a gate message
about a commit that had nothing to gate.

Ask git whether the staged paths actually differ instead. Two conjuncts
are load-bearing: the `length === 0` short-circuit keeps the
all-missing-paths case exact (a pathspec-less `diff --cached` would test
the whole index, so unrelated staged work would suppress the guard), and
`!isMergeInProgress` keeps a merge from being abandoned — during a merge
git refuses a partial commit, so the pathspec describes nothing about what
would land.

Nine regression cases in tests/commands.test.cjs cover the brief's six
acceptance criteria plus the merge interaction. Against the pre-fix build
exactly one fails; the other eight pin behaviour that was already correct.

Residual sibling of #2608/#2693, which covered `git add` failing; this
covers `git add` succeeding and contributing nothing.

* fix(#3776): exempt a cherry-pick too, but never a revert

The empty-diff guard must not decide from a pathspec git will not honour.
That test was merge-only; git refuses a partial commit during a cherry-pick
for the same reason, so the guard would have fired there and reported a
silent `nothing_to_commit` where the pre-fix code surfaced git's refusal.

The three sequencer states do not agree, so this is driven rather than
reasoned by analogy (git 2.54):

  MERGE_HEAD        fatal: cannot do a partial commit during a merge.
  CHERRY_PICK_HEAD  fatal: cannot do a partial commit during a cherry-pick.
  REVERT_HEAD       permitted; behaves like an ordinary commit.

REVERT_HEAD is therefore deliberately excluded: enumerating it alongside the
other two — the obvious move — would suppress this fix during a revert and
reintroduce the very misreport it removes. Both new states are pinned by a
test, and the revert arm fails against the pre-fix build exactly as AC1 does.

`canScope` keeps its narrower merge-only test on purpose; widening it would
change pre-existing cherry-pick behaviour, which is outside this fix.

* fix(#3776): probe the working tree, not the index

`git commit -- <paths>` is a PARTIAL commit: it records the working-tree
content of those paths and ignores what is staged. The guard was probing
`git diff --cached` — the index — which answers a different question than
the commit asks.

Driven on the same path, in this order: `git add` an unmodified file, then
write to it, then probe.

  git diff --cached --quiet -- p   rc 0   ("nothing staged")
  git diff --quiet HEAD -- p       rc 1   ("the tree differs")
  git commit -m m -- p             committed the new content

So a working-tree write landing between the `git add` above and the probe —
another process in a shared checkout, which this project explicitly supports
— would let the guard report `nothing_to_commit` for a call that would have
recorded that content. Probing `HEAD` asks the question the commit answers.

Not reachable through this function single-threaded, because the staging
loop re-adds every named path immediately beforehand, so index and working
tree agree at the probe. The change is correctness by construction rather
than a fix for an observed miscommit.

An unborn HEAD makes `diff HEAD` fatal; that falls through to the commit as
any other probe error does, and is now pinned by a test — the first commit
in a repo must not be swallowed by an empty-diff guard.

Both added probes are now gated on the guard being able to fire at all, so
an unscoped commit and an `--amend` pay for neither.

Found by adversarial pre-filing review; the index/worktree distinction was
not something my own path-shape probes could have surfaced.

* test(#3776): pin the assume-unchanged boundary; correct a stale comment

`git update-index --assume-unchanged` makes `git add` stage nothing and makes
BOTH diff forms — `--cached` and `HEAD` — report no difference, so no
diff-based guard can see a change to such a path. `git commit -- <path>` is
the odd one out: it reads the working tree directly and records it.

So a modified assume-unchanged path now reports `nothing_to_commit` where it
previously committed. That is the answer consistent with this function's own
staging step, which honoured the flag one loop earlier — but it is a
behaviour change, and it belongs on the record as a decision rather than
surfacing later as a surprise.

Also corrects an AC3 comment still describing the `diff --cached` whole-index
form that the previous commit replaced.

* chore(#3776): set changeset fragment pr to 3859

The fragment carries the PR's own number, which is unknowable before the PR
exists. Backfilled post-create; the repo's changeset lint rejects the `pr: 0`
placeholder.

* fix(#3776): do not read an unanswered sequencer probe as "no merge"

`execGit` surfaces a spawn timeout as `exitCode: 1` (`_spawnResult`:
`result.status ?? 1`) — the same code `rev-parse --verify` returns for a ref
that does not exist. So the MERGE_HEAD and CHERRY_PICK_HEAD probes could not
tell "not in that state" from "never answered", and the empty-diff guard read
both as "not in that state". That is the one path in #3776 that did not fail
toward the previous behaviour: a timeout during a real merge decided
`nothing_to_commit` from a pathspec git will not honour and left the merge
unconcluded, where before it was a loud `commit_failed`.

Treat an unanswered probe as "assume the partial commit would be refused" —
which falls through to `git commit` and lets git speak for itself.

Routed into `partialCommitRefused` only, deliberately never into
`isMergeInProgress`. That flag also feeds the pre-existing `canScope`, and
widening it there is worse than the misreport it fixes: with `canScope` false
the commit runs bare, and a bare commit during a merge is PERMITTED — git
concludes the merge with the whole index under a message naming one file.
Driven: the whole-flag form reports `committed` where this form reports
`commit_failed`, and it drops the pathspec on the ordinary timeout, re-opening
the #2112 scope leak.

* fix(#3776): pin the empty-diff probe against diff-only configuration

`git diff` is porcelain and honours settings `git commit -- <paths>` does not,
so an unpinned probe let a caller's configuration decide whether the guard
fires. Driven against git 2.54, each with the paired `git commit -- <path>`
confirmed to record the change the unpinned probe reported as absent:

  diff.ignoreSubmodules=all      a gitlink bump is invisible to the probe
  .gitmodules  ignore = all      the same, and it needs NO local config at
                                 all — it is checked in, so it arrives with
                                 a clone
  diff=<driver> + textconv       two different blobs converge to one text,
                                 so the probe sees no change; no submodule
                                 involved

`--ignore-submodules=dirty` rather than `=none`, because `dirty` is what a
partial commit of a submodule path actually means: it records the gitlink,
which moves only when the submodule's HEAD does. Under `=none` a merely dirty
submodule work tree reports a difference the commit would not record, sending
an empty call back to `git commit` — the same misreport, re-entered from the
other side. `dirty` still overrides both `diff.ignoreSubmodules` and a
checked-in `.gitmodules` `ignore`, so the gitlink vectors stay closed.

`--no-ext-diff` is deliberately absent: `--quiet` short-circuits ahead of an
external diff driver, so `diff.<driver>.command` cannot invert the probe
(driven: rc 1 with and without the flag).

* docs(#3776): disclose the two outcome changes the changeset omitted

The body listed what stays unchanged and never named the arms whose
user-visible outcome moves, so neither would have reached the changelog:

  - a modified path under `git update-index --assume-unchanged` now reports
    `nothing_to_commit` where it was previously committed. `git add` already
    honoured the flag one loop earlier; the guard reports what staging did.
    Documented in a code comment and pinned by a test since the first round,
    but absent from the fragment.
  - naming a submodule whose work tree is dirty while its recorded commit has
    not moved now reports `nothing_to_commit` rather than `commit_failed`,
    because nothing would have landed. New in this round, from the
    `--ignore-submodules=dirty` pin.

* test(#3776): use helpers.cleanup() for the submodule fixture teardown

`local/no-raw-rmsync-in-tests` rejects a bare `fs.rmSync` in a test: the
helper carries the Windows-EBUSY retry budget (`maxRetries`/`retryDelay`)
that a raw call does not, and a submodule work tree is exactly the shape
that holds handles open on Windows.

Caught by CI, not locally — the round ran the two affected suites but not
`npm run lint:ci`, so the repo's own rule never fired until the push. The
chain now exits 0 locally against this tree.

* fix(#3776): never drop a named assume-unchanged path

`--assume-unchanged` is the one state where `git diff` and
`git commit -- <paths>` genuinely disagree: `git add` stages nothing,
both diff forms report no difference, and `git commit -- <path>` still
reads the working tree and records it. The empty-diff guard therefore
reported `nothing_to_commit` about content the caller named in `--files`
and git would have written.

commit is made — so suppressing its misreport must not be paid for by a
silent drop. Same rule the timeout routing already follows: a fix for a
misreport may not cost content.

The guard now asks `git commit --dry-run --porcelain` whether the commit
would record anything, and stands aside on rc 0. That is the same
decision the real commit makes, so there is no second implementation of
it to drift. It does not run the `pre-commit` hook (driven: a rejecting
one neither fires nor writes a marker), which is what matters — a firing
`pre-commit` is the whole of #3776. It is NOT hook-free in general: git
2.54 fires `post-index-change` here, so a repo using that hook sees it
once for the probe and once for the commit. Stated rather than claimed
away.

Asking git rather than reconstructing its answer was reached by
measurement. Comparing `git hash-object` against `HEAD:<path>` was tried
and is wrong three ways, each a silent drop of named content: it misses a
mode-only change (`chmod +x` leaves the blob identical while the commit
records `100755`); it cannot hash a submodule path at all (`fatal: Unable
to hash sub`, while the commit advances the gitlink); and the path it
needs must be parsed out of `ls-files` output, which `core.quotePath`
renders as `"caf\303\251.md"` by default. Each has its own arm, and the
non-ASCII arm pins `core.quotePath` so it cannot go vacuous.

Falling through on the `ls-files` tag alone — without asking whether
anything would land — is also wrong: an UNMODIFIED assume-unchanged path
would reach `git commit`, which with any unrelated modified file present
prints `no changes added to commit`, a string the fallback does not
match, and returns `commit_failed`. That is #3776 re-entered from the
other side, the same shape `--ignore-submodules=none` would have
re-entered it. Pinned by its own arm.

The `ls-files` read is an optimisation, not a gate: it keeps the dry run
off the hot path when no assume-unchanged entry is present, and when it
cannot answer the dry run simply runs, because the dry run needs nothing
from it. Failing closed there would drop content and failing open would
re-enter #3776 — both are wrong answers to a question that can be asked
directly.

`--skip-worktree` is not a second instance. A present, modified one exits
1 from `git add` and fails closed as `staging_failed` above the guard; an
absent one is skipped before `git add` runs (#2014) and is answered by
the `stagedPaths.length === 0` arm, exactly as it was pre-fix. Both
shapes pinned, because the shorter claim ("never reaches the guard") is
too strong.

* test(#3776): register fixture teardown so a failed assertion cannot leak

`bumpedSubmodule()` creates its sub-repo as a SIBLING of `tmpDir`, and
the unborn-HEAD arm creates `fresh` outside it too, so the describe's
`afterEach(() => cleanup(tmpDir))` reaches neither. Both were cleaned by
a trailing statement in the test body, which any failing assertion above
it skips — leaking a git repo into the temp root.

`bumpedSubmodule()` now records the path and a describe-scoped
`afterEach` drains it, which covers all three of its callers at once;
the unborn-HEAD arm takes `t.after`, the form already used elsewhere in
this file.

Negative-controlled both ways with a deliberate assertion failure
injected into the dirty-submodule arm, under an overridden TMPDIR:
before, one `*-sub` repo survives the run; after, none.

* fix(#3776): never read an unanswered dry-run probe as "nothing to record"

The `git commit --dry-run --porcelain` probe that decides the assume-unchanged
boundary is the one probe in the guard whose rc 0 is the reassuring answer, so
it inverts the diff probe's safety: `execGit` collapses a spawn timeout (or any
spawn error) to `exitCode: 1`, byte-identical to git's own "nothing to record",
and the guard then reported `nothing_to_commit` about content named in
`--files` that git was never asked to write. Same conflation the sequencer
probes already defend against.

Only a CONFIRMED rc 1 with no spawn error closes the path now; a timeout, a
spawn error, or rc 128 falls toward the real commit, where git speaks for
itself. Five arms in tests/commit-files-pathspec.test.cjs pin it (posix +
windows timeout shapes, rc 128, the ls-files optimisation's own timeout, and a
negative control on an unmodified path); the injection helper gains an optional
`matchArg` so the dry run can be targeted without intercepting the real commit.

Also corrects the comment that claimed both sequencer probes are gated on
`guardApplies` — the MERGE_HEAD probe predates this fix and is unconditional.

* fix(#3776): probe with --no-verify so a hook-firing git cannot close the guard

Round 4, review finding 3 (Minor). The `git commit --dry-run --porcelain`
probe's safety rested on an empirical claim about one git version: that
`--dry-run` does not run `pre-commit`. git 2.54 satisfies it, but the failure a
differing version would produce is silent and lands in exactly #3776's own
configuration.

A `pre-commit` that fires and rejects exits 1 — the same code git returns for
"nothing to record" — so the closure would read it as a CONFIRMED empty answer,
drop the content the caller named in `--files`, and report `nothing_to_commit`.
That is #3776 re-entered through the probe the fix added.

`--no-verify` forecloses it structurally rather than documenting the version
dependency. Driven on git 2.54: rc-identical in both directions (rc 0
would-record, rc 1 nothing) with and without the flag, so it is behaviour-
neutral where the version already agrees.

Two claims deliberately NOT widened: `--no-verify` does not suppress
`post-index-change`, which still fires on this call with or without it (driven
both ways); and the real `git commit` is untouched — #3776 is a bug about a
hook's message reaching the caller wrongly, never a licence to skip hooks.

The new arm pins the FLAG rather than an outcome, because the outcome it
protects is unobservable on a git that already declines to run the hook. It is
a seam assertion over the argv the guard actually issued, not a source grep.

* test(#3776): pin all-missing --files during a merge or cherry-pick

Round 4, review finding 1 (Major) and finding 7 (Nit, its coverage half). The
review asks for the `stagedPaths.length === 0` disjunct to be gated on
`!partialCommitRefused`, or for the combination to be documented and tested.
Documented and tested — the gating is refused, with cause.

The premise is confirmed: the state is reachable exactly as described, and
during a merge the `nothing_to_commit` report does not tell the caller the merge
is still open. The prescription is not. With every named path missing,
`stagedPaths` is empty, so `canScope` is false and the fall-through reaches a
BARE `git commit`, which git PERMITS during a merge and which then CONCLUDES it.

Driven, git 2.54, through cmdCommit with the prescription applied:

  cmdCommit(cwd, 'add the thing', ['.planning/never-produced.md'])
  -> { "committed": true, "hash": "8e6bf45", "reason": "committed" }
     MERGE_HEAD gone; HEAD is a 2-parent merge commit recording
     .planning/shared.md with the caller's resolution content.

So the gating trades a report that writes nothing for one that silently writes
the whole index under a message naming a path that does not exist, and reports
success. That is the same trade the timeout routing already refuses one block
up, which is why the sequencer states gate the DIFF branch only.

The behaviour is also pre-existing and unchanged by this PR: at 86452da7 the
identical short-circuit sat ABOVE the MERGE_HEAD probe, so it never consulted
the sequencer either. The residual — a merge held open behind a
`nothing_to_commit` report — is offered as a separate issue alongside the three
already deferred, not folded into this fix.

These are behaviour pins, not regression tests: they pass at base and red on the
gated implementation (both arms, verified).

* docs(#3776): state the git-version provenance once, not at three claims

Round 4, review finding 6 (Nit). The guard carries ~159 comment lines around 32
lines of executable logic, and the review's specific complaint is that "driven
against git 2.54" is repeated near-verbatim in three places, which makes the
decision tree harder to scan.

Hoists the provenance to a single block header and reduces the three repeats to
the observation each actually carries. One claim keeps its version explicitly
and now says why: the `--no-verify` reasoning is version-SENSITIVE rather than
merely version-observed, so it is the one place the version is load-bearing
instead of incidental.

The behavioural matrix stays inline rather than moving to an ADR or a doc block.
Every claim in it is a constraint on the four flags immediately below it, and the
value of having it here is that the next reader who wants to "simplify" one of
those flags meets the driven counter-example in the same screen. Splitting the
constraint from the code it constrains is how the flags get dropped.

Comments only. No behaviour change; suite and lint:ci unchanged.

* docs(#3776): correct two driven figures in the new guard commentary

Both found by this round's own pre-push adversarial review, and both re-driven
before adopting.

1. The empty-paths rationale said a bare commit during a merge produces a
   "three-parent commit". It produces a TWO-parent merge commit. Three was the
   token count of `git rev-list --parents -n1 HEAD` (commit + two parents) read
   as a parent count. `git cat-file -p HEAD | grep -c '^parent '` returns 2.
   This round's commit message for the pins already said two, so the tree
   contradicted itself.

2. The `post-index-change` disclosure said a repo using that hook "sees it once
   for the probe and once for the commit". Driven with a counting hook: git
   fires it TWICE per `git commit --dry-run`, and twice again for the real
   commit — 2/2/2 across the flagged probe, the unflagged probe and the real
   commit. The disclosure understated the cost by half in both halves.

Comments only. No behaviour change; suite 351/351 and lint:ci unchanged.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-01 13:59:11 -04:00
BeeHiggs
41466e8e88 fix(#4023): preserve decimal phase ids in init progress ordering and smart-entry output (#4110)
* test(#4023): reproduce decimal phase-id coercions

* fix(#4023): preserve decimal phase ids in progress signals

* test(#4023): align phase token contract expectations

* chore(#4023): point the changeset at PR #4110

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-01 13:58:20 -04:00
Tom Boucher
80d2236ca5 fix(#4112): drop redundant $(( subshell wrapping in pause-work.md Context Detection (#4140)
* test(#4112): add failing-first regression coverage for pause-work.md Context Detection

Extracts and executes the shipped Context Detection bash block to prove the
$(( arithmetic-context misparse (SC1102/SC1106/SC2205) as a hard syntax
error, plus boundary/independence coverage for the phase/spike/sketch/
deliberation resolution the fix must preserve.

* test(#4112): fix false-green regression test — assert dash/sh failure, not bash -n

The prior version asserted bash -n syntax validity and wrapped bash -c
execution, neither of which detects this bug: bash/zsh silently fall back to
a working command-substitution reading of the ambiguous $(( construct, and
bash -n never evaluates arithmetic-context content at parse time. A gsd-test
run against the still-broken file passed 41594/41594 with this test in
place — a false green. POSIX sh/dash does not implement that fallback and
throws a real "Syntax error: Missing '))'", which is what this version
asserts against.

* fix(#4112): remove leaked tool-call tags corrupting the regression test file

A subagent's Write introduced trailing </content>/</invoke> markup at the
end of tests/pause-work-context-detection.test.cjs. gsd-test's own
tests/portability-rule-disable-ban.test.cjs (which parses every test file)
correctly caught this as a parse error. Strips the garbage lines; no
behavioral change.

* fix(#4112): drop redundant $(( subshell wrapping in pause-work.md Context Detection

phase=/spike=/sketch= opened with $((, which POSIX sh/dash parses as
arithmetic expansion and rejects (Syntax error: Missing '))') since the
enclosed text isn't valid arithmetic. bash/zsh silently retry it as command
substitution, which masked the defect there. Dropping the inner grouping
parens (matching the deliberation= line already in the same block) makes
the construct valid under bash, zsh, and dash alike, with identical
fallback-to-empty-string behavior when nothing matches.

Once the surrounding syntax parses, ShellCheck can now also analyze the
inner ls -lt calls it previously couldn't see past the parse failure,
newly surfacing the same pre-existing SC2012 ("use find instead of ls")
suggestion already baselined for the file's other ls usage. Baseline
updated to reflect the true current count; no ls-vs-find behavior change
made, as that is a separate, pre-existing, out-of-scope question.

* fix(#4112): verify shell discrimination instead of trusting the binary name

Code review finding: falling back from dash to sh could silently produce a
non-discriminating test on a host where /bin/sh is bash-compatible (e.g.
macOS) — such a shell never throws on the broken construct either, so the
test would pass whether the bug were present or not. resolvePosixShell now
verifies the candidate actually rejects a known-ambiguous $(( snippet
before using it, and skips with an explicit reason when none does. The
gsd-test Linux bench (dash as /bin/sh) is unaffected either way.

* fix(#4112): route regression test through process-seam, splitLines, cleanup

npm run lint:ci flagged 7 violations gsd-test's node:test run doesn't check:
local/no-adhoc-markdown-parsing (single fence-spanning regex),
local/no-crlf-fragile-split (bare \n split), local/no-unbounded-spawn (x4,
hand-rolled execFileSync with no timeout), and local/no-raw-rmsync-in-tests.
Rewrites the test to route every subprocess call through
tests/helpers/process-seam.cjs's runHook() (bounded by construction),
tests/helpers.cjs's cleanup() for temp-dir removal, and a line-by-line fence
scan using the text-lines.cts splitLines() seam, matching this repo's
established extraction idiom (tests/no-hardcoded-home-gsd-tools.test.cjs).
No behavioral change to what is asserted.

* docs(#4112): add changeset for the pause-work Context Detection fix

* docs(#4112): backfill changeset PR number (pr:0 -> pr:4140)

---------

Co-authored-by: sim <sim@local>
2026-09-01 13:03:33 -04:00
Tom Boucher
cad70f4f3e fix(#4120): replace shellcheck npm dep with dependency-free downloader (#4121)
* fix(#4120): replace shellcheck npm dep with dependency-free downloader

The `shellcheck` devDependency (added in #4109) pulled in decompress@4.2.1
for archive extraction, which carries an unpatched CRITICAL zip-slip
vulnerability (GHSA-mp2f-45pm-3cg9, CVSS 9.1) plus two moderate findings.
decompress's latest published version IS the vulnerable one -- no patched
release exists upstream, so npm audit fix cannot resolve this by upgrading.

Removes the shellcheck package entirely and replaces its role with
scripts/lib/shellcheck-fetch.cjs: a small downloader using only Node's
built-in https/zlib plus a hand-written tar-entry reader, fetching a pinned
koalaman/shellcheck release directly from GitHub releases. The reader never
uses an archive-supplied name as a filesystem path (the exact defect class
decompress had) -- it only returns the matched entry's bytes; the caller
writes those bytes to a path it constructs itself. Bounds the download with
a 30s-per-hop timeout, consistent with the ShellCheck subprocess's own
timeout. Covers linux/darwin on x86_64/aarch64, matching this repo's actual
CI (lint-tests runs only on ubuntu-latest) and local dev needs; Windows
fails with a clear, honest error rather than silently misbehaving.

Adds tests/lint-workflow-shellcheck-fetch.test.cjs covering the tar-parser
(unit cases plus a fast-check property test per CLAUDE.md's parser-testing
requirement), a security behavioral pin confirming traversal-style entry
names are treated as opaque strings never filesystem paths, and boundary
coverage for the redirect-following logic's MAX_REDIRECTS limit
(limit-1/limit/limit+1, via an injectable transport, no real network I/O).

npm audit: 0 vulnerabilities (was 1 critical + 5 moderate). The lint script
reproduces the identical result against the current tree:
"212 pre-existing finding(s) from baseline, 0 new" -- no behavior
regression, no baseline changes needed.

* fix(#4120): register shellcheck-fetch.cjs with the installer

scripts/lib/shellcheck-fetch.cjs shipped without being added to
GSD_SCRIPTS_LIB_FILES in bin/install.js, which would have left it orphaned
on uninstall and broken the golden install-tree fixtures for every runtime.
Adds the entry and regenerates the 19 affected fixtures via
npm run gen:install-tree.

* docs(#4120): add changeset for the decompress CVE fix

---------

Co-authored-by: sim <sim@local>
2026-08-31 22:31:23 -04:00
Cody Anderson
8c9265d4e5 fix(#3724): warning-only Dimension 3b findings no longer force the revision loop (#3758)
* fix(#3724): stop advisory Dimension 3b findings from forcing the revision loop

Dimension 3b (undeclared/temporal coupling, #1954) is spec'd "never a
blocker" but tagged severity: warning — the tier plan-phase's revision
loop counts as must-fix — and the planner is never taught the rule, so
every multi-wave phase touching shared mutable state replans at least
once, and intentionally coupled plans re-flag identically every
iteration to the stall prompt.

Three coordinated changes:
- gsd-plan-checker: retag 3b to severity: info, the tier
  references/revision-loop.md already exempts by design; recognize a
  coupling_justified frontmatter declaration in the Do-NOT-flag list so
  deliberate pairs converge. Additions are offset by trimming 3b
  motivation prose — the checker sits 45 bytes under its LARGE hard cap.
- plan-phase step 12: INFO-only accept — an issues block with zero
  BLOCKER/WARNING entries accepts the plan and surfaces the advisories
  instead of re-entering the revision loop. Real blockers and warnings
  still gate unconditionally.
- gsd-planner: slim pointer in assign_waves to the new
  progressive-disclosure reference gsd-core/references/planner-coupling.md
  (the planner sits 19 chars under its own cap), which carries the
  shared-mutable-state rule and the coupling_justified escape hatch so
  first-pass plans avoid the finding when the coupling is unintentional.

Documented the coupling_justified field in docs/reference/plan-md.md.
Growth acks per #2914; inventory manifest and install-tree fixtures
regenerated for the new reference file.

Closes #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin Dimension 3b at severity: info

The severity retag makes the old assertion (severity: warning) stale;
lock the advisory tier from both directions — info must be present,
warning must not — so a future edit cannot silently re-arm the
revision-loop trigger.

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* chore(#3724): changeset fragment for PR #3758

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* docs(#3724): roster planner-coupling.md in docs/INVENTORY.md

The new reference was enumerated in the manifest and all 19 install-tree
fixtures but missing its row in the Modular Planner Decomposition table —
the roster half the manifest-sync test cannot check. (Review Blocker.)

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): cover all four acceptance criteria (review round 1)

- plan-checker-coupling: the 3b severity assertion is now a PARITY check
  deriving the exempt tier from revision-loop.md's flow instead of
  hardcoding info — editing either side alone reds the suite. New
  describe pins the other three criteria: plan-phase's INFO-only accept
  clause (proven failing-first), the BLOCKER + WARNING count staying
  intact, the coupling_justified Do-NOT-flag exemption + fix_hint, and
  the planner pointer + planner-coupling.md content.
- ack fragment: $comment's plan-phase figure corrected to +79B; the 2775
  pin note carried forward into the gsd-planner.md entry, updated for
  upstream's #3761/#3764 Rule-paragraph anchor (which this diff leaves
  verbatim).

The parallel-dependent-plans re-anchor this commit originally carried was
superseded by upstream #3764 during review; this branch no longer touches
that file.

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 2 — align the stance enumeration, complete the template contract

MAJOR: <adversarial_stance>'s severity enumeration gains the INFO bullet so it
agrees with Dimension 3b's 'ALWAYS INFO' mandate instead of contradicting it.
Funded by extracting the inline <examples> block to the new progressive-
disclosure reference gsd-core/references/plan-checker-examples.md (@-inlined
from the same spot; #1949 precedent), which also restores the 3b motivation
clause round 1 traded away (Nit 4) and nets the agent file SMALLER than base
(49107 -> 48486) — the extraction the byte pressure was owed.

MINOR: gsd-core/templates/phase-prompt.md now carries coupling_justified, and
the field's shape becomes one 'plan-id: reason' string per coupled peer so a
plan justified against two peers can express it; docs/reference/plan-md.md's
Type column names the shape.

NIT: the 3409 ack's plan-phase entry no longer calls the #1168 workflow
ratchet an 'XL tier'.

Acks and derived artifacts updated accordingly (checker entry removed — a
shrink needs no ack; INVENTORY roster row + regen:derived for the new file).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): derive the 3b negative severity assertion (review round 2)

Every severity token in the 3b span must BE the tier revision-loop.md exempts,
replacing the hardcoded severity:warning negative — if the loop's exemption
ever moves, the failure names the real conflict instead of blaming the agent
file with a mutually-unsatisfiable pair.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): refit the planner coupling pointer under the char cap

Upstream #3299 (PR #3390) grew agents/gsd-planner.md to 49146 chars at the
base, leaving 5 chars of headroom where the +16-char pointer was measured
against 13 more. The pointer prose shortens to 'Non-file coupling:' —
49150 chars, back under the strict 49152-char cap — and the ack figures
follow. The @-path the tests pin is unchanged.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): re-home the plan-phase ack after the #3823 spent-fragment sweep

Upstream #3078/#3823 deleted all fully-spent ack fragments, including
3409-unreachable-guard-arms.json, which carried this PR's plan-phase.md
+79B append. Per the collision remedy that sweep added: take the deletion
and home the still-live entry in this PR's own fragment. Figures
re-measured at this merge base (90871 -> 90950 LF bytes).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): absorb the spent #3172 plan-phase fragment into this PR's ack

Upstream #3825 shipped 3172-stated-failing-direction.json naming only
plan-phase.md, now spent at the base — colliding with this PR's live
plan-phase entry. Per the #3003 pattern the fully-spent single-path
fragment is deleted and this fragment stays the path's one source;
figures re-measured at this base (93073 -> 93152 LF bytes).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 3 — true up the ack figures, restore the wave comment

The fragment's absolute sizes are re-measured and anchored to base
e40e9670 (planner 47259 -> 47330 chars, checker 45537 -> 44916 B,
plan-phase 91186 -> 91265 LF bytes), with a note that absolutes rot as
next moves — the deltas are the durable claims. The round-1 removal of
the '# Implicit dependency: files_modified overlap forces a later wave.'
pseudocode comment offset headroom base drift had already returned, so
it is restored (findings 2-3). Changeset gains the (#3724) backlink
(finding 4).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 4 — close the verify-work surface, harden the boundaries

BLOCKER: verify-work.md's verify_gap_plans is the second multi-plan
consumer of the checker's sentinels, and its ISSUES FOUND handler entered
revision_loop with zero severity parsing — the guaranteed replan #3724
fixed in plan-phase, alive on the gap-closure surface. The handler now
counts BLOCKER + WARNING and accepts INFO-only returns with advisories
displayed. The checker's INFO stance bullet is reworded to the claim that
is true everywhere ('revision gates count only BLOCKER + WARNING').

Minor 1: plan-phase's iteration_count >= 3 arm recounts severities, so an
INFO-only third check accepts instead of halting on a '0 issues remain'
user gate. Minor 2: the coupling_justified exemption now requires the
entry to NAME the other plan, closing the blanket-suppression reading.
Nit 1: INVENTORY row states the extraction buys cap headroom, not context.
Nit 2: the advisory display gains a concrete format on both surfaces.

Ack fragment re-anchored at base ddde001a: verify-work.md +264B (new
entry), plan-phase.md +395B, checker still net negative (-512B).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin the verify-work accept and the iteration-cap boundary (review round 4)

Two wiring assertions: verify_gap_plans' ISSUES FOUND handler gates on
BLOCKER + WARNING and accepts INFO-only blocks, and plan-phase's
iteration_count >= 3 arm recounts severities instead of gating advisories
— the limit+1 boundary of the gate this PR fixes.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 5 — fail closed at the gates, surface the advisory

Blocker 1: the checker's step-10 status rule routes an INFO-only result to
## ISSUES FOUND (with a new ### Advisories (info) template section and a
severity-aware recommendation) so the orchestrator receives the block and
displays the advisory instead of silently accepting a bare PASSED.

Blockers 2+3: all three gate surfaces (plan-phase step 12 both arms,
verify-work verify_gap_plans) carry one canonical clause verbatim — an entry
whose severity is missing or unrecognized counts as a BLOCKER (fail closed) —
making the accept condition an explicit-INFO whitelist while keeping
issue_count coherent for stall math.

Major 1: the INFO stance bullet scopes its claim to the plan-phase and
verify-work gates (quick mode's loop still revises on any ISSUES FOUND).
Major 2: INVENTORY row and ack $comment state the extraction's real trade
(readability, +0.6 KB eager runtime context), not a cap remedy.
Minor 1: plan-md.md marks coupling_justified as prompt convention, unvalidated.
Nit 1: ack absolutes re-anchored at base 1e67ec97; checker now +120B and acked.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin the round-5 contract — fail-closed parity, INFO-only return shape

New: three-surface verbatim parity test for the fail-closed clause (Blockers
2+3); checker return-contract test for the INFO-only ## ISSUES FOUND route and
advisories section (Blocker 1). All seven newly pinned tokens are absent at
f3a5682d, so each new assertion fails pre-fix.

Updated: accept-clause regexes track the explicit-INFO whitelist wording;
the severity sweep scopes to the span's fenced yaml examples via
yamlSeverityTiers (round-5 Minor 3, applied to the blocker negative too);
the iteration-cap comment states it is a prose pin, not an executed boundary
check (Minor 4); splitLines call sites document the line-pin coupling (Nit 2).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): adopt next's line wrap in the 3b motivation clause — drops a wrap-only hunk from the diff

Byte-identical content; the wrap difference was an artifact of the round-1
base adaptation predating upstream's #3003 landing.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 7 — gate every checker consumer, not just the two audited ones

Blocker: quick/steps/plan-checker-loop.md (issue-named in #3724) gets the
same canonical fail-closed clause and explicit-INFO whitelist accept as
plan-phase/verify-work — an INFO-only result proceeds instead of entering
quick mode's revision loop.
Major: import.md plan_validate handles the checker return by severity
(INFO-only never blocks an import) and is added to agent-contracts.md's
consumer enumeration, which had omitted it.
The checker's INFO stance bullet drops the quick-mode carve-out — the claim
is universally true again now that every consuming gate is severity-aware.
Minor: an applied coupling_justified exemption is surfaced as its own info
advisory so a stale one-sided declaration stays observable.
Nit: plan-phase's revision-iteration Display line is explicitly conditioned
on not having already proceeded to step 13.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM
Emitted-Drift-Ack-Growth: import.md — #3724 round 7: the plan_validate step's checker-return handler becomes severity-aware — counts BLOCKER + WARNING failing closed and accepts an explicitly-INFO-only return with advisories displayed instead of blocking the import

* test(#3724): pin the round-7 surfaces — five-gate parity, quick/import accepts, exemption visibility

The verbatim fail-closed parity test extends to quick/steps/plan-checker-loop.md
and import.md plan_validate; new assertions pin quick mode's INFO-only proceed,
import's never-blocks accept, import.md's presence in agent-contracts.md's
consumer row, and the surfaced coupling_justified exemption advisory. All four
newly pinned token families are absent at the pre-fix head, so each new
assertion fails first.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-31 19:08:11 -04:00
Tom Boucher
e242d85c92 fix(#4109): rewrap unquoted DISPATCH_SLUGS-shaped consumption to survive zsh (#4116)
* test(#4109): add zsh regression coverage for gate-check and invoke_reviewers dispatch

Extends tests/review-plan-coverage-manifest.test.cjs to cover the two
DISPATCH_SLUGS consumption sites the existing #3301 coverage-check tests
don't reach: the write_reviews gate-check block (ALL_LANES_SKIPPED /
TOTAL_LANE_FAILURE counting) and invoke_reviewers' dispatch + join loops.
Both extract the real shipped bash from review.md and run it under bash and
zsh, same convention as the existing coverage-check rows. Expected RED under
zsh pre-fix, GREEN post-fix.

* fix(#4109): rewrap unquoted DISPATCH_SLUGS-shaped consumption to survive zsh

Bash word-splits an unquoted scalar on IFS by default; zsh does not (no
setopt SH_WORD_SPLIT anywhere in these files), so a scalar accumulated from
multiple space-separated tokens and then consumed via bare `for x in $VAR`
collapses onto one bogus iteration under zsh whenever it holds 2+ tokens.

Fixes all four review.md sites reported by #4108/#4109 (invoke_reviewers
dispatch + join loops, the gate-check block, and the #3301 coverage-check
block), plus the identical pattern found by a repo-wide sweep in
complete-milestone.md, code-review.md, pr-branch.md, sync-skills.md, and
execute-phase/steps/per-plan-worktree-gate.md. Each site is rewrapped in
unquoted command substitution (`$(printf '%s' "$VAR")`), which re-splits
identically under both shells regardless of SH_WORD_SPLIT — the same
mechanism that already made the accumulator-building loops in these files
shell-safe.

* test(#4109): warn loudly when the zsh probe fails instead of skipping silently

detectShells() drops the zsh test lane whenever a live zsh probe fails (e.g.
on a CI runner without zsh installed), and a dropped lane reads identically
to a passing one in the suite's own output — exactly the blind spot that let
#4109's bug class ship undetected. Both copies of this helper (this file and
its byte-identical duplicate in review-build-prompt-optional-sections.test.cjs)
now emit a greppable console.warn when the probe fails, so a run without zsh
reads as "zsh coverage unknown" rather than "all lanes green".

* ci(#4109): gate workflows/*.md's embedded bash blocks in lint:ci

Adds scripts/lint-workflow-shellcheck.cjs, wired into npm run lint:ci, so
this bug class can't land undetected a third time. Two independent checks
run over every ```bash block in gsd-core/workflows/**/*.md:

- ShellCheck (new `shellcheck` devDependency, downloads and caches the real
  koalaman/shellcheck binary) catches the general unquoted-expansion/
  word-splitting family in argument position. 212 pre-existing findings
  across the tree are absorbed into scripts/lint-workflow-shellcheck-baseline.json
  (matched on {file, code, message}, not line number, so unrelated edits
  elsewhere in a file can't spuriously un-baseline anything) — fixing all of
  them is out of scope for this issue; only NEW findings fail the build.
- A custom structural check specifically for #4109's own shape: ShellCheck
  does not flag a bare `for x in $VAR` word-list — it treats that as an
  intentional idiom under any ruleset (confirmed empirically). This check
  does, and gates the build on any occurrence, with zero tolerance (no
  baseline) since every known site was already swept and fixed on this
  branch.

* test(#4109): update pr-branch cherry-pick loop test anchor for the zsh fix

extractPickLoop()'s PICK_LOOP_MARKER located the create_pr_branch cherry-pick
loop by the literal substring "for HASH in $INCLUDED_COMMITS", which no
longer appears verbatim after this issue's fix rewrapped that loop in
unquoted command substitution. Updates the anchor (and its error message,
now derived from the same constant instead of duplicating stale text) to the
new literal form. No behavior change to the extraction logic itself.

Emitted-Drift-Ack-Growth: review.md — #4109's fix adds explanatory comments at 4 sites; net code is functionally equivalent, comment expansion grows the file
Emitted-Drift-Ack-Growth: complete-milestone.md — #4109's fix adds explanatory comments at 3 sites documenting the zsh word-splitting divergence
Emitted-Drift-Ack-Growth: code-review.md — #4109's fix adds an explanatory comment documenting the zsh word-splitting divergence
Emitted-Drift-Ack-Growth: pr-branch.md — #4109's fix adds explanatory comments at 3 sites (TRANSIENT_DIRS, INCLUDED_COMMITS, FILTER_PATHS)
Emitted-Drift-Ack-Growth: sync-skills.md — #4109's fix adds explanatory comments at 2 sites (CREATE_LIST/UPDATE_LIST, REMOVE_LIST)

* test(#4109): cover lint-workflow-shellcheck parsers + add timeout

Adds tests/lint-workflow-shellcheck.test.cjs covering the hand-rolled
parser/logic functions exported by scripts/lint-workflow-shellcheck.cjs
(stripCommandSubstitutions, stripShellComments, substitutePlaceholders,
extractForLoops, findBareForLoopSplits, findingKey, partitionAgainstBaseline)
that shipped with zero coverage — CLAUDE.md requires at least one fast-check
property test for parsers, included here (bare/braced forms always flagged
naming the variable; quoted/substituted/literal forms never are).

Also bounds runShellcheck's subprocess with a 60s timeout: the `shellcheck`
npm package's own API has no timeout option and internally blocks on a
synchronous spawnSync, so this reimplements the binary resolve/download step
via the package's own exported config/download and calls spawnSync directly
with a native timeout. And guards the CLI entry point with
`if (require.main === module)`, matching this repo's sibling dual-purpose
lint scripts — without it, requiring the module for its exported functions
(as the new test file does) also triggered a live ShellCheck run as a side
effect.

* docs(#4109): add changeset for the zsh word-splitting fix

* chore(#4109): backfill changeset PR number (#4116)

---------

Co-authored-by: sim <sim@local>
2026-08-31 18:08:06 -04:00
BeeHiggs
1a358ce0fd feat(#2761): bracket-tolerant read path — roadmap/validate/verify/state recognize bracket ids (epic #612 PR-2) (#2867)
* feat(#2761): gated heading-intro selection + one bracket identity grammar

Foundation. Two owner-level changes plus a federated convention resolver; no
reader consumes them yet.

1. GATED SELECTION, not an ungated widening.

   Widening every heading matcher requires the claim "no legacy ROADMAP contains
   a `[CODE.MM]` bracket followed by a digit", and that is false:
   `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and
   `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each
   as a phase — moving phase_count and total_phases and adding W006 on projects
   that never opted in. No narrowing rescues it: the premise is about documents
   we do not control.

   `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the
   pattern SOURCE at construction time. A project whose resolved
   `phase_id_convention` is not exactly 'bracket' compiles the same source string
   it compiled before. `baseline` is explicit because whether a site spells the
   any-bracket prefix or a bare `Phase\s+` is a fact about that site's history:
   handing the wider grammar to a bare site retro-grants tolerance it never had,
   in both directions — warnings appear, and a warning that fires today vanishes.

   Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through
   the base alternative, which captures nothing, so a reader saw no bracket, fell
   back to the legacy token rule, and counted a labeled icebox heading while
   excluding the label-less one beside it — two derivations of one ROADMAP
   disagreeing.

2. ONE bracket identity grammar, one width rule.

   The milestone width is reconciled with the emit validator: pad2 output, so
   two digits or 3+ with no leading zero. Earlier spellings diverged in both
   directions — admitting `002`, which the validator rejects, and a bare `0` pad2
   never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED
   a milestone no phase heading could then resolve into, recreating the
   on-disk-count fallback this epic removes. An unpadded bracket is now uniformly
   malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its
   directories is the surfacing signal.

   The milestone field is boundary-anchored, so a malformed run cannot match by
   its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays
   case-insensitive because readers compile `/i`, but identity helpers match
   `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:`
   failed every sentinel test. The qualified key shares the width, the `(?=-|$)`
   boundary and the single-sub-phase shape of the directory token, because
   phaseTokenMatches returns unconditionally on a qualified hit: a key matching a
   directory isPhaseDirName rejects would be a final wrong answer.

3. resolvePhaseIdConvention federates workstream -> root exactly as
   config-loader does — including that root is a fallback only when a WORKSTREAM
   is active, so a project-scoped directory stands alone. loadConfig cannot serve
   this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and
   this key is not among them. It governs the bracket-selection reads ONLY.

PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing
consumes it, and it is superseded rather than redefined.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#2761): roadmap.cts selects its heading grammar from the convention

Six matchers build their intro through the gated selector, and
cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each
resolve the convention ONCE per command and thread it down.

Three sites take the any-bracket baseline (they already tolerated
`[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`).
Handing the wider grammar to a label-only site retro-grants tolerance it never
had — and not only by adding matches: on a legacy repo an unchecked
`- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that
fires today.

Sentinel handling under bracket ADDS a rule rather than replacing one: a
bracketed heading is a sentinel when its bracket milestone is reserved
(`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog
convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a
mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly
the content this epic targets — add entries to the progress denominator. The
captured id is folded before the identity test, so a lowercase
`### [gsd.999] 07:` is excluded too.

The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention
once and hands it to all four of its heading/checklist patterns, but the single
`phaseTokenMatches` call that decides `disk_status`, `plan_count`,
`summary_count`, `has_context` and `has_research` was left two-argument — so
every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with
zero counts, on the PR's own headline verb, while the SAME build resolved those
same directories correctly in three other places on the same repo (W006/W007 via
phaseTokenFromDir, `state json` via the milestone filter, and the W021
milestone-complete read through this very helper's three-argument form). It
failed ONLY for the directory shape the convention exists to name: a
mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is
why nothing caught it. Measured, bracket vs its flat-legacy twin:
`[["01","no_directory",0,0],["02","no_directory",0,0]]` against
`[["01","complete",1,1],["02","planned",1,0]]`.

The oracle is the twin, computed in the same test run, plus exact literals —
`grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix
nor a future regression had any gate at all.

Disclosed: a ROADMAP written in bracket form before config.json is switched
reads as empty rather than mis-counted. Silent invisibility during the migration
window is the deliberate trade against claiming phases on projects that never
opted in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#2761): validate.cts selects its grammar; gated directory recognition

The W006/W007 feeders take the resolved convention as a threaded parameter.
These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them
where an ungated widening does the most damage: `### [RFC.2119] 5:` enters
roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on
disk" on a project that never opted in.

buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket
headings. Surfaced rather than filtered in place because roadmapPhases feeds both
a membership check and a missing-directory warning, and only the latter should
ignore an icebox item.

That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is
a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying
suppression on the token alone let an icebox heading silence a REAL phase that
happens to share its number — a false negative strictly worse than the warning it
removed. A token is suppressed only when no non-sentinel heading bears it.

Directory recognition is added as gated FUNCTIONS beside the exported RegExp
constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is
string-indistinguishable from the letter-prefixed-decimal family this repo
documents as ambiguous, and folding a branch in changes those constants' answers
on exactly that family. A RegExp constant has nowhere to attach a gate.

The recognizer mirrors the emit grammar and delegates the token to the canonical
owner, so recognizer and resolver agree on rejected input as well as accepted.
Both functions throw on a non-string, matching the call pattern they replace.

buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin
and like the sibling checklist scan in roadmap.cts, and for the reason that one
states: the bracket id has to ride along or the sentinel filter is blind to
`- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every
checklist token REAL, and the occurrence-aware un-suppression loop then deleted
the icebox token the HEADING scan had correctly marked sentinel — so `validate
consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE
ROADMAP shape where an icebox appears as both a bold bullet and a detail
heading. `validate health` stayed silent on that same repo, so the two verbs
disagreed — which is the disagreement `sentinelPhases` exists to close.

Both directions are pinned, because the failure mode of a careless fix here is
the opposite one: a real phase sharing a sentinel's token must still warn. It
does, in all four shapes that attack it (sentinel heading + real bullet,
lowercase sentinel, sentinel after the real heading, colon-less bullet).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2761): count bracket headings, and retire them, in both derivations

Both `total_phases` derivations select their grammar from the resolved
convention, in one commit — cmdStateSync already carries the comment that it
mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)",
so teaching one and not the other ships that divergence.

The #1514 retirement filter widens WITH the counter it protects. The canonical
gesture strikes the checklist BULLET and leaves the detail heading intact, so a
bracket-form retirement went undetected and the phase stayed in the denominator
forever. That is half a fix alone: the retired key is compared against
phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves
land here.

Under bracket the sentinel token rule composes as the full engine set {0, 999},
so this counter agrees with `roadmap analyze`, which has always excluded both —
otherwise the two derivations report different numbers for one ROADMAP and the
changeset's "excluded from every count" is false as written. The LEGACY path
keeps its pre-existing 999-only rule: widening it there would move legacy totals,
so the two stay split off the bracket path exactly as they are today.

The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not
the frontmatter total_phases. Sync's own counter never reaches that field — the
read derivation writes it — so asserting the frontmatter after a sync measures the
read path twice and lets a mutation to the write-path guard survive.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read

The shipped milestone-prefixed W021 gate keeps its ROOT-only config read,
verbatim base semantics. Federating it silently moved a legacy convention's
answer in BOTH directions on workstream repos — a W021 that fires at base
vanishing, and one that is silent at base firing. resolvePhaseIdConvention
governs the new bracket-selection reads only.

B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it
with an empty config) but selects its grammar from the convention. Inferring
'bracket' from the shape of a matched bracket ran a repo-failing check against a
legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution
widens with the heading read, so a bracket repo whose phases are on disk stays
silent, and a bracket sentinel is not reported as unstarted.

checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so
fenced examples cannot warn and heading level is structural. Its scope rules each
close a way it silently did nothing or fired wrongly: only a genuine MILESTONE
heading opens or closes a section (a `### Notes` used to reset scope and disable
both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed
phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a
phase; the full h2-h6 range is processed. Its section recognizer shares the one
milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the
id grammar and a section to the section grammar at once, silently re-scoping
every warning after it.

validate consistency suppresses bracket sentinels in its missing-directory
warning — the two verbs disagreed, health suppressing via notStartedPhases while
consistency did not. The legacy reading is untouched, including its pre-existing
wart that `### Phase 999:` still warns there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2761): scope the milestone by its bracket; select the disk-side filter

Two roadmap-parser reads, both of which made a bracket project's totals track
the disk instead of the ROADMAP.

The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name,
no version — but scoping matched STATE's `milestone: v2.0` STRING against a
heading, so the canonical form matched nothing and total_phases fell back to the
directory count. The rule was re-derived in THREE places: extractCurrentMilestone
plus two `milestoneBounded` guards; fixing one left the others falling back
regardless, so they are now one gated helper. It matches the CANONICAL padded
spelling only — accepting `0*N` bounded a milestone whose phases were invisible,
which un-suppressed a progress percent computed off an unscoped disk count.

getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a
bracket ROADMAP it collected nothing, so the filter degraded to pass-all and
buildStateFrontmatter counted every other milestone's directories — making the
bracket convention strictly worse than the M-NN one it supersedes on the property
that matters most: totals must track the ROADMAP, not the disk.

The DIRECTORY side of that same filter is selected with it. Teaching only the
heading scan was half a fix and a worse one: `milestonePhaseNums` became
non-empty, so the pass-all degrade stopped firing, but no bracket directory could
satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the
custom-id match captures the project code `GSD`, and stripProjectCodePrefix does
not strip a dotted prefix). Every bracket directory was rejected, and
completed_phases / total_plans / completed_plans / percent all collapsed to 0
while `state sync` went on writing a percent off the unfiltered disk — `state
json` reporting 0% on the same repo, in the same second, that STATE.md's body
called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and
total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)`
floors it at the ROADMAP count no matter how many directories are rejected.

The dir side matches on the milestone-QUALIFIED id, delegated to the owner's
gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B
puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one`
share the token `01` and only the qualified key separates them. The qualified ids
are kept in their own set — a hyphen in `milestonePhaseNums` would flip
`roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo
— and the branch is ADDITIVE: on a miss it falls through to the three legacy
checks, so a bracket project carrying legacy-shaped directories reads unchanged.

Both are resolved lazily and gated, so the legacy path pays neither a config read
nor a second scan and cannot change answer. The scoping call is also GUARDED:
resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a
GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only
planningDir call in extractCurrentMilestone sits inside the STATE-read try, so
the function returned normally on such an environment; an unguarded one here let
that escape and broke the never-throws invariant that getRoadmapPhaseInternal and
getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about.
Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the
workstream-name policy and GSD_PROJECT throws identically at base — but reachable
by any in-process embedder, which is precisely who that invariant is for. The
filter's own resolve call was already inside its try and is unaffected.

The milestone-qualified key is formed only for a token that is itself a bracket
phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration
heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to
`GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02:
the `-01` truncated, both such headings collapsing to one key, and the heading
claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting
`GSD.02-01-one`, the one it does. The guard drops those headings back to the
unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned
against the milestone-prefixed reading of the same ROADMAP, which is
base-identical on this shape.

Scoped precisely, because the fixture moves one number that the guard does not
touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket
heading COUNT this PR exists to add, not the splice — measured identical with and
without the guard, and identical to what the canonical `### [GSD.02] 01:`
spelling does on the same fixture (both read 2 with zero directories on disk,
where base reads 0). The claim is base-equivalent ACCEPTANCE, not a
base-equivalent reading.

One consequence is stated rather than fixed: a heading whose token carries a
hyphen still puts that hyphen into milestonePhaseNums and so still flips
`roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it
is what keeps the shape base-equivalent; excluding the token would have moved
answers versus base on malformed input. The comment at the qualified-set
declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out
of that flag's input, not hyphens in general.

The oracles ship with it, and they are the five numbers, not the one: the parity
gate now asserts total_phases, completed_phases, total_plans, completed_plans AND
percent, on both derivations, on two fixture shapes (one milestone; two
milestones with stale prior-milestone directories on disk). The oracle is the
flat-legacy twin, built in the same test run and compared number for number,
plus exact literals so a shared wrong answer cannot pass.

The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could
not serve, because buildStateFrontmatter's #2445 de-dup key captures only a
directory's leading integer and collapses `02-01-one` / `02-02-two` /
`02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's
[3,2,3,2,67], identically at base and before this fix, and structurally
unreachable from the bracket key space. That reasoning is only sound while it
stays true, so a characterization test holds the M-NN reading down on the two
numbers that do not depend on which directory wins the mtime race. Widen the
de-dup key and it fails, instead of quietly invalidating the changeset's
disclosure.

Also adds the call-site pin. The structural table pins transcription against the
selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping
verify.cts's milestone-complete site to the wider baseline grants a
fires-on-every-repo check tolerance it has never had, and every behavioural test
still passed. The pin reads the shipped sources and asserts the mode at each of
the 14 sites, count-exact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): pin the bracket read surfaces in the parity gate

This gate exists because #2043 fixed one bug across five hand-edited copies of a
rule and #2232 was the residual that survived, because a later reader could not
tell the copies were one rule. PR-2 adds two consumers, so they belong here.

Surface 7 — the heading read and the directory read must agree about WHICH phase
a `MM-<seg>` pair names, across the shared width corpus, and the bracket and
legacy spellings of one heading must yield the same token.

Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on
ACCEPTED input was already pinned; agreement on REJECTED input is where they
actually diverged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2761): changeset

Disclosures for the PR body (deliberate, not defects):

- phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and
  cannot serve as the convention resolver however the file is federated. This PR
  ships its own workstream->root resolver; adding the key and its value enum is
  later-slice work.
- Convention matching is strictly === 'bracket'. A misspelled value reads as
  not-configured and the project keeps legacy behaviour silently.
- An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing,
  bounds nothing, sections nothing, and is not a phase id. W005 on its
  directories is the surfacing signal.
- WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical
  inputs, none of which toDir can emit and none of which had a bracket caller at
  base:
    isSentinelPhaseId('GSD.0-01',    'bracket')  true  -> false
    isSentinelPhaseId('GSD.0999-01', 'bracket')  true  -> false
    getMilestoneFromPhaseId('GSD.2-01',   'bracket')  'v2.0' -> null
    getMilestoneFromPhaseId('GSD.002-01', 'bracket')  'v2.0' -> null
  The canonical pad2 sentinel spelling `[GSD.00]` still tests true.
- FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x ->
  milestone null) is preserved". After the unification that holds for the
  canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours;
  flagging the tension rather than editing it.
- The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is
  a sentinel when its bracket milestone OR its token is reserved. Under bracket
  the state-side token rule is the full {0, 999} set so both derivations agree;
  the LEGACY path keeps its pre-existing 999-only rule, unchanged.
- validate consistency's legacy reading is untouched, including the pre-existing
  wart that `### Phase 999:` warns there while validate health suppresses it.
- find-phase still cannot resolve a bracket phase directory. phase-locator.cts is
  outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls
  phaseTokenMatches without a convention, so whichever slice lands second must
  thread it through.
- Four of the five bracket readers scan raw ROADMAP content, so a bracket heading
  inside a fenced code block is read as a phase. Pre-existing for the legacy
  spelling; parity, not a new class.
- roadmapPhaseLookupSources gained no bracket source: nothing emits a
  milestone-qualified query into it yet.
- roadmap validate remains a separate, unfederated convention reader.
  Pre-existing and base-identical, but two verbs can disagree about the active
  convention on one project.
- _diskScanCache keys on cwd while the values it caches are now
  convention-dependent. Not reproducible through the CLI; pre-existing for the
  workstream dimension, widened here. Stated as inconclusive.
- A ROADMAP written in bracket form before config.json is switched reads as empty
  rather than mis-counted — the deliberate migration-window trade.
- THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that
  divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter
  applies the milestone filter; cmdStateSync does its own fs.readdirSync and never
  calls it, so on a repo carrying prior-milestone directories the read path
  reports the SCOPED percent and the sync body reports the WHOLE-DISK one.
  Measured on the true base build (d04592de), flat-legacy spelling, 3 in-scope
  phases with 1 complete plus 2 stale prior-milestone dirs: `state json`
  [3,1,3,1,33], sync body 60%. The bracket twin of that repo now reads the same
  two numbers — 33 and 60. Scoping the sync counter would move every legacy
  repo's percent, which a bracket read-path PR must not do. The gate pins both
  sides, so the mirror cannot silently become a one-sided fix.
- THE PARITY ORACLE IS THE FLAT-LEGACY TWIN, NOT THE M-NN ONE, and that is a
  measurement finding rather than a preference. buildStateFrontmatter's #2445
  de-dup key captures only a directory's LEADING integer, so the M-NN dirs
  `02-01-one` / `02-02-two` / `02-03-three` all key to `2` and two of the three
  are dropped before they are ever counted: base reads [3,0,1,0,0] where the
  flat-legacy twin of the same repo reads [3,2,3,2,67]. Present identically at
  base and at HEAD, untouched here, and structurally unreachable from the bracket
  key space — `GSD.02-01-one` does not match that pattern at all, so every bracket
  directory keys to its own name. The source line already carries a
  `phase-id-owner:` sanction recording the divergence. Mirroring it under bracket
  would mean manufacturing a collision that cannot occur, so the gate compares
  against the flat-legacy spelling, which is uncontaminated. This paragraph is
  itself pinned: a characterization test holds the M-NN reading on the two
  numbers that do not depend on which directory wins the mtime race, so widening
  the de-dup key in a later slice fails the suite rather than silently making
  this disclosure false.
- THE `phaseTokenMatches` CALL-SITE CENSUS, stated so the remaining gaps are
  auditable rather than implied. 13 call sites outside the owner (phase-id.cts).
  THREE are three-argument: verify.cts:2229 (the W021 milestone-complete read,
  already was), roadmap.cts:436 (`roadmap analyze`'s directory lookup, threaded
  by this PR) and roadmap-parser.cts:792 (the disk-side milestone filter, added
  by this PR). The other TEN are two-argument and stay that way — phase.cts ×5
  (220, 277, 444, 585, 1547), phase-locator.cts:62, smart-entry.cts:243,
  init.cts:1414, milestone.cts:551 and verify.cts:2467. All ten are untouched by
  this PR and base-identical.
  One of them sits in a file this PR DOES edit, so it is named rather than left
  to a reader's grep: verify.cts:2467, `verify schema-drift <phase>`. Measured on
  a bracket repo across base / pre-fix branch / this HEAD, all three agree on all
  three argument forms — `verify schema-drift GSD.02-01` and `… 01` both report
  "Phase directory not found" on every build, and `… GSD.02-01-one` resolves on
  every build through the exact-directory-name fallback. So the user-visible
  shape of what stays broken is: a bracket phase is addressable there by full
  directory name only, exactly as at base. Threading the convention into a
  function this PR never touched, in the last round before ship, is the wrong
  trade; it is where the same one-argument fix goes next, alongside
  milestone.cts:551 and init.cts:1414.
- A BRACKET HEADING WHOSE TOKEN CARRIES A HYPHEN (`### [GSD.02] Phase 02-01:`,
  a mid-migration spelling) forms NO milestone-qualified key, and therefore
  scopes through the unqualified legacy path — base-equivalent ACCEPTANCE, which
  is the claim, and not a base-equivalent reading: `total_phases` on that shape
  moves 1 -> 2 for the same reason it moves on the canonical `### [GSD.02] 01:`
  spelling, because counting bracket headings is what this PR does. Such a token
  still flips `roadmapUsesHyphenedIds`, as it also does at base. The comment at
  the qualified-set declaration now claims only that narrower, true thing.

The `pr:` field carries the sub-issue number as a placeholder — it must be
updated to the real PR number when the PR is opened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2761): point the changeset at PR #2867

* test(#2761): fast-check properties for the convention-selection layer

CONTRIBUTING.md mandates a generative property test for parser/bijective-contract
changes; PR-2 shipped six example-based files and none. This adds the missing
layer, scoped to what PR-2 actually contracts — WHICH pattern each reader
compiles, decided by the resolved `phase_id_convention` — rather than restating
PR-1's grammar round-trip properties, which already live in
tests/adr-612-bracket-grammar.test.cjs.

Four properties: P1 an opted-in repo reads the ADR-canonical label-less bracket
heading/dir and a non-opted-in repo is byte-blind to the identical input; P2
every non-bracket convention agrees with the hand-transcribed BASE source over
generated content, including bracket-DOTTED legacy prose (`[RFC.2119] 5:`) that
must never be claimed as a phase; P3 nine per-field mutations are rejected and
the one case variation folds instead; P4 both sides of a phase comparison derive
the same key under the same convention.

Generators template every input from raw primitives — nothing is seeded through
renderPhaseId/toDir, the p2() tautology that made #2258 round 1's property test
structurally unable to find B1. Domain reaches past 99 into the 3+-digit branch
(round 2's numArb-capped-at-99 miss), forces sub-phases in at weight, and pins
both sentinel milestones.

Falsified against the COMPILED lib, not the source: five deliberate mutants
(gate never fires; gate always fires; milestone width widened to \d+; the #612
convention forwarding dropped from phaseKeyFromDir; extractPhaseToken's bracket
branch ungated) each fail the specific property that should catch them —
16/2, 16/2, 15/3, 17/1, 17/1 pass/fail — and the lib restores byte-identical.

An earlier draft of P2 held vacuously: its base regex omitted the markdown
furniture the selected one carried, so every realistic `### Phase NN:` line
matched neither side. The gate-always-fires mutant did not kill it. Both are now
compiled through one function, and that mutant kills P2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): adversarial malformed bracket tokens across the tolerant readers

The existing boundary coverage stopped at shapes the emit grammar rejects
(unpadded `[GSD.3]`, wrong-case, `12A`). It never exercised a STRUCTURALLY
broken token — a non-numeric milestone, a bracket that never closes, a bracket
nested in another — which is the input a tolerant reader is most likely to
half-read, and the one the PR's own regex commentary is explicit about.

read-tolerance (roadmap heading scan + validate's dir and variant builders):
ten malformed headings, each asserted to be read as a phase by NO convention and
to give the opted-in repo the same answer as the legacy one; the corpus driven
through `roadmap analyze` end to end; malformed DIRECTORY names asserted
unrecognized and non-throwing on all four conventions; and the two variant
builders asserted to agree, since a widening that reaches only one splits
`validate consistency` from `validate health` (the #3242 Bug B shape).

coherence (verify.cts W021): the same six broken shapes asserted to raise no
W021 of their own AND not to re-scope the W021 that follows them — the G2
failure mode reached from a different shape, where a heading that is not a phase
but IS read as a section silently moves later warnings onto the wrong milestone.

Both files gain a pathological-input time bound. Nested quantifiers over a long
unclosed bracket are the classic ReDoS shape and two commits on next (#2828,
#2944) were CodeQL-flagged for exactly that, so the bound is asserted rather
than argued from reading the pattern. The probes themselves parse no regex.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2761): document "bracket" as a phase_id_convention value

The row listed only `"milestone-prefixed"` and `null`, so after two shipped
slices (#2258 grammar, this PR's read path) the convention had no documented
enum value. CONFIGURATION.md is also a top-10 historical co-changer of both
src/verify.cts and src/state.cts and was absent from this PR.

The row states the boundary rather than the ambition: `"bracket"` changes the
READ path only, there is no migrator and no emit yet, and a project on any other
value compiles the patterns it compiled before. That keeps the docs honest for
the two releases before PR-3 and PR-4 land, instead of describing a convention a
user cannot yet migrate to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2761): retype the changeset Added, drop the docs-exempt marker

`Fixed` was wrong by CONTRIBUTING.md's own definition — a fix restores
documented behavior, and bracket read tolerance is the second slice of a
capability that did not exist before #2258. The type also carried a
`docs-exempt` marker, and `Fixed`/`Security` are exempt from the docs-required
lint, so the typing had the effect of routing around a gate this change should
pass. It now passes it: `lint-docs-required` returns ok_docs_updated on the
CONFIGURATION.md row added in the previous commit.

Body gains one sentence pointing at that row and restating that `"bracket"` is a
read-path opt-in until the migrator and write path land.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): pin the version-less bracket milestone scoping gap; narrow the claim

The changeset asserted that milestone scoping "recognises the ADR-canonical
`## [GSD.02] Foundation` heading and applies to the phase DIRECTORIES too." A
CLI probe on that exact heading form falsifies the second half: with no `vN.N`
in the milestone heading the directory side does not scope, and directories from
BOTH the prior and the later milestone are admitted. Measured 4 dirs counted
where the milestone declares 2.

Every bracket fixture in the suite writes `## [GSD.02] v2.0: …`, so nothing
covered the form the ADR actually specifies — and the state.cts doc comment
calls that version-less form canonical.

Mechanism, in extractCurrentMilestone: the bracket scope branch selects the
right currentSection, but `preambleCutoff` keys off a pattern requiring a
version or status emoji, so a version-less roadmap falls back to the current
milestone's own offset and every PRIOR milestone lands in the preamble — whose
phase-stripping regex only strips `Phase N:`-labelled headings, so bracket phase
headings survive it. Independently, `computeSectionEnd` accepts a boundary only
on a version/emoji heading, so the section runs to EOF and every LATER milestone
is swept in. Two sites, bidirectional.

Not fixed here: it changes milestone scoping, which is shared with the legacy
path. Five characterization tests pin today's reading plus a versioned CONTROL
proving the version string is the only difference, and the changeset sentence is
narrowed to what the code does. The DEFECT assertions are written to be
INVERTED by the fix, not deleted — that inversion is its regression proof.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2761): re-anchor the branch's own cross-file line citations after the rebase

Three of this branch's code comments cite sibling call sites by line number, and
the rebase onto 178ec000 moved two of the three targets:

  roadmap-parser.cts  validate.cts:210 -> :218   (const g = capturing ? 1 : 0)
                      state.cts:1715   -> :1752  (const bg = … 'bracket' ? 1 : 0)
  roadmap.cts         verify.cts:2229  -> :2355  (phaseTokenMatches 3-arg form)

`state.cts:1715` had drifted 37 lines and now lands on the retirement skip, not
the capture-offset idiom the sentence is about — the citation read as evidence
for a claim the cited line does not support.

planning-workspace.cts's `config-loader.cts:618/:649` was checked and is still
correct; left alone.

Comment-only. Build, drift guard and the bracket suites re-run unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2761): scope the version-less bracket milestone heading too (B1)

computeSectionEnd and the preambleCutoff scan in extractCurrentMilestone
(roadmap-parser.cts) only recognized a milestone boundary heading that
carried a vN.N token or a status emoji. The ADR-canonical bracket
heading (## [GSD.02] Foundation) carries neither, so on that shape
computeSectionEnd fell through to content.length (sweeping every LATER
milestone into scope) and preambleCutoff fell back to the current
milestone's own offset (leaking every PRIOR milestone's bracket phases
into the preamble, whose Phase-N: strip regex never matches them).

Under the bracket scope branch, both sites now also accept a
`#{1,2}\s+\[CODE.MM\]` boundary, built from phase-id.cts's BRACKET_ID_SRC
(single owner of the bracket-id grammar) rather than a re-typed literal.
`#{1,2}` is the deliberate discriminator: a bracket PHASE heading is
level 3 and shares the same `[CODE.MM]` prefix, so a `#{1,3}` boundary
would swallow it too. Reachable only when bracketScopeConvention ===
'bracket' was already resolved (i.e. the bracket scope branch actually
fired), so version-bearing/emoji headings and non-bracket conventions
take the exact pre-existing code path byte-identically — confirmed by
the full adr-612 suite staying green.

Inverts the four DEFECT assertions in the
"#612 PR-2 CHARACTERIZATION: a version-less bracket milestone does not
scope" describe block (tests/adr-612-bracket-phase-counting.test.cjs)
into their regression-proof form, per the block's own doc comment, and
reframes the describe title/comments accordingly. Corrects the
.changeset/2761-bracket-read-tolerance.md fragment, which described the
directory-side version-less gap as an open, un-closed bound.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): thread sentinelPhases into validate health's W006 loop (B2)

cmdValidateHealth's W006 loop (src/verify.cts) destructured only
roadmapPhases from buildRoadmapPhaseVariants, not sentinelPhases —
unlike cmdValidateConsistency, which already skips sentinelPhases with
the identical guard a few hundred lines up. A heading-only bracket
icebox/pre-milestone entry ([GSD.999] / [GSD.00]) therefore gained a
false W006 "no directory on disk" from validate health while validate
consistency correctly stayed silent on the very same ROADMAP — the two
validators contradicting each other.

Threads sentinelPhases through and skips it before the existsOnDisk
check, mirroring the consistency guard exactly. Gated the same way
sentinelPhases already is (empty unless phase_id_convention is
'bracket'), so a legacy repo's W006 reading — including its own
pre-existing wart where a legacy `### Phase 999:` still warns on both
verbs — is untouched; confirmed by the existing "INHERITED WART,
unchanged" test staying green.

Adds the paired-agreement regression test (#612 PR-2 B2 describe block
in tests/adr-612-bracket-read-tolerance.test.cjs): a sentinel-only
bracket roadmap must produce no missing-directory warning from EITHER
validator, plus a CONTROL proving a real phase with no directory still
warns on both. Confirmed red (health false-W006) against the pre-fix
code before applying the fix.

Corrects the .changeset/2761-bracket-read-tolerance.md fragment, which
described the asymmetry as already closed and in the wrong direction
(it credited validate health with already staying silent, when health
was the one falsely warning).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#2761): pin mixed-shape preamble cutoff and boundary heading levels

Closes two self-flagged coverage gaps in the B1 fix (commit 08d5b0c4)
ahead of adversarial review. No src change — all three new tests are
green against the code as committed.

1. earliest-of-either preambleCutoff comparison: only exercised where
   the version/emoji match and the bracket match happen to land on the
   same heading. Adds the mid-migration mixed shape (version-bearing
   PRIOR + version-less CURRENT) and asserts scoping outcomes (accepts
   booleans + total_phases), not internals.

2. `h.level <= 2` conjunct in computeSectionEnd: provably redundant
   whenever the selected milestone heading is level 2 (every existing
   fixture), since `h.level > level` alone already implies it there —
   a mutant deleting the conjunct would have survived every prior test
   in this file. Adds a level-3 CURRENT-heading fixture (with a real
   PRIOR milestone so the preamble side-channel can't independently
   rescue the truncated phases) that makes the conjunct's deletion
   test-visible, confirmed by hand-mutating a throwaway copy of the
   compiled output (never touching tracked src or the real build) and
   observing the assertion flip. Also pins a level-1 companion case
   (#{1,2} tolerance, not just level 2).

NOT included here: the other mixed-shape direction (version-less PRIOR
+ version-bearing CURRENT) turned out to be a genuine, currently-unfixed
gap — reported separately rather than silently patched or weakened, per
instruction not to touch src while a probe run is in flight.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): engage bracket boundaries when the current milestone heading is version-bearing (B3, self-caught)

Found during round-2 self-verification of B1 (commit 08d5b0c4), while
closing the mixed-heading-shape coverage gaps flagged in my own review
notes. The B1 fix resolved `bracketScopeConvention` only inside the
`if (headingMatches.length === 0)` gate that also drives SELECTION's
own bracket fallback (which heading counts as "current"). That gate is
correct for selection, but `bracketScopeConvention` also feeds
computeSectionEnd's and preambleCutoff's boundary detection further
down — which accidentally inherited selection's gate instead of having
its own.

Trigger shape: the CURRENT milestone heading is itself version-bearing
(`## [GSD.02] v2.0: Current Milestone`), so the primary version-string
match succeeds immediately — headingMatches.length !== 0 from the very
first check — and the entire bracket-resolution branch was skipped. A
sibling milestone (PRIOR or LATER) that is version-less then got
neither the version/emoji boundary rule (it has none) nor the bracket
boundary rule (never resolved), reproducing the original #612 defect
(total_phases falling back to the whole-disk count) through a
structural shape B1's own fixtures never exercised — every one of them
is uniformly version-bearing or uniformly version-less across all
three milestones, never mixed with CURRENT specifically being the
version-bearing one.

Fix: resolve `bracketScopeConvention` unconditionally, decoupled from
`headingMatches.length`. SELECTION is deliberately left untouched — the
`if (headingMatches.length === 0 && bracketScopeConvention === 'bracket')`
fallback that picks which heading is "current" keeps its original gate
byte-for-byte (confirmed by diff: that line is unmodified). Only the
convention *resolution* moved out from behind it, so boundary detection
can consult it regardless of which branch selected the heading. The
extra `resolvePhaseIdConvention` call this now costs on every
invocation (previously paid only when the version match found nothing)
is the accepted cost: a non-bracket repo still resolves to something
other than 'bracket' (or null on a poisoned env, caught exactly as
before), so `bracketMilestoneHeadingRe` stays null and every downstream
branch is byte-identical to today — confirmed by the full adr-612 +
roadmap-parser + state + verify + health-validation suite staying green
(1260/1260) and the all-version-bearing/legacy fixtures showing no
behavior change.

TDD: tests/adr-612-bracket-phase-counting.test.cjs describe block
"#612 PR-2 B3: bracket boundaries engage even when CURRENT is
version-bearing but a sibling is not" — 4 tests, confirmed red against
pre-fix code (leak-in booleans true/true, total_phases 4) before this
change, green after.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): reject same-milestone continuation headings as boundaries (B1)

Gate-2 adversarial review Blocker 1: the B1/B3 boundary fired on ANY
as the one currently selected — a version-less checklist/detail split
(`## [GSD.02] Foundation (Phase Details)`, or an ad-hoc continuation
heading) truncated the current milestone's own section instead of
being recognised as a continuation of it. The `(Phase Details)`
re-append only searches VERSION-STRING matches, so a version-less
continuation heading was cut out and never re-appended — a confidently
wrong, non-degraded phase count for a still-incomplete milestone
(repro8 case 1: 1/1/100 instead of 2/1/50; repro5: same, on a fully
version-less roadmap with no sibling milestones at all).

Introduces one shared helper, isBracketMilestoneBoundary(headingText,
level, selectedBracketId), used by both computeSectionEnd and the
preambleCutoff bracket scan, replacing the ungated `h.level <= 2 &&
bracketMilestoneHeadingRe.test(...)` inline check. `selectedBracketId`
(case-folded via phase-id.cts's foldBracketId, matching the branch's
own fold-before-identity convention) is derived from `selected[0]`,
which is the full matched heading line on BOTH selection paths
(version-string and bracket-fallback), so one extraction covers both.

Level cap stays at `level > 2` for now (temporary — ADR-612's content
discriminator replaces it in the next commit); same-milestone rejection
is the change this commit is scoped to.

DEVIATION from the reviewed plan, caught empirically: applying the
same-milestone rejection at the preambleCutoff site (as literally
specified) regressed an existing pin ("boundary heading level: a
level-1 CURRENT milestone heading also scopes correctly") and a
fenced-heading case (repro10 A3) — because preambleCutoff's job is
"where does the earliest milestone-shaped heading sit, scanning from
the TOP of the document," and the selected heading's own occurrence is
always a correct answer to that question regardless of same-id-ness;
rejecting it let the earliest-of-either comparison fall through to a
stray LATER heading instead. `selectedBracketId` is threaded through as
`null` at the preambleCutoff call site for this reason — bracket-shaped
(and, from the next commit, phase-tail) discrimination still applies
uniformly at both sites; only the same-milestone component is
call-site-specific, since it encodes a "keep scanning past this
heading" instruction with no counterpart in a top-of-document search.

Tests: new describe block "#612 PR-2 B1 round-2: a same-milestone
continuation heading is not a boundary" — RED-turned-GREEN fixtures for
repro8 case 1 and repro5, plus PINs for repro8 case 3 (trailing
different-id icebox still terminates) and repro10 A1 (all-version-
bearing + icebox + Phase Details stays exactly 2/1/50 — no double-count
from the same-milestone exclusion interacting with the pre-existing
detailsMatch re-append). syncedTotal()/syncedPercent() assertions
omitted from the repro10 A1 pin: that fixture carries dirs outside the
current milestone, which exposes the SEPARATE Major 1 defect
(cmdStateSync's body percent from an unfiltered disk scan) — asserted
once Major 1 is fixed, not here.

Full suite green (796/796 across the targeted adr-612 + roadmap-parser
+ state files); node scripts/lint-phase-id-drift.cjs clean; eslint
clean on both changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): bracket boundary discriminates by content, not heading level (B2)

Gate-2 adversarial review Blocker 2: three sites disagreed about which
heading levels are a bracket milestone. The selector
(roadmap-parser.cts's bracket-fallback SELECTION branch,
`^#{1,3}\s+\[CODE.MM\]`) and `isMilestoneBounded` (state.cts) both
admit level 1-3, but isBracketMilestoneBoundary's level cap only
admitted level 1-2 (`h.level <= 2`, from the B1 commit). A `###`-level
bracket milestone heading was therefore SELECTED and BOUNDED but never
TERMINATED: computeSectionEnd ran with level=3, a level-3 SIBLING
milestone survived the pre-existing `h.level > level` (not-deeper)
filter, failed the version/emoji test (version-less), then failed
`h.level <= 2` — falling through to `return content.length` and
sweeping the sibling milestone's own phases into the current one.
Reproduces trek-e's original #612 defect verbatim ("a safe degrade
became a confidently-wrong persisted number") on a heading level the
selector and bounding predicate both already admit (repro2 case C:
4/75% instead of 2/100%; mechanism confirmed directly via repro7 —
extractCurrentMilestone returned the whole 214-byte document).

ADR-612 Decision 1 (docs/adr/612-bracket-phase-id-convention.md:56)
specifies the discriminator as CONTENT, not level: "a phase heading is
a bracket followed by a digit-then-colon ([GSD.02] 05:); a milestone
heading is a bracket followed by a name." Replaces the `level > 2`
rejection with BRACKET_PHASE_TAIL_RE — built by interpolating
phase-id.cts's single-owner phaseHeadingPrefixSrcFor(ANY_BRACKET,
'bracket', false) plus the digit + optional-tag + colon tail every
phase-heading counter in this file already spells, not a re-typed
grammar — and widens the level check to a depth-sanity cap of 3
(mirroring the selector's own `#{1,3}` ceiling; NOT itself a
phase/milestone discriminator). Covers the dotted sub-phase heading
form (`[GSD.02] 05.03:`) via the same `[\w][\w.-]*` token, pinned by a
new fixture — the shape where a regex slip in the tail grammar would
hide.

preambleCutoff's own raw-scan regex is widened from `^(#{1,2})` to
`^(#{1,3})` in lockstep: the outer pattern's level ceiling must track
the helper's cap, or a level-3 PRIOR milestone heading is invisible to
that scan and its own phase heading leaks into the preamble
un-stripped (a real double-count this widening closes, verified
against repro2 case C directly).

The existing "boundary heading level: a level-3 CURRENT milestone
heading still scopes correctly" pin (3e562f12) now passes via a
DIFFERENT mechanism than before — its own neighbours are version-
bearing, so it previously passed via the version/emoji rule (the level
cap was never actually exercised by that fixture, per the round-2
review's own finding); with the content discriminator, the SAME
fixture's level-3 phase headings are now correctly excluded because
they are phase-tail-shaped, not because they are too deep. A
deliberate mechanism change, confirmed by re-running that test green
after this commit.

Also updates the "every selector call site declares the right
baseline" governance pin (adr-612-bracket-heading-selection.test.cjs):
BRACKET_PHASE_TAIL_RE is a new, legitimate ANY_BRACKET call site in
roadmap-parser.cts (always passing the literal 'bracket' convention,
since its only caller is already gated on bracketBoundaryActive) —
EXPECTED count bumped 1->2, with a matching BASE_SITES transcription
entry (identical src to every other ANY_BRACKET site, since the
function is pure).

Tests: new describe block "#612 PR-2 B2 round-2: the bracket boundary
is a CONTENT discriminator, not a level cap" — RED-turned-GREEN for
repro2 case C (exact total AND truthful percent, since
isMilestoneBounded already returns true at #{1,3}) and repro7's
mechanism, a PIN for the dotted sub-phase form, and a re-pin of repro8
case 3 (icebox) under the new mechanism.

Full suite green (907/907 across the targeted adr-612 + roadmap-parser
+ state + phase-id files); node scripts/lint-phase-id-drift.cjs clean;
eslint clean on all changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): fence-aware preamble cutoff on the bracket branch (Blocker 3)

Gate-2 adversarial review Blocker 3: preambleCutoff's bracket scan used
a raw content.match/matchAll — blind to fenced code blocks — while its
sibling computeSectionEnd (a few lines above it) already consumed
tokenizeHeadings(content), which strips fences. The two halves of one
boundary semantic disagreed about what a heading is.

A fenced markdown example in the preamble containing a bracket heading
(ADR-612's own docs do exactly this) was textually the earliest
`#{1,3} [CODE.MM]` match: preambleCutoff landed INSIDE the fence,
`preamble = content.slice(0, preambleCutoff)` ended with an unclosed
opener, and the unbalanced fence then blinded
getMilestonePhaseFilter's own tokenizeHeadings(scope) call — every
heading in the returned scope vanished, phaseCount degraded to 0, and
the pass-all filter admitted every directory on disk (repro11's
mechanism, confirmed directly: fence count 1/odd, tokenizeHeadings(scope)
-> only "Roadmap"). Regression vs round-1, which had no bracket pattern
to blind and so fell back to the correct heading (repro12 bracket row:
2/1/50 at round-1, 4/3/75 at HEAD).

Fixed by hoisting one tokenizeHeadings(content) call
(currentMilestoneHeadings) shared by computeSectionEnd and the
preambleCutoff scan, which now iterates that same fence-aware token
list instead of a raw regex. HeadingToken.text is already hash-stripped
and trimmed, so isBracketMilestoneBoundary needs no `^#{1,3}\s+`
re-derivation at this site (that spelling would not match h.text — a
note the round-2 review called out explicitly, confirmed while
porting). selectedBracketId stays `null` here, unchanged from the B1
commit's same-milestone-exclusion reasoning.

DISCLOSED, not fixed (explicitly out of scope per the round-2 review's
own minimal-fix note): the LEGACY (non-bracket) anyMilestonePattern
raw-match path shares the identical fence-blindness hazard and stays
byte-identical — a bracket repo whose preamble has a fenced
VERSION-BEARING heading still has the legacy raw-match win the
earliest-of-either min() (repro12's LEGACY control: 4/3/75, unchanged
across base/round-1/HEAD/this commit). Pinned here so a future reviewer
files this as a known, pre-existing gap rather than a new regression.

Tests: new describe block "#612 PR-2 Blocker 3 round-2: preambleCutoff
is fence-aware (bracket branch only)" — RED-turned-GREEN for repro12's
bracket row and repro11's mechanism (fence balance + non-degraded
phaseCount + correct per-directory admission), a PIN for repro12's
LEGACY control (the disclosed gap, explicitly unchanged), and a PIN for
repro10 A3 (a fenced heading INSIDE the current section must still not
terminate it).

Full suite green (1072/1072 across the targeted adr-612 + roadmap-
parser + state + phase-id + markdown-sectionizer files); node
scripts/lint-phase-id-drift.cjs clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): scope cmdStateSync's disk scan by milestone under bracket (Major 1)

Gate-2 adversarial review Major 1: `state sync` wrote a Progress
PERCENT computed from an UNFILTERED whole-disk scan, beside the
milestone-scoped total_phases/completed_phases it writes into the same
STATE.md via the refreshed frontmatter (syncStateFrontmatter ->
buildStateFrontmatter, which has always applied getMilestonePhaseFilter
for the READ path). cmdStateSync's own `fs.readdirSync` chain (the
WRITE-path scan) never called the milestone filter at all, unlike
buildStateFrontmatter's identical-purpose scan. One command therefore
wrote two contradictory numbers into one file: on the ADR-canonical
version-less bracket fixture (4 dirs, 3 complete; asserted milestone =
2 phases, both complete), base wrote total_phases:2/completed_phases:2
(correct, from the READ derivation) alongside body Progress 75% (wrong
— from the unfiltered WRITE derivation; repro3).

Fixed by threading `getMilestonePhaseFilter(cwd)` through the same
`.filter()` chain buildStateFrontmatter already applies, gated on
`syncConvention === 'bracket'` (falling back to a pass-all predicate
otherwise) — so totalDiskPlans/totalDiskSummaries/diskCompletedPhases/
syncTotalPhases become milestone-scoped under bracket, byte-identical
under legacy.

DEVIATION (approved, stated plainly): an earlier phrasing of this fix
called for mirroring buildStateFrontmatter's filter UNCONDITIONALLY.
Implemented GATED instead — an unconditional filter would ALSO move
every LEGACY repo's persisted percent, since the milestone-scoping-vs-
whole-disk divergence this closes is engine-wide, not bracket-specific.
The gate keeps legacy byte-identical, which is the binding constraint:
this is a bracket read-path PR, not a legacy behavior change.

Nit 2 (informational, no code change): 10 calls to
extractCurrentMilestone on a legacy repo cost 10 config.json
existsSync + 10 readFileSync (0 before B3); accepted, unmemoized cost,
unaffected by this commit.

Also folds in two minors from the round-2 review:
- Corrects .changeset/2761-bracket-read-tolerance.md: the sibling-
  exclusion sentence now states it holds at any heading level 1-3 and
  across a milestone split over two headings (true again now that
  Blockers 1 and 2 are fixed); the percent sentence states plainly that
  `state sync`'s body percent is now milestone-scoped under bracket,
  and unaffected under legacy.
- Records the read/write scoping divergence at currentMilestoneRawRanges
  (src/roadmap-parser.cts) in a comment: it did not receive B1/B2's
  bracket boundary fixes, currently harmless (its only consumer falls
  back to whole-content mutation, and every mutation there is still
  Phase-labelled-only, not bracket-widened), but live the moment the
  write path is bracket-widened — flagged so a future PR closes it in
  lockstep with that work, not after.

Tests: 6 pre-existing tests in tests/adr-612-bracket-phase-counting.test.cjs
needed fixture updates, not logic changes — they used the default
single directory (`GSD.02-01-setup`, phase "01"), which the SENTINEL/
retirement/mixed-heading fixtures in those tests never declare as a
real phase (only 04/05/06/999/etc are declared). Before this fix,
cmdStateSync's unfiltered scan counted that off-roadmap directory
anyway; after this fix the milestone filter correctly excludes it,
which for several of these fixtures made `state sync` a no-op (the
computed 0% coincided with STATE.md's initial template default) and
broke `syncedTotal()`/`syncedPercent()`'s ability to observe anything.
Updated each to pass an EXPLICIT directory naming one of the fixture's
REAL declared phases, preserving each test's original numerator/
denominator intent. One test — "shape 2 WRITE" — was substantively
rewritten: it was a CHARACTERIZATION of the Major 1 bug itself ("the
DISCLOSED legacy gap, mirrored — not closed"), and now correctly pins
bracket closing to 33% (agrees with the read path) while legacy stays
at the disclosed 60% (unchanged, deliberately, per the gating decision
above).

Full suite green: `npm test` 1449/1449 (0 fail, 0 skipped, 0 todo,
single-shard "all" run — includes issue-2765-brace-expansion-lockfile
passing); `npm run lint:ci` clean (0 errors; 2 pre-existing timing-
assertion warnings in files this PR does not touch); node
scripts/lint-phase-id-drift.cjs clean; node scripts/changeset/lint.cjs
ok.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): preambleCutoff identity is offset- and child-aware (round-3 Blocker 1)

Gate-2 round-3 re-verify Blocker 1 (NEW): the round-2 B1 deviation
(39c42a89) threaded `selectedBracketId` as the real value at
computeSectionEnd but as bare `null` at the preambleCutoff scan. The
deviation's rationale — "the selected heading's own occurrence is
always a correct earliest answer" — was right, but `null` disables the
same-milestone check for EVERY candidate, not just the selected one.
Any bracket-shaped heading earlier than the selected milestone was
accepted as a boundary regardless of identity: a same-id checklist/
overview heading preceding the version-bearing selected heading (cases
A, B — the version lands on the LATER half of a split, or a plain
overview heading with no "(Phase Details)" spelling), or a DIFFERENT-id
bracket-shaped PROSE heading with no children of its own sitting above
the current milestone's content (case D — `## [ADR.612] Heading
convention used by this roadmap`). In every case the region between
that false boundary and the real sectionStart was silently dropped —
a completed phase vanished and `state sync` persisted a confident 0%
where base and round-1 both correctly wrote 50%. Regression vs base
AND round-1 (not merely "under-fixed", per the round-3 review's own
severity note).

Fixed with two changes, both scoped to the preambleCutoff scan only
(computeSectionEnd already threads the real `selectedBracketId` and is
untouched):

(a) `h.offset === sectionStart` now bypasses BOTH the same-milestone
    check inside isBracketMilestoneBoundary (passing the REAL
    `selectedBracketId` for every other candidate) and the new child
    rule below — the selected heading's own position is definitionally
    the correct answer, so neither discriminator should run against it
    (rejecting it would mean rejecting the heading against ITSELF).
    Closes cases A and B — verified by the reviewer's own one-liner,
    reproduced here.

(b) New `bracketHeadingHasMatchingChild`: an otherwise-accepted
    candidate (bracket-shaped, not phase-tail-shaped, not the same id
    as the selected milestone) must ALSO have a next-strictly-deeper
    heading carrying its OWN bracket id to count as a boundary. This is
    what a genuine sibling milestone has (its own phase children share
    its bracket id — `## [GSD.01] Setup` / `### [GSD.01] 01: …`) and an
    unrelated bracket-shaped prose heading does not. A candidate with
    no such child at all (childless — e.g. an empty prior milestone, or
    one immediately followed by a same-or-shallower heading) degrades
    to NOT a boundary — over-inclusive, the safe direction: its own
    heading text stays in the preamble, contributing nothing to any
    phase count (not phase-shaped). Closes case D, which (a) alone does
    not — verified: without this rule, `[ADR.612]`'s prose heading is
    indistinguishable from a genuine prior sibling at this site.

As a side effect, also neutralizes Nit 2 (a colon-less `[GSD.02] 05`
heading spuriously terminating the preamble): a colon-less bracket
heading is not phase-tail-shaped so isBracketMilestoneBoundary alone
would accept it, but it is — precisely because it is malformed/
incomplete rather than a real milestone — childless, so the child rule
rejects it too. Pinned.

Known interaction with the fence-blind SELECTION path (disclosed by
the reviewer, not introduced here, tracked for the next commit): when
`sectionPattern` selects a FENCED version-bearing heading (an
extremely pathological shape — a fenced example whose text happens to
match STATE's asserted version), no token exists at `sectionStart`, so
the `h.offset === sectionStart` bypass never fires and the loop falls
through to the ordinary same-id / child-rule checks. This composes
with the round-3 Major 1 fix (next commit) rather than introducing a
new defect — SELECTION itself is untouched by any of this — but is
worth stating plainly rather than rediscovering.

Tests: new describe block "#612 PR-2 Blocker 1 round-3: preambleCutoff
identity is offset- and child-aware" — RED-turned-GREEN for cases A, B
(rv-attack1) and D (rv-attack1b) with syncedTotal()/syncedPercent()
assertions (the persisted 0% is the point), a PIN for a genuine prior
sibling with real children (still excluded), a PIN for a childless
prior sibling (degrades to not-cutting, over-inclusive/safe), a PIN
for the colon-less Nit 2 shape, and the reviewer's rv-mech1 mechanism
re-run as a proper test (scope now equals the full input document,
phaseCount 2, both dirs accepted).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 937/937 pass (930 baseline + 7 new). node
scripts/lint-phase-id-drift.cjs clean; eslint clean on both changed
files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): fence-aware version/emoji half of preambleCutoff on the bracket branch (round-3 Major 1)

Gate-2 round-3 re-verify Major 1 (NEW): ff6bf0a8 (round-2 Blocker 3)
made the BRACKET half of preambleCutoff's "earliest milestone-shaped
heading" search fence-aware, but left the VERSION/emoji half a raw
`content.match` even on the bracket branch. A fenced VERSION-BEARING
example heading in a bracket repo's preamble (ADR-612's own docs
illustrate the LEGACY heading shape exactly this way, inside a fenced
authoring-guide block) was still textually the earliest match for that
raw regex, winning the min() and un-suppressing a wrong persisted 75%
that base correctly suppressed (rv-attack3c fixture C1: base
suppressed the percent entirely — `isMilestoneBounded` false — HEAD
wrote 75% where truth is 50%).

Fixed by deriving the version/emoji half from the SAME fence-aware
`currentMilestoneHeadings` token list as the bracket half, on the
bracket branch only — the exact `/^Phase\s+\S/i` / `/v\d+\.\d+|✅|📋|🚧/i`
pair `computeSectionEnd` already uses against `h.text`. The non-bracket
(legacy) path is untouched: it keeps the raw `content.match`, byte-
identical to before, including its own fence-blindness (repro12's
LEGACY control, pinned unchanged in the round-2 Blocker-3 test block —
not re-pinned here to avoid duplicating an already-covered assertion).

Not rated Blocker (per the review) because it is not a regression vs
round-1 and the fixture (a version-BEARING fenced example in a bracket
repo) is rarer than the already-fixed bracket-heading case; still
fixed now rather than disclosed, per this arc's own precedent (every
prior "disclose instead of fix" call in this PR has been overturned on
re-review).

Tests: new describe block "#612 PR-2 Major 1 round-3: preambleCutoff's
version/emoji half is fence-aware on the bracket branch" — RED-turned-
GREEN for case C1 (readTotal + syncedPercent, so the persisted 75% is
directly observed, not just the read-path total), PIN for case C2 (the
already-fixed fenced-bracket-heading shape, unchanged).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 939/939 pass (937 + 2 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean on both changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): correct changeset claims + mark runtime-gated BASE_SITES row (round-3 minors)

Gate-2 round-3 re-verify Minor 2: three changeset sentences in
.changeset/2761-bracket-read-tolerance.md were overstated in a new
direction after round-2:

- The split-milestone claim ("across a milestone split over two
  headings") was true only when the version-bearing heading came
  FIRST (repro8 case 1); false when it came LATER (round-3 Blocker 1
  case A). Now restated to say plainly "with the version-bearing
  heading in EITHER position" — true again now that round-3's Blocker
  1 fix lands earlier in this range.
- The "counted from the phases... rather than from every directory on
  disk" claim was false on cases A/B/D (a strict subset of the
  milestone's own phases). Restated as "ALL of the phases... not a
  subset", and extended to state that an unrelated bracket-shaped
  heading with no phase children of its own (case D's `[ADR.612]`
  shape) does not truncate the milestone either — true now, not before.
- "Each widened read is SELECTED by the project's phase_id_convention"
  was literally false for BRACKET_PHASE_TAIL_RE, which is RUNTIME-gated
  (via its only caller, isBracketMilestoneBoundary, itself only
  consulted when bracketBoundaryActive) rather than selector-gated.
  Restated behaviourally: "every widened read ENGAGES only when the
  project's resolved phase_id_convention is bracket" — true for both
  gating mechanisms, so it no longer implies a selector call this site
  does not make.

Minor 1: the STRUCTURAL IDENTITY test's BASE_SITES row for
BRACKET_PHASE_TAIL_RE (added in the B2 commit) asserts a property of
`phaseHeadingPrefixSrcFor` — the function — not of the call site; it
would pass unchanged even if the site were deleted. Safety at that
specific site rests entirely on a runtime gate the test cannot see.
Added `runtimeGated: true` to the row and threaded it into the
generated test's own title (`… [runtime-gated, not selector-covered]`),
so the gating mechanism is visible in test OUTPUT, not only in a source
comment that could drift silently.

No production code changed. Targeted suite green (48/48 in the
affected file); `node scripts/changeset/lint.cjs` ok; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): harden round-3 preamble cutoff — subtree child scan + level cap

Team-lead review of f87bba0e found two edges in the round-3 preamble-
cutoff code, both hardened here.

AMENDMENT 1 — bracketHeadingHasMatchingChild (2e06aef5) checked only
`headings[index + 1]`, the IMMEDIATE next heading, not the candidate's
whole subtree. A genuine prior sibling milestone whose section opens
with a non-bracket subsection before its first phase heading
(`## [GSD.01] Setup` / `### Notes` / `### [GSD.01] 01: Old`) was
therefore wrongly rejected as a boundary — its real phase heading sits
TWO headings deep, not one — leaking its entire section into the
preamble unstripped.

CONFIRMED RED, not merely theoretical (built and ran the fixture
against f87bba0e before touching the fix, per instruction): scope
membership DOES drive the disk-side filter on this shape.
`GSD.01-01-old`'s directory was wrongly admitted into the CURRENT
milestone's filter via the leaked heading's qualified key
(`GSD.01-01`) — 3/2/67% where truth is 2/1/50%.

Fixed by scanning the candidate's full SUBTREE: continue past a
non-matching deeper heading instead of returning false on the first
one; only a same-or-shallower heading actually closes the subtree and
yields "no match found". A candidate whose entire subtree closes with
no same-id hit (including a genuinely childless one) still degrades to
`false` — over-inclusive, safe, unchanged from before.

AMENDMENT 2 — c483552a ported the version/emoji half of preambleCutoff
to the token-based scan with no level cap; the raw
`content.match(anyMilestonePattern)` it replaced was anchored
`^#{1,3}\s+`. A level-4+ version-bearing heading in the preamble
(`#### v2.0 notes`) therefore won the scan on the bracket branch where
the raw pattern — and the legacy path, unaffected — ignores it
outright. Fixed with `if (h.level > 3) continue;`, mirroring the
depth-sanity cap isBracketMilestoneBoundary already applies to the
bracket half of this same scan.

Tests: new describe block "#612 PR-2 round-3 hardening: subtree child
scan + level cap on preambleCutoff" —
- RED-turned-GREEN for the Notes-intervening fixture: exact 2/1/50 (was
  3/2/67), plus the disk-filter observable (`GSD.01-01-old` now
  correctly excluded).
- PIN for the level-4 preamble heading: the scope now PRESERVES the
  heading's text (was silently dropped before this fix — harmless in
  this minimal fixture's total_phases specifically, since the dropped
  text carries no phase-shaped content, but a real correctness gap
  against the raw pattern's own ceiling) — asserted via scope content,
  not total_phases, since that number is invariant here either way.
- PIN for the LEGACY control on the same level-4 shape — unchanged,
  confirming the raw content.match path is untouched.

Re-verified the existing genuine-prior-sibling and childless-sibling
pins (round-3 Blocker 1 commit) still pass under the subtree scan —
both fixtures' outcomes are unchanged since their same-id hit (or its
absence) was already at the first deeper heading.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 942/942 pass (939 + 3 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean on both changed files. Full `npm test` + `npm run
lint:ci` deferred to the team lead's own run per instruction.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): bracketHeadingHasMatchingChild requires a same-id PHASE child (round-4 Blocker 1)

Gate-2 round-4 re-verify Blocker 1 (NEW): the round-3 hardening's
subtree scan (fbfd0fca) proved SAME-ID-NESS but never asked whether
the matching child was PHASE-shaped. Case F1 re-opens round-3's case D
one heading later: `## [ADR.612] Heading convention` is followed by
its OWN sub-heading `### [ADR.612] Examples` — same bracket id as the
candidate, but MILESTONE-shaped (a name, no digit-then-colon), not a
phase. Same-id-ness alone satisfied the subtree scan and re-cut the
preamble at exactly the shape the round-3 hardening was written to
close.

Failing input: `## [ADR.612] Heading convention` / `### [ADR.612]
Examples` (prose) / `### [GSD.02] 01: One` (the current milestone's
own first phase, now unreachable) / `## [GSD.02] v2.0: Foundation` /
`### [GSD.02] 02: Two`. Truth 2/1/50. HEAD read 1/0/0, and `state sync`
reported "nothing to do" (exit 0, `{synced:true,changes:[]}`) because
its wrong 0% happened to equal the STATE.md seed — a half-done
milestone read as untouched with no write-path signal at all.

Fixed with the reviewer's one-conjunct addition: a same-id child only
counts if it is ALSO phase-tail-shaped (`BRACKET_PHASE_TAIL_RE`) — the
same single-owner discriminator `isBracketMilestoneBoundary` already
uses one level up for the identical distinction (phase vs milestone),
reused here rather than re-derived. This is exactly what the
changeset's own wording already claimed ("no phase children of its
own") — the code now matches the sentence rather than the other way
around.

Docstring updated at the function itself: the rule is "same-id PHASE
child", not "same-id child".

Tests: new describe block "#612 PR-2 round-4 Blocker 1: the same-id
child must be PHASE-shaped" — RED-turned-GREEN for F1 with
syncedTotal()/syncedPercent() (the persisted 0% — and the
report-nothing-to-do write-path silence — is the point), PINs for F11
(colon-less same-id child) and F11b (bullet-only phase list): both
correctly stay excluded either way, and the leak the phase-shape
requirement newly creates for these two shapes is INERT — a colon-less
heading forms no qualified key (getMilestonePhaseFilter's own
phase-heading pattern requires the colon too) and a bracket bullet
never matches the legacy-only BULLET_PHASE_LINE_PATTERN — confirmed
directly via getMilestonePhaseFilter, not merely inferred. Re-verified
the four existing child-rule pins (F2 subtree-closure, F3 deep-nested
same-id, F4 level-4 same-id, F5 childless-at-EOF) are unaffected by the
phase-shape requirement, since every one of them already used a
colon-bearing same-id child.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 946/946 pass (942 + 4 new). Zero drift measured across the full F1-F12
corpus except F1 itself (F9/F10/F12 remain red, deferred to the
separate Major 1 fix). node scripts/lint-phase-id-drift.cjs clean;
eslint clean on both changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): fence-aware phase counting, milestone bounding, and bracket-fallback selection (round-4 Major 1)

Gate-2 round-4 re-verify Major 1 (NEW): roadmapPhaseCount is a
fence-blind raw `.exec()` over the scope string, duplicated in TWO
independent copies (buildStateFrontmatter's read path, cmdStateSync's
write path). With the bracket alternative now compiled into it (#612),
a fenced EXAMPLE phase heading in the preamble inflates total_phases
and persists a wrong percent that base got right. Two further
fence-blind sites participate: isMilestoneBounded (a raw
`.test(roadmapRaw)`) and the bracket-fallback SELECTOR inside
extractCurrentMilestone (a raw `content.matchAll`, only reachable when
version-string selection finds nothing).

Failing inputs:
- F10 (clean isolate, version-bearing selection): a fenced
  `### [GSD.02] 05: Example phase` in the preamble inflates
  total_phases 2->3, persisting 33% where truth is 50%. LEGACY control
  on the same shape is correct on every build — not a pre-existing
  hazard being inherited, bracket-only.
- F9 (version-less selection): a fenced example carrying the project's
  OWN milestone id additionally confuses the bracket-fallback selector
  (the fenced heading gets SELECTED), compounding with the same
  fence-blind counter. base suppressed the percent; round-1 and HEAD
  both wrote 33%.
- F12 (isMilestoneBounded isolate): the ONLY `[GSD.02]` heading in the
  document is inside a fence, and the asserted milestone genuinely has
  no section at all — HEAD persisted 67% where base correctly
  suppressed the percent (the milestone is absent from the roadmap).

Fixed at the CONSUMER level, not the producer — extractCurrentMilestone's
returned scope string is deliberately UNCHANGED, since every other
consumer of that string needs its full content fidelity and legacy
identity forbids touching the shared string (this branch's own
precedent, ff6bf0a8/c483552a, was producer-level; here the ruling is
consumer-level because the string is shared far more broadly than the
two round-3 fixes' narrower producer edits):

(a) New `countRoadmapPhaseHeadings` (src/state.cts, immediately above
    extractRetiredPhaseNumbers) — ONE shared implementation for both
    call sites, replacing two independently-maintained copies. BRACKET
    convention counts via `tokenizeHeadings(scope)` at levels 2-4,
    testing each heading's hash-stripped text directly — fence-aware by
    construction, since tokenizeHeadings never produces a token for a
    fenced line. LEGACY convention keeps the exact pre-existing raw
    `.exec()` loop, byte-for-byte. A pre-existing, deliberately
    PRESERVED asymmetry between the two original call sites — the read
    path always excluded a bare `/^999\b/` token, the write path never
    did — is threaded through as an explicit
    `includeUnconditional999Check` parameter per call site, so sharing
    the implementation does not silently unify (and thereby move)
    either total.
(b) isMilestoneBounded's bracket branch now scans
    `tokenizeHeadings(roadmapRaw)` for a matching heading (level <= 3)
    instead of a raw regex test. Legacy version-string branch untouched.
(c) The bracket-fallback SELECTOR now builds its candidate set from
    `tokenizeHeadings(content)` instead of `content.matchAll`,
    reconstructing a match-shaped array so every downstream consumer of
    `headingMatches` sees the identical shape the raw-regex path always
    produced. This is the ONE site in this entire arc where SELECTION
    itself changes — selection SEMANTICS are otherwise unchanged (same
    pattern, same first-match-wins by document order); only the
    candidate set is now fence-aware. Pinned that unfenced selection is
    byte-identical.

Zero drift measured across the full historical corpus (repro2-13,
rv-attack1/1b/3c, rv-mech1, rv2-amend1/2, and F1-F11b) except the three
target fixtures.

Tests: new describe block "#612 PR-2 round-4 Major 1: four fence-blind
sites on the bracket path" — RED-turned-GREEN for F10 (with
syncedTotal()/syncedPercent()) and F9 (both layers), PINs for F10's
LEGACY control and F10c (non-phase-shaped fence, unaffected either
way), RED-turned-GREEN for F12 (asserts the percent KEY is absent from
`state json`'s output and that `state sync`'s body stays at its
unmodified seed — the persisted-suppression signal, not merely a
total_phases number), and an explicit PIN that unfenced bracket-fallback
selection (first real milestone-shaped heading wins, no fences
involved) is unaffected.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 952/952 pass (946 + 6 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean on all three changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): correct docstring overstatement + stale consumer-count sentence (round-4 minors)

Gate-2 round-4 re-verify Minor 1 + Nit 1. No production code changed.

Minor 1: bracketHeadingHasMatchingChild's own docstring said a
rejected (no-same-id-PHASE-child) candidate's degrade "contributes
nothing to any phase count" — true of the candidate's OWN heading
text, but not of its SUBTREE, which is what actually stays in the
preamble. F7 (`## [GSD.01] Setup` / `### [GSD.07] 01: Foreign`) shows
a DIFFERENT-id bracket PHASE heading inside a rejected candidate's
subtree DOES form a qualified key and CAN admit a foreign directory —
3/2/67%, stable across base, round-1 and HEAD (base via its own
pass-all degrade). Not a regression, still the declared over-inclusive
/ never-under-inclusive safe direction — the comment now says that,
with F7's numbers cited, at the call site that actually decides
`isBoundary` (roadmap-parser.cts's preambleCutoff loop) rather than
only at the helper's own definition.

Minor 2 (changeset) — VERIFIED, no wording change needed: re-ran F1,
F9, F10, F12 at this HEAD. The "no phase children of its own... does
not truncate it either" sentence (naming the `[ADR.612]` shape
directly) is now literally true — F1 reads 2/1/50. The "counted from
ALL of the phases... not a subset" sentence is now true on every
measured shape — F1/F9/F10 all read 2/1/50, F12 correctly suppresses
the percent. No carve-out for F9/F10 is needed since round-4 Major 1
(3be5c412) closes both; per the fix-round instruction to "only carve
out anything genuinely left," nothing is.

Nit 1: tests/adr-612-bracket-heading-selection.test.cjs's runtimeGated
row claimed `BRACKET_HEADING_INTRO_RE` has "no other consumers" — true
when round-3's f87bba0e wrote it, stale since 2e06aef5 (round-3's own
earlier commit) had already added two more uses inside
bracketHeadingHasMatchingChild. Corrected to state the true count
(three consumers) and re-confirm the conclusion is unaffected: all
three are still nested inside the same bracketBoundaryActive runtime
gate, and BRACKET_HEADING_INTRO_RE is built from BRACKET_ID_SRC, not
phaseHeadingPrefixSrcFor, so it was never a selector site regardless of
consumer count.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 952/952 pass (comment-only changes, no count movement). eslint
clean; node scripts/changeset/lint.cjs ok.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): restore the bracketId guard in countRoadmapPhaseHeadings (round-5 Blocker 1)

Gate-2 round-5 re-verify Blocker 1 (NEW, introduced by 3be5c412): the
merge that created the shared countRoadmapPhaseHeadings helper dropped
the `bracketId && ` guard both original inline loops carried before
calling isSentinelPhaseId. Every other isSentinelPhaseId call site in
src/ (roadmap-parser.cts, roadmap.cts, validate.cts x2, verify.cts)
keeps the guard; state.cts's shared counter was the only one of seven
without it.

When the phase-heading-intro grammar's LEGACY alternative matches (a
`### Phase 00:` heading in a `phase_id_convention: "bracket"` repo —
the mid-migration shape this PR exists for), the bracket capture group
is `undefined`, so the unguarded call became
`isSentinelPhaseId("undefined-00", 'bracket')` — measured TRUE, so
phase 00 (and 000, 0a, 0.5, 999.1 — any token whose splice with the
literal string "undefined" happens to fall in a sentinel range) was
silently dropped from the denominator. `getMilestonePhaseFilter` (which
still carries its own guard) counts the phase and admits its directory
regardless, so the filter and the counter disagree — a half-done
milestone reads as 100% complete, persisted.

One-line fix, restoring the guard every sibling call site already has:

    if (bracketId && isSentinelPhaseId(`${bracketId}-${token}`, 'bracket')) continue;

Line count: `git diff --stat src/state.cts` -> 1 file changed, 1
insertion(+), 1 deletion(-).

Tests: new describe block "#612 PR-2 round-5 Blocker 1:
countRoadmapPhaseHeadings restores the bracketId guard" — RED-turned-
GREEN for G3 (3/2/67, was 2/2/100) and G3d (the mixed bracket+legacy
mid-migration shape, same numbers) with syncedTotal()/syncedPercent(),
PIN for G3's legacy control (unaffected), PIN for G3b (isolates the
counter with no directory to admit), PIN for G3c (legacy 01/02 only,
no sentinel-shaped token present).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 957/957 pass (952 + 5 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): extractRetiredPhaseNumbers is fence-aware on the bracket path (round-5 Major 1)

Gate-2 round-5 re-verify Major 1 (NEW) — a FIFTH fence-blind site on
the bracket path, missed by 3be5c412's own enumeration of "four".
extractRetiredPhaseNumbers' line scan (`scope.split(/\r?\n/)`) has no
fence awareness. This PR compiles the bracket alternative into
`introSrc` ("the retirement filter has to widen with the counter it
protects" — the function's own pre-existing comment), so a FENCED
authoring EXAMPLE showing the #1514 retirement gesture in bracket
spelling is now indistinguishable from a real one: it retires a
genuine phase, shrinking the denominator and persisting a confident
100% where base correctly read 50%.

Fixed with the SAME consumer-level ruling this arc has used at every
other fence-blind site, reusing markdown-sectionizer's existing
exported `stripFencedCode` rather than hand-rolling a second fence
parser (single-owner rule) — retirement lines are BULLETS, not
headings, so `tokenizeHeadings` doesn't serve here; `stripFencedCode`
is the general-purpose fence stripper the tokenizer itself is built on.
Gated on `convention === 'bracket'`; the LEGACY line scan stays the raw
`scope` string, byte-identical — its own fenced-example hazard is
pre-existing (wrong at base too) and out of scope.

Line count: `git diff --stat src/state.cts` -> 1 file changed, 9
insertions(+), 2 deletions(-) — one import added, four lines inside
the function (a comment + the `scanScope` computation + the changed
`.split()` call).

Tests: new describe block "#612 PR-2 round-5 Major 1:
extractRetiredPhaseNumbers is fence-aware on the bracket path" —
RED-turned-GREEN for G2 (fenced example in the preamble) and G2b (the
same example placed INSIDE the milestone section, ruling out a
preamble-scoping artifact — the site itself was fence-blind wherever
the fence sits), PIN for the LEGACY control (unchanged, pre-existing,
out of scope — base is wrong on this shape too).

Also folds in the "four fence-blind sites" correction: 3be5c412's
commit message and any restatement of it should read FIVE going
forward; this commit's own message states the count correctly.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 960/960 pass (957 + 3 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): state sync's counter excludes the bracket 999 icebox token (round-5 Major 2)

Gate-2 round-5 re-verify Major 2 (NEW) — `includeUnconditional999Check`
left `state json` and `state sync` reporting different totals for one
bracket repo. Under bracket, READING-B puts the sentinel in the
bracket, so `isSentinelPhaseId("GSD.02-999", 'bracket')` is false —
the `/^999\b/` TOKEN rule is the only thing excluding a
`### [GSD.02] 999:` icebox heading, and it ran on the read path
(buildStateFrontmatter, `true`) and on getMilestonePhaseFilter
(unconditional), but not on cmdStateSync's own counter (`false`). One
`state sync` call could leave a single STATE.md with its own
frontmatter (percent 50, from the read-path re-sync inside
writeStateMd) and body (percent 33, from the write-path counter that
alone still counted the icebox heading) disagreeing — falsifying this
PR's own stated invariant that sharing countRoadmapPhaseHeadings made
"the two counters must see the same phases" structural.

Functional change is one argument, exactly as specified: the write
call site now passes `syncConvention === 'bracket'` instead of the
literal `false`. `syncConvention === 'bracket'` is `false` for every
non-bracket value, so the LEGACY path resolves to the exact same
`false` it always did — this file's own pre-existing, deliberately-
unchanged read/write divergence on that path is untouched. The READ
site (`:1860`) is NOT touched — its historical behaviour applied
`/^999\b/` to legacy and bracket alike, so changing it would move
legacy READ totals, exactly the class of mistake this arc's own
Blocker 1 (this round) was.

Line count: the functional change is ONE argument
(`false` -> `syncConvention === 'bracket'`); the surrounding comment
was rewritten because the previous one asserted the now-superseded
behaviour ("preserving this file's pre-existing... divergence... a
bare 999 token is not excluded here") and leaving it would mislead the
next reader — not a structural change.

Tests: new describe block "#612 PR-2 round-5 Major 2: state sync
excludes the bracket 999 icebox token like the read path" —
RED-turned-GREEN for G1, asserting `state sync`'s body percent equals
`state json`'s own percent (both 50, not 33 vs 50), PIN for G1's
legacy control (33 vs 50 unchanged, deliberately).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 962/962 pass (960 + 2 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): restore indent parity at the two line-start-anchored bracket-heading reconstructions (round-5 Minor 1)

`HeadingToken.offset` (tokenizeHeadings) is the LINE-START character offset,
not necessarily the `#` character's own offset — a ≤3-space-indented ATX
heading has both. Two round-4 reconstructions built on tokenizeHeadings
inherited this gap and accepted indented headings their raw, line-start-
anchored predecessors (`^#{1,3}\s+\[...`) never matched:

  1. roadmap-parser.cts's bracket-fallback SELECTOR (extractCurrentMilestone,
     ~line 377) — an indented, version-less `[GSD.02]` milestone heading
     could be reconstructed into headingMatches, then mis-parsed downstream
     (selectedBracketId null, level fallback to 1), leaking a SIBLING
     milestone's phases into the counted scope (G6: HEAD read 3/2/67 instead
     of 2/1/50 — a real phase heading's own directory belonging to the NEXT
     milestone got counted).

  2. state.cts's isMilestoneBounded (~line 1571) — the same gap let an
     indented-only `[GSD.02]`-shaped heading wrongly bound a milestone
     absent from the roadmap, un-suppressing a percent that should stay
     suppressed (mirrors round-4's F12 fenced-only case, but via indentation
     instead of a fence).

Fix: one added conjunct per site — `content[h.offset] === '#'` (source-named
`roadmapRaw` in state.cts) — filtering to tokens whose LINE-START offset IS
the `#` character, i.e. exactly the set the raw line-start-anchored regex
would ever have matched. Restores byte-for-byte raw parity; no other logic
in either function changes. computeSectionEnd and the preamble version/
emoji-token scan are untouched, as instructed — they consumed tokenizeHeadings
output before this arc and are out of scope here.

Also corrects roadmap-parser.cts's now-provably-false docstring claim that
`h.offset` is unconditionally "the same `#`-character coordinate space
`content.match().index` used" — true only for the survivors of the new
filter, not for every token tokenizeHeadings produces.

Line count: the FUNCTIONAL change is exactly 2 lines (one added `&&` conjunct
per call site — `git diff --stat` on the two source files shows 20
insertions/7 deletions, but only those 2 lines change behavior; the rest is
docstring/comment rationale, per this round's "state the line count" ask).

TDD: both fixtures verified RED at HEAD before this commit, GREEN after,
via /tmp/pr612rev/rv5-attack.cjs G6 and a locally-authored isMilestoneBounded-
isolating probe (G6 alone doesn't distinguish the two sites — its unindented
phase headings already satisfy isMilestoneBounded's loose prefix regex
either way, so a second, indentation-only fixture was needed to prove that
site's fix is not a no-op; verified by temporarily reverting just that one
conjunct, confirming 100%-wrongly-bounded RED, then restoring it, confirming
suppressed-percent GREEN).

New tests (adr-612-bracket-phase-counting.test.cjs):
  - RED (G6): indented version-less bracket milestone heading — 2/1/50, not
    the pinned-before-fix 3/2/67.
  - PIN (G6c unindented control): identical document, no indent — 2/1/50
    unaffected on every build.
  - RED (isMilestoneBounded site, indented-ONLY): mirrors round-4's F12
    shape (fenced-ONLY → indented-ONLY) — percent stays suppressed instead
    of the pinned-before-fix wrongly-bounded 100%.

Full G1-G10 (rv5-attack.cjs) + G2b/G3c/G3d/G6c (rv5b.cjs) re-verified
zero-drift against TRUTH after this change. Targeted suite (adr-612-*,
roadmap-parser, state, verify): 962 -> 965 (+3), 0 fail.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): changeset discloses counting-set narrowing + correct stale consumer-count sentence (round-5 minors)

FIX 5 (Minor 2): .changeset/2761-bracket-read-tolerance.md did not disclose
that bracket phase-heading counting is narrower than the raw-regex
predecessor in two ways the review's G4/G5 fixtures surfaced: a level-5
heading (`##### [GSD.02] 05: ...`) is no longer counted (the counter's
tokenizeHeadings scan caps at level 4, matching the selector/isMilestoneBounded
ceiling), and a space-less heading (`###[GSD.02] 05: ...`) is no longer
counted (CommonMark requires ≥1 space/tab after the hashes, which
tokenizeHeadings correctly enforces and the old raw regex did not). Both
G4 and G5 moved from round-1's wrong (inflated) values back to base's
original values as an incidental side effect of routing through
tokenizeHeadings — never a deliberate feature of this PR, and previously
undocumented. One clause added to the existing run-on paragraph; no other
wording in the changeset touched.

FIX 6 (Nit 1): tests/adr-612-bracket-heading-selection.test.cjs:76 —
`BRACKET_PHASE_TAIL_RE` has TWO consumers as of round-4's 65d257ce
(isBracketMilestoneBoundary's own use, plus bracketHeadingHasMatchingChild's
same-id-PHASE-child conjunct), not the "no other consumers" the comment
claimed. Same correction pattern round-4 already applied to this row's
BRACKET_HEADING_INTRO_RE neighbor (that sentence's own staleness was fixed
in 4d7184b8): note the true consumer count, confirm both stay nested inside
the same bracketBoundaryActive runtime gate (verified at
src/roadmap-parser.cts:178 and :250, both reached only through the
`if (bracketBoundaryActive)` block starting at :529), and record which
commit and which fix introduced the drift. Comment-only; no assertion
logic changed.

Line count: 2 files, 8 insertions / 2 deletions total — one added clause
in the changeset (1 line changed) and one comment block replacing the
single stale line in the test file (6 comment lines replacing 1).

Verified: targeted suite (adr-612-*, roadmap-parser, state, verify)
unchanged at 965/965 pass (comment/prose-only diffs, no test count change).
`node scripts/changeset/lint.cjs` and `npx eslint` on the touched files both
clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): restore the bracketId guard on the bracket-only /^0\b/ sibling rule (round-6 Blocker 1)

Round-5's Blocker 1 was `isSentinelPhaseId("undefined-00")` — the merge
that introduced the shared counter dropped the `bracketId &&` guard on the
sentinel check. This is the identical failure one line further down, in the
sibling rule this PR itself added: when the phase-heading grammar's LEGACY
alternative matches (`### Phase 0:` in a `phase_id_convention: "bracket"`
repo — the mid-migration shape this PR exists for), `bracketId` is
`undefined`, and the unguarded `/^0\b/` fires on the bare token anyway.
Neither the LEGACY branch of this same function nor
`getMilestonePhaseFilter` has a `/^0\b/` rule at all, so the filter counts
the phase and admits its completed directory while the counter refuses to
count its heading — a milestone with an unstarted phase 02 persists as a
confident 100%.

`/^0\b/` matches `0` and `0.5` (word boundary before the `.`) but not `00`
(no boundary between the two zeros), which is exactly why round-5's G3/G3d
fixtures (`### Phase 00:`) never tripped this one — same defect class,
different token spelling.

Fix: one word, mirroring the guard round 5 restored two lines above —
`if (/^0\b/.test(token)) continue;` -> `if (bracketId && /^0\b/.test(token)) continue;`.

Expected and intentional side effect: `roadmap analyze` and `state json`
now disagree again on this bracket-repo shape (analyze phase_count=2, json
total_phases=3) — exactly as they already do under the legacy convention
today (verified via /tmp/pr612rev/rv6c.cjs on both conventions). That is
the counter regaining agreement with `getMilestonePhaseFilter` (the tighter
constraint — it is what actually decides `completed_phases`), not a new
break; the counter/filter disagreement is what was wrong.

Out of scope, deliberately NOT fixed here (Minor 1, disclosed via a PIN
test only): the bracket-SPELLED `### [GSD.02] 0:` shape has the same
counter/filter disagreement, but reads 2/2/100 on base too — never closed
by any build in this arc, so it is a pre-existing gap rather than a
regression this commit could introduce.

Line count: src/state.cts is exactly 1 insertion / 1 deletion (one word,
`bracketId && ` prepended to the existing condition).

TDD: T0 (`Phase 0:`) and T05 (`Phase 0.5:`) verified RED at HEAD before this
commit (json 2/2/100, sync body 100%) via /tmp/pr612rev/rv6b.cjs, GREEN
after (3/2/67 on both derivations, matching TRUTH). Zero-drift verified by
diffing the FULL corpus (rv5-attack, rv5b, rv4-attack, rv-attack1,
rv-attack1b, rv-attack3c, rv-mech1, rv2-amend1, rv2-amend2, rv6-attack
H1-H19, rv6b) between a pre-fix and post-fix build: the only differing
lines in the entire diff are T0, T05, and H14 — the three target fixtures.

New tests (adr-612-bracket-phase-counting.test.cjs):
  - RED (T0) + PIN (T0L legacy control)
  - RED (T05) + PIN (T05L legacy control)
  - PIN (B0, bracket-spelled `0` — base-parity characterization, Minor 1,
    not fixed this round)
  Round-5's own G3 test (`### Phase 00:`) already serves as the T00 pin —
  `/^0\b/` never matched `00`, so it is unaffected and untouched.

Targeted suite (adr-612-*, roadmap-parser, state, verify): 965 -> 970 (+5),
0 fail. `lint-phase-id-drift.cjs` and `npx eslint` on touched files clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): changeset re-attributes the phase-count level cap + discloses the bare-0 carve-out; correct two stale test comments (round-6 Minor 1 + Nits)

FIX 2 (Minor 1 — disclosure only, NOT a code fix): the bracket-SPELLED
`### [GSD.02] 0:` / `0.5:` shape has the same counter/filter disagreement
round-6's Blocker fixed for the legacy spelling, but it reads 2/2/100 on
BASE too — never closed by any build in this arc, so it is a pre-existing
gap rather than a regression this round could introduce. Two changes:

  - Pinned as a base-parity characterization: new PIN test (B0) in
    04b5a95f's describe block asserting the exact unchanged value with a
    comment stating why it's deliberately not touched.
  - .changeset/2761-bracket-read-tolerance.md: one qualifying clause added
    to the "counted from ALL of the phases … not a subset" sentence — a
    bare `0`/`0.x` phase token (however spelled) keeps `roadmap analyze`'s
    own pre-existing sentinel reading under bracket too, so it stays
    excluded from these counts. Carried over, not newly introduced by this
    PR — `roadmap analyze` has always read it this way.

FIX 3(a) — same changeset paragraph misattributed the phase counter's
2-4 level cap to "the milestone-boundary machinery in this PR" (the
selector, `isMilestoneBounded`, the preamble scan) — that machinery caps
at `###` (level ≤3), not 2-4. The 2-4 cap belongs to
`getMilestonePhaseFilter`'s own phase scan. Re-pointed the attribution;
the paragraph's earlier, correct ≤3 claim (selector/isMilestoneBounded)
is untouched.

FIX 3(b) — tests/adr-612-bracket-heading-selection.test.cjs:76-82 (added by
1395bd89, round-5's own Nit-1 correction) claimed both
`BRACKET_PHASE_TAIL_RE` consumers are reached "only through the
`if (bracketBoundaryActive)` block starting at :529". Verified at HEAD:
`bracketHeadingHasMatchingChild` is (only caller :593, inside that block).
`isBracketMilestoneBoundary` is not — it has two callers, an inline
`bracketBoundaryActive &&` conjunct at `:473` (inside `computeSectionEnd`)
and the `:529` block at `:567` — the very fact this row's own earlier
"exactly two callers" paragraph already stated correctly. Corrected the
mechanism claim; the CONCLUSION (every consumer is still gated on the same
flag) is unchanged, exactly as round-5's own Nit-1 fix left round-4's
conclusion unchanged when it corrected the consumer count.

FIX 3(c) — tests/adr-612-bracket-phase-counting.test.cjs:2311,2340 still
said "four fence-blind sites" after round-5 (357ba671) added a fifth
(the retirement scan) in its own block below. Reworded both the section
comment and the `describe()` label to read as historical scoping of
round-4's own fix ("the four sites known at round 4 … a fifth was found at
round 5, see its own block below") rather than a live exhaustive claim.

Line count: 3 files, changeset 1/1, heading-selection test 17/8,
phase-counting test 6/2 (comment/prose-only; B0's own PIN test landed in
04b5a95f alongside T0/T05 since it was investigated as part of that same
Blocker's defect class, not in this commit — noted here for the record).

Verified: targeted suite (adr-612-*, roadmap-parser, state, verify)
unchanged at 970/970 (comment/prose-only diffs, no test count change).
`npx eslint` and `node scripts/changeset/lint.cjs` both clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): W021's bracket remediation hint stops pointing at a command that hard-errors (round-7 Minor 2)

`checkBracketCoherence`'s W021 (added by 94abf5df) attached the fix hint
`Run \`gsd-tools roadmap upgrade --convention bracket\` to migrate
(dry-run by default)` to every bracket-convention W021 — both sub-checks
(`missing-bracket` and `mismatch`) share the single `addIssue` call at
src/verify.cts:2255-2256. But `roadmap-command-router.cts:204` throws
unconditionally for any `--convention` value other than
`milestone-prefixed`:

  $ gsd-tools roadmap upgrade --convention bracket
  Error: Only --convention milestone-prefixed is supported

This contradicts the PR's own two disclosures: the changeset ("`bracket`
is a READ-path opt-in until the migrator and write path land") and
docs/CONFIGURATION.md:188 ("There is no bracket migrator and no bracket
emit yet"). Reachability is the exact mid-migration repo this PR targets —
any bracket project with one un-migrated heading gets an unfollowable
instruction on every `validate health`.

Fix: one string. Replaced the hint with what a user can actually do today —
manually align the heading's bracket milestone to its section — and named
the tracked future landing (#612 PR-3) instead of a command that errors.
The milestone-prefixed sibling hint at src/verify.cts:2240 (a different,
already-functional convention/command pair — verified against
roadmap-command-router.cts:65) is untouched.

Line count: src/verify.cts is exactly 1 insertion / 1 deletion (one string
literal).

No test in the suite previously asserted this string's content
(`grep -rn "upgrade --convention bracket" tests/` was empty), so the
unfollowable hint shipped unpinned — `lint-fix-has-regression-test.cjs`
would not have caught a string-only change without a new test. Added one:
asserts the fix string both (a) does not match `--convention bracket`
(the specific pinned-before-this-fix hazard) and (b) equals the new string
exactly, covering the invariant a future edit must not re-break: no
unsupported `--convention` value named in remediation text users are
expected to run verbatim.

Verified end-to-end (not just unit-level) via /tmp/pr612rev/rv7f.cjs: the
W021 issue's `fix` field now reads the new string; `gsd-tools roadmap
upgrade --convention bracket` (and its --dry-run variant) still correctly
hard-error — that command remains unsupported, only the hint text changed.

Targeted suite (adr-612-*, roadmap-parser, state, verify): 970 -> 971 (+1),
0 fail. `lint-phase-id-drift.cjs` clean; `npx eslint` on touched files:
0 errors (1 pre-existing unrelated no-elapsed-assertion warning, same file,
same line this arc's prior rounds already disclosed).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): correct the bare-0 changeset disclosure + disclose W021's second sub-check (round-7 Minor 1 + Nit 1)

FIX 1 (Minor 1) — round-6's bare-0 changeset clause (1395bd89) was wrong in
three falsifiable ways:

  (a) Its own illustration (`### Phase 0:`) is exactly the LEGACY spelling
      04b5a95f made COUNTED. The residual exclusion after that fix applies
      only to the BRACKET-spelled token (`### [GSD.02] 0:`) — the clause
      named the wrong shape as its example.
  (b) "excluded from these counts" over-scoped the carve-out. The phase-0
      directory is admitted into `completed_phases`/`total_plans` in BOTH
      spellings (measured: B0 at HEAD has `completed_phases=2`,
      `accepts {"GSD.02-0-bootstrap":true}`). The exclusion that survives
      lives in `total_phases`/`phase_count` (heading counting) only.
  (c) The pre-existing "so `roadmap analyze` and `state json` report the
      same number" sentence (present since before this arc's bracket work)
      is now false for the bare-0 LEGACY-spelled shape: 04b5a95f's own
      commit message discloses this exact re-divergence as expected —
      `roadmap analyze` phase_count=2 vs `state json` total_phases=3 on T0
      — matching the disagreement legacy already carries today. The
      changeset still asserted unconditional agreement.

Rewrote both sentences: dropped the `### Phase 0:` example, scoped the
carve-out to `total_phases`/`phase_count`, named the bracket-spelled token
as the one that keeps analyze's sentinel reading, stated plainly that the
legacy-spelled form in a bracket repo IS counted (the mid-migration guard),
and qualified the "report the same number" claim to the `999` token, with
the bare-0 legacy-spelled shape named as the one exception and why.

FIX 3 (Nit 1) — `checkBracketCoherence` has always had two sub-checks
(its own docstring: "Two sub-checks, both surfaced as W021") but both the
changeset and docs/CONFIGURATION.md:188 described only the `mismatch`
sub-check. The `missing-bracket` sub-check — which fires on every
legacy-spelled heading in a bracket repo, the noisier of the two on a
mid-migration project — was undisclosed in both places. One clause added
to each.

Line count: 2 files, 1 line changed each (both are single-paragraph/
single-row files; git diff --stat reports 1/1 per file though several
distinct clauses were edited within that one line each).

Verified: `node scripts/changeset/lint.cjs` clean. Targeted suite
(adr-612-*, roadmap-parser, state, verify) unchanged at 971/971
(prose-only diffs, no test count change, no code touched).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): resolvePhaseIdConvention's docstring no longer claims loadConfig drops the key

Upstream #2997 (aa7697fe, landed in `next` during this PR's final
verification) added `phase_id_convention: get('phase_id_convention') ?? null`
to `_baseConfig` in config-loader.cts — `loadConfig` now surfaces
`phase_id_convention` in its resolved config. This function's docstring
gave "loadConfig merges against CONFIG_DEFAULTS and drops keys it does not
know, and `phase_id_convention` is not among them" as the reason for
reading config.json directly instead of calling `loadConfig(cwd)`. That
rationale is now stale against live `next`.

Corrected the comment to state the two reasons that actually survive #2997:

  1. The workstream->root federation this function performs is a standalone
     resolution run against a GIVEN cwd, not necessarily the same base a
     `loadConfig(cwd)` call elsewhere in the codebase would federate from.
  2. Convention-ENUM validation is still #612 PR-4 work — this function
     returns the raw string unvalidated, exactly as the now-surfaced
     resolved key would.

Noted that #2997 surfacing the key makes consuming it from resolved config
(instead of re-reading config.json here) a natural PR-4 consolidation —
not this PR's scope. The "cycles were never the obstacle" close and every
other paragraph in the docstring (federation rationale, SCOPE note) are
untouched; they still hold.

Comment-only — no code behavior changed. This worktree's own history does
not contain aa7697fe (git merge-base --is-ancestor confirms neither branch
is an ancestor of the other; not rebasing per instruction), so this is a
textual correction against a documented external fact, not a functional
sync with upstream.

git diff --stat:
 src/planning-workspace.cts | 17 ++++++++++++-----
 1 file changed, 12 insertions(+), 5 deletions(-)

Verified: `npm run build:lib` clean, `node scripts/lint-phase-id-drift.cjs`
clean, `npm run lint:ci` exit 0 (same 2 pre-existing unrelated eslint
warnings as every prior round this arc, 0 errors; lint-fix-has-regression-test
PASS). Full suite not re-run per instruction (comment-only diff).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): resolve phase_id_convention against the caller's workstream

`resolvePhaseIdConvention` took no workstream, so it resolved from
`planningDir(cwd)` — which falls back to `GSD_WORKSTREAM` only when its `ws`
argument is `undefined`. Every caller that passes a workstream by ARGUMENT
(they cannot set the env var per iteration) therefore read the convention from
the ROOT config while reading that workstream's ROADMAP.

Two reproduced consequences:

- A workstream that explicitly declares its own `phase_id_convention` had it
  ignored. Flipping ONLY the root config between bracket and milestone-prefixed
  changed which milestone that workstream extracted.
- `--workstream foo` and `GSD_WORKSTREAM=foo` disagreed on the same repo: the
  arg form fell through to the root config, the env form did not.

`resolvePhaseIdConvention(cwd, ws?)` now forwards `ws` to `planningDir`, and
both roadmap-parser call sites pass theirs — `extractCurrentMilestoneScoped`
(which already reads STATE from `planningDir(cwd, ws)`) and
`getMilestonePhaseFilter`'s lazy branch. `undefined` keeps the env fallback, so
convention-less call sites are byte-identical; `null` still means "explicitly no
workstream". The `undefined` vs `null` discriminator on `phaseIdConvention` is
untouched — only the base the `undefined` branch resolves FROM moves.

workstream-inventory's two sites (`countRoadmapPhases`, `inspectWorkstream`)
passed a literal `null` where they meant `undefined`, pinning every workstream
to the legacy grammar. On a bracket workstream whose milestone declares 3
phases that returned phaseCount 0 and fell back to the on-disk directory count;
it now returns 3. A non-bracket workstream is unchanged (legacy control pinned).

The workstream -> root federation is preserved: a workstream that declares no
convention still inherits the root, as config-loader does. That inheritance is
pinned so the fix cannot be over-applied into isolation.

* fix(#2761): classify missing phase details per occurrence, not per token

Under READING-B a phase's sentinel status lives in the BRACKET, so
`[GSD.999] 01` and `[GSD.02] 01` share a token and are not the same phase.
Both sides of the missing-detail check were keyed by the bare token anyway, and
both produced false negatives in `missing_phase_details`:

- The checklist scan built a token -> bracket-id Map, FIRST-WINS. Of two entries
  sharing a token, whichever the author wrote first classified both. With
  `- [ ] **[GSD.999] 01: Icebox**` above `- [ ] **[GSD.02] 01: ...**` the real
  phase inherited the icebox's sentinel verdict and vanished from the report;
  swapping the two bullets — same document, same phases — reported it. Pinned
  with a test asserting BOTH orders.

- The detail set was `new Set(phases.map(p => p.number))`, also token-keyed, so
  `[GSD.02] 01`'s heading marked token `01` present and satisfied
  `[GSD.03] 01`, which has no heading anywhere. Order-independent, same class.

Both sides now key on an occurrence key: the owner's `bracketQualifiedKey`
(fold- and padding-insensitive, so `[gsd.2] 01` and `[GSD.02] 01` are one
phase), falling back to a fold-normalized composite for the two shapes that
grammar refuses — a token carrying its own hyphen, which splices to an id whose
trailing segment the qualified-key grammar truncates (the hazard
`getMilestonePhaseFilter` guards with its own `!token.includes('-')`), and any
id it does not accept. With no bracket id the key IS the bare token, so the
legacy path keeps its exact keys and dedupe order.

The emitted value is unchanged — `missing_phase_details` stays an array of bare
tokens, matching `phases[].number`. Only the classification moved to the
qualified key, so two brackets' `01` both missing report `01` once instead of
one silently covering for the other.

* test(#2761): replace the two wall-clock ReDoS assertions with algorithmic bounds

Both guards asserted elapsed wall-clock time — `Date.now()` against a 20s
ceiling in the coherence suite, `process.hrtime.bigint()` against 1s in the
read-tolerance suite. Those measure the host machine rather than the SUT and
flake on a loaded CI runner (RULESET.TESTS.no-timing-assertion). Per
RULESET.TESTS.delete-bad-tests they are REPLACED, not skipped, and the
behavioral property each one guarded is preserved.

The property is "the widened bracket patterns do not backtrack
catastrophically", which is a claim about growth, so it is now stated by
scaling the input instead of by timing it:

- validate health runs the pathological unclosed bracket at 4,000 and 16,000
  characters and must return the same correct result (no W021) at both.
- The four reader entry points run five ReDoS shapes at 5,000 and 20,000 and
  must name exactly the phases each shape should name at each size, with the
  variant cardinality unchanged across the two.

Catastrophic backtracking is superlinear, so a regression cannot complete the
4x leg under any ceiling, while a bounded matcher is indifferent to the
scaling. The one attack that is well-formed-but-oversized now pins its reading
precisely (`['1'.repeat(n)]`) rather than being lumped in with the malformed
ones. A positive control asserts the readers still extract a well-formed
bracket heading, so "names no phase" cannot pass by the readers being inert.

`{ timeout }` is a hang backstop, not an assertion: it turns a runaway into a
deterministic failure instead of a suite that never returns.

* fix(#2761): give the bracket grammar one owner and teach the drift guard to see it

The bracket milestone-intro grammar was re-typed verbatim in three readers —
roadmap-parser's bracket-fallback selector, state's `isMilestoneBounded` and
verify's `checkBracketCoherence` — which is exactly what #2761's own gate
forbids ("no token literal outside src/phase-id.cts"). `check:phase-id-drift`
reported clean the whole time: its detector only ever knew the phase-NUMBER
token grammar, so the bracket class `[A-Z][A-Z0-9_]*` was invisible to it.

Ownership. `src/phase-id.cts` now exports the class as
`BRACKET_PROJECT_CODE_SRC` and the intro in the two shapes its readers need:
`bracketMilestoneIntroSrcFor(milestone)` (pinned to one milestone) and
`BRACKET_MILESTONE_INTRO_CAPTURING_SRC` (milestone captured). The pinned
builder owns the pad2 spelling rule too — "canonical spelling only, not `0*N`"
was previously restated in prose beside each copy, a convention two files had
to keep agreeing on by hand. `BRACKET_ID_SRC`, `BRACKET_ID_PREFIX_RE` and
`BRACKET_QUALIFIED_KEY_RE` now compose the class rather than re-spelling it.
All three call sites consume the owner; the regex sources are byte-identical to
what they spelled, asserted against hand transcriptions of the pre-fix lines.

Guard. `scripts/lint-phase-id-drift.cjs` gains a bracket rule
(`findBracketGrammarDrift`), wired into `scanRepo` and tagged `kind`. It
deliberately does NOT copy the token rule's `line.includes(CANON_REF)` escape:
that escape is line-level, and verify's copy referenced the owner for the
MILESTONE field on the same line as the re-typed PROJECT-CODE class — so a
bracket rule with that escape would have kept passing on the very site under
review. Partial ownership is the drift; only a dedicated `// phase-id-owner:`
comment suppresses it.

Proof, end to end: planting the shipped verify.cts literal back into src/ makes
`npm run check:phase-id-drift` exit 1 naming `[bracket] src/verify.cts:1475`;
restoring it returns the gate to ok.

Tests. phase-id-drift-guard carries all three shipped literals as negative
fixtures, the same-line-owner-reference case, the case-widened evasion variant,
the sanction rules, and a temp-tree scan proving the rule is wired into
scanRepo rather than merely exported. continuation-grammar-parity drives the
pinned and capturing shapes over a 12-entry corpus and requires the same
verdict from both plus the same captured milestone — widening either alone
fails there.

* test(#2761): drop the out-of-scope source-grep exemption from the selector pin

The baseline-selector pin in tests/adr-612-bracket-heading-selection.test.cjs
read `src/*.cts` with readFileSync and regex, claiming the no-source-grep
escape with a source-text-is-the-product reason. CONTEXT.md's documented scope
(RULESET.TESTS.no-source-grep.exemption) reserves that escape for tests whose
subject is a runtime CONTRACT FILE — STATE.md, config.toml, hooks.json, agent
.md — and `src/*.cts` is none of those.

It was also broader than it looked: eslint-rules/no-source-grep.cjs matches the
marker in ANY comment in the file, so one block's claim disarmed the rule for
the whole ~700-line suite.

The pin itself is worth keeping — the BASELINE ARGUMENT at each call site is a
fact no behavioural test can recover (flipping verify's milestone-complete site
from LABEL_ONLY to ANY_BRACKET grants a tolerance it has never had, and every
behavioural test still passes), so pinning it does require reading the authored
source. That reading moved to `scripts/lint-phase-id-drift.cjs` — the seam's
own guard, where source scanning is sanctioned (`warn` scope) and already
happens for the grammar rules — as `countSelectorBaselines` /
`scanSelectorBaselines`, which return a structured census. The test asserts on
the returned data and touches no file text.

CONTEXT.md is unchanged: the exemption scope was not widened to fit the test.
No marker remains in the suite, so the rule is live across all of it again, and
eslint passes with the escape removed rather than relocated. A floor assertion
pins that the census actually found the five consumers, so the "no other file
consumes the selector unpinned" check cannot pass on an empty scan.

* docs(#2761): rewrite the changeset as a lead + deltas instead of one paragraph

The fragment was a single ~6,400-character paragraph that opened on internal
mechanics, buried the user-visible change, and named an internal test path
(tests/adr-612-bracket-phase-counting.test.cjs) that means nothing to a reader
of the CHANGELOG.

It now leads with what a user sees — bracket-style phase IDs are recognized on
the read path across roadmap, validate, verify and state — followed by compact
bullets for the behavioral deltas, the opt-in caveat, and the upstream
consequences. Both merge-added disclosures are kept: the #3185 legacy-sentinel
Phase-0 delta with the reason the bracket counter keeps the narrower rule, and
the enumerator's convention-argument default flip with the four newly-scoped
read surfaces and the note that archival and milestone-completion paths are
unchanged. Two deltas from this review round are stated as their own bullets:
per-occurrence classification in `missing_phase_details`, and workstream-scoped
convention resolution including the workstream-rollup count change.

The internal test path is gone and the body is down to ~3,850 characters. The
bullets render as sub-bullets under the CHANGELOG entry and the `(#NNNN)`
suffix still lands on a trailing paragraph rather than mid-list.
`npm run lint:changeset` passes.

* test(#2761): use helpers.cleanup in the drift-guard temp-tree test

`local/no-raw-rmsync-in-tests` rejects a bare `fs.rmSync` in tests — the shared
helper carries the Windows-EBUSY retry budget. The temp-tree scan added with
M3's guard coverage test used the raw call.

* docs(#2761): correct the occurrenceKey comment on qualified-key normalization

The comment claimed `bracketQualifiedKey` is "fold- and padding-insensitive, so
`[gsd.2] 01` and `[GSD.02] 01` are one phase". The fold half is right; the
padding half is not. `BRACKET_MILESTONE_NUMERIC_SRC` is `(?:[1-9]\d{2,}|\d{2})`,
so `bracketQualifiedKey('GSD.2-01', 'bracket')` returns null and that input
takes the composite fallback instead.

No behavior change — the two key spaces use different separators and cannot
collide — but the claim was checkable and wrong. Restated: the owner case-folds,
and padding-tolerance is not a property it has or needs, because each accepted
milestone has exactly one canonical spelling and `[GSD.2]` is malformed rather
than an alternate spelling of `[GSD.02]`.

* ci(#2761): give the coverage-gate job the heap floor the shard jobs already have

The gate's report steps parse ~2.5GB of merged raw V8 dumps in one
process at the runner's implicit ~4GB ceiling, so pass/fail comes down
to GC timing (run 31338081337 OOM'd; next passes the same volume at
28s). Deterministic locally: crashes at --max-old-space-size=4096,
completes in 11s at 8192. PR shard data is within 0.15% of green next
runs — the load is pre-existing, only the ceiling was missing. #2952
set 6144 on the shard-collection step; the gate job was missed.

* Revert "ci(#2761): give the coverage-gate job the heap floor the shard jobs already have"

This reverts commit 03c74a97c911e9ddc53675aded9c9bbb89f14cd1.

* fix(#2761): thread the resolved convention into cmdStateValidate's phase-directory lookup

#3208 replaced cmdStateValidate's `startsWith` prefix test with the canonical
`phaseKeyFromDir(...) === selectedPhaseKey` comparison. That is the right
surface, and it is why the lookup now needs the resolved `phase_id_convention`,
which the rewrite does not pass.

`phaseKeyFromDir` deliberately refuses to read a bracket directory without an
explicit signal (ADR-2121: a bracket dir is string-indistinguishable from the
legacy letter-prefixed-decimal family), so un-threaded it returns the whole dir
name as the key — `GSD.02-05-real-work` -> `GSD.02-5-REAL-WORK` — while the
STATE side is the bare `05` that `parsePhaseFromProse` yields. Both sides of one
comparison derived under different conventions is #2562's defect class, and this
file's other three `phaseKeyFromDir` call sites already thread against it.

Observable: a bracket repo whose phase directory plainly exists reported
`valid: false` and "no phase directory matches phase 05", and the drift scan
(plan-count mismatch, verification status) never ran. The pre-#3208 `startsWith`
missed the same directory but skipped silently, so this is a visible-failure
regression on bracket repos, not a new miss.

Non-bracket conventions are byte-identical by construction: `extractPhaseToken`
branches only on `=== 'bracket'`, so null / 'milestone-prefixed' / unresolvable
compile the same path as the un-threaded call. The flat-legacy twin assertion
pins that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2761): note the cmdStateValidate convention threading in the changeset fragment

The fragment described the bracket read path as of the pre-merge branch. 504c64ff
added a shipped behaviour change — `state validate` now resolves bracket phase
directories — that the fragment did not mention, so the rendered changelog would
have under-described what ships.

Body text only; `type` and `pr` are untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): invert the inherited-wart characterization — #3225 fixed it upstream

Required by the merge of next @ 86101ee6; test-only, no src delta.

This branch disclosed rather than fixed a pre-existing upstream wart: on a
LEGACY repo, `### Phase 999:` warned from `validate consistency` while
`validate health` suppressed it — the two verbs contradicting each other.
The case was pinned as a characterization test ("INHERITED WART, unchanged")
precisely so it would INVERT if upstream ever closed it, rather than rot
silently.

ae7dc529 (#3225) closed it, by adding the `isSentinelPhaseId` guard to this
very loop. So the assertion inverted on the merge — as designed. Flipped to
assert the FIXED behaviour rather than deleted: it is the negative-space
proof that this branch's `sentinelPhases` guard never had to grow a legacy
reading of its own, and it reds if a future conflict resolution keeps our
guard while dropping upstream's.

Added a scope control alongside it (a legacy NON-sentinel `### Phase 09:`
with no directory still warns), so deleting the loop outright cannot pass.

Union proven load-bearing in BOTH directions against the merged tree —
neither guard subsumes the other:
  - drop `isSentinelPhaseId(p)` (ours only) -> 1 red, exactly this case.
    Note upstream's own #3225 tests stay GREEN there: they cover the
    disk-side loops (sentinel dir on disk) and the gap-numbering filter,
    not the ROADMAP-side loop. This case is that site's only coverage,
    which is the second reason to keep it rather than delete it.
  - drop `sentinelPhases.has(p)` (upstream only) -> 4 red bracket suites
    (icebox-not-missing, health/consistency agreement, occurrence-aware
    suppression, checklist-index suppression).

tests/adr-612-bracket-read-tolerance.test.cjs 79 -> 80, all green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): pin the three re-homed W006/W007 reads the merge left unfalsifiable

#3309/#3310 deleted the helpers this PR threaded (collectDiskPhases,
collectDiskPhaseEntries, collectArchivedPhaseDirNames,
forEachArchivedPhaseToken) and rebuilt their reads inside
src/planning-snapshot.cts + src/health-diagnostic-rules/*.cts. The convention
threading moved with them in the merge commit — but a mutation sweep over the
re-homed sites found three where reverting the convention argument changed real
CLI output and NOT ONE existing test went red. The pre-migration sites were
covered indirectly, through helpers that no longer exist, so the coverage did
not survive the relocation even though the behaviour did.

Each case is pinned at the CLI with its flat-legacy twin as the byte-identity
control, plus a non-vacuity control in the opposite direction:

  archivedPhaseTokens      revert -> "Phase 05 in ROADMAP.md but no directory on
  (planning-snapshot.cts)            disk" for a phase whose only directory is
                                     under .planning/milestones/v1.0-phases/
  roadmapPhaseCheckboxes   revert -> "Phase 09 in ROADMAP.md but no directory on
  (planning-snapshot.cts)            disk" for an unstarted `- [ ]` phase
  W007's extractPhaseToken revert -> "Phase GSD.02-77-orphan exists on disk"
  (roadmap-disk-consistency)         instead of "Phase 77 exists on disk"

Red/green: each single-line revert reds exactly its own case and nothing else;
every legacy control stays green in all three runs.

NOT pinned, and disclosed as droppable rather than given an unfalsifiable test:
buildValidPhaseSet's extractPhaseToken(dir, convention) in W002. That rule
unions disk tokens with roadmapDeclaredPhases and archivedPhaseTokens, and any
STATE.md reference a bracket disk token would rescue is already rescued by the
ROADMAP half — probed directly, the argument makes no observable difference. It
is threaded because it restores collectDiskPhases(planBase, convention)'s
derivation exactly, not because a test needs it.

Test-only; revertable independently of the merge commit.

* test(#2761): pin the two reads the #3165 extraction left unfalsifiable

DROPPABLE, offered as such. Test-only; the merge commit is correct without it.

#3428's extraction gave `collectAnalyzePhases` TWO call sites — the scoped
milestone window and the truncated-window recovery path. The merge threads
`convention` into both and swaps `detailKeys` with `phases` across the recovery.
Neither of those was falsifiable by the suite as it stood:

  - passing a NULL convention at the FALLBACK site, with the scoped site still
    threaded, is green across all 518 tests of the bracket, roadmap and
    milestone-window files. Every pre-existing bracket assertion reaches the
    enrichment through the scoped window, so none of them can observe the
    fallback's reading at all;
  - dropping `detailKeys = fallbackScan.detailKeys;` — the line this merge
    authored — is likewise green, because no fixture that reaches the recovery
    path carries a checklist bullet, so `missing_phase_details` is `null`
    either way.

That is the same hole class lap 3 closed in 9914359c: an argument whose revert
changes real CLI output with zero reds.

Pinned at the CLI on the shape that reaches the fallback on a bracket repo — a
MID-MIGRATION ROADMAP: bracketed ACTIVE milestone, its phase-detail sections
and checklist bullets sitting after an intervening CLOSED legacy milestone (so
the window closes over prose only), plus one legacy `### Phase N:` section of
its own. Six tests:

  - a NON-VACUITY control that removes the phase directories, so #3428's own
    precondition fails and the result is the empty one the recovery exists to
    replace — this is what proves the rest read the FALLBACK's output rather
    than the scoped scan's;
  - the directory assertion (`disk_status`/`plan_count`/`summary_count` for
    canonical `{CODE}.{MM}-{PP}-slug` dirs) + a flat-legacy twin asserting
    byte-equal shape, and that the twin is the right answer rather than a
    shared wrong one;
  - `missing_phase_details: null` for phases the recovery just found, and its
    companion direction — a bullet with no heading anywhere is STILL reported,
    so a fix that merely suppressed the report does not pass;
  - a non-bracket repo taking the same path unaffected.

Red/green, each mutation reverted afterwards:
  - fallback call site -> `null` convention: exactly 2 reds, both directory
    assertions in this block; 307 other tests green, including upstream's own
    #3428 tests in tests/milestone-window-single-owner.test.cjs.
  - `detailKeys` swap deleted: exactly 2 reds, both `missing_phase_details`
    assertions; 414 other tests green.
  - `matchPhaseDirs` 3rd argument dropped (the lap-2 site): 4 reds — the 2
    already on record plus this block's 2, which is the point: one seam, now
    reached by two paths.

DISCLOSED IN THE TEST BODY WITH MEASURED OUTPUT, not fixed: `hasPhaseEntries`
(src/roadmap-parser.cts) is convention-blind, so on a PURE-bracket ROADMAP the
window classifies COMPLETE and #3428's recovery is gated off entirely. That
document returns `{"scope":"complete","phase_count":0,"next_phase":null,
"phases":[]}` — the "genuinely empty milestone" answer, the indistinguishability
#3184 introduced `scope` to remove — where the flat-legacy twin returns
`{"scope":"truncated","phase_count":2,...}`. Widening it changes the value of an
upstream-owned output field on bracket repos, so it is a behaviour slice (the
PR-2.5/PR-4 convention-less-readers question), not a merge-round change. The
mid-migration shape pinned here is the reachable half.

* fix(#2761): mirror the bracket terminator in the milestone-scope write guard

parser's terminator vocabulary — "a level 1-3 heading that is not a Phase
heading and carries a milestone signal". On this branch that vocabulary is
convention-SELECTED: `computeBracketSectionEnd` adds `isBracketMilestoneBoundary`
as a terminator arm, and the ADR-canonical `## [GSD.09] Hidden` carries NO
vN.N token, NO ✅/📋/🚧/🔄 marker, and not the word "Milestone". Left
unmirrored, the guard accepted exactly the description it exists to reject.

NOT a defect on clean next — a bracket heading terminates nothing there. The
branch widens the terminator set, so the branch owns the mirror.

MEASURED at the CLI seam before the fix (bracket fixture, one milestone,
one phase):

  $ gsd-tools phase add $'Sneaky\n## [GSD.09] Hidden'   -> exit 0, written
  $ gsd-tools phase add 'Innocent follow up'            -> exit 0, written
  $ gsd-tools roadmap milestone-scope
    { "scope": "complete", "phases": ["01","1"], "phase_count": 2 }
  $ grep '^### Phase' .planning/ROADMAP.md
    ### Phase 1: Sneaky
    ### Phase 2: Innocent follow up      <- in the document, out of the window

After the fix the first add exits 1, ROADMAP.md is byte-unchanged and no
phase directory is created. The legacy twin (same text, no
`phase_id_convention`) still exits 0 — opt-in only, base behaviour preserved.

SHAPE
- `findMilestoneScopeHeadingLines(text, convention)` — REQUIRED, the same
  tripwire `scanMilestonePhaseIds` carries in the merge commit and for the
  sharper reason: a blind call here fails OPEN (the guard quietly ACCEPTS a
  window-narrowing description), so a future call site must fail to COMPILE.
  Census is one caller, which pays nothing for it. A non-bracket value takes
  the pre-existing path byte-identically. The bracket arm routes through
  `isBracketMilestoneBoundary`, the same single-owner phase-vs-milestone
  discriminator `computeBracketSectionEnd` consults; no second bracket-heading
  grammar is spelled here.
- `selectedBracketId` is deliberately `null`, so the same-milestone
  CONTINUATION exemption never fires and a value naming the ACTIVE milestone
  is flagged too. That is this predicate's third stated conservatism and
  rests on its own existing argument: which milestone is active is a property
  of the document at write time, not of the text being validated, and
  over-rejecting is one-directional.
- `assertDescriptionPreservesMilestoneScope` takes `cwd` and resolves through
  the branch's tolerant try/catch — this guard runs BEFORE `loadConfig` and
  before the ROADMAP existence check, so an unresolvable convention must
  degrade to the legacy vocabulary, never turn a rejection into a crash.
- The error's marker list gains the bracket form only when the project is on
  the bracket convention.

RED/GREEN
- Revert the bracket arm -> 2 reds (phase add; insert + add-batch), the
  non-bracket control and both no-false-positive cases stay green.
- Revert the probe threading in the merge commit -> 1 red (the probe case).
- 6 new cases in the #3262 file, each with a non-bracket or
  no-false-positive control: fenced bracket milestone heading and bracket
  PHASE heading are both non-violations.

Droppable: revert this commit and the merge stands on its own; the branch
then ships the gap as a disclosure instead of a fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changeset): narrow the enumerator claim to the directory set — per-entry rendering in progress/stats/init-manager is display-slice scope

Round-7 review confirmed 3 of the 4 surfaces named by the closing claim
parse each directory or heading with legacy-only patterns that live in
files this PR does not touch (commands.cts:1766, commands.cts:2116,
init.cts:2241). The claim now states exactly what this slice delivers:
the scoped directory set. Their per-entry conversion is display work,
deferred to the epic's display PR with the statusline/progress-card
gates it belongs beside.

* fix(#2761): de-accident the unmatched-milestone bracket fixture, re-pin to #3480's withhold contract

tests/adr-612-bracket-phase-counting.test.cjs:515 ("a milestone that
does NOT match STATE is not scoped in") asserted total_phases === 0.
Since today's merge brought in 70b5c1a1 (#3354/#3480, already in
`next`), buildStateFrontmatter withholds total_phases (omits the key,
read back as null) for a milestone that is genuinely sectioned but
matches no ROADMAP heading, instead of substituting a computed number
— the fixture's `state json` read now returns null, failing the
strictEqual(0) assertion.

The original single-section fixture's 0 only ever survived by
accident: hasMilestoneSectioning requires >=2 milestone-vocabulary
headings to call a ROADMAP sectioned, and its isPhaseHeading helper
recognizes only the legacy `Phase N:` text form — so the bracket
phase heading `### [GSD.03] 09: Not this milestone` (title containing
the word "milestone") miscounted as a second milestone heading,
tipping hasMilestoneSectioning true and routing to the disk-count
branch. With that miscount removed, the same one-section fixture
reads 1 (the non-matching milestone's phase count leaking into v2.0's
total) — proving the pinned 0 was never validating the scoping this
test claims to exercise.

Replaced the fixture with two genuine milestone sections (GSD.03,
GSD.04), neither matching STATE's v2.0, with phase titles carrying no
incidental vocabulary — the real #3354 shape. Re-pinned the assertion
to null, matching upstream's own tested contract (tests/state-
document.test.cjs, "#3354 with nothing stored, the key is omitted
rather than written from the dir count").

Verified: file green 3/3 runs (104/104), full state-suite regression
771/771 green. The isPhaseHeading gap that made the old fixture
accidental is a real, pre-existing, upstream-owned limitation
(hasMilestoneSectioning only guards >=2-section conflation, not a
single non-matching section leaking through) — not introduced by

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): exempt two self-authored-fixture regexes from #3441's unbounded-quantifier rule

Both sites parse the STATE.md the test itself just wrote to its own
tmpdir — fixed-size fixture output, not adversarial or document-scale
input. Exempted with the justification-comment pattern the tree's own
tests use for this exact case (settings-integrations, copilot-install,
research-agent-profiles). The sites predate the rule; #3441 landed on
next this morning and this branch picked it up in the catch-up merge.

* fix(#2761): fix forward two next-movement test regressions from the round-11 rebase

origin/next's #3573 newly threads STATE's stored `milestone:` value into
`state json`'s buildStateFrontmatter call (previously always `undefined` on
that read surface). Two pre-existing PR-2 fixtures reach code paths that
call never exercised before that merge:

- RED (repro2 case C): a version-less, all-bracket-id ROADMAP now hits
  getMilestonePhaseFilter's pre-existing (unaffected by this branch) row-5
  `versionResolved && !headingFound => SCOPE.UNSCOPED` rule, which withholds
  `progress.percent` even though the scoped total_phases/completed_phases
  are still correct. That row-5 rule is load-bearing for six other
  version-less-document pins in this same file; narrowing it broke seven of
  them in testing, so production code is untouched here. Reassert the
  test's real claim (scoping, via total_phases/completed_phases) and
  disclose the now-withheld percent instead of silently dropping it.

- PIN (repro12 LEGACY control): a fenced-example-vs-real version heading
  fixture. #3573 routes this read through sliceMilestoneWindow (fence-aware)
  instead of the legacy anyMilestonePattern raw scan (fence-blind) this pin
  was disclosing as out of scope, so the fixture no longer exercises the
  blind path — total_phases moves from the whole-doc 4 to the correctly
  scoped 2. Re-pinned with an upstream-attributed comment, matching this
  file's existing house style for prior origin/next movements (see the
  sibling "PIN (G3 LEGACY control)" comment).

Both are test-only fix-forwards: no production code changed. The remaining
12 failures in this file at 27363e9d0 (dropped seam commits' VERIFICATION.md
naming + fixture updates) are pre-existing and deferred to the stacked
follow-up per the task's own scope.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): thread phaseIdConvention explicitly at the two write-adjacent enumerator call sites (round-11 BLOCKER)

listMilestonePhaseDirs' phaseIdConvention param lost its `= null` default
(phase-locator.cts), so an omitted convention now means "resolve from
config" instead of "explicitly not bracket" — a deliberate flip, but
milestone.cts and state.cts were not in the PR-2 diff and both call the
enumerator without threading it:

- milestone.cts cmdMilestoneComplete (the single #3597 shared derivation
  feeding the stats loop, --dry-run preview, and the real archive/rename
  pass) now resolves phase_id_convention once and threads it explicitly,
  so a bracket project's `milestone complete` archives its real
  bracket-declared phase directories instead of silently inheriting
  whatever the enumerator's lazy default resolves to.
- state.cts cmdStateUpdateProgress's own enumerator call threads the same
  resolved convention. Empirically this does not change the #3217 withhold
  gate (scope is assigned before headingConvention resolves in
  getMilestonePhaseFilter, so it's convention-independent either way) or
  the reported percent (already correctly threaded via
  computeUpdateProgressPreview -> buildStateFrontmatter); it closes a
  second, silently-resolved answer to the same question this file's own
  ONCE-and-THREAD rule (~:2300) already states as policy.
- state.cts's other listMilestonePhaseDirs call site (the state-sync
  scope-only read, ~:4791) is left unthreaded on purpose, with an inline
  note explaining why: only `.scope` is consumed, and `.scope` is set
  before convention resolution in getMilestonePhaseFilter, so it cannot
  disagree with a threaded convention.

Both enumerated sets are pinned by new tests (round-11 BLOCKER block in
tests/adr-612-bracket-phase-counting.test.cjs), including a mutation-style
regression check on the milestone.cts fix (forcing the un-threaded default
back on makes the pinned test fail, catching a regression to the
pass-all-degrade legacy reading).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): correct the changeset's false archival claim and amend ADR-612 for the round-11 M2(2) split scope

The changeset (.changeset/2761-bracket-read-tolerance.md) asserted "The
archival and milestone-completion paths are unchanged" — false: both
paths reach the widened enumerator, and this PR now threads their
convention explicitly (previous commit). Replaced the closing paragraph
with an accurate description of what changes for a bracket project at
those two call sites, and notes both enumerated sets are pinned by tests.

ADR-612 (docs/adr/612-bracket-phase-id-convention.md) amended per M2(2):
the 2026-08-03 PR-2/PR-4 boundary proposal is added in-body (PROPOSED,
not stamped — mechanics per docs/contributor-standards.md:143 reserve ADR
ratification to maintainers), rescoped to what actually ships in PR-2
(#2761) now that the round-11 M1 split moved the completion-seam
threading (isPhaseArtifact/scopeToPhase, phase-id.cts:964-1090) out into
a separate, stacked follow-up PR referenced generically via epic #612:

- state.cts read-tolerance (both total_phases derivations, the #1514
  retirement filter) moves into PR-2's row, alongside milestone.cts,
  reflecting the round-11 fix above.
- the write-observability point is restated in terms of what actually
  ships: explicit threading at two named call sites, pinned by tests,
  rather than silent inheritance.
- the completion-seam threading is named and explicitly excluded from
  PR-2's scope, with a new ratify item (5) and a Negative consequence
  bullet covering the sequencing cost.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): repair round-11 response gaps — disk-side sentinel bug, test-fixture bug, scanMilestonePhaseIds caller drift

Four independently-verified gaps in the round-11 repair response:

1. isSentinelPhaseId (src/phase-id.cts) treated a bare, untagged phase
   directory under phase_id_convention: "bracket" (`0-bootstrap`, no
   `{CODE}.{MM}-` prefix) as sentinel milestone 0 by falling through to the
   legacy leading-int rule when the bracket-tag match failed. This silently
   dropped a real, on-disk, milestone-declared phase directory from
   `listMilestonePhaseDirs` (phase-locator.cts:424, the only unguarded call
   site) and undercounted completed_phases/percent. Mirrors the carve-out
   already present on both heading-side counters (state.cts's
   countRoadmapPhaseHeadings guards its bare-0 exclusion with `bracketId &&`;
   roadmap-parser.cts's scanMilestonePhaseIds composes the bare-token rule as
   999-only) — under bracket convention, milestone 0 is expressed only via an
   explicit bracket tag, so an untagged leading 0 is a real phase token.

2. tests/adr-612-bracket-phase-counting.test.cjs's writeProject fixture
   hardcoded `01-VERIFICATION.md` for every "complete" phase dir regardless
   of the dir's real phase token. Under legacy convention this file already
   correctly failed #3511's strict isPhaseArtifact match for any dir other
   than phase 01 (the fixture never modeled what it claimed to); under
   bracket convention the pre-existing ambiguity fail-safe admitted it
   regardless. The two readings' disagreement was mistaken for a missing
   completion-seam threading (isPhaseArtifact/scopeToPhase convention
   awareness, correctly split out to the stacked #3644 per round-11's M1).
   Replaced the hardcoded name with verificationNameFor(dir), deriving the
   real per-phase filename from the production normalizePhaseName the same
   way cmdScaffold does — the 7 failing "flat-legacy-twin" / CHARACTERIZATION
   assertions pass on #2867's own code with no seam threading required, and
   the genuine seam-dependent cross-phase-stray-exclusion test the fixture
   fix would otherwise have hidden lives on the stacked branch instead.

3. tests/roadmap-parser.test.cjs's two #3577 tests still called
   scanMilestonePhaseIds with the old single-Set return shape; this PR
   changed it to `{ ids, qualifiedIds }` for every other caller but missed
   these two, which don't touch #2761/bracket code at all. TypeErrors at
   runtime, not silently-wrong assertions. Updated both call sites.

4. .changeset/2761-bracket-read-tolerance.md gains a paragraph disclosing
   fix 1 above, so the changeset stays accurate to what actually ships (the
   round-11 BLOCKER was exactly this changeset going stale once).

Verified: tests/adr-612-bracket-phase-counting.test.cjs +
tests/continuation-grammar-parity.test.cjs + tests/roadmap-parser.test.cjs =
354/354. adr-612-{coherence,grammar,heading-selection,read-tolerance,
selection.property}.test.cjs + collision-characterization = 287/287
(CONFIRMED-CLEAN set, unaffected). Full unfiltered suite run separately.

PR #3643 (round-11's M1 split, opened draft per that round's explicit
requirement) was auto-closed by this repo's own draft-PR policy 11 seconds
after opening — draft PRs are unconditionally closed here. Reopened as
non-draft #3644 (same branch, same commits, DO NOT MERGE / stacked-on-#2867
marker in the body, enhancement template) since the repo's own bot confirms
non-draft contributor PRs are tolerated even off-template.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): narrow the state-update-progress pin claim to what mutation testing actually proved, add the test that covers the rest

Round-11 BLOCKER response gap (verified, not a guess): the changeset and
ADR-612 amendment point 2 both claimed the round-11 tests pin
`cmdStateUpdateProgress` so "a future change to the enumerator's default
cannot silently move ... what state update-progress reports without failing
a test." Mutation testing disproves this. Reverting BOTH the state.cts
explicit `phaseIdConvention` thread AND simulating the phase-locator.cts
pre-#612 hardcoded-null default (`phaseIdConvention: null` at that one call
site) leaves the existing PIN test ("writes a real percent for a
bracket-scoped milestone") green, because that percent comes from
`computeUpdateProgressPreview` -> `buildStateFrontmatter`, a separately and
already-correctly-threaded derivation the state.cts inline comment at ~:965
already candidly documents.

What the reverted thread DOES change, empirically, is `phaseDirs`/`totalPlans`
— the enumerated `.value` this call site feeds into the #3233 zero-plans
no-op gate a few lines below. That is the one place a regression at this call
site is observable in the command's output. Added a test that pins exactly
that: a bracket milestone whose declared phases carry no plans on disk,
alongside a decoy directory that plainly does not belong to the milestone
(no bracket tag, no phase token) but does have a plan. Correctly scoped, the
decoy is excluded and the #3233 no-op fires (`updated: false`). Degraded to
the pass-all legacy reading, the decoy is swept in, `totalPlans` flips
nonzero, and the no-op never fires (`updated: true`).

Mutation-tested against both scenarios:
- Reverting ONLY the state.cts explicit thread (falls back to `undefined`,
  which `getMilestonePhaseFilter` resolves via the identical
  `resolvePhaseIdConvention(cwd, undefined)` call the explicit thread also
  makes): new test stays green — confirms the single-hunk thread really is
  pure single-derivation hygiene, exactly as the existing inline comment
  claims, for this test too.
- Forcing `phaseIdConvention: null` at that call site (the combined-revert
  scenario the round-11 mutation testing actually exercised): new test FAILS
  (`updated: true, percent: 0` instead of the expected `updated: false`).
  Restoring the real code makes it pass again.

Narrowed the changeset and ADR-612 point 2 prose to match: `cmdMilestoneComplete`'s
enumerated set is pinned against a pass-all-degrade regression (genuine,
already verified in round 11); `cmdStateUpdateProgress`'s own call site is now
pinned against the #3233 zero-plans no-op specifically, not against the
reported percent, which the prose no longer claims. Added a one-line pointer
to the new test in the state.cts inline comment. No production behavior
changes — test and documentation only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(#2761): reconcile ADR PR-2 module map

* fix(#2761): restore #3639's dir-aware W007 sentinel exclusion lost in the rebase replay

The rebase replayed this file's pre-#3639 patch over next, reverting the
isSentinelPhaseId(token) -> isSentinelPhaseDir(dirName) fix: the extracted
token is milestone-stripped, so a bracket sentinel (GSD.999-07-icebox,
GSD.00-01-backlog) was invisible to the id predicate and false-fired W007.
Restores upstream's call and comment verbatim; the branch's convention-aware
token remains for the diagnostic message only.

Caught by next's own #3639 regression tests (CI shard 2). Local: the file's
18/18, health-diagnostic-rules 165/165, full npm test exit 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MeAsZNfhiUQioFPA4ygEGS

* docs(#2761): ratify scope and correct surface claims

* test(#2761): strengthen legacy selection and drift claims

* docs(#2761): correct the enumerator claim, scope list, and ratification receipt

Three documentation corrections, none touching production code.

The changeset claimed the shared phase-directory enumerator "now defaults its
convention argument to 'not yet resolved' rather than 'resolved, and not
bracket'". That is false: `phase-locator.cts:375` still destructures
`phaseIdConvention = null`, and the lazy resolve-from-config fires only on
`undefined` (`roadmap-parser.cts:941`, `:1928` — whose own comment records that
"explicit null still means 'resolved and non-bracket'"). Measured: only 4 of 17
`listMilestonePhaseDirs` call sites thread a resolved convention (`milestone
complete` and `state`'s three). The changeset went on to name `progress`,
`stats`, `phase list` and the init manager view as now receiving a correctly
scoped set — those are precisely callers that omit it. Replaced with what the
code does, and the deferral stated plainly.

The ADR's PR-2 row omitted `scripts/lint-phase-id-drift.cjs` (+141/-14) and
`scripts/lint-phase-enumeration-drift.cjs`, both changed by this PR. An
under-claim rather than an over-claim, but the row is the epic's
scope-of-record.

Ratify item 5 asserted maintainer ratification while citing only the review that
raised the question. It now cites the review that granted it (#2867 review
`5012940978`, 2026-08-24) and quotes its terms, so the claim carries its receipt.

Found by an adversarial review pass over the round-12 diff. (#2761)

* fix(#2761): drop the no-op --json that strict argv now rejects

Surfaced by the rebase onto next, not by a change in this PR's subject.

#3884 ("failure is a value — strict argv", e20744eac) made an unrecognized
flag a hard error: `state validate --json` now exits 1 with
`unknown flag "--json"; accepted: --strict` on stderr and EMPTY stdout,
where the token was previously accepted and ignored. `state validate` never
had a `--json` flag — JSON is its only output shape — so the argument was a
silent no-op from the start.

The helper parsed that empty stdout, so all five subtests in the
"state validate resolves bracket phase DIRECTORIES" suite failed identically
with `SyntaxError: Unexpected end of JSON input` at the JSON.parse, masking
what they actually assert.

Dropping the token restores the same envelope the helper already parsed. No
assertion changes. Verified against the fixture the suite builds: exit 0, and
the output carries the S005 plan-count warning the drift assertions require
with no S004 phase-directory warning — i.e. the bracket directory resolves,
which is the behaviour these tests exist to pin. 94/94 in the file.

* fix(#2761): exempt the phase-counting suite from the docs-guard lane

Surfaced by the rebase onto next. #3753 (107eb8c1d) added
lint-docs-guard-registration, which requires every test file that reads a
docs/ path to be either registered in the docs-guard lane or carry an
explicit marker. It flags adr-612-bracket-phase-counting.test.cjs, which
reads no docs/ path at all.

The file's only docs/ occurrence is the ADR-612 Decision 1 citation in a line
comment; every read call it makes targets a tmpdir .planning fixture. It trips
Detector 3, whose DOCS_TEMPLATE_LITERAL_RE sees an odd prose backtick in the
comment block above that citation as opening a template literal and reads the
span between them — citation included — as a docs/ path expression. That
detector documents this trade in its own header: it favours recall, and says a
false positive costs one docs-guard-exempt marker with a reason. The baseline
records the same class ("comment-only mentions") for 46 of its entries.

Registration was the wrong side of the trade here: it would run this suite in
the docs-guard lane on docs/ changes whose content it never reads.

Three pieces, matching what the gate requires and what its 54 existing entries
already do:
  - the marker in the file's header window, written without backticks so it
    cannot itself disturb the parity tracking findExemption does;
  - the basename in DOCS_GUARD_EXEMPT_BASELINE, since the ratchet fails a new
    marker until the baseline is deliberately updated — the reviewable diff is
    the point;
  - the FIX 3 per-file fingerprint, derived with the lint's own
    extractDocsPathReferences rather than retyped, so the exemption fails loudly
    if the set of docs/ paths this file mentions ever changes.

Gates: lint-docs-guard-registration 0 violations, ci-docs-guard-registry 51/51.

* docs(#2761): correct the bare-0 comment and state what opting into bracket costs

Round-13 review items Minor 2 and Minor 3, both still live on the previous head.

Minor 2 — src/phase-id.cts. The comment block claimed "Bare `0` is admitted
alongside it because a 0.x sentinel is a legitimate identity that predates
padding", which contradicted both the shipped constant and its own next
paragraph. BRACKET_CANONICAL_NUMERIC_SOURCE is `(?:[1-9]\d{2,}|\d{2})`;
measured against it, `0` is rejected while `00`, `05`, `99`, `100` and `999`
are admitted. The paragraph four lines below already recorded the removal
("the earlier `(?:\d{2,}|0)` ... admitted ... a bare `0` that pad2 never
produces"), so the block asserted and denied the same fact. The stale sentence
is replaced with what ships: `00` is the backlog sentinel's canonical padded
identity, `\d{2}` already covers it, and nothing needs the unpadded spelling.
No behaviour change — the constant is untouched.

Minor 3 — docs/CONFIGURATION.md. The phase_id_convention row described the
read-path widening but never said what a repo GIVES UP by opting in. Added:
on a bracket repo a heading whose bracket is followed directly by a digit is
read as a phase heading, so `### [RFC.2119] 5:`, `### [v1.0] 2024:` and
`### [ADR.612] 3:` — legal prose headings under any other convention — are
claimed as phases and move phase_count, total_phases and W006. Those three
shapes are the ones phase-id.cts's own selector comment names as the reason
the widened read is selected at construction time from this value rather than
applied everywhere; the row now carries that trade instead of only its
upside.

Gates: tsc --noEmit exit 0, npm run lint:ci exit 0, npm run test:unit
33890/33891 with the single failure being emitted-attribution's base drift
against a next that moved after the rebase (133/133 against the rebase base;
this push re-bases onto the current tip, which resolves it).

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-31 14:03:48 -04:00
0xdhx
ef9ce3e598 fix(#3702): count asterisk, plus and ordered markers as deferred-items entries (#3739)
* fix(#3702): deferred-items counts `*`, `+` and ordered markers as list items

`deferred-items.md` has no template and no mandated shape, but its parser
recognised only the `- ` hyphen marker. Asterisk bullets, plus bullets and
dot-terminated ordered lists — all lists in CommonMark and GFM — contributed
ZERO entries on both the headless and the heading-delimited path, and a mixed
file dropped its non-hyphen entries while keeping their hyphenated siblings,
under-reporting without ever looking empty.

The restriction was a regex literal inherited from the Gaps seam, where the
template genuinely mandates the hyphen YAML-lite form; nothing in the module's
stated rationale distinguishes `*` from `-`.

Widened on the deferred path only:
- `splitGapsEntriesCore`'s entry opener, `extractGapEntryFields`' line-0 strip
  and `rawGapEntryText`'s line-0 strip take a `BulletMarkers` parameter that
  DEFAULTS to the hyphen-only set, so `## Gaps` keeps its template-mandated
  grammar byte-for-byte and the module still has exactly one grouping pass.
- `splitDeferredHeadingEntries`' body-bullet test, `stripLeadingBulletMarker`
  and `acknowledgeDeferredItem`'s status-field regexes move in lockstep —
  widening what OPENS an entry without widening what is STRIPPED before field
  extraction would surface an entry that can never resolve.

Unchanged, and pinned by tests: prose-only and bare headings still contribute
nothing ("prose is not an item"); a table under a leaf heading still yields
exactly its rows, since table lines are skipped before the body-marker flag can
be set and a `|` row is not a list marker; the paren-terminated ordered form
`1)` is out of this fix's scope.

* docs(#3702): changeset fragment (pr: 0 placeholder pre-create)

* fix(#3702): widen the forensic-audit prose entry rule to match the parser

Sibling site of the same defect class, found by a defect-class sweep of the
deferred-items consumers. `/gsd-progress` check 7 does NOT go through
`gsd-tools query` — it globs `deferred-items.md` and has the model read entries
by a prose rule that mandated "one entry per top-level `- ` line". Left as-is,
the marker widening would hold on the CLI path while the one consumer that
bypasses the parser kept reporting "No unresolved deferred items" for a file
written with `*`, `+` or an ordered marker: the same false negative, surviving
in the only place the fix could not reach by code.

Also pass DEFERRED_BULLET_MARKERS explicitly where the heading path extracts
fields. It was already correct — stripLeadingBulletMarker pre-strips the widened
set from every line, so the default hyphen strip is a no-op there — but relying
on that leaves a detection site and a strip site nominally on different marker
sets, which is exactly the asymmetry the BulletMarkers doc comment warns about.
Explicit is local; inferred is a trap for whoever edits the strip next.

Out of scope, noted rather than fixed: forensic-audit.md globs only
`.planning/phases/*/` and so misses archived milestone phases that
`scanDeferredItems` covers. Pre-existing, a different defect, and not this
issue's ruling.

* docs(#3702): note the milestone-close halt for heading-shape non-hyphen files in the changeset

A heading-delimited deferred-items.md written with */+/ordered markers
previously parsed to zero and closed silently; it now yields entries whose
heading shape acknowledgeDeferredItem refuses, halting complete-milestone
until hand-edited. User-visible, so the fragment states it.

* chore(#3702): set changeset fragment pr to 3739

* fix(#3702): CR-normalise the heading path and the acknowledge writer (review B1, M4, m2)

B1 — `splitDeferredHeadingEntries` stored RAW lines; on a CRLF file every
body line but the last still carried its `\r`, the `$`-anchored marker
strip failed on it, the marker survived into field extraction and the
field was lost — a `**Status:** resolved` that was not the file's final
line resurfaced its entry as open. The heading path now stores CR-stripped
lines like the headless path already did, and the strip regex tolerates a
trailing CR on its own. Round 1's CRLF test put `**Status:**` on the last
line, the one position `collectSection`'s `.trimEnd()` had already
de-CR'd; the new tests put it first and mid-body.

M4 (pre-existing on `next`) — `acknowledgeDeferredItem` found the status
line on a CR-stripped copy but rewrote the raw line with a `$`-anchored
`.*`, which cannot consume `\r`; `replace` returned its input, and the
writer reported `ok` over byte-identical content. The rewrite now runs on
a CR-stripped line. The comment that claimed `.*$` consumed the `\r` is
corrected — it was the bug, stated as the design.

m2 — the indent probe for an inserted `status:` line ran on the raw line
and fell back to indent 0 on CRLF; it is CR-stripped too.

* fix(#3702): derive every deferred-items marker regex from one source (review M3, N1, N2)

M3 — round 1 carried the marker alternation in FOUR places: the
`BulletMarkers` pair and two inline literals inside
`acknowledgeDeferredItem`, under a doc comment saying the interface
existed so a detection site and its strip site could not drift. All four
now derive from `DEFERRED_MARKER_ALT`; drift is impossible rather than
discouraged. A parity test pins the vocabulary against
`markdown-sectionizer`'s `iterateBullets` on everything the two grammars
are meant to agree on, and names the two points they deliberately differ.

N1 — the ordered marker is `\d{1,9}\.` (CommonMark §5.2), not `\d+\.`.

N2 — the marker is followed by `[ \t]`, not `\s`, which also accepted
`\r`; the tab remains accepted (CommonMark-legal) and the divergence from
`iterateBullets`' literal space is pinned rather than papered over.

The four regexes are exported for the parity test only.

* fix(#3702): an ordered marker opens an entry only from `1.` or inside a run (review B2, m1)

B2 — `\d+\.` alone read ordinary prose as a list: "2026. was a bad year
for this module" and, under `### Notes`, "3. is the number of retries we
settled on." both opened an entry on round 1, the second straight through
the "prose is not an item" contract that round's AC4 claimed to preserve.
CommonMark §5.3 faces the same ambiguity when an ordered list would
interrupt a paragraph and resolves it by requiring the list to start with
1; `matchListOpener` applies that rule wherever an ordered marker is seen,
with the run carried per list (headless) or per leaf-heading body. Numbers
after the first are ignored, as CommonMark ignores them. Stated cost,
pinned: a hand-numbered list starting at 2 reads as prose — every ordered
record in the #3702 scan starts at 1.

Both reviewer cases are pinned as prose; the ruling's `1. alpha / 2. beta`
shape still counts.

m1 — the 9-digit boundary is pinned at both sides (`999999999.` opens,
ten digits is not a marker), and the 3-vs-4-space indentation cliff is
pinned as deliberately NOT applied: the parser is indent-lenient because
surfacing a questionable hand-written entry beats dropping a real one.

* fix(#3702): thematic breaks close the list and fenced code never opens an entry (review M1, M2)

M1 — `- - -` was a phantom `"- -"` entry on base; widening the marker set
added `* * *` and `+ + +` to the class, and `* * *` is the separator an
author writing in the `*` style is most likely to use. A CommonMark §4.1
thematic break (plus the `+ + +` gesture, which is the same garbage as an
entry name) now closes the open entry on the headless path and is dropped
from the body on the heading path — neither an item nor a continuation.

M2 — neither splitter was fence-aware, so `+ `-prefixed diff lines and
`1.`-numbered repro steps inside a code block counted as entries; #3702's
wild records carry exactly those blocks. Both splitters now classify lines
by the sectionizer's own `scanFencedBlocks` (so `~~~`, indented and
unterminated fences behave as `stripFencedCode` would): fence content
never opens an entry, is continuation inside an open one — keeping the
span invariant `acknowledgeDeferredItem` re-verifies — and is discarded
before the first.

* test(#3702): range the #2287 deferred-items property over marker × shape × line ending (review B3)

The `#2287` property hard-coded `- ` and filtered `\r\n` out of its
arbitraries, so the widened marker set — an enumerated domain, exactly
what a property is for — was never under it. It now ranges over
`{-, *, +, ordered}` × `{headless, heading}` × `{LF, CRLF}`, with the
heading shape placing `**Status:**` first or last: the review's
prescription (markers × line endings) would not have reached B1, which
lives on the heading path only, so the shape axis is the load-bearing
addition. Ordered entries are numbered from 1, so the B2 run rule is
under the property too.

A second property drives `acknowledgeDeferredItem` over every unresolved
headless entry across the same marker × line-ending grid — the one that
reaches M4 (a CRLF rewrite that reported `ok` and wrote nothing) and m2.

* test(#3702): pin the milestone-close halt on a heading-delimited `*`/`+`/`1.` file (review m3)

A heading-delimited `deferred-items.md` written with a non-hyphen marker
previously parsed to zero entries and let `complete-milestone` close
silently; it now yields entries whose heading shape `acknowledgeDeferredItem`
refuses, which the milestone loop turns into `record_ack_failure` → exit 1.
The loop is prose in a workflow, so the test drives the two CLI calls it
makes: `audit-open --json` must list the entry, and
`audit-open acknowledge --text <the audit's own text>` must refuse with the
heading-delimited message and write nothing.

* docs(#3702): changeset and forensic-audit prose carry the round-2 grammar

The changeset names the CRLF fixes, the ordered start-at-1 rule, thematic
breaks and fences. The `/gsd-progress` forensic-audit step is the one
prose parser of this file and must state the same grammar the code has.

* fix(#3702): round-review refinements — run ends at a paragraph, rejected ordinals unstripped, breaks at any indent, fenced fields, `## Gaps` scope

Findings from the pre-push adversarial review of round 2, each pinned:

- An ordered run ENDS at a paragraph that follows a blank line (CommonMark
  §5.3); a non-indented line with no blank before it is lazy continuation
  and keeps the run open. `1. a` / blank / `paragraph` / blank / `5. x` is
  one entry, not two.
- The heading path strips the marker off every body line before field
  extraction (#3457); a line whose ordinal `matchListOpener` REJECTED must
  not be stripped, or "3. status: resolved" as prose loses its `3. ` and
  reads as a resolved field. `splitDeferredHeadingEntriesDetailed` now
  carries a per-line opener flag and only accepted openers are stripped —
  in headless regions of a heading-shaped file too.
- A thematic break is recognised at any indent, matching the parser's
  indent-lenient reading of items; `    * * *` was a phantom `* *`.
- Fenced lines carry no FIELDS either: a `status: resolved` quoted inside a
  code block no longer resolves its entry on either path.
- Block structure (breaks, fences) is a property of the GRAMMAR, carried as
  `BulletMarkers.blockStructure`: the deferred set opts in, the Gaps set
  does not, so `## Gaps` is byte-for-byte on its `next` behaviour — the
  round-2 M1/M2 change had reached it through the shared splitter.

* test(#3702): the property exercises the rejected-ordinal branch; the N2 control is independent

Round review: the widened #2287 property numbered every ordered run from 1
and so never generated an ordinal the start-at-1 rule rejects — it could
not tell round 1 from round 2 on B2. Each entry may now carry a decoy prose
line beginning with a non-1 ordinal, placed where it cannot end a run
(before the first headless entry; first in a heading body), followed by a
`status: resolved` that must never become a field; and a decoy-only
heading body must yield no entry.

The N2 assertion accepted a tab, which round 1's `\s` accepted too, so a
`[ \t]` → `\s` revert alone stayed green. NBSP, form-feed and vertical-tab
are now asserted refused — the assertion that fails on that revert on its
own, and the disclosure that `[ \t]` narrows what round 1 accepted.

* fix(#3702): the splitter records its own opener flags; an opener clears the blank-line memory

Round-review continuation, two state defects in the ordered-run logic:

- `blankSeen` survived the headless splitter's opener branch, so an opener
  followed by a lazy continuation line read as "paragraph after a blank" and
  ended the run — `1. a` / blank / `2. b` / lazy / `3. c` folded `c` into `b`.
  The opener branch now clears it.
- The heading path re-derived per-line opener flags for headless regions
  without the paragraph reset, re-accepting a rejected `3. status: resolved`
  under a stale run and stripping it into a field. `GapsEntrySpan` now
  carries the flags the splitter itself computed, and the heading path reads
  them; the re-derivation is deleted.

* fix(#3702): ordered-run memory is per indent — nested runs resolve, nested ordinals never inherit the top-level run

Round-review continuation 2: nested openers consulted the TOP-LEVEL run
flag and never wrote their own, so a nested `1. / 2.` run under a hyphen
entry rejected its `2. status: resolved` (round 1 resolved it), while a
nested `3. status: resolved` under a nested `- ` bullet inherited an open
top-level run and was stripped into a false field.

`OrderedRuns` keys the memory by indent: a new opener at indent d resets
every deeper level, a paragraph after a blank at indent d ends the runs at
d and deeper, a thematic break or a heading clears all. Both splitters use
it; the top level still decides entry boundaries, nested levels decide
only which continuation lines are accepted openers for field stripping.
Pinned for LF and CRLF.

* fix(#3702): run levels — one top level at or above the base, CommonMark column indents, a fence ends its level's runs

Round-review continuation 3:

- A dedenting top-level list (`    1.` / `  2.` / `3.`) lost its entry
  boundaries: the exact-indent run lookup rejected the shallower ordinals
  before the boundary check ran. Every indent at or shallower than the
  list's base is now ONE level, in both splitters.
- `indentOf` counted characters, so a tab and a space aliased to one level
  and `\t1. nested` / ` 2. status: resolved` resolved falsely. Indent is now
  measured in CommonMark columns (§2.2: a tab advances to the next multiple
  of 4), for the run level and the entry-boundary check alike.
- A nested run survived a fenced block. A fence is a non-list block: its
  opening delimiter ends the runs at its level and deeper, exactly as a
  paragraph after a blank does.

* fix(#3702): the indent measure is grammar-scoped — Gaps keeps next's character count

`blockStructure: false` promised the Gaps grammar byte-for-byte parity with
`next`, but the CommonMark-column indent measure added for the deferred
grammar was shared by the whole splitter core, so tab-indented Gaps input
changed entry boundaries in BOTH directions:

  `\t- a` / `  - b`  — next folded into one entry, HEAD split into two
  `  - a` / `\t- b`  — next split into two,      HEAD folded into one

`indentWidth` now keys the measure on the grammar: columns for the deferred
set, raw character count for Gaps. The opt-out covers indent semantics, not
only fences and thematic breaks.

Four cases pin both halves — the two flipped Gaps pairs, the two Gaps pairs
that never moved, and the same tab/space pairs on the deferred path returning
the opposite (column-measured) verdict by design.

* fix(#3702): the acknowledge path reads and writes through one classifier

Round 3, Blockers 1 and 3, and Minors 7 and 8 — one mechanism, so one commit.
Every consumer of an entry's lines now reads the splitter's own per-line
verdict instead of a re-derivation of it.

B1. Round 2 widened the WRITER's status-line finder to the deferred marker set
while `extractGapEntryFields` still de-bulleted line 0 only. A nested
`  * status: pending` was therefore selectable by the writer and invisible to
the reader: acknowledge rewrote it in place, returned `ok`, and the item stayed
outstanding on every later audit. Measured against a `next` build, `*`, `+` and
`1.` each resolved on base and stopped resolving at round 2's head — a
regression, not a gap in new behaviour. The hyphen form of the same shape was
already broken on `next` and is fixed here too: one classifier cannot be right
for three markers and wrong for the fourth.

`parseGapEntryFieldLine` is now the single place a line is classified as a
field, and it reports the offset at which the VALUE begins. The rewrite happens
at that offset rather than through a second regex, so a line the classifier can
select is one whose rewrite it has already located — the selection and the
rewrite cannot disagree. Both `DEFERRED_STATUS_FIELD_RE` and
`DEFERRED_STATUS_REWRITE_RE` are deleted rather than widened. A read-back guard
returns `rewrite_not_readable` rather than `ok`; it is unreachable by
construction today and is the fail-loud floor under the next divergence.

B3. This is the end state the round-3 review prescribed on both #3739 and
#3773: #3773's shared classifier, parameterised by this PR's marker set, with
this PR's two status regexes deleted. #3773 lands first. Its hyphen-only strip
is consistent with `next`'s hyphen-only splitter today, so the writer/reader
divergence is created by THIS merge, which is why widening every consumer
belongs to the PR that widens the domain.

m7. The heading path marker-stripped its lines before calling the reader, so
the reader's fence scan ran over text the splitter never saw: `- ```sh` is an
ordinary bullet to the splitter but strips to a fence opener, and a
`**Status:** resolved` after it was suppressed as fence content — a resolved
entry resurfaced as open. Stripping now happens inside the reader, after the
fence scan.

m8. `rawGapEntryText` stripped a marker off line 0 unconditionally, but on the
heading shape line 0 is the heading TEXT: `### 1. Race in the writer` was
silently renamed to `Race in the writer`, and the name is the key acknowledge
matches on. Line 0 is stripped only when the splitter accepted it as an opener.

Also removed: `splitDeferredHeadingEntries`, whose sole caller only null-checked
it (round 3, M4 — the claim was zero callers, which was wrong; the wrapper's
`.map` was waste at the one call site), and `stripLeadingBulletMarker`, which
this change leaves with no callers at all. The export surface narrows to the two
splitter regexes the behavioural parity test reads (M6).

[PEER-ASK pr-order-12d5]
q: Reviewer blocked both on merge order. I'm declaring #3773 lands first and
   building the end-state shape into #3739 now (both my status regexes
   deleted). Does that match your plan?
reply: CONFIRMED - same order, derived independently. #3773 cannot carry the
   fold: `DEFERRED_BULLET_MARKERS`/`BulletMarkers` have zero occurrences at
   `next` (verified), so the prescribed end state is not executable inside
   #3773 without absorbing this PR's work.
deadline: 03:55 UTC (answered before it)
fallback: declare #3773 first, adopt end-state shape in #3739, push+comment
decision: proceeded as stated; #3773 lands first, this PR carries the widening
   of every consumer.

Refs #3740

* test(#3702): pin the detect/strip symmetry, and drop a white-box test that could not reach it

Round 3, Blocker 2 and Minors 6 and 9.

B2. The regression shipped green because no fixture put a marker on a nested
status line. Four markers x {nested status line}, each asserting the entry
READS BACK as acknowledged rather than that acknowledge merely reported `ok` —
reporting `ok` over a line the reader skips is the whole defect. Plus the bare
capitalised `Status:` case (the reader stores it case-sensitively, so the
writer must not select it), and an idempotence test, which is the failure the
defect actually produced: the item resurfaces, is acknowledged again, and never
settles.

Each of these was run against the pre-fix build first: all five fail there and
pass here. Two further assertions in the block are labelled CONTROL because
they held pre-fix — they guard the new offset-based rewrite and the opener-flag
threading against regressing, and calling them regression tests for a reported
defect would overclaim.

M6. The round-2 parity test asserted that four writer-side regexes embedded the
same source string. That is true of a detect/read asymmetry too, so it could
not have caught B1 — and two of the four regexes were widened into `export =`
purely to let it read them. Replaced with a behavioural test that drives the
real seam: every marker that opens an entry must also resolve it through
acknowledge. The structural assertion is kept for the two splitter regexes,
which really are two copies of one alternation.

m9. `expectedResolved` was computed and immediately voided; the loop beneath it
already asserts both polarities.

m7/m8 coverage lands here too: a bullet whose content is a fence opener must
not suppress the entry's fields, and a heading beginning with a list marker
must keep it in the entry name.

* docs(#3702): document the deferred-items entry shape where the file is written

Round 3, Major 5, and #3702's own item 2. The widened grammar was documented in
the reader (`forensic-audit.md`) but not at the write site, where
`executor-examples.md` still said only "log to deferred-items.md" — so the
question the issue actually raised, which shapes count, remained unanswered
anywhere a human writes the file.

States what opens an entry (`-`, `*`, `+`, and `1.` when the list starts at
`1.`), that `1)` is not a marker here, that a separator closes the list and
fenced content is never an entry or a field, and that an entry without an
explicit `status: resolved` stays open by design.

* chore(#3702): regenerate the changeset through the generator

Round 3, Minor 10. The fragment was hand-named against 64 generated names on
`next`, and its body ran ~250 words against CONTRIBUTING's one-sentence form.
Regenerated via `npm run changeset`, which is also what the random three-word
name is for: concurrent PRs never collide.

* fix(#3702): the fence gate lives on the seam both sides call, not just the reader

Found by the pre-push adversarial review of this round, and it is a regression
this round introduced rather than a pre-existing one.

`extractGapEntryFields` applied `fencedLineSet` before classifying; the
acknowledge writer's status-line search did not. So a `status:` line inside a
fenced block was SELECTED by the writer and SKIPPED by the reader — the write
produced a line nothing reads, the read-back guard refused it, and the entry
became impossible to acknowledge at all: `audit acknowledge` raised an internal
error and `complete-milestone` halted on it.

Measured, `- alpha` / fence / `  status: pending` / fence:

  next          ack=ok                   -> reads back "acknowledged"
  round-2 head  ack=ok                   -> reads back ""      (the B1 defect)
  before this   ack=rewrite_not_readable -> refuses entirely   (worse than next)

`entryFieldLines` is now the seam — per line of an entry, the field it declares
or `null`, fences included — and the reader and the writer both go through it.
That makes "the writer cannot select a line the reader will not read back"
structural rather than asserted, which is what the previous commit's message
claimed while a second read-side filter still lived outside the classifier.

Two comments corrected with it. The read-back guard is NOT "unreachable by
construction": this round shipped a reachable path to it, which is precisely
what an invariant asserted in a comment is worth. And the M6 replacement test
put its marker only on the entry opener, so it passed against the defective
build — the exact weakness it was introduced to fix in round 2's test. It now
marks the nested status line too, and fails pre-fix like the rest.

Round-3 tests against the pre-fix build: 10 of 12 fail there, and the 2 that
hold are labelled CONTROL because they guard this round's new code rather than
pin a reported defect.

* fix(#3702): one end-of-file CRLF algorithm, adopting #3773's with its B4 closed

Round-4 M1. Two open PRs shipped two different answers to "what line ending
does an entry that ENDS THE FILE get?", and the review's ruling was that the
disagreement needs one answer, not two. Neither shipped answer was that one.
Measured on builds of both heads:

  case                                     #3739 r3   #3773   here
  undelimited single entry, CRLF preamble    pass      FAIL    pass
  LF-dominant list, one stray CRLF at EOF    FAIL      pass    pass
  (the other five)                           pass      pass    pass

This PR's content.endsWith('\r\n', matchIndexInContent) reads the terminator of
the PREVIOUS line, so it propagated an isolated CRLF into an LF-dominant list --
refuted by #3773's own LF-dominant fixture, ported here. Withdrawn.

#3773's crlfAtEof asks the right question -- does anything before the entry,
within scope, contradict CRLF -- and fails closed. But its scope goes EMPTY for
an undelimited single-entry list, because the entry-list region runs from the
first entry's start to the insertion point and those coincide; crlfAtEof('') is
false by its own before.length > 0 guard, so 'preamble\r\n\r\n- alpha' gained a
bare \n in a CRLF document. That is #3773's B4, verified by driving its head.

Adopted here with the scope widened to everything preceding the insertion point
where the preferred region is empty, rather than asserting LF from no evidence.
That only ever loosens a scope carrying zero information, and the predicate
stays fail-closed over the wider one. An entry at offset 0 of an undelimited
document has no evidence under either scope and stays LF.

Tests: 10 added. Negative control, driven -- 1 of the 10 fails against this
branch's own pre-fix head (the stray-CRLF fixture); B4 fails against #3773's
head; the remaining 8 are the scope counterexamples ported with the function,
which were regression pins in #3773 and are guards here. Each still kills a
simpler algorithm: drop any one and a refuted scope passes again.

Four deferred-items suites 450/450, 0 skipped. npm run lint:ci exit 0.

* fix(#3702): drop the unreachable rewrite_not_readable guard (B3)

Round-4 B3: the status had zero test coverage in either file. The review
offered two branches -- drive it from a test, or delete it and stop carrying an
untested terminal status. Taking the second, with the reason stated rather than
assumed.

Why it cannot be driven. Round 3 added the guard after a fenced `status:` line
proved the writer could select a line the reader would not read back. Round 3
then closed that divergence STRUCTURALLY, by routing the writer's line selection
and the reader's field extraction through one entryFieldLines seam. The guard
now detects a state construction prevents: 21 document shapes were driven
against it -- fence openers on the bullet line for every marker in the widened
set, duplicate and triplicate status lines, bolded and nested variants, fences
between duplicates -- and none reached it. The only seam that would is routing
the internal call through the module's exports so a test could stub it, which
reshapes production surface for a test.

Why leaving it undriven is not free. RULESET.TESTS.mutation-score runs Stryker
incrementally over changed files at an 80% threshold and says to treat a
surviving mutant as a failing test specification. An undriven `if` on a changed
file is exactly that, on both the condition and the .toLowerCase() comparison.

What this gives up, stated rather than hidden: if a future change re-splits the
writer's selection from the reader's extraction, acknowledgeDeferredItem returns
ok over an item that stays outstanding -- the original #3702 defect class. One
correction to the review's framing: match_verification_failed does NOT backfill
it. That check runs BEFORE the write and compares the matched span to the
target, so it cannot see a post-write read-back failure. The protection against
re-splitting is the shared seam and the round-3 tests that pin it, not a runtime
assertion. A comment at the removal site records all of this.

Removing it also drops the union member from both files, which resolves the PR
body's internal contradiction (it claimed no type-signature changes while adding
one) and the duplicate-status surface #3773 collides on.

No test changed behaviour: 450/450 across the four deferred-items suites, 149/149
across the audit suites, npm run lint:ci exit 0 -- the same figures as before the
removal, which is itself the evidence that nothing exercised the branch.

* fix(#3702): the deferred fence gate is indent-unbounded, like the rest of the grammar (M2)

Round-4 M2. scanFencedBlocks is CommonMark, which caps a fence delimiter's
indent at three spaces -- a fourth makes it an indented code block instead. This
grammar had already opted out of that cliff for entry openers ([ \t]*) and for
THEMATIC_BREAK_RE (^[ \t]*), but not for fences. So a fence at four spaces was
not a fence to the gate, and a `status: resolved` line inside it RESOLVED the
entry containing it.

That is not an exotic shape. A fenced block written under a NESTED bullet sits
at four spaces, so ordinary hand-written deferred-items.md files reach it.
Driven before the fix at indents 4, 5, 8 and a leading tab: all four silently
resolved. It is the #3702 silent-resolution defect class in a new place.

gsd-core/references/executor-examples.md, added by this PR, states flatly that
"nothing inside a fenced code block is an entry or a field". The review offered
fixing the parser or bounding that claim in three places. Fixing it -- the claim
is the one users will rely on, and the grammar had already chosen unbounded
indent everywhere else.

NO second fence dialect (the rule blankIndentedFenceDelimiters states). The
classification is still done by scanFencedBlocks, the one exported CommonMark
state machine, over a de-indented VIEW of the same lines. Run lengths, backtick
vs tilde, closer-must-match-and-not-trail, info-string rules and the
unterminated-at-EOF case remain that engine's answers. Indent is the only
dimension hidden from it, and it is exactly the dimension this grammar has
already declared it does not measure. Index alignment is 1:1 -- map preserves
length -- so every returned line index still addresses the original line.

Scope is the deferred grammar only. Both marker-parameterised call sites gate on
markers.blockStructure, which the Gaps set does not set, so Gaps reaches an empty
set. Verified, not asserted: the 47-fixture Gaps differential (marker x
line-ending x separator x fence x break x key-shape x list-shape) is
BYTE-IDENTICAL across this change, 8033 bytes both sides.

Tests: 14 added, of which 8 fail against the pre-fix source and pass here; the
other 6 are the deliberate controls -- indents 0 through 3, which must NOT move,
and the Gaps opt-out guard.

Four deferred-items suites green; the 58 suites touching uat/deferred/sectionizer
run 6045 tests with an IDENTICAL failing set before and after this change (17
pre-existing environment failures -- installs and an unpinned GSD_EMITTED_BASE;
emitted-attribution passes 259/259 in isolation with its base pinned). lint:ci
exit 0.

* fix(#3702): changeset, both prose parsers, and the minors (M3, M4, m1-m3, m5, n1-n2)

M3 -- the changeset omitted a user-BREAKING change. Measured against next: a
heading-delimited deferred-items.md written with `*`, `+` or `1.` went from
"0 entries, so complete-milestone has nothing to acknowledge and closes" to
"1 entry, the CLI writer refuses the heading shape, ACK_FAILURES accumulates,
exit 1". The `-` form already halted and is unchanged. That is release-note
material: a close that used to succeed now fails, and the correct response is to
fix the file, not revert. Also names the fence-indent fix below, and adds #3740
so #3773's issue is attributed here as it is absorbed.

M4 -- gsd-core/workflows/progress/steps/forensic-audit.md is a SECOND,
model-executed parser of the same grammar, and prose cannot carry a parity test.
Its widened text stated the start-at-1 rule, fences and separators but not the
`1)` exclusion nor the nine-digit ordinal cap, both enforced in code with pinned
tests. Both stated now, along with the round-4 fence-indent rule. (No ack
fragment: the size ratchet's currentSizes does a NON-recursive readdirSync of
gsd-core/workflows and agents, so a file under workflows/progress/steps/ is
outside its scope -- verified by reading the helper, not by the green.)

n1 -- executor-examples.md documented that the BOLDED status key is matched
case-insensitively and left the bare key's rule to inference. Driven: bare
`Status: resolved` is NOT read, so the entry stays open with no warning, while
`**Status:**` is. Stated explicitly, with the digit cap and the any-indent fence
rule (n2).

m1 -- boundary coverage was 2/3. limit (999999999.) and limit+1 (1234567890.)
were pinned; limit-1 (12345678.) added, per RULESET.TESTS.boundary-coverage.

m2 -- THEMATIC_BREAK_RE and the tab-expanding indent counter are hand-rolled
CommonMark rules with no in-repo peer to compare against, so the parity
assertion is against the SPEC: eight positive and five negative fixtures, plus
the two DELIBERATE divergences pinned as deliberate (`+` is a separator here but
not in CommonMark, because `+` is a list marker in this grammar and `+ + +`
would otherwise be a phantom entry; indent is unbounded). One fixture was
initially wrong -- `-- -` IS a CommonMark break, since the spec allows free
spacing between the three characters -- and the parser was right.

m3 -- the result union is hand-duplicated in audit.cts as part of a deliberate
structural view of uat.cjs, so the fix is not to delete a copy but to make drift
observable. Every REACHABLE status is now driven from a fixture; four of the six
(ambiguous, unsupported_heading_shape, already_resolved, match_verification_failed)
had no assertion anywhere in the suite before this. match_verification_failed is
still undriven and the test says so rather than omitting it.

m5 -- DECLINED, with the measurement. The review is right that `(\s*)` in the
opener and `/^[ \t]*/` in the reader disagree about \f, \v and NBSP, but its
prescribed narrowing was implemented, driven and REVERTED: as shipped, an entry
indented with any of those surfaces, parses its status field, acknowledges, and
reads back acknowledged -- a complete round-trip. Narrowing turns all three into
SILENTLY DROPPED entries, which is the #3702 defect class itself and the opposite
of this file's stated fail-safe rule. A latent inconsistency in the safe
direction is not worth a live regression in the unsafe one. Pinned by three
round-trip tests so the prescription cannot be re-applied silently; if it is ever
closed, the direction is to make the readers agree with the opener, not to make
the opener reject lines it accepts today.

Four deferred-items suites 475/475, 0 skipped. lint:ci and lint:changeset exit 0.
The 47-fixture Gaps differential is byte-identical at 8033 bytes.

* fix(#3702): the pinned `## Gaps` phantom now cites its issue (m4)

Round-4 m4. The second assertion in the Gaps byte-for-byte test pins a real
defect as expected output: a spaced hyphen thematic break in `## Gaps` is read
as an ITEM, so `- - -` surfaces a phantom open gap named `- -`. Reproduced on
pristine next at 389bc86e0 across nine separator shapes -- every spaced hyphen
form is affected, `---`/`----`/`* * *`/`___` are not, and the dividing line is a
space after the first hyphen (the Gaps opener is /^(\s*)(-)\s/ with no
thematic-break concept at all).

Filed as open-gsd/gsd-core#3898. The pin stays: scope-limiting Gaps is the point
of the blockStructure opt-out, and this assertion is the only thing that would
notice the Gaps path moving. What was missing was the tracking -- a pinned defect
with no issue behind it reads as intended behaviour to the next reader. The
comment now says which it is and what the expectation becomes when #3898 lands.

* fix(#3702): an unterminated fence runs to the end of its entry, never past it (B1, B2)

Round 4 de-indented every line before `scanFencedBlocks`, so a fence
opened at any indent — and `scanFencedBlocks` runs an unterminated
fence to end-of-document — so one stray delimiter swallowed every entry
after it into the entry before it. `- a` / blank / four-space ``` /
blank / `- b` yielded ONE entry where `next` yields two: a widening
that made an already-counted item vanish, on the mixed-file shape #3702
exists to close. Reproduces at indent 0 as well.

The bound is the entry. CommonMark closes a fence with its container
and a container at the next item at its level; this parser extends
that to a document-level stray delimiter, where CommonMark would
swallow to EOF and the fail-safe rule (surface, don't drop) will not.
`scanFencesFrom` reports the unterminated opener and the walk supplies
the bound — the next line shaped like a top-level item — then RESCANS
from it, so a later delimiter is read on its own terms. Still one
fence dialect: every block boundary is `scanFencedBlocks`' answer.
Entry-scoped `fencedLineSet` (the field reader) already ran an
unterminated fence to the end of its lines, so reader and walk agree
by construction.

Tests: the M2 pin that asserted `[]` for a stray fence before an item
flips (the item counts); the round-4 "runs to end-of-file, exactly as
CommonMark says" test is retitled — its assertion stands because the
bound is the entry — and extended with the next entry; a new block pins
the review reproduction at both indents, a terminated deep fence still
gating, the gated status inside the bounded fence, the rescan case, and
the heading-tokenizer caveat (at indent 0 the tokenizer applies
CommonMark's own fence rule, so a heading after a stray delimiter is
body text there, exactly as on `next`).

Reverted in isolation against the final tree: 3 named tests fail.

* fix(#3702): `0.` starts an ordered list (M1)

The start-at-1 rule applied unconditionally dropped ONLY the first item
of a `0.`-numbered list — the run then started at `1.` — which is the
mixed under-report that looks like a clean parse. CommonMark §5.2
permits any 1-9-digit start and a `0.` list is ordinary; a sentence
opening with "0." is not a shape anyone writes. The threshold is now
`> 1`. The cost is restated accurately in the doc comment and pinned:
a list starting at 2 or more, at a paragraph position, reads as prose
until its first `0.`/`1.` line — the prefix, not the whole list.

Boundary tests at the threshold itself: `0.`, `1.`, `2.` starts, `00.`/
`01.`, and the prefix-loss case. Reverted in isolation: 1 named test
fails.

* fix(#3702): a non-1 ordinal is an item wherever a list is already open at its level (M2)

The per-indent run memory recorded whether the previous opener was
ORDERED, so a bullet item closed the run and `1. a` / `- b` / `5. c`
folded `5. c` into `b` — another mixed-file under-report. In CommonMark
`5. c` there opens a fresh ordered list (start=5): a non-1 start is
refused only where it would interrupt a PARAGRAPH (§5.3), and after a
list item it interrupts nothing. `ListRuns` now records "a list is
open here"; the start rule applies where no list is open at the line's
level — the positions a sentence can occupy — so the round-2 B2 pins
(doc start, after a heading, after a paragraph) hold unchanged.

Two round-2 pins move with it, both CommonMark-backed: `1. alpha` /
`- beta` / `2. gamma` is three items, and a nested `3. status:` after
a nested bullet is a nested item (a field line, as `- status:` would
be); the "rejected ordinal is not stripped" pin is re-anchored at a
paragraph position, where it still holds. Reverted in isolation: 4
named tests fail.

* docs(#3702): the two runtime-loaded docs state the grammar the parser ships (B3, M4 parity)

`executor-examples.md` (the write-site doc) and `forensic-audit.md`
check 7 (the model-executed parser) both asserted "never silently drops
a possibly-open item" over a grammar that dropped three measured shapes.
Both now carry the round-5 grammar — `0.`/`1.` starts, a non-1 ordinal
inside an open list, an unclosed fence ending with its own entry — and
the fail-safe sentence is kept with what it does NOT cover named
beside it: a fenced line, a separator, and an ordered list numbered
from `2.` upward at a paragraph position, and nothing else.

* docs(#3702): changeset reflects the merged contract

The "Breaking, and deliberate: … HALTS complete-milestone" paragraph
described a refusal that #3781 removed from `next`; a heading-shaped
file written with a newly recognised marker now surfaces its entries
and `complete-milestone` acknowledges them in place. The fragment cites
#3702 alone — #3740 and #3775 closed on `next` through #3940 and #3989;
this PR's shared reader/writer classifier subsumes both fixes rather
than closing either issue. The round-5 grammar (ordered start, unclosed
fence bound) is stated in the user-facing sentence.

* fix(#3702): the heading-shape insert lands on a line the reader reads, and keeps a closing `#` sequence

Found by the round's pre-push adversarial review. An entry whose body ends
in a fenced block — closed, or unclosed and therefore running to the
entry's end — received `status: acknowledged` AFTER its last non-blank
line, i.e. as fence content: the writer returned `ok` and the reader
never saw the marker, the item stayed outstanding. That is the #3702
class itself (a write nothing reads), on the shape #3781 just opened.
The insert now walks back over blank AND fenced lines, classified by the
reader's own `fencedLineSet`, so the marker lands on a line the reader
reads; pinned as round-trips for an unclosed fence, a closed fence, and
a pending entry ending in an unclosed fence before a heading. Reverted
in isolation: the round-trip test fails.

Separately, the leaf line-0 rewrite (`### status: open ###`) dropped the
closing `#` sequence; it is kept now. Cosmetic, pinned.

* docs(#3702): the prose parser states the bare-key case rule; both docs say what an unclosed fence does, no more

`forensic-audit.md` check 7 called `status: resolved` case-insensitive
where the code reads a bare key lower-case only (the bolded form in any
case; the value case-insensitively) — `executor-examples.md` already said
so, the model-executed parser did not. And both docs claimed "a stray
delimiter cannot hide the entries after it", which overstates B1: an
UNCLOSED fence ends with its entry; a closed pair of delimiters is a
fence, whatever sits between them, as CommonMark reads it. Found by the
round's pre-push review.

* fix(#3702): a heading whose text is a fence delimiter is a heading, not a fence

Second finding of the round's pre-push review, one door over from the
first: for a leaf headed `### ```` (or `~~~`) the entry-level fence scan
read line 0 — the heading TEXT, not a Markdown line — as a fence opener,
so every body line was fenced: the reader read no field under it, and
the writer's marker (placed by the same scan) landed on a line nothing
reads — `ok`, item outstanding. `entryFencedLines` now owns the entry's
fence view for reader and writer alike, and a leaf's line 0 never opens
a fence (the leaf tell is `openerFlags[0] === false`; a pending or
headless entry's line 0 is a marker line, never a delimiter). Pinned for
both delimiters, read and write; reverted in isolation the pin fails.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-31 13:44:47 -04:00
Tom Boucher
6beaa66b25 enhance(#3304): gate re-verification blockers on deterministic evidence (#4085)
* test(#3304): add failing-first suite for the convergence evidence gate

Content-assertion suite for the Step 7 re-verification evidence gate
(agents/gsd-verifier.md / gsd-core/references/verifier-evidence-gate.md).
Committed before the implementation to prove RED via gsd-test.

* enhance(#3304): gate re-verification blockers on deterministic evidence

Step 7's anti-pattern scan re-runs at full, unbounded scope on every
re-verification pass, independent of the must-haves established in Step 2.
A blocker it finds — other than the self-evidencing debt-marker check —
previously reverted a completed gap-closure round and started another
--gaps cycle on nothing more than the verifier's own new judgment call,
with no bound on how many times that could repeat.

A Step 7 blocker now blocks unconditionally in re-verification mode only
if it is a carried-forward gap (present in the prior VERIFICATION.md's
gaps: list) or the flagged file was git-modified since the prior pass
(a regression; fails closed toward blocking when history is unresolvable).
Otherwise it predates the gap-closure round unflagged and needs
deterministic evidence — a named test run red, or another concrete
reproducible artifact — to stay blocking. Unevidenced, it downgrades to
a new advisory: frontmatter list and report section instead of setting
status: gaps_found, and never reverts a completed must-have.

Maintainer approval was narrowed to this evidence condition only,
explicitly rejecting the broader "advisory whenever untraceable to a
requirement/decision/prior-gap" proposal — implemented and pinned by
tests/verifier-evidence-gate.test.cjs and documented as rejected in
gsd-core/references/verifier-evidence-gate.md so it can't silently
re-expand.

Closes #3304

* fix(#3304): correct window-truncation and indentation bugs in evidence-gate tests

gsd-test's GREEN checkpoint caught 3 real bugs in the test file itself
(not the production prose): a {0,600} match window was shorter than the
724-char paragraph it was scanning (the "exclude from Step 9 Rule 1"
phrase starts at offset 662), and two regexes assumed no indentation
after a markdown list-continuation line break. All three phrases are
confirmed unique across agents/gsd-verifier.md, so the windowed
submatches are replaced with direct whole-string assertions instead of
just widening the window.

Also acknowledges the deliberate byte growth in agents/gsd-verifier.md
that the differential-attribution check (ADR-2719) correctly flagged.

Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the #3304 re-verification evidence gate (Step 7 rule, Advisory bucket, advisory: frontmatter, report section); 1488 bytes, still within the LARGE-tier 48 KiB cap (48751/49152).

* docs(#3304): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-30 16:51:52 -04:00
Tom Boucher
62b0d939b6 feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey (#4083)
* feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey

Add an optional `timeoutConfigKey` field to the reviewer lane descriptor,
resolved in `resolveLanePlan` at invocation time and falling back to the
frozen `timeoutFloorMs` when unset or invalid, in the same spirit as the
existing `promptBudgetKey`/`modelConfigKey` fields. All 12 shipped lanes
declare `review.timeouts.<slug>` on both surfaces (the descriptor and their
capability.json manifest), validated by capability-validator.cjs.

For the antigravity lane, the native `agy --print-timeout` flag — previously
a second hardcoded literal (`540s`) independent of the outer cap — is now
derived from the same resolved outer timeout in `antigravityArgv`, preserving
the existing 60-second buffer relationship (ADR-2782 D6: the outer bound is
declared data, the inner one is handler-owned).

The antigravity default timeoutFloorMs stays at 600s per the maintainer's
disposition; users raise it through the new config key instead.

* docs(#3274): document review.timeouts.* and extract resolveTimeoutMs helper

Address code-review findings on the timeoutConfigKey change: extract the
inline timeout-resolution logic into a named, exported, directly-tested
resolveTimeoutMs helper (matching the file's existing configString/
normalizeHost convention); document the new review.timeouts.* federated
config keys in docs/CONFIGURATION.md, docs/reference/capability-manifest.md,
and docs/how-to/ship-a-reviewer-lane.md; add the changeset fragment.

* fix(#3274): resolve native antigravity timeout in resolveLanePlan, not the runner

gsd-test caught two design mistakes in the prior commits:

1. SpawnPlan.argv is documented and tested as fully resolved by
   resolveLanePlan (model/effort/output/prompt already folded in) — leaving
   the antigravity '{{nativeTimeout}}' marker unresolved until the runner's
   antigravityArgv violated that contract and broke tests that read
   plan.argv directly (tests/antigravity-reviewer.test.cjs,
   tests/review-default-reviewers-workflow.test.cjs). Fix: '{{nativeTimeout}}'
   is now a fifth ARGV_PLACEHOLDER member, resolved by resolveLanePlan itself
   via the new nativeTimeoutToken() helper, exactly like the other four.
   antigravityArgv reverts to its pre-#3274 four-argument form. Also missed
   updating capabilities/antigravity/capability.json's invoke.args to match
   the descriptor, which broke the manifest/descriptor parity test.

2. tests/reviewer-config-federation.test.cjs enforces a deliberate, narrow
   invariant (#3691 narrows #2797): qwen, cursor, and coderabbit — the three
   lanes with neither a model flag nor a host — may own no config key beyond
   their own prompt-budget key. Adding review.timeouts.<slug> to all 12 lanes
   violated it. Fix: those three keep timeoutConfigKey: null and own no
   review.timeouts.* key, matching their existing modelConfigKey: null. The
   other 9 lanes are unaffected.

* chore(#3274): backfill changeset PR number (pr:0 -> 4083)

---------

Co-authored-by: sim <sim@local>
2026-08-30 14:26:02 -04:00
Tom Boucher
6eea00b707 enhance(#3301): tell reviewers the plan ids and total count, grade coverage (#4084)
* test(#3301): add failing-first plan coverage manifest tests

Failing-first regression tests for the plan-id manifest, the updated Review
Instructions, and the mechanical per-reviewer coverage check, ahead of the
review.md implementation. RED baseline before the fix lands.

* test(#3301): raise allow-test-rule-refs unverified ceiling for new marker

Adding tests/review-plan-coverage-manifest.test.cjs's source-text-is-the-product
marker grows the unverified-exemption pool by one (282 -> 283), the same
documented growth path scripts/lint-allow-test-rule-refs.cjs's own failure
output names. Confirmed clean via 'npm run lint:allow-test-rule-refs' locally.

* feat(#3301): tell reviewers the plan ids and total count, grade coverage

build_prompt now derives a plan-id manifest from each *-PLAN.md filename
(stripping the -PLAN.md suffix) and appends it, with the total plan count,
to both gsd-review-instructions.md and gsd-review-prompt.md. The Review
Instructions prose requires one heading-verbatim section per id before any
cross-plan or overall-risk content.

write_reviews grades each dispatched lane's real (non-stub, non-empty)
review against that same manifest and records an optional plan_coverage:
frontmatter block, present only when a lane is incomplete. The match
escapes regex metacharacters in the id and excludes a preceding/trailing
hyphen or word character as a boundary, closing the two traps named in the
issue (a decimal phase like 12.6 satisfied by 12X6-01; a threat id like
T-04-07 registering as coverage of plan 04-07). CodeRabbit is exempt, since
it never receives the source-grounding prompt carrying the manifest.

This closes the gap where a review that silently covers only some plans in
a multi-plan phase is indistinguishable from one that covers all of them.

* docs(#3301): add changeset fragment

* test(#3301): use t.after() instead of try/finally for cleanup

CONTRIBUTING.md bans try/finally inside test bodies. Code review caught
this in the new coverage-manifest test file; switch every fixture-cleanup
site to the approved t.after() pattern.

* test: use t.after() instead of try/finally in #3300's build_prompt tests

Pre-existing try/finally-for-cleanup pattern in this file (landed for
#3300) violates CONTRIBUTING.md's explicit ban on try/finally inside test
bodies. Surfaced incidentally while reviewing #3301's diff, which cites
this file as its extraction-pattern precedent; fixed inline per the
no-defer rule rather than deferred to a separate PR.

* test(#3301): anchor coverage-check extraction on the fence line, not prose

`.plans-manifest.md` also appears in write_reviews' own prose ahead of the
```bash fence, so indexOf found that occurrence first and the
backward-walk-to-fence-open landed on the earlier, unrelated gate-check
block instead. gsd-test caught this: coverage-check tests expecting a real
verdict got null, because the wrong block ran and never writes
.plan-coverage-<slug>.json. Anchor on the fence-only bash assignment line
instead.

Emitted-Drift-Ack-Growth: review.md — #3301 adds the plan-coverage manifest and mechanical coverage check to build_prompt/write_reviews.

* fix(#3301): route id escaping through the canonical pattern seam

ADR-3212 (epic #3212) consolidated ~44 hand-rolled regex-escape copies into
one owner, src/pattern.cts's escapeRegex, specifically to stop this exact
class of duplication. My coverage-check node -e script hand-rolled the
identical metachar-escape regex — invisible to eslint-rules/no-adhoc-regex-escape.cjs
only because it lives inside a workflow markdown file, not a .cts/.cjs
source file the shape-matching guard scans. Require the compiled seam
(gsd-core/bin/lib/pattern.cjs) instead, matching the established
node -e-requires-a-compiled-lib idiom already used elsewhere in this
workflow (code-review.md's code-review-flags.cjs/code-review-depth.cjs
calls). Verified both named traps from the issue still resolve correctly
under escapeRegex's RegExp.escape-backed implementation, which differs in
escaped-text shape (hex-escapes hyphens/leading chars) but not match
result.

* test(#3301): run coverage-check block with cwd at the repo root

The block's node -e now requires ./gsd-core/bin/lib/pattern.cjs, a path
relative to the repo root (correct for production, which always runs
from there). The test harness ran it with cwd at the fixture's own temp
dir instead, so the require failed. Add an optional cwd param to
runScript (default: root, unchanged for the plan-copy-block tests) and
pass the real repo root for every coverage-check call site. Manually
verified end-to-end before spending another remote run: the extracted
block now produces the expected {complete:true} verdict.

* docs(#3301): backfill changeset pr number (pr:0 -> pr:4084)

---------

Co-authored-by: sim <sim@local>
2026-08-30 14:25:24 -04:00
Tom Boucher
86452da7cb fix(#4070): reserve shard 1's aux-suite cost out of the LPT unit-test packer (#4072)
* test(#4070): failing-first regression for shard1 aux-suite budget imbalance

- selectShard has no way to reserve virtual weight on a bin, so the LPT
  unit-test packer cannot account for shard 1's fixed aux-suite cost
  (integration/security/install/slow all pinned to shard 1/3).
- test.yml wires no such reserve into the workflow.
- ci-test-job-timeout-budget.test.cjs's LANE_COSTS entry for job `test`
  was stale (7m12s from run 30677442953, predating the aux-suite growth);
  corrected to the real evidence cited in #4070 (13m48s / cancelled at
  ~14m51s), which now honestly fails the file's own 1.5x headroom policy
  against the current 15-minute cap.

All three are expected RED on this commit; see
.gsd/bug/fix-4070-shard1-aux-suite-budget/50-test-matrix.md.

* fix(#4070): reserve shard 1's aux-suite cost out of the LPT unit-test packer

selectShard now accepts an optional initialWeights array giving one or more
bins a virtual head start before any file is placed, so LPT converges each
bin's FINAL total (assigned weight + head start) toward equal instead of
raw assigned weight alone. test.yml wires RUN_TESTS_SHARD_RESERVE=1:77 into
the full-scope unit-test step (gated on matrix.scope == 'full', so the
unrelated windows lane is unaffected) -- 77 weight units is the empirical
conversion of the aux suites' ~220s measured fixed cost, derived against
the real tests/test-timings.json (see the diagnosis artifact for the full
computation).

Also corrects two pieces of now-stale bookkeeping this issue exposed:
- ci-test-job-timeout-budget.test.cjs's LANE_COSTS entry for job `test`
  carried a 7m12s figure that predated the aux-suite growth; corrected to
  the real pre-fix evidence (13m48s / cancelled at ~14m51s, issue #4070),
  which requires raising timeout-minutes from 15 to 21 (1.5x headroom over
  the real worst-case measurement) to satisfy the file's own policy.
- test.yml's job-header and matrix comments, which still claimed the aux
  suites cost "~1m35s combined" (they now measure ~216-224s).

Closes the gap the previous commit's failing-first tests proved: selectShard
had no reserve-capacity mechanism and test.yml wired none in.

* fix(#4070): correct the reserved-weight property bound; cover main()'s reserve bounds check

Isolated code review found a genuine gap and gsd-test's real run confirmed a real
test bug it exposed:

- The fast-check property "no shard exceeds average(+reserve) + heaviest file" was
  falsified by gsd-test itself (weights=[1,1,1], total=2, reserve=6 on bin 0):
  selectShard is correct, the BOUND was wrong. A reserve large enough that its bin
  never receives a real item stays at exactly that reserve forever -- no amount of
  routing real items elsewhere can dilute a fixed head start below itself -- so the
  true bound is max(reserve, the classic Graham term), not the Graham term alone.
  Verified the corrected bound against the exact counterexample plus 20,000
  additional random trials (zero violations) before re-running gsd-test.

- Isolated review (MAJOR): the shard-total bounds check on RUN_TESTS_SHARD_RESERVE
  and its console.error fallback in main() were untested end-to-end --
  parseShardReserve itself has no concept of the shard total, so only main()
  enforces that guard, and nothing exercised it through the subprocess seam. Added
  an E2E harness test that sets RUN_TESTS_SHARD_RESERVE to an out-of-range index
  via the real CLI, asserts the fallback warning fires, AND asserts the resulting
  file selection is byte-identical to a no-reserve control run against the same
  injected timings table -- proving the reserve was actually ignored, not just
  that a warning printed.

* chore(#4070): backfill changeset PR number

pr:0 -> pr:4072

* fix(#4070): strip leaked RUN_TESTS_SHARD_RESERVE from the harness test's child env

Real GH Actions CI on this PR (run 33288554040, ubuntu shard 2/3) failed 7
tests in the shard-partitioning describe block, all with the same symptom:
`run-tests: no tests in suite "all"` where a real file count was expected.
gsd-test's own dockerized bench run never showed this, and ubuntu shards 1/3
and 3/3 (which run the same test.yml step) passed clean -- the discrepancy is
the tell: only shard 2/3 happened to schedule this specific test FILE for
that run, and the outer CI job's own environment is where the leak lives.

Root cause: test.yml's "Run unit tests" step now sets
RUN_TESTS_SHARD_RESERVE=1:77 (this issue's own reserve mechanism) on the
OUTER job that runs `npm run test:coverage:unit:raw -- --shard N`. The
harness test file's runHarness() helper spawns run-tests.cjs as a CHILD of
that same job via `{...process.env, ...extraEnv}`, so every pre-existing
--shard test in this describe block silently inherited the ambient reserve
-- even though none of them know it exists. A reserve of 77 weight units
utterly dwarfs the ~0.3 total weight of the 9-file synthetic fixtures these
tests use (none are in the real timings table, so all fall back to the same
tiny median weight), so shard index 1 is routed zero files every time --
exactly the observed "no tests" failures, and exactly the skewed 5/4 split
observed on the shard-2 test that expected a plain 3/3/3 round-robin.

Reproduced locally end to end (set RUN_TESTS_SHARD_RESERVE=1:77, spawn the
old runHarness against a synthetic 9-file fixture, --shard 1/3 -- reproduces
the exact "no tests in suite \"all\"" stderr) and confirmed the fix (env
stripped unless a test opts in via extraEnv, as the #4070 E2E bounds-check
test already does) resolves it, before re-running gsd-test.

This is a genuine bug this PR introduced -- a new ambient env var that a
pre-existing subprocess-spawning test helper didn't know to isolate against
-- not a pre-existing flake and not resource contention.

---------

Co-authored-by: sim <sim@local>
2026-08-30 00:43:02 -04:00
Tom Boucher
4d70b4dc43 fix(#4068): add --merge-async to c8 coverage-merge invocations (#4069)
* test(#4068): failing-first regression guard for c8 --merge-async flag

Adds a config-invariant test asserting test:coverage:unit and
test:coverage:report pass --merge-async to c8, guarding against the
coverage-merge OOM (release.yml finalize dry-run, exit 134/SIGABRT)
regressing a third time (prior stopgap: #199).

Also commits the sourced research memo backing the diagnosis.

* test(#4068): rename regression test to avoid lint-test-file-count collision

coverage-merge-async-flag.test.cjs's effective prefix (coverage-merge-
async-flag) matched the unrelated gsd-core/bin/lib/coverage.cjs module's
2-file cap under lint-test-file-count.cjs's startsWith bucketing, failing
FAIL_EXCEEDS_LIMIT (confirmed via the RED gsd-test run on 7b6e6ca9e).
Renamed to c8-merge-async-flag.test.cjs -- no colliding prefix.

* fix(#4068): add --merge-async to c8 coverage-merge invocations

The `test:coverage:unit` script OOM-crashed (exit 134, SIGABRT) in the
release.yml finalize dry-run of 1.12.0: all 1785 unit tests pass, then
c8's report/merge phase crashes ~76s later against the 6144 MB heap
ceiling. Root cause (verified against this repo's pinned c8@11.0.0
source, node_modules/c8/lib/report.js): the default sync merge path,
Report._getMergedProcessCov(), loads every raw V8 coverage file for the
whole run into memory as one array before merging. This is a recurrence
of #199 (same OOM class at ~466 tests, "fixed" by raising the heap
ceiling) -- the suite outgrew 6144 MB as it grew to 1785 tests, and
test.yml's own coverage-gate job already independently hit and
stopgap-fixed the identical class once (test.yml:614-620, 4096->8192 MB).

c8 ships the upstream fix for exactly this: --merge-async (c8 v7.14.0,
already inside the pinned c8@^11.0.0 range -- no dependency bump),
which switches to Report._getMergedProcessCovAsync(), reading and
merging one raw file at a time instead of loading them all at once.
The merge arithmetic (mergeProcessCovs()) and --all's zero-coverage-file
inclusion are unchanged between the two paths, so the coverage gate's
accuracy is unaffected.

Verified empirically (not just by source inspection): against a 223 MB
synthetic raw-coverage corpus derived from this repo's actual 235
gsd-core/bin/lib/**/*.cjs files (same order of magnitude as the 358 MB
figure documented in test.yml), c8's real Report class crashes with the
same "Ineffective mark-compacts near heap limit" signature via the sync
path at a 300 MB heap ceiling, while the async path completes at a flat
~40-50 MB peak down to a 100 MB ceiling. Full repro steps and numbers:
.gsd/bug/fix-4068-coverage-merge-oom/10-diagnosis.md.

Only `test:coverage:unit` and `test:coverage:report` are changed.
`test:coverage` (line 150) is a non-CI dev-convenience script, never
invoked as a command by any workflow. `test:coverage:unit:raw` (line
153) already runs --reporter none, which skips the merge/report phase
this flag affects, entirely -- adding it there would be a no-op.

Fixes #4068

* fix(#4068): add changeset fragment for the coverage-merge OOM fix

The changeset fragment was generated locally (npm run changeset) but
never committed -- isolated code review caught the gap (no
.changeset/*.md on the branch, confirmed via git diff --name-status).

* chore(#4068): backfill changeset PR number

pr:0 -> pr:4069

---------

Co-authored-by: sim <sim@local>
2026-08-29 22:03:28 -04:00
Tom Boucher
8793d307f6 fix(#4060): drive repo-baseline lint check in-process, not via a subprocess timeout race (#4065)
* test(#4060): failing-first regression for repo-baseline subtest timeout race

The "repo baseline passes" subtest in lint-allow-test-rule-refs.test.cjs
drives the script under test via a spawnSync subprocess with a fixed
30s timeout, which races the script's real wall-clock completion
against unbounded CI-load contention -- it has already died at this
race twice (#4060, and once before at a lower bound). Rewrites the
subtest to call the script's `main` directly, in-process, removing the
subprocess timeout race entirely. This commit only changes the test
(main is not yet exported), so it fails first with
`TypeError: scriptUnderTest.main is not a function`.

* fix(#4060): export lint-allow-test-rule-refs main() for in-process drive

The "repo baseline passes" subtest previously drove this script via a
spawnSync subprocess with a fixed 30s timeout, racing the script's
real completion time against unbounded CI-load contention -- it has
now died at that race twice (#4060, and once before at a lower
bound). A fixed timeout racing unbounded contention has no value that
is both tight and safe, so raising it again would not fix the
mechanism, only its odds.

Parameterizes main() to accept an explicit argv (defaulting to
process.argv.slice(2) only when omitted, so the CLI entrypoint is
unaffected) and exports it, so the test can call it directly,
in-process -- removing the subprocess and its spawnSync timeout kill
race entirely for this one row.

* fix(#4060): capture stderr too in the in-process repo-baseline subtest

Code-review finding: the in-process rewrite captured only console.log,
but main()'s real failure path throws a bare, messageless ExitError --
all diagnostic detail goes to process.stderr.write. The old
subprocess-based assertion embedded both stdout and stderr in its
failure message; this silently dropped that debuggability. Captures
process.stderr.write the same way (restored in finally) and surfaces
both streams in assertion failure messages and in a wrapped re-thrown
error on an unexpected throw from main().

---------

Co-authored-by: sim <sim@local>
2026-08-29 18:20:24 -04:00
Tom Boucher
6e681cb36c fix(#3901): commit-docs-guard suites isolate children from a global core.hooksPath (#4063)
* test(#3901): the guard suite must survive a developer's global core.hooksPath (failing first)

* fix(#3901): the commit-docs-guard suite isolates children from the host git config

A developer's GLOBAL core.hooksPath (~/.gitconfig) applies to every
fresh repo, so the guard's correct refusal — it will not install a hook
git would never run — failed this suite's beforeEach and 18 tests read
as a commit-docs-guard regression on any machine that centralizes
commit hooks. CI never saw it (runners have no global git config).

before() pins GIT_CONFIG_GLOBAL to an empty file for the whole suite
(runGsdTools and the git helpers propagate process.env to every child).
A LOCAL repo value cannot isolate this: any non-empty local value trips
the same refusal, and an empty local value makes rev-parse --git-path
hooks resolve to ./ instead of .git/hooks. The regression test pins the
seam: the isolation file is armed and empty, a child git resolves no
hooksPath from the host, and the beforeEach enable installed the hook
at the repo-local default path. A child given an explicitly hostile
GIT_CONFIG_GLOBAL still refuses — that is the guard being correct.

* fix(#3901): review fold-ins — isolation hoisted to all three guard suites; changeset

Adversarial review: the pin covered suite A only — the B (#3588 B1-B15)
and D (real git commit wiring) suites run enable against fresh repos
too and failed identically on hostile-global machines; the issue's 18
failures could not have come from A's five tests alone. The pin is
extracted to isolateGlobalGitConfig() and armed in all three suites
(B9's LOCAL hooksPath test is unaffected — the pin neutralizes only the
global scope; D's premise is the hook firing from .git/hooks, which the
isolation preserves).

* chore(#3901): backfill changeset PR number (4063)

* fix(#3901): drop the now-unused before import — max-warnings 0 fails CI

---------

Co-authored-by: sim <sim@local>
2026-08-29 18:20:14 -04:00
Tom Boucher
73919526e9 fix(#3902): packaging guard resolves npm 12's pack --json shape — stays armed on Node 26 (#4064)
* test(#3902): packListToPathSet must resolve npm 12's object-keyed pack --json shape (failing first)

* fix(#3902): resolve both npm pack --json shapes — the packaging guard stays armed on Node 26

npm 12 (bundled with Node 26) emits an object keyed by package name
where npm <=11 emitted an array; parsed[0] was undefined, the before()
hook threw, and all 6 tests in the packaging guard — including the two
'does NOT ship' guards that stop repo-only CI tooling leaking into the
published package — went dark for contributors on Node 26. CONTRIBUTING
pins Node 24 as the floor but requires Node 26 compatibility for code
AND tests; this was the test's parse, not the code.

packListToPathSet now resolves either shape and fails loud on anything
else — a silently-empty set would be the exact dark-guard failure mode
this guard exists to prevent.

* fix(#3902): review fold-in — a multi-key object fails loud

Single-package assumption documented and enforced: npm 12 emits exactly
one key for this repo (no workspaces — verified); if workspaces were
ever adopted, first-key selection would silently validate one package's
file list and recreate the exact silent-disarm this fix kills.

* chore(#3902): changeset fragment (pr number backfilled after PR creation)

* chore(#3902): backfill changeset PR number (4064)

---------

Co-authored-by: sim <sim@local>
2026-08-29 18:17:24 -04:00
Tom Boucher
686c05a740 fix(#4058): raise MAX_ACK_TRAILERS from 64 to 128 (#4059)
* fix(#4058): raise MAX_ACK_TRAILERS from 64 to 128

64 had no principled derivation (no git trailer limit, no CI resource
bound) and a wide-touching maintenance PR can legitimately accumulate
more than 64 distinct emitted-drift acknowledgments, tripping the cap
and failing Required tests even though nothing is actually wrong.

The cap's underlying purpose is unchanged: parseAckTrailers() still
throws rather than truncates on overflow, and still forces pruning of
stale acknowledgments at some ceiling. Only the ceiling moves.

* chore(#4058): backfill changeset PR number (4059)

---------

Co-authored-by: sim <sim@local>
2026-08-29 17:47:34 -04:00
Tom Boucher
90fff40f5a fix(#3900): overlay walker skips non-regular files — a repo-root socket no longer fails 32 tests (#4062)
* test(#3900): a unix socket under the repo root must not break the overlay (failing first)

* fix(#3900): the overlay walker skips non-regular files (sockets, FIFOs, device nodes)

The walker classified every not-a-directory entry as a file, so a unix
socket anywhere under the repo root (an MCP server, a language-server
daemon, a watcher) made copyFileSync throw ENXIO — the reporter
measured 32 local test failures across 6 files, all in before() hooks,
none about the code under test; a FIFO would block until a writer
appeared. The hard-link path does not save it: the overlay sits under
tmpfs, so cross-device EXDEV always falls through to the copy.

One guard at the classification site (srcStat.isFile() before the
copy/link branches) — sockets, FIFOs, and device nodes are not
repository content; CI never saw this because a fresh clone has no
working-tree daemons, which is exactly why it ate local debugging
time. Also adds opts.root (test-only root override, default REPO_ROOT)
so the regression tests build the overlay from a tiny synthetic root
in milliseconds instead of paying a full repo walk.

* chore(#3900): changeset fragment (pr number backfilled after PR creation)

* fix(#3900): review fold-ins — sibling cold-fixture guard, changeset format, root doc

buildColdInstallTree enumerated the live repo's hooks/ with a NAME-only
filter — the same class #3900 documents: a daemon socket or FIFO inside
hooks/ would reach cpSync and throw ENXIO (or block). Dirent type guard
added. The changeset gains the mandated bold user-visible lead, and
opts.root is documented at its definition (mirroring cold-runtime-lib-
fixture's repoRoot convention).

* chore(#3900): backfill changeset PR number (4062)

---------

Co-authored-by: sim <sim@local>
2026-08-29 17:24:40 -04:00
Dennis Kim
8487f0ed42 enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage

- pin configured, absent, and malformed branch-list behavior
- require opposite CLI and execute warning outcomes

* feat(01-01): warn on configured protected branches

- resolve the base branch union configured protected branch names
- expose exact boolean CLI comparison output for workflow callers
- keep execute-phase warning advisory and within its byte budget

* test(01-01): add failing protected branch config coverage

- cover valid list persistence and null unset
- reject hostile shapes while preserving the prior value

* feat(01-01): validate protected branch configuration

- register git.protected_branches as a canonical config key
- require a non-empty array of non-blank branch names

* test(01-02): add failing ship protected-branch controls

- Execute both workflow warning blocks with exact predicate arguments
- Require true and false results to produce opposite warning outcomes
- Preserve the none-strategy feature-branch offer contract

* feat(01-02): warn at ship on protected branches

- Reuse the typed protected-branch predicate in ship preflight
- Keep raw base resolution for PR targeting and advisory branch creation
- Prove execute and ship warning blocks with opposite-result controls

* test(01-02): add failing protected-branch docs parity

- Require the canonical schema key in both English config references
- Pin the non-empty string-array type and absent default
- Require synchronized multi-branch examples and advisory semantics

* feat(01-02): publish protected branch configuration contract

- Document the optional non-empty string-array field in both references
- Explain resolved-base union and absent-field compatibility
- Keep execute and ship warnings advisory under branching_strategy none

* fix(01): CR-01 honor active workstream branch policy

* fix(01): WR-01 assert protected config path selection

* docs: add changeset fragment for #3648

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx

* fix(#3648): resolve base_branch precedence inversion and round-1 findings

Blocker 1/2: production config resolution was flat-first, so a project
that migrated to git.base_branch but still carried a stale flat
base_branch got the old value back. Add base_branch to
normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos
pattern: canonical nested wins) and route readEffectiveGitConfig's
test seam through the same normalization so it can't silently diverge
from production again. Adds a regression test with both keys set that
fails without the fix.

Blocker 3/4/5: restore the handle_branching case-selector prose and
"none" contract sentence that #3389's tests anchor on, and revert the
unrelated prose/comment compaction in the same step — both were
drive-by edits outside #3552's scope.

Also addresses review majors/minors: delete readConfigBaseBranch and
readConfigProtectedBranches (dead in production, only self-tested);
--is-protected now fails closed (reports protected) instead of
silently answering false when the base branch can't be verified;
trim configured protected-branch names; fix HOME-without-USERPROFILE
vacuous isolation on Windows; correct the drift-ack's byte accounting.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N

* test(#3648): add failing legacy-key hoist safety coverage

Round-2 review found normalizeLegacyKeys block 5 records a normalization
carrying the DISCARDED flat value on the canonical-wins branch. Probing
that turned up a second, unreported defect in the same helper shape:
blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no
object guard, so a config whose section key holds a string is spread into
index keys —

  {"git":"main","base_branch":"release"}
    -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}}

The resolved value is accidentally still correct, so nothing fails and no
diagnostic fires. But normalizations.length > 0 sets configDirty, and
config-loader then serializes that shape back into the user's
config.json — a read that silently corrupts config.

The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"}
explicitly; this is the input it would have caught.

Covers both defects across blocks 1 and 5, with object/array/null
negative controls that must stay green in both phases, and a fast-check
property over arbitrary `git` values.

* test(#3648): pin fail-closed handling of malformed protected_branches

Replaces the test that pinned the fail-OPEN behaviour. The old
assertion — ['develop', 42] yields isProtected === false for 'develop' —
locked in the exact failure #3552 exists to close: config-set validation
is bypassable by a direct edit of .planning/config.json, so a user who
believes 'develop' is protected got a silent false and no warning.

It was also inconsistent with the fail-CLOSED direction twelve lines
away, where an unverified base reports protected and writes a
diagnostic. A protection predicate must not have two opposite failure
directions depending on which input is bad (#3648 review Blocker 3).

New coverage: a bad element drops only itself, a non-array contributes
no names, an empty list is well-formed rather than malformed, and
--is-protected surfaces the rejection. Both negative controls — a clean
list reports nothing rejected and writes no diagnostic — must stay green
in either phase, so the reject channel cannot fire unconditionally.

* fix(#3648): drop only invalid protected_branches and report them

Partition git.protected_branches instead of discarding the whole list on
one bad element, and carry the rejections out through
ProtectedBranchStatus so --is-protected can name them on stderr. Valid
names keep protecting; the user finds out the rest were ignored.

A non-array value still contributes no names — a bare string is not a
list of branch names — but is now reported rather than swallowed. An
empty array stays silent: declaring no extra protected branches is a
valid choice, not a misconfiguration.

writeDiagnostic is hoisted out of the unverified-base branch since both
arms now use it.

* test(#3648): prove the predicate diagnostic survives both call sites

The workflow bash stub now emits a stderr diagnostic the way the real
command does, which is what makes a swallowed `2>/dev/null` visible to a
test — previously the stub was silent on stderr, so discarding it changed
no observable behaviour and the call sites could drop the explanation
undetected.

Adds the Minor 2 binding check as well: ship must expose the predicate
result as IS_PROTECTED rather than only echoing a warning, asserted by
running the extracted bash and reading the bound value, not by grepping
the workflow source.

Both tests carry opposite-outcome controls — an empty diagnostic must
leave the text absent, and a false predicate must bind false.

* fix(#3648): surface the predicate diagnostic and bind ship's result

Drop `2>/dev/null` from the --is-protected call at both call sites. The
fail-closed explanation and the new rejected-entry warning both go to
stderr, so discarding it left the user with a bare "protected branch"
warning on a branch that is not protected and no way to tell a real
match from a degraded-git guess. `git branch --show-current` keeps its
own redirect — that one is genuine noise.

ship.md binds IS_PROTECTED and its prose now branches on the variable,
so the following steps have evaluable state instead of having to infer
it from warning text in tool output.

execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth
319 bytes (was 331 before the redirect came out). Baseline re-verified
against the current rebase base by blob id; the ceiling check passes
with 755 bytes of margin.

* test(#3648): restore negative space for the readFile config seam

The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm
it pinned survives verbatim in readEffectiveGitConfig's readFile branch —
the JSON.parse catch, the non-object guard, the git-section object guard,
.trim() and blank-string rejection — and the four surviving readFile
injections were positive-path only. protected_branches was never driven
through this seam at all.

Restores nine cases against the seam, including protected_branches
partitioning, plus a control proving loadConfig still wins when both
seams are supplied.

Records honestly what the suite pins. Mutating the built lib shows
.trim() is KILLED, while the non-object guard and the blank-string
rejection SURVIVE — both are unreachable through this entry point for
the same reasons the deleted suite documented against its own
equivalents: a JSON-parsed non-object carries no relevant own-property
either way, and a blank value is rejected a second time downstream by
the resolver's truthiness check. They stay as defence-in-depth and are
labelled known-unkillable rather than left looking like coverage this
suite does not provide.

* test(#3648): distinguish detached HEAD from a missing branch argument

`args[1] ?? ''` collapsed two different situations into one: a detached
HEAD, where `git branch --show-current` legitimately prints nothing, and
the flag being called with no argument at all. Both answered false, so
the right outcome arrived by an unintentional path and a caller bug was
indistinguishable from normal operation.

Asserts the detached case stays silent and the missing-argument case
reports, with a control that the two diagnostics differ.

* fix(#3648): report a missing --is-protected branch argument

Answer false either way, but say so when the flag arrives with no
argument. A detached HEAD passes an explicit empty string and stays
silent, since that is a normal state rather than a misconfiguration.

* docs(#3648): state exact-name matching and per-entry rejection

isProtected is exact string equality, so a git-flow project must
enumerate every release/* and hotfix/* by name. #3552 only asked for an
integration-branch field, so the implementation satisfies the letter of
the issue while leaving its git-flow motivation partly unserved — say so
where users will meet it rather than leaving them to discover it.

Also documents the Blocker 3 behaviour change: an invalid entry is
ignored with a warning naming it and the remaining names still apply.

Both statements land in docs/CONFIGURATION.md and
gsd-core/references/planning-config.md, and the config-field-docs parity
test asserts each in both so the two cannot drift.

* refactor(#3648): extract isValidProtectedBranches for cross-surface pinning

The `git.protected_branches` check inside `cmdConfigSet` and the resolver's
per-entry filter in `git-base-branch.cts` are deliberately different shapes —
all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot
fail the guard open. Nothing structural keeps their two definitions of "usable
branch name" in step.

Lifting the write-side check into a named, exported predicate lets a property
test ask both surfaces about the same value and assert they agree, which is the
fast-check gap the round-2 review flagged. No behaviour change: the predicate is
the same expression, called from the same place.

* fix(#3648): stop --is-protected rewriting the config it is asking about

`gsd_run query git.base-branch --is-protected` runs on every execute-phase and
every ship. It resolved config through `loadConfig`, whose normalize-then-write
path rewrites `.planning/config.json` whenever any legacy key normalizes — so a
boolean question was silently editing the user's checked-in config. This PR had
widened the trigger by adding a fifth normalization block (top-level
`base_branch` -> `git.base_branch`), making it fire for exactly the projects the
feature targets.

`loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution
is unchanged, only the two write-back side effects are suppressed. The predicate
passes `persist: false`; the ~30 other callers are untouched, so a legacy config
is still migrated by ordinary use.

Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and
reflows whitespace even when the values are equivalent. Three tests, each with
its own control: the end-to-end CLI leaves the file byte-identical while still
answering `true` from the legacy key (proving the config WAS read); an ordinary
persisting load of the same fixture DOES change the bytes (proving the fixture
is live rather than inert); and `persist:false` vs default over one directory
returns deep-equal config while differing on the write. Reverting the one-line
`persist: false` fails the first of those and only that one.

Also from the review:

- `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through
  the same precedence authority production uses". It does not, and cannot — it
  reproduces two of production's steps over a single file. The comment now names
  what the seam covers and what it does NOT (root/workstream deep merge, builtin
  and global defaults, federated merge), and the seam now applies production's
  flat-then-nested lookup so it stops disagreeing about a surviving flat key.

- The missing-argument diagnostic promised "answering false", which the
  fail-closed guard on the same call can contradict by printing `true`. It now
  states what it did with the argument and leaves the answer to stdout.

* test(#3648): re-pin block 5 on #3760's refusal contract

#3767 landed on next while this PR was in review and fixed the non-object
config-section defect properly: a present-but-non-object section now BLOCKS its
own migration — value preserved, no Normalization pushed, refusal reported via
`skipped[]` — rather than being rebuilt from a plain-object view. That supersedes
this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread
but still dropped the section value silently, and which the round-3 review
correctly called out as destruction in place of corruption. The rebase drops that
commit and routes block 5 through the upstream helper.

This file's tests asserted the superseded design, so they are rewritten to pin
block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's
suite was written — against the contract that now governs it: ordinary hoist into
an absent/null/object section, canonical-nested-wins, and refusal for each of
string/number/boolean/array sections with the exact `skipped` entry.

Two controls keep it from passing vacuously: the refusal must be scoped to block
5 (an unrelated block still normalizes in the same call), and a property over
arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually
exclusive per key, that a refusal leaves both the section and the legacy key
untouched, and that a hoist manufactures no index key the input did not carry.

* docs(#3648): correct the Git Query and Config Loader module contracts

CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct
`.planning/config.json` read. Since this PR it is the EFFECTIVE configuration
resolved by the Config Loader — a materially different authority, carrying the
root/workstream deep merge, flat-then-nested lookup and builtin/federated
defaults. The `--is-protected` predicate, `git.protected_branches`, and the two
invariants that distinguish the predicate from the plain query (fails closed on
an unverified base; must not write) were undocumented entirely.

The Config Loader entry now states that loading is not side-effect-free by
default and documents `options.persist`.

docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and
no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write`
was run and produced no diff: the manifest indexes roster NAMES, not row prose,
so a description edit cannot move it.

Also closes the global-defaults minor: `git.protected_branches` is inert in
`~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key
appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is
section-wide and predates this PR, so the fix is to state the scope where users
meet it rather than to quietly extend the resolution set for two new keys.

* fix(#3648): close four defects found by the round-4 external review

Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially
against this branch. Four findings reproduced against source; each is fixed with a
failing-first test and a control, and each fix was verified by reverting it and
watching exactly the intended test fail.

1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was
   only half closed. `loadConfigResolved` re-enters itself with a bare
   `{ workstream: null }` when a workstream has no config.json of its own, and
   that literal discarded every other option — so the recursive pass ran at the
   DEFAULT persistence and rewrote the ROOT config. Reproduced: with
   GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected`
   rewrote `.planning/config.json` despite `persist:false`. Both recursions now
   forward `options` and override only `workstream`; the explicit override still
   wins the hasOwnProperty check, so spreading cannot let `workstreamContext`
   reintroduce a workstream.

2. Both workflow call sites failed OPEN, and aborted under `set -e` (both
   reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty
   string when the query fails, so `[ "$X" = true ]` was simply false: no
   warning, no trace — a silent hole in the guard whose only job is to warn. The
   bare assignment also aborted the step under `set -e`. Both sites now degrade
   VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the
   check did not run. Deliberately not fail-closed — claiming "protected" on no
   evidence would warn on every branch whenever gsd-tools is unavailable.

3. `isValidProtectedBranches` and the resolver disagreed on a sparse array
   (antigravity). `.every()` skips holes; the resolver's `for...of` yields
   `undefined` for them, so `["main", , "develop"]` was accepted by config-set
   and rejected by the resolver. The cross-surface property passed only because
   `fc.array` cannot generate a hole. The predicate now indexes, and the
   generator punches holes so that axis is actually falsifiable. JSON cannot
   express a hole, so this is unreachable in production — but two definitions of
   one predicate must not contradict each other.

4. A top-level `protected_branches` silently outranked `git.protected_branches`
   (antigravity). Routing the key through `get(key, {section, field})` gave it
   flat-then-nested precedence, which is back-compat for keys
   `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has
   no legacy form, so that invented an undocumented alias. It now resolves
   nested-only through a new `getNested`, in production and in the test seam.
   `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's
   refusal path can leave behind — and a control pins that distinction.

Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed
only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`);
a git command that runs and exits non-zero counts as a clean negative, so a cwd
that is not a repository answers `false`, not `true`. Verified pre-existing on
next @ 738f42f4, so the documentation was over-claiming rather than the code
regressing — but an over-broad contract is exactly what the module docs must not
carry.

Both workflow byte figures re-derived after the call-site change:
execute-phase.md 92356 -> 92865 (+509), ship.md 36784 -> 37227 (+443).

* test(#3648): pin git config read parity

* docs(#3648): document git query contracts

* fix(#3648): expose protected branch default

* test(#3648): snapshot planning tree for read-only query

* test(#3648): pin planning snapshot stray-write detection

* fix(#3648): resolve merge conflict from #3078's ack-fragment sweep

next swept the fully-spent 2818/3003 ack fragments this branch had
appended to (#3078, a84f7563). Rebased onto upstream/next and took
the deletions on both, then moved the #3552 append into a new
fragment of its own.

Rebasing onto the current base also left execute-phase.md only 34
bytes under the frozen ADR-857 Phase 6 margin ceiling (93400 bytes) —
intervening next PRs consumed the rest while this PR was in review.
Extracted the "none" arm's protected-branch-warning bash block into
gsd-core/workflows/execute-phase/steps/protected-branch.md (content
unchanged, matching the existing steps/ extraction pattern used
elsewhere in this file) so the inline growth is a one-line pointer
instead of the full block. 93366 -> 93385 bytes (+19), 15 bytes
inside the ceiling.

* fix(#3648): drop stale ack entry for the new step file

The extracted execute-phase/steps/protected-branch.md needed no
acknowledgment of its own — the differential-attribution check flagged
the entry as stale once the build ran, so removed it and kept the two
growth entries (execute-phase.md, ship.md) that actually needed one.

* fix(#3648): follow the step-file reference in the bash-extraction test helper

extractProtectedBranchWarningBash() read the "none" arm's bash block
directly out of execute-phase.md. That block now lives in
execute-phase/steps/protected-branch.md (byte-ceiling extraction);
the helper follows the step-file reference and extracts from there
when no inline block is found, so the three execute-phase tests that
execute this bash for real keep exercising the actual behavior.

* fix(#3648): regenerate INVENTORY-MANIFEST.json and satisfy the CRLF-fragile lint rule

- gen-inventory-manifest.cjs --write to pick up the new
  execute-phase/steps/protected-branch.md entry (already covered by
  docs/INVENTORY.md's generic workflow_steps wildcard row, so no
  INVENTORY.md edit is needed).
- Reworked the step-file-reference lookup in
  extractProtectedBranchWarningBash() to avoid a bare-\n regex split
  on file content (local/no-crlf-fragile-split), using the same
  line-array scan the function already uses elsewhere.

* fix(#3648): regenerate golden install-tree fixtures for the new step file

npm run gen:install-tree, adding gsd-core/workflows/execute-phase/
steps/protected-branch.md to all 19 runtime install-tree fixtures.
CI's tests/golden-install-tree.test.cjs caught this on push — I'd
verified the differential-attribution and INVENTORY-MANIFEST checks
but missed this separate golden-fixture check for the new file.

* fix(#3648): add the canonical gsd_run preamble to the new step file

CI's runtime-launcher-parity suite requires exactly one canonical
resolver preamble in every workflow .md that calls gsd_run. The
inline "none"-arm block never needed one (execute-phase.md already
carried a preamble elsewhere in the same file), but the extracted
execute-phase/steps/protected-branch.md is now its own file with no
preamble of its own. Ran node scripts/sync-runtime-launcher.cjs to
insert it (execute-phase.md itself is untouched — still 93385 bytes,
inside the ADR-857 ceiling).

That preamble defines its own gsd_run(), which shadows the mock
tests/git-base-branch.test.cjs injects for the three #3648 tests that
execute this bash for real — without stripping it, those tests reached
the real gsd-tools.cjs on the machine running them instead of the
test's fixture. Preamble correctness is already covered by
tests/runtime-launcher-parity.test.cjs, so extractProtectedBranchWarningBash()
now strips the preamble line before handing the block to the harness;
it only needs to exercise the #3552 warning logic.

* fix(#3552): address PR 3648 review feedback on protected branch warnings

- Fix execute-phase handle_branching branching_strategy=none instruction
  to "Read and execute execute-phase/steps/protected-branch.md"
- Use io.error(..., ERROR_REASON.USAGE) for cmdGitBaseBranch usage errors
- Align git.protected_branches schema default to (none) without fallback []
- Relocate CONTEXT.md forward-referencing sentence into module body
- Sanitize control and ANSI characters in renderRejected diagnostics
- Clean up out-of-scope whitespace hunks in gsd-tools.cjs

Emitted-Drift-Ack-Growth: execute-phase.md — #3552: execute-phase handle_branching adds a pointer to execute-phase/steps/protected-branch.md for branching_strategy=none so the protected-branch check executes while keeping execute-phase.md within the ADR-857 Phase 6 margin ceiling (93400 bytes). 93392 bytes, 8 bytes inside the ceiling.
Emitted-Drift-Ack-Growth: ship.md — #3552: ship preflight step 3 now asks the same typed git.base-branch --is-protected predicate as execute-phase, binding IS_PROTECTED and warning without refusing execution or blocking the branching_strategy=none feature-branch offer; it degrades visibly (rather than silently reading an empty result as "not protected") when the query itself fails to run. 36841 bytes, well inside the XL cap (98304, tests/workflow-size-budget.test.cjs).

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-29 17:00:52 -04:00
0xdhx
472f585f7c fix(#3726)!: require --confirm before milestone complete mutates (#3774)
* fix(#3726): require --confirm before milestone complete mutates

`milestone complete <version>` is a one-way door — ROADMAP.md and
REQUIREMENTS.md archived, every phase directory in the milestone MOVED,
STATE.md rewritten — and ran unconditionally on first invocation through
every invocation path, including `query milestone.complete <version>`,
whose `query` meta-prefix reads as a read-only namespace but performs no
filtering (#167's invocation-compatibility shim + #3243's dotted-form
normalization).

The gate lives on the destructive command itself, not on the `query`
prefix (the prefix is an intentional invocation mechanism, not a
permission boundary — restricting it would break dozens of shipped
workflow callers). Without --confirm and without --dry-run the command
now refuses via error() before reading anything beyond its arg checks,
so an unconfirmed invocation is a guaranteed no-op on disk. --dry-run
still previews with no confirmation needed and is now documented in the
usage block (it was only documented for the sibling archive-quick).
--force keeps its narrow meaning — bypassing the TRUNCATED-scope and
unstarted-phase guards — and does not double as the mutation opt-in.
--confirm follows the existing `phases clear --confirm` idiom in the
same module.

complete-milestone.md's two invocations pass --confirm (the workflow has
gathered explicit user intent by that step). Existing tests get
--confirm appended — pre-change behavior is exactly confirmed behavior —
and a #3726 regression block covers: refusal + full-tree byte-identity
on both invocation forms, --force not satisfying the gate, --dry-run
still passing without confirmation, and --confirm proceeding. The
refusal tests fail against pre-fix code (negative control run).

Fixes #3726

* docs(#3726): document the --confirm requirement in CLI-TOOLS and COMMANDS

Cross-AI review of the fix diff (codex, pre-create) caught three shipped
doc sites still instructing the now-refused bare invocation: the
CLI-TOOLS.md milestone-complete synopsis + flag table, and COMMANDS.md's
two guard-override instructions (`--force` alone now refuses without
--confirm). Localized CLI-TOOLS copies already lag the English synopsis
(no --force/--dry-run either) and follow the translation pipeline, not
this fix.

* chore(#3726): set changeset fragment pr to 3774

* test(#3726): confirm-gate CI repairs — QA scenario caller + growth ack

Two CI reds from the --confirm gate, both this branch's own misses:

- tests/qa/scenarios/milestone-rollover.json invoked `milestone complete
  1.0 --force` as a JSON arg-array fixture — a caller shape the test
  sweep (which grepped runGsdTools/runSdkQuery in tests/*.cjs) never
  enumerated. Adds --confirm; the scenario's boundary-crossing contract
  is otherwise untouched.
- complete-milestone.md's +420-byte --confirm note trips the
  emitted-attribution growth ratchet. Acknowledged as a #3726 append to
  the existing complete-milestone.md entry in
  3409-unreachable-guard-arms.json (two ack sources may never name the
  same path, per that fragment's own precedent).

Local: lint-emitted-drift-ack ok; loop-walk.qa 115/115 green sandboxed.

* docs(#3726): CLI-TOOLS.md guard-override sentences say --force --confirm

Review Major 1: the truncated-window and unstarted-phase guard paragraphs
still told the reader to "Pass `--force` to override", which now refuses
(--force alone does not satisfy the confirmation gate), while the flag
table 470 lines later said the opposite. Mirror the docs/COMMANDS.md pair
so the file no longer contradicts itself.

* docs(#3726): synopsis renders --confirm and --dry-run as alternatives

Review Nit 1: `milestone complete <version> --confirm [--dry-run]` read as
"a dry run still needs --confirm", the opposite of AC 3. Render the pair
as `(--confirm | --dry-run)` in the CLI-TOOLS.md synopsis and the usage
docblock, and let the flag rows carry the rule.

* test(#3726): pass --confirm in base-added milestone fixtures; re-file the growth ack

Rebase onto next (26 commits) surfaced three tests the gate now refuses:
the #3685 write-flag contract pair in tests/milestone.test.cjs and the
`milestone complete` boundary fixture in tests/state-contract.test.cjs
all invoke the command bare. Each now passes --confirm (a mutating run is
exactly what they assert on).

The +420 byte complete-milestone.md growth ack rode on
3409-unreachable-guard-arms.json, which #3078 swept from next as fully
spent — hence the modify/delete conflict. Re-filed under a fresh fragment
named for this issue, never resurrecting the swept one.

* test(#3726): pin the present-but-falsy arm of the confirmation gate

Review Minor 1: the boundary triple covered absent and present but not
present-but-falsy. The gate is an exact-token match, so --confirm=false
and --confirm=0 refuse today — pinned (canonical + query forms, whole
.planning/ tree byte-identical) so a future `=`-aware or prefix-matching
parser cannot silently turn --confirm=false into a confirmed run of an
irreversible command.

* test(#3726): drop --confirm from dry-run-only invocations

Review Nit 2: --confirm was mass-appended to 14 pre-existing --dry-run
invocations that never needed it, so each stopped standing as incidental
proof that a preview needs no confirmation. Reverted to the pre-PR form;
the dedicated AC-3 test carries the explicit assertion.

* docs(#3726): sync the localized CLI-TOOLS synopsis with the confirm gate

REQ-I18N-02 (docs/features/internationalized-documentation.md) requires
translations to stay synchronized with the English source. The four
localized CLI-TOOLS.md guides still advertised a bare
`milestone complete <version>`, which now exits 1. Render the English
synopsis verbatim — `(--confirm | --dry-run)` plus the `[--force]` and
`[--archive-quick]` flags the translations had also fallen behind on.

* test(#3726): drop --confirm from the remaining preview-only invocations

Round 2 reverted the --confirm appends on --dry-run-only invocations in
tests/milestone.test.cjs, but four more sat in two files the sweep missed:
tests/milestone-archive.test.cjs (three) and
tests/milestone-window-single-owner.test.cjs (one).

Each is a preview run whose whole purpose is to document that a preview
mutates nothing, so `--dry-run ... --confirm` contradicted the semantics
the test exists to pin. Dropping the token restores each as incidental
proof that a preview needs no confirmation; the dedicated AC-3 test keeps
the explicit assertion.

No assertion added, relaxed, or removed — the change is four tokens.

* chore(#3726): migrate the emitted-drift ack from a fragment to a commit trailer

#3954 (ADR-3942) moved emitted-drift acknowledgments out of
tests/emitted-drift-acks/ and into git commit trailers, and the fragment
directory no longer exists on next. The reason this PR's fragment carried
moves verbatim into the Emitted-Drift-Ack-Growth trailer on this commit;
the fragment file is removed rather than resurrected.

Emitted-Drift-Ack-Growth: complete-milestone.md — #3726: +420 bytes (40186 -> 40606). The archive_milestone step's two `milestone complete` invocations now pass the required --confirm flag (the command refuses to mutate without it — the archive is irreversible), with a note explaining the flag and pointing at --dry-run for previews. Deliberate runtime-loaded workflow text for the new gate, not converter drift.

* fix(#3726): name --confirm in the version-required refusal

The documented arg-discovery path (gsd-tools.cjs top-level usage: invoke
the command without args and the error lists what is required) stopped at
`version required for milestone complete (e.g., v1.0)` — one required
argument short. Discovering --confirm took a second round trip through the
gate. The refusal now reads `… — and --confirm to mutate`, pinned by a test
that also asserts the version-less invocation leaves .planning/ untouched.

* test(#3726): pin the milestone complete docs against a silent regression

The changeset is `type: Fixed`, which the docs-required lint exempts, so
nothing in CI would notice a later edit that reinstated the bare-`--force`
override prose or dropped `--confirm` from the synopsis. Four tests in
tests/milestone.test.cjs now pin: the synopsis line in docs/CLI-TOOLS.md
and its four localized mirrors; the `--confirm` flag row; both
guard-override instructions in docs/CLI-TOOLS.md and docs/COMMANDS.md,
by guard name (a substring match on each instruction's `--force
--confirm` text); and — as an identity ratchet over the
milestone-complete sections — every `--force` sentence or clause that
lacks `--confirm`, so a new bare instruction in its own sentence or
clause fails whatever its wording. Named residual: a bare instruction
spliced into the same clause as a compliant one coalesces with it and
passes the ratchet; the by-name pins are what keep the four known
instructions from losing the pairing that way. The file is registered
in scripts/docs-guard-registry.cjs so the pin runs on the PR that
changes those docs, not only after merge.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-29 17:00:45 -04:00
Tom Boucher
370cfc6680 enhance(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% (#4043)
* feat(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90%

Adds two new mechanisms plus an audit-coverage extension:

- scripts/lib/ci-job-timing.cjs: shared elapsed-vs-cap arithmetic
- scripts/ci-check-job-near-cap.cjs: in-job advisory near-cap check,
  wired into test/test-full/mutate/smoke as each job's last step
- scripts/ci-timeout-report.cjs + .github/workflows/ci-timeout-report.yml:
  scheduled REST-API poll that appends new records to
  tests/ci-timeout-budget-history.jsonl and opens a small data-only PR
- tests/ci-test-job-timeout-budget.test.cjs: extended to cover mutate
  (mutation.yml) and smoke (install-smoke.yml), which previously had no
  headroom-factor gate coverage at all

Does not change any timeout-minutes value, shard composition, or shard-1
contents — those stay maintainer policy calls per the issue's own scope.

* fix(#4036): address two-orthogonal-review findings

- Parity tests guarding the two hand-duplicated literals this design
  cannot single-source through GH Actions YAML: CI_JOB_TIMEOUT_MINUTES
  vs each job's own timeout-minutes, and ci-timeout-report.cjs's
  JOB_RULES name-prefixes vs each job's actual name: template.
- Thread run.event through as runEvent on every persisted record, so
  PR-context and push-context install-smoke timings (genuinely
  different matrix shape) are distinguishable in the history rather
  than silently conflated under one job name.
- Replace the Windows near-cap start-time step's ambiguous PowerShell
  +/>> precedence with GitHub's documented string-interpolation form.
- Move github.run_id out of direct ${{ }} shell interpolation into an
  env: var in the new scheduled workflow, per this repo's own
  expression-injection-safe convention.

* test(#4036): regenerate golden install-tree fixtures for scripts/lib/ci-job-timing.cjs

npm run gen:install-tree — scripts/ ships wholesale into the installed
package (per ADR/known-defect precedent from #4012's own PR history: a
new scripts/lib/*.cjs file needs its golden entry regenerated or every
runtime's install-tree test fails). Confirmed via gsd-test: this was the
sole cause of the first real verification run's 25 failures (all in
tests/golden-install-tree.test.cjs, one per runtime). Top-level
scripts/*.cjs files (ci-check-job-near-cap.cjs, ci-timeout-report.cjs)
are not individually tracked in these fixtures — consistent with every
other existing top-level scripts/*.cjs file, so no entry was expected
or added for those two.

* fix(#4036): register new lib file with installer, fix H1 shell policy

- bin/install.js: add ci-job-timing.cjs to GSD_SCRIPTS_LIB_FILES (a
  hand-maintained registry, not generated — tests/install.test.cjs
  asserts every scripts/lib/ file is enumerated here)
- test.yml: replace the two OS-conditional "Record job start time"
  step pairs (test + test-full jobs) with a single unconditional
  `node -e` step. The prior pair's Windows variant declared an
  explicit shell: pwsh, which scripts/workflow-policy.cjs's H1 checker
  statically flags against every OS a job's matrix can realize,
  independent of the step's own if: gate. A single Node one-liner
  needs no shell override at all — it's syntactically valid and
  behaves identically under bash, zsh, and pwsh — which is both H1
  compliant and removes the last OS-specific shell syntax from this
  change entirely.

Both defects were found by a real gsd-test run, not local gates —
lint:ci and build:lib were clean throughout because neither the
scripts/lib/ install-manifest parity check nor the H1 shell-policy
baseline runs as part of lint:ci; both are gsd-test-only suites.

* docs(#4036): how-to for reading CI timeout budget signals

The phase-gate docs check correctly flagged the enablement sequence as
3 real steps (read the near-cap warning, find the accumulated trend
file, pick the right maintainer lever) — a reference table can't carry
a sequence. Adds docs/how-to/read-ci-timeout-signals.md, indexed from
docs/README.md.

* chore(#4036): backfill changeset PR number (4043)

---------

Co-authored-by: sim <sim@local>
2026-08-29 16:13:15 -04:00
Tom Boucher
b431ae9f0d fix(#3898): a separator-shaped line in ## Gaps is skipped, not made an entry (#4057)
* test(#3898): a spaced-hyphen thematic break in ## Gaps is not an entry (failing first)

* fix(#3898): a separator-shaped line in ## Gaps is skipped, not made an entry

splitGapsEntriesCore's opener regex (/^(\s*)-\s/) matched a spaced
hyphen thematic break, fabricating a gap named '- -' with result
'unknown' — surfaced by audit-uat as an outstanding finding that
cannot be cleared by editing any entry, because there is no entry,
only the separator the author wrote for readability. Five shapes were
affected (- - -, -, and wider/indented variants); the unaffected ones
(---, ----, * * *, ___) were safe only by accident — the path never
matched them, not because it understood breaks.

A line whose content after the opening marker is solely hyphens and
whitespace (with at least one further hyphen) is now skipped entirely —
neither an opener nor a continuation. Deliberately option 2 from the
issue, not a full thematic-break concept: a break does not close the
Gaps list, entries after it keep parsing, and the deliberately-frozen
byte-for-byte Gaps behavior changes ONLY for documents carrying such a
separator (which previously produced a phantom). A real entry whose
truth begins with a hyphen (- truth: "-5 error budget...") is untouched
— its remainder contains non-hyphen characters.

* fix(#3898): review fold-ins — span-contiguous skip, property coverage

The skip is narrowed to where the phantom came from: a separator-shaped
line BETWEEN entries (nothing open, or it would open a top-level entry).
One landing strictly inside a live entry (indent > baseIndent) folds
back as a continuation line, so entry lines and the GapsEntrySpan agree
byte-for-byte — the span invariant and the #3805 ack writer's identity
re-verification both hold (the review traced the unconditional skip to
a match_verification_failed refusal in that corner). Adds the parser-
convention property test (arbitrary hyphen counts/indents/spacings) and
a span-contiguity pin.

* chore(#3898): changeset fragment (pr number backfilled after PR creation)

* chore(#3898): backfill changeset PR number (4057)

---------

Co-authored-by: sim <sim@local>
2026-08-29 15:51:58 -04:00
Tom Boucher
44ddc6dc46 fix(#3865): init.* phase queries accept --phase <N> as the positional alias (#4054)
* test(#3865): init.* phase queries accept --phase as positional alias (failing first)

* fix(#3865): init.* phase queries accept --phase <N> as the positional alias

The phase-taking init.* queries read their phase token at args[2]
blindly: '--phase 60' made the literal '--phase' the phase, which the
locator can never resolve — the reported incident answered a well-formed
phase_found:false, plan_count:0 with exit 0 for a phase holding seven
committed plans (ADR-3473 §8.4's later strict validation turned the
paired form into a usage error instead, still not the alias).

normalizePhaseAlias (shared by all eight phase-taking handlers — the
issue's four plus phase-op, review, discuss-phase-assumptions, todos,
the same class) rewrites '--phase N'/'--phase=N' into the caller-owned
positional slot before flag parsing, so handlers and strict validation
see exactly the argv the positional form produces. A valueless --phase
is a usage error naming the flag. Any other flag-shaped args[2] now
resolves to undefined (the commands' designed use-the-current-phase
input) instead of passing flag text down as a phase name.

isFlagToken exported from command-arg-projection (single owner of the
flag-shape predicate).

* fix(#3865): review fold-ins — honest no-position-given comment, fail-closed throws

The helper's comment claimed flag-shaped args[2] resolves to 'the
commands' designed use-the-current-phase input' — no such cross-module
behavior exists: execute-phase/plan-phase/verify-work usage-error
'phase required', the find-based queries answer phase_found:false, and
todos drops its area filter. Reworded to what actually happens, so the
contract-grade comment cannot mislead a future edit. Adds the
parseNamedArgsOrExit-style throw after each error() call as a
fail-closed backstop against a returning fail().

* chore(#3865): changeset fragment (pr number backfilled after PR creation)

* chore(#3865): backfill changeset PR number (4054)

---------

Co-authored-by: sim <sim@local>
2026-08-29 15:22:34 -04:00
Tom Boucher
d9e906a744 fix(#3864): smart-entry classify() matches the verif* status stem (#4052)
* test(#3864): smart-entry classify must match the verif* status stem (failing first)

* fix(#3864): classify() matches the verif* status stem, aligning with normalizeStateStatus

"verified"/"verification" contain no "verify" substring, so the
exact-word branch never matched: a STATE.md declaring status: verified
fell through to situation "unknown" (or idle-stranded on a clean tree
with unpushed commits — differently wrong, which made it look
intermittent). state-document's normalizeStateStatus already matches
any verif* stem; the classifier now uses the same stem, so the two
owners agree on the invariant. verify_failed is tested earlier in
classify(), so failed verification still wins. Negative controls
(completed/executing/planning/paused) unchanged and pinned.

* test(#3864): review fold-in — the real handler-written verification status pinned both ways

'Phase complete — ready for verification' (state-document.cts's own
Status default) carries the verif stem but names completion: at 5/5 it
must stay complete (isComplete beats the stem), mid-project it must
route to verify-pending (pre-fix it fell through to unknown/idle-
stranded — a real behavior change beyond the literal verified repro,
sanctioned by the issue's own verification→verify-pending table).

* chore(#3864): changeset fragment (pr number backfilled after PR creation)

* chore(#3864): backfill changeset PR number (4052)

---------

Co-authored-by: sim <sim@local>
2026-08-29 15:07:14 -04:00
Tom Boucher
4048ba2e80 fix(#3860): Quick Tasks lookup accepts milestone-suffixed headings; schema-aware section selection (#4050)
* test(#3860): quick-tasks heading tolerance for milestone suffixes (failing first)

* fix(#3860): prefix-match Quick Tasks heading + schema-aware section selection

The exact-anchored predicate (^quick tasks completed$) never matched a
milestone-suffixed heading ('Quick Tasks Completed (v1.1+)'), so every
/gsd-fast append AND every reset failed with QUICK_TASKS_SECTION_ABSENT
before the columns were ever looked at — the reported table was already
canonical. A heading is a section label, not data: the predicate is now
prefix-anchored with a word boundary (Completedness still excluded),
hoisted to one shared isQuickTasksHeading beside QUICK_TASKS_SECTION_ABSENT
so the two call sites cannot drift.

Section selection is now schema-aware (the issue's deliberate-choice
ask): among ALL matching sections, the first whose table parses with a
recognized Quick Tasks schema wins — a legacy (v1.0) table first in
document order no longer shadows a usable (v1.1+) one below it. When no
section is usable, the FIRST is returned so the error names the real
problem (unrecognized schema) instead of a false 'no section'.

Fixes one over-assertion in the new tests (v1.1+'s own next ordinal IS 2;
pins v1.0 rows byte-identical instead).

* fix(#3860): review fold-ins — level-bounded bodies, splice guard, convention pin

Adversarial review caught that collectSections ends a candidate's body
only at the next MATCHING heading: an intervening '## Deferred Items'
table (canonical STATE.md layout — templates/state.md puts it after the
Quick Tasks section) would be swallowed into the body, and the append's
last-table-line scan would splice the quick-task row into that WRONG
table — silent corruption on the primary /gsd-fast path. Each candidate
is now re-collected through collectSection with an offset-precise
predicate, restoring the pre-#3860 level-bounded stop (next same-or-
higher heading) for both probing and splicing.

Also pins the both-schema-valid tie-break (document order, newest-on-top
— the layout the issue itself demonstrates; this layer has no
active-milestone signal) with a test, guards the Deferred Items layout
with a dedicated splice-target test, and updates resetQuickTaskRows's
doc comment to name the new pipeline.

* chore(#3860): changeset fragment (pr number backfilled after PR creation)

* chore(#3860): backfill changeset PR number (4050)

---------

Co-authored-by: sim <sim@local>
2026-08-29 14:52:24 -04:00
Tom Boucher
fd63889d1f fix(#3854): write normalization preserves tight multi-line lists (#4049)
* test(#3854): write normalization must preserve tight multi-line lists (failing first)

* fix(#3854): no blank before a bullet whose previous line is an indented continuation

_normalizeMd's 'separate a list from a preceding paragraph' rule inserted
a blank before any bullet whose previous line wasn't a bullet — but an
INDENTED CONTINUATION of the previous multi-line item also isn't a
bullet. Every .md write (phase.complete in the report, but any write
through platformWriteSync) therefore converted tight lists to loose
ones: +61 blank lines on the reporter's 1015-line ROADMAP, one before
each bullet following a wrapped item. Tight and loose lists render
differently, so this was a rendering change plus misleading diff noise;
one-shot (idempotent afterwards), which is why integrity checks on
headings/content passed.

The guard is the mirror image of the after-a-bullet rule two lines
below, which already excludes indented next lines. Paragraph→list and
heading→list separations — the rule's purpose — are pinned unchanged by
the new suite.

* fix(#3854): review fold-ins — ceiling tracks next's 281 + this branch's marker (282), header/require nits

The ceiling is not ratcheted but must track the tree: origin/next raised
it to 281 (sibling branch's marker file); this tree adds one more
(shell-command-projection-md-normalize), so 282/282. Also fixes the test
header's stale pre-rename filename and hoists the inline require to the
file's single import.

* chore(#3854): changeset fragment (pr number backfilled after PR creation)

* chore(#3854): backfill changeset PR number (4049)

---------

Co-authored-by: sim <sim@local>
2026-08-29 14:37:21 -04:00
Tom Boucher
529480b4a5 fix(#3895): delete the mempalace-curator's model frontmatter pin — the fleet's only hardcoded model (#4048)
* test(#3895): no shipped agent may hardcode a model frontmatter pin (failing first)

* fix(#3895): delete the mempalace-curator's model frontmatter pin — the fleet's only hardcoded model

Exactly one of the 34 shipped agents carried 'model: sonnet' in its
frontmatter; every other agent resolves through the model-profile
system. The ship:post dispatch (#2684) resolves per-hook and — per
#2517 — deliberately OMITS model= on inherit so the agent inherits the
orchestrator's model; the frontmatter pin intercepted that inherit
case, silently forcing sonnet where all 33 siblings would inherit, and
operators could not durably remove it (install rewrites live copies
wholesale).

Deleting the line changes nothing for default profiles — the catalog
entry (model-catalog.json agents.gsd-mempalace-curator:
golden/balanced sonnet, budget haiku) preserves today's behavior —
while restoring model_overrides and inherit authority. Pinned by a new
agent-frontmatter guard: no shipped agent may hardcode a model pin,
and the catalog entry must keep existing so the pin's deletion can
never orphan the agent.

* chore(#3895): changeset fragment (pr number backfilled after PR creation)

* chore(#3895): backfill changeset PR number (4048)

---------

Co-authored-by: sim <sim@local>
2026-08-29 14:11:39 -04:00
Tom Boucher
400db94e02 fix(#3894): quick path honors workflow.research_before_questions; key resolves from global defaults (#4047)
* test(#3894): research_before_questions must resolve globally and order quick.md (failing first)

* fix(#3894): quick path honors workflow.research_before_questions; key resolves from global defaults

Two layers, one key. The quick workflow ran its discussion phase
before its research phase unconditionally — neither quick.md nor its
steps ever read workflow.research_before_questions, though the key is
documented, schema-registered, /gsd-settings-writable, and honored by
/gsd-discuss-phase and /gsd-new-project. A gray-area answer given
without research is then written to <quick_id>-CONTEXT.md as a locked
decision downstream agents are told not to revisit — an evidence-free
choice made unfalsifiable (the reporter's #3714 misresolution).

- quick.md Step 4 now carries the same research-before-questions check
  the two honoring paths make: when enabled, research-phase executes
  before discussion-phase; false/unset keeps the written order. Both
  sections stay section-manifest gated.
- src/config-loader.cts forwarded workflow.post_planning_gaps from
  ~/.gsd/defaults.json but silently dropped this key — same file, same
  nesting, one resolved and one didn't. Now forwarded with the same
  flat + nested-alias fallback shape, added to the resolution-keys
  lockstep canary and the #3532 shadowed-warning set (nested alias
  reporting generalized over both keys).

Emitted-Drift-Ack-Growth: quick.md — #3894: +Step 4 ordering rule (the research-before-questions check the discuss-phase and new-project paths already make); a real behavioral gate, not incidental bloat.

* fix(#3894): review fold-ins — gate the CONTEXT.md reference, colon slash-forms

- quick/steps/research-phase.md directed the researcher subagent to read
  <quick_id>-CONTEXT.md under DISCUSS_MODE with no existence hedge — but
  under the new ordering (research BEFORE discussion) the file cannot
  exist yet when the researcher is dispatched. The reference now says
  read-only-if-present with the #3894 reason; the alignment purpose
  still applies on the default ordering.
- quick.md's new rule used the hyphen slash forms (/gsd-discuss-phase,
  /gsd-new-project); source artifacts under gsd-core/workflows must
  author the colon form the install-time converters key on — the same
  file already uses /gsd:new-project and /gsd:quick elsewhere.

* docs(#3894): planning-config row names the flat CONFIG_DEFAULTS alias

config-field-docs requires every CONFIG_DEFAULTS key to appear in the
doc; the row documented the canonical namespaced form only. Adds the
same alias sentence post_planning_gaps's row carries, plus the #3894
quick-path note.

* chore(#3894): changeset fragment (pr number backfilled after PR creation)

* chore(#3894): backfill changeset PR number (4047)

---------

Co-authored-by: sim <sim@local>
2026-08-29 13:51:42 -04:00
Tom Boucher
192eb1dfbd fix(#3886): git commit timeout reported as commit_timeout; stale lock surfaced; 30s band (#4046)
* test(#3886): a timed-out git commit reports commit_timeout, not commit_failed (failing first)

* fix(#3886): git commit timeout reported as commit_timeout; 30s band; stale-lock surfaced

cmdCommit's git commit invocation did not distinguish a spawnSync
timeout from a real non-zero exit (#2608 fixed this for the staging
loop only): a slow pre-commit hook crossing the 10s cap was
SIGTERM'd mid-hook and reported as reason commit_failed with whatever
partial stderr git had flushed (in the reporter's case an incidental
CRLF warning), while the kill left a stale .git/index.lock blocking
the next attempt.

All three commit sites now check isSpawnTimeout before the
nothing-to-commit/ordinary-failure branches: cmdCommit reports
reason commit_timeout + timed_out:true and names the stale lock's
path (surfaced, not auto-deleted — deleting a lock a live git holds
is destructive; the caller recovers deliberately); the subrepo
counterparts do the same within their per-repo result / rollback
error. The commit calls also move to the 30s band the push call
already uses — husky+lint-staged alone idles ~4s on Windows before
any task runs.

* fix(#3886): review fold-ins — git-path lock resolution, shared band constant, executor contract row, precedence pin

- The stale-lock path is resolved via git rev-parse --git-path
  index.lock, never a literal .git/index.lock join (#3588 row 8's
  class: a linked worktree's .git is a FILE, so the literal path cannot
  exist there while the real lock — under <gitdir>/worktrees/<name>/ —
  blocks the next commit; this repo leans on linked worktrees).
- COMMIT_TIMEOUT_MS hoisted; all three sites and their messages build
  from it (the subrepo variant also regains the stdout fallback the
  primary site had).
- agents/gsd-executor.md's commit-result contract gains the
  commit_timeout row with the OPPOSITE retry advice from
  staging_timeout (remove the stale lock, then retry once) — an
  executor matching the doc previously had no handling for the new
  reason.
- Precedence pin: a timeout whose partial output contains 'nothing to
  commit' must still read as a timeout (branch-reorder mutant).

Emitted-Drift-Ack-Growth: gsd-executor.md — #3886: +commit_timeout row to the commit-result contract with the retry guidance OPPOSITE staging_timeout's (remove the stale lock, then retry once); the executor previously had no handling for the new reason.

* chore(#3886): changeset fragment (pr number backfilled after PR creation)

* chore(#3886): backfill changeset PR number (4046)

---------

Co-authored-by: sim <sim@local>
2026-08-29 13:27:14 -04:00
Tom Boucher
0bf778c352 fix(#3849): phase allocation counts numbers held by sibling git worktrees (#4042)
* test(#3849): phase allocation must skip numbers held by sibling worktrees (failing first)

* fix(#3849): phase allocation counts numbers held by sibling git worktrees

Both allocators (cmdPhaseAdd, cmdPhaseAddBatch) chose max+1 over numbers
gathered from ONE checkout — headers, bullets (add only), on-disk dirs.
Every sibling git worktree carries its own .planning/ on its own branch,
so a phase minted there was invisible and the same number was allocated
twice (the reported incident: two Phase 441s, one with six written
plans, surfaced a day late by human memory).

New shared horizon collectSiblingWorktreePhaseNums: one
git worktree list --porcelain, then per sibling — phase-dir names (the
cheap scan that would have caught the incident) and the WHOLE sibling
ROADMAP.md headers (a row can predate its dir; milestone-scoping would
be wrong — a number used under any milestone on another branch is
taken). Widen, never refuse: unreadable sibling / no .planning / not a
git repo / git unavailable each contribute nothing and allocation is
unchanged. Reuses isSentinelPhaseId and the allocators' own patterns.

Secondary (#1229 never reached batch): cmdPhaseAddBatch now also scans
roadmap bullets — a bullet-only 'Phase N' row was invisible to batch
allocation, exactly the condition #1229 was filed for.

Also fixes the two new tests' result-key access (output.phases, not
output.results).

* fix(#3849): review fold-ins — subprocess band, bounded test git, linked-worktree fixture

- execFileSync options now match the repo's git band (10s window,
  windowsHide, 4MiB maxBuffer) — a spurious 4s timeout silently reverted
  to the pre-fix collision.
- test git helper bounded (15s) per local/no-unbounded-spawn.
- new fixture: allocation FROM a linked worktree counts the main
  checkout — the incident's actual topology direction.
- fixture-setup rmSync carries the sanctioned lint-disable (setup, not
  teardown; cleanup() still owns directory removal).

* fix(#3849): exempt the sibling-worktree scan from the enumeration-drift guard

collectSiblingWorktreePhaseNums reads a SIBLING checkout's phases dir —
a different question from the cwd-scoped listMilestonePhaseDirs the
guard routes everything to (which cannot see another worktree's
.planning at all). Function-scoped, per ADR-3180 Decision 4(a): any
other re-derivation in phase.cts is still caught. The GREEN bench
caught the omission.

* chore(#3849): changeset fragment (pr number backfilled after PR creation)

* chore(#3849): backfill changeset PR number (4042)

---------

Co-authored-by: sim <sim@local>
2026-08-29 11:35:19 -04:00
Tom Boucher
519ac23ebb fix(#3839): hook tables say PreToolUse (validate-commit) and SessionStart (session-state) (#4041)
* test(#3839): docs hook tables must match surface registrations (failing first)

* docs(#3839): hook tables say PreToolUse for validate-commit, SessionStart for session-state

gsd-validate-commit.sh is registered PreToolUse (src/runtime-hooks-surface.cts;
its exit-2 block IS the contract — a post-tool hook cannot prevent a commit)
and gsd-session-state.sh is registered SessionStart (session orientation, not
post-tool tracking). Both rows said PostToolUse in ARCHITECTURE.md and the
three INVENTORY locales; the issue asked for a neighbouring-row scan, which
is how the session-state row was found. All other rows in the four tables
verify against the surface.

* fix(#3839): review fold-ins — 10 more wrong rows in ko-KR/pt-BR/zh-CN, parser authority + drift pins

Adversarial review found the same two wrong rows shipped in five more
files the issue's table missed (ko-KR ARCHITECTURE+INVENTORY, pt-BR
ARCHITECTURE+INVENTORY, zh-CN ARCHITECTURE) — all fixed; DOC_TABLES now
covers all ten shipped tables. The parity parser unioned only the Kimi
mirror list, silently exempting agent-isolation-guard (registered via
the dynamic preToolEvent push): probes are now parsed too, with bare
hook names resolved against hooks/ ground truth and dynamic event
variables resolved to their canonical (non-Gemini) events; an exact-set
pin replaces the loose size guard. allow-test-rule marker carries the
issue ref; unverified-ceiling 280→281 (audited: the new marker is
legitimate — the suite reads product docs whose text is the contract).

* fix(#3839): register the hook-table parity suite in the docs-guard lane

The new suite reads ten docs/ paths, so lint-docs-guard-registration
requires it in the docs-guard registry — the first GREEN bench run
caught the omission (the RED run's docs-guard failures were the same
signal, previously misread as marker fallout).

* chore(#3839): changeset fragment (pr number backfilled after PR creation)

* chore(#3839): backfill changeset PR number (4041)

---------

Co-authored-by: sim <sim@local>
2026-08-29 11:11:24 -04:00
Tom Boucher
3c08315a5e fix(#3827): new-mode routing gate — classify approval no longer authorizes scaffold writes (#4037)
* test(#3827): new-mode routing approval gate before roadmapper (failing first)

* fix(#3827): new-mode routing gate — classify approval no longer authorizes scaffold writes

The route_new_mode step delegated to gsd-roadmapper (PROJECT.md,
REQUIREMENTS.md, ROADMAP.md, STATE.md creation + commit) with no gate:
the discovery gate approved classification only, and the zero-conflict
branch said 'proceed to routing silently'. Merge mode previews its diff
and gates via approve-revise-abort; new mode now shows the exact
destinations and requires Create planning setup | Keep synthesized
intel only | Abort, per the skill contract's routing-gate requirement.
The keep-intel-only choice is the analysis-only path the issue asks
for: no roadmapper, no destination writes, intel preserved; finalize
labels it 'new (intel only)' and points at /gsd:new-project instead of
plan-phase. Ambiguous gate answers re-ask once, then treat as Abort —
never infer Create. Zero-conflict wording is mode-aware: silence about
conflicts is not authorization to write.

Emitted-Drift-Ack-Growth: ingest-docs.md — #3827: +routing gate display block, AskUserQuestion contract, three disposition branches (incl. the intel-only no-write path), ambiguous-answer rule, intel-only finalize line, mode-aware zero-conflict pointer; a real behavioral gate, not incidental bloat.

* chore(#3827): changeset fragment (pr number backfilled after PR creation)

* chore(#3827): backfill changeset PR number (4037)

---------

Co-authored-by: sim <sim@local>
2026-08-29 09:41:00 -04:00
Tom Boucher
331747ea99 fix(#3817): count the truncation remainder — display truncates, counting must not (#4034)
* test(#3817): audit-open counts must include the truncation remainder

* fix(#3817): count the truncation remainder — display truncates, counting must not

* chore(#3817): changeset fragment (pr number backfilled after PR creation)

* chore(#3817): backfill changeset PR number (4034)

---------

Co-authored-by: sim <sim@local>
2026-08-29 08:53:21 -04:00
Tom Boucher
ac0eed1267 Merge pull request #4015 from open-gsd/fix/3889-instrument-chunk-timeout 2026-08-29 08:02:14 -04:00
Tom Boucher
213a2fff63 chore(#3813): delete the caller-less listMilestoneArchiveDirs seam; #1883 contract now pins the live path (#4029)
* test(#3813): pin the #1883 unreadable-milestones contract on the live planning-snapshot path

* fix(#3813): delete the caller-less listMilestoneArchiveDirs seam; #1883 contract now pins the live path

* chore(#3813): changeset fragment (pr number backfilled after PR creation)

* chore(#3813): backfill changeset PR number (4029)

* chore(#3813): docs-exempt marker — internal dead-code removal

---------

Co-authored-by: sim <sim@local>
2026-08-29 07:44:22 -04:00
Tom Boucher
b811ea16fc fix(#3807): advance-plan refuses an ambiguous multi-Phase Current Position (#4028)
* test(#3807): advance-plan must refuse an ambiguous multi-entry Current Position

* fix(#3807): refuse an ambiguous multi-Phase Current Position before advancing

* chore(#3807): changeset fragment (pr number backfilled after PR creation)

* chore(#3807): backfill changeset PR number (4028)

---------

Co-authored-by: sim <sim@local>
2026-08-29 03:50:03 -04:00
Tom Boucher
3a4c3cb83e fix(#3805): audit-uat honours the audit_acknowledged marker via the shared predicate (#4025)
* test(#3805): audit-uat must honour the audit_acknowledged marker

* fix(#3805): route audit-uat's UAT and VERIFICATION scans through the shared acknowledged predicate

* chore(#3805): changeset fragment (pr number backfilled after PR creation)

* chore(#3805): backfill changeset PR number (4025)

---------

Co-authored-by: sim <sim@local>
2026-08-29 02:51:54 -04:00
Tom Boucher
f4fefb0bef fix(#3804): audit-uat enumerates all three phase-archive layouts (#4022)
* test(#3804): audit-uat must see all three phase-archive layouts

* fix(#3804): enumerate all three phase-archive layouts (flat, workstream-archived, workstream-active)

* chore(#3804): changeset fragment (pr number backfilled after PR creation)

* chore(#3804): backfill changeset PR number (4022)

---------

Co-authored-by: sim <sim@local>
2026-08-29 01:14:58 -04:00
sim
c1a2c33886 perf(#4012): two subprocess spawns, not five
My instrumentation suite tipped ubuntu shard 1/3 over its 15-minute job cap.
Measured, not guessed: baseline on next (c8f08b61f) passed that shard in
14m21s — 39 seconds of headroom — and the chunk-timeout instrumentation suite
I added costs 31,621ms, about 80% of what was left. The job ran 15m09s and
GitHub reported the cap as a cancel, which failed the required-tests rollup.

Cause is spawn count, not test content. The suite booted scripts/run-tests.cjs
five separate times — three success-path, two timeout-path — and each boot
globs the suite and spawns node --test children before any assertion runs. The
direct-require tests beside it cost milliseconds.

Now two spawns: one success run backing the elapsed-timing, temp-dir-non-leak
and no-EINVAL assertions, and one timeout run backing the in-flight-file
naming and the diagnostic wording. Every assertion is kept verbatim; only the
per-assertion subprocess boot is gone. Modelling per-boot overhead from the
measured total puts the saving around 18s, but that is an estimate derived
from one data point, not a measurement — the real number comes from CI.

The chunk timeout stays at 2000ms. It is already the lowest value used
anywhere in this file, and lowering it further would race a loaded CI box that
has to boot node --test, register the hang, and observe the kill inside the
window.

Worth recording separately: that lane has 39 seconds of margin on next, so it
is one cliff away from this happening to whoever adds the next test. The
per-chunk timing this PR adds is what made the attribution possible at all —
chunk 1/5 alone is 263s of a 900s budget, which was previously invisible.

Verification runs on the remote runner.

Refs #4012
2026-08-29 01:01:36 -04:00
Tom Boucher
80de48c319 enhance(#3914): every phase records a truthful guard ledger (#4018)
* fix(#3914): retire n/no-process-exit where its successor governs

Epic #3889 criterion 5 — no phase closes with a guard added and its
predecessor left standing — is violated in the tree by the epic that wrote it.

local/require-registered-exit was registered on gsd-core/bin/**/*.cjs and
scripts/**/*.cjs, while n/no-process-exit stayed 'error' over a nine-glob block
covering those same two. Only the hooks 'off' exemption ever came down; the
predecessor's registration never did. Both rules have been enforcing the same
property on the same surfaces since P6.

Narrowed, not deleted. Seven of those nine globs have NO successor —
eslint-rules/, bin/lib/, pi/, examples/, vscode/, .kilo/, .opencode/ — so
deleting the rule outright would silently drop enforcement on all seven. That
is the inversion this epic has already hit three times: removing a coarse guard
because a narrower one exists somewhere it does not reach. Flat config is
last-match-wins and both successor blocks come after the nine-glob block, so
'n/no-process-exit': 'off' in exactly those two retires the predecessor
precisely where the successor governs and nowhere else.

The successor is strictly more precise: it permits process.exit only inside
terminateNow in cli-exit.cts, the single sanctioned terminator (ADR-3889 §3),
where n/no-process-exit permits none and would flag terminateNow's own
generated copy.

Asserted at the consumer's altitude via ESLint.calculateConfigForFile on real
paths, with the positive control that matters: n/no-process-exit is still
'error' on six of the seven successor-less globs, so a future edit that turns
this into a blanket disable goes red. bin/lib/ has no file in this checkout and
is reported as untested rather than given an invented path. Severity is
normalized across the string/numeric/array forms the API can return, and the
normalized value asserted — not truthiness.

Verified by running calculateConfigForFile myself on both superseded globs and
four controls before trusting the test.

Found and fixed inline: the change made an eslint-disable directive at
gsd-tools.cjs:257 partially unused, which --max-warnings 0 rejects; narrowed to
the one rule still in force.

Verification runs on the remote runner.

Refs #3914

* docs(#3914): the epic added three guards, it did not remove one

The audit reconciled the epic ledger against what actually landed. The net is
+3, not -1: four lint:generated-sync --check arms (gen-scripts-cli-exit,
gen-hooks-cli-exit, gen-exit-code-registry, gen-exit-code-docs) plus one rule,
against two retirements.

An epic whose thesis was consolidation ended with a larger guard surface than
it started with. The additions are each defensible; the claim that the total
fell was never true.

Two of the three prior errors in this amendment are mine. It said "Net -1 by
count" above terms reading -1 -1 +1 +1 +1, which sums to +1 — an arithmetic
error in the paragraph directly below the sentence arguing that an ADR about
honest accounting must not pad its own ledger. And the term list omitted two of
the four --check arms, which is what turns that +1 into the real +3.

Recorded rather than quietly rewritten. This ledger has now been wrong three
times — the original -2, the -1 that replaced it, and #3914's own table, which
states -1 above terms summing to 0 — and a written claim nobody checked against
the thing it describes is the exact failure this epic exists to close.

Refs #3914

* fix(#3914): make the successor actually supersede before retiring the predecessor

An isolated security review found that the previous commit turned off a guard
that was still doing work. Reproduced by executing both rules against a
fixture, not inferred:

  const exit = 'exit';
  process[exit](1);

n/no-process-exit flags it; local/require-registered-exit did not, because it
early-returned on callee.computed. So retiring the predecessor on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs un-guarded that shape on precisely
the two globs this epic's exit contract cares most about.

This is the third time in this epic I have removed a coarse guard on the claim
that a narrower one covered it, without checking construct-level parity — after
the allowlist key-to-prefix-to-exact-membership sequence and the band
ranges-to-categories one. The rule is the same every time: a narrower guard
supersedes a coarser one only where it demonstrably reaches at least as far,
and "demonstrably" means executing both against the constructs, not reading
either.

The successor now resolves computed property access for the statically
determinable cases — a string Literal, and an Identifier bound once to a string
Literal, resolved through scope — and leaves genuinely dynamic properties
alone so the rule does not over-fire. Measured after the fix: plain
process.exit flagged, process['exit']() flagged, process[exit]() flagged,
process[globalThis.k]() not flagged. That makes it a strict superset of the
predecessor on these globs, since process['exit']() was caught by NEITHER rule
before.

The second finding is worse than the first, because it was reasoning rather
than oversight. My justification comment claimed n/no-process-exit "would flag
terminateNow's own generated copy here". It would not — that file is in the
global ignore list, so neither rule ever lints it. There was no conflict to
resolve; I wrote a rationale I had not checked, in a change whose entire
subject is written claims nobody verified. Both comment blocks now state the
real basis.

The tests that should have caught this asserted only rule SEVERITY per glob and
never construct REACH, which is exactly how a coverage hole passed. A parity
matrix now pins all five shapes, including a RED/GREEN regression pin against
an inlined reproduction of the pre-fix rule — inlined rather than loaded from
HEAD, because HEAD resolves to the fixed commit under the remote runner and
would silently stop testing anything.

Verification runs on the remote runner.

Refs #3914

* fix(#3914): the two exit rules are complementary — keep both

Reverts this branch's retirement of n/no-process-exit. The premise was wrong
twice, and the second review proved the change itself was wrong.

I claimed local/require-registered-exit was a strict superset on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs. Measured, successor vs predecessor:

  function f(exit) { process[exit](1); }         0  vs  1
  let exit='exit'; exit='exit'; process[exit]()   0  vs  1
  const { exit } = ...; process[exit](1)         0  vs  1

plus for-of bindings, let-then-assign, var redeclaration, catch params, and an
undeclared global named exit. The predecessor matches any identifier NAMED
exit however it is bound; the successor resolves only a string literal or a
single-write const. It never was a superset — I asserted the relationship after
fixing one construct and did not re-check the rest.

The justification was independently false: all three generated cli-exit copies
are in the global ignore list, so n/no-process-exit was never flagging
terminateNow. There was no conflict to resolve. I wrote a rationale I had not
verified, in the phase whose subject is written claims nobody checked.

So criterion 5 does not apply to this pair. They are not predecessor and
successor — they are complementary, each catching constructs the other misses.
The epic's criterion assumed a replacement relationship that does not exist
here, and retiring either rule loses real coverage. The ADR ledger now says so
with the measured shapes.

What survives is the genuine improvement: the computed-property strengthening.
local/require-registered-exit now catches process['exit'](1) and optional-chain
terminators like process?.[k]?.(1), which NEITHER rule caught before, while
correctly ignoring a genuinely dynamic property so it does not over-fire.

The parity tests are rewritten to assert what is true rather than what I wanted
to be true: a bidirectional matrix where each rule is shown catching shapes the
other misses. The previous matrix tested only the four shapes where the
successor wins, which is precisely why the regression shipped — a test set
selected to confirm the thesis.

Also corrected: a stale ADR sentence claiming a third wrong ledger version that
does not exist (the table it described now reads +3 over terms summing to +3),
and a changeset whose stated motivation was the false generated-copy conflict.

Verification runs on the remote runner.

Refs #3914

* fix(#3914): the exemption term was a no-op — the net is +4

Fourth correction to this ledger, and a fourth error of the same kind.

Every version counted removing the n/no-process-exit 'off' entry from the hooks
block as -1. Measured: calculateConfigForFile returns undefined for that rule on
hooks/**. It was never registered there, and no broader block sets it globally,
so the 'off' entry overrode nothing and removing it changed no enforcement at
all. A no-op removal, not a guard removal — the same category error as counting
baseline acknowledgement entries: a thing that is not a guard, in guard units.
It is misattributed too; that block came down in d98b55562 (#3910), already on
next before this branch existed.

So the epic added FOUR guards, not three.

This surfaced from a test of mine that overclaimed. I asserted n/no-process-exit
was error on "all nine CommonJS/hook globs" — but hooks is not one of the nine,
and the rule resolves to undefined there. Fixing the test to match reality is
what exposed the ledger term, which is the argument for tests that assert
identity rather than a comfortable shape.

The hooks state is now pinned explicitly rather than glossed: n/no-process-exit
unregistered, local/require-registered-exit error. It is mildly surprising and
therefore worth a test.

Also updates a pre-existing test that documented the old name-based-only
boundary as intentional. The computed-property strengthening deliberately moves
that boundary — process['exit'](0) was caught by NEITHER rule before — so the
test now asserts the new contract and cites the ADR, rather than being left to
fail or the rule weakened to satisfy it. A contract change should read as
deliberate in the test that pins it.

Verification runs on the remote runner.

Refs #3914

* chore(#3914): backfill changeset pr number to 4018

---------

Co-authored-by: sim <sim@local>
2026-08-29 00:53:33 -04:00
Tom Boucher
51ca9f39ba fix(#3801): register inline_plan_threshold in the defaults manifest and correct the docs (#4019)
* fix(#3801): register inline_plan_threshold in the defaults manifest (default 2) and correct settings-advanced

* chore(#3801): changeset fragment (pr number backfilled after PR creation)

* chore(#3801): backfill changeset PR number (4019)

* test(#3801): parse the defaults table with the shared markdown-table parser

---------

Co-authored-by: sim <sim@local>
2026-08-28 23:48:42 -04:00
Tom Boucher
ac7587287b fix(#3812): document how Current Position actually resolves a duplicate field (#4017)
* docs(#3812): say that Current Position is single-valued, and pin the behavior that makes it true

#3812 shipped CLOSED with half its acceptance unmet. #3873 delivered cardinality for FRONTMATTER
keys - current_phase/current_plan render as optional at docs/reference/state-md.md:89,91, covered by
tests/gen-state-md-docs.test.cjs:374. The issue's actual ask was the ## Current Position BODY
section, and that never landed. Surfaced by an /adr-phase-coverage audit of epic #3473; the issue
was reopened rather than noted.

The section now states three things: every field is single-valued, the section is overwritten rather
than appended to, and a duplicate resolves to the FIRST occurrence with no warning - so a line
appended in good faith is silently ignored rather than winning. Progress history belongs in
## Performance Metrics, two headings down, and the text now points there.

The third claim is a behavioral promise about the reader, so it was VERIFIED BY EXECUTION before
being written rather than inferred from the issue title:

  stateExtractField(<"Phase: 1 of 5 (First)" ... "Phase: 9 of 9 (Appended later)">, "Phase")
    -> "1 of 5 (First)"

The mechanism is state-document.cjs:405 - the plain-line pattern ^<field>:[ \t]*(.+) carries flags
im with NO g, so String.match returns the first hit. Writing "first wins" without running it would
have repeated the exact error I had to retract twice in this epic already.

A test pins the reader, not the prose. Three rows in tests/state.test.cjs: T1 (load-bearing) asserts
the duplicated case resolves first; T2 asserts the ordinary single-field case still works, so a fix
that only functions when duplicated cannot pass; T3 puts a Plan: line BETWEEN the two Phase: lines
and asserts it resolves independently - negative space, because a reader returning the first line of
the SECTION rather than the first matching FIELD would satisfy T1 alone. Proven to discriminate: a
last-match variant returns "9 of 9 (Appended later)" and T1 reds.

No assertion checks that the document contains a sentence. That is what local/no-source-grep exists
to stop, and it would pin wording that is allowed to improve. The point of the test is that if that
regex ever gains g and a last-match walk, the test fails - instead of the documentation quietly
becoming a lie with nothing to notice.

Prose only, no new heading. docs-state-md-locale-parity compares heading-level sequences by LCS
rather than text, so added paragraphs cannot fail it while an added HEADING would fail all four
locales. The constraint is structural, not stylistic - confirmed by running that comparison after
the edit.

The four locale copies are translated rather than left stale. They are not gate-enforced for prose,
so "nothing fails" was available and is not the same as correct: leaving four documents asserting
something the English one now contradicts is a correctness problem. Code spans and the anchor link
stay untranslated - they name real tokens.

The whole approach rests on one fact, checked first: ## Current Position at :196-208 sits OUTSIDE
every generated marker region (:81-104, :138-151), so a hand edit survives --write. Re-confirmed
after all five edits - gen-state-md-docs --check reports all 6 targets up to date. Had that been
false the fix would have belonged in the generator, and a hand edit would have been silently
reverted.

One real gate failure fixed inline rather than reported: the new test's comments referenced
docs/reference/state-md.md, which was not in that file's registered exempt-docs paths, and
lint-docs-guard-registration failed lint:ci correctly. Registered.

Known limit, named rather than folded in: gsd-tools validate/health still do NOT warn on a
duplicated Phase:. #3812 records that as a "consider", not a requirement, and confirms none of the
nine rules in src/health-diagnostic-rules/{state-consistency,phase-structure}.cts counts
occurrences. Documenting the silent first-match is the delivered scope; making it loud is new scope
and stays unclaimed.

Closes #3812

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3812): the rule I documented was false — replace it with the measured one

An isolated review returned two blockers. Both mine, and the first is the worse
kind: I wrote a falsifiable rule into a reference page and got it wrong.

1. "A duplicate resolves to the FIRST occurrence" is FALSE.

   stateExtractField (src/state-document.cts:401-419) tries BOLD `**F:**` across
   the whole input, THEN plain `^F:`, THEN a pipe-table row. Form precedence
   beats document order. Measured against the built reader, all intra-section:

     Phase: A (plain)  /  **Phase:** B (bold, later)   -> B    LATER WINS
       Phase: A (indented) / Phase: B (plain, later)   -> B    LATER WINS
     Phase: A (plain)  /  | Phase | T (table) |        -> A    first wins

   My original verification tested plain-versus-plain, saw first-wins, and
   generalized to all forms. Measuring one case and claiming the general rule is
   the same error I have had to retract twice already in this epic.

   It is also worse than silence. The sentence told authors an appended line is
   safely ignored; a bold line appended "for emphasis" silently overrides the
   original. Someone trusting the doc would have corrupted their own state file.
   And #3812 never asked for a resolution rule - it asked for single-valued,
   overwrite-not-append, and where history goes. The rule was my unrequested
   addition.

   Replaced with the measured truth: resolution is by FORM (bold anywhere, then
   plain at line-start, then table row), and only WITHIN the winning form does
   the first occurrence win. Both consequences stated plainly - a higher-ranked
   form wins regardless of position, and an indented `Phase:` is invisible to the
   plain form. All five claims in the new paragraph verified by execution before
   being written, including the two I had wrong.

2. The tests tested the wrong case and passed for the wrong reason.

   T1/T3 put the second `Phase:` under `## Somewhere else` - the INTER-section
   case, which #2956 already fixed by scoping. #3812 says verbatim that #2956
   "fixed the inter-section case and never addressed intra-section duplication",
   so the case the new prose describes was untested, and the fixtures passed
   because of section scoping rather than field resolution. They also called bare
   stateExtractField rather than the production chain, T2 could not discriminate
   first from last at all, and no fixture mixed forms - which is precisely why the
   false claim survived to review.

   Rewritten as four rows, all intra-section, all through the real
   stateCurrentPositionSlice -> stateExtractField path: plain-then-plain (first
   wins within a form), plain-then-bold (the bold LATER value wins - the row whose
   absence let the false claim ship), indented-then-plain (indented invisible),
   and sibling-field independence. Each proven to fail against a reader that
   disagrees.

3. Two dead anchors. pt-BR and zh-CN linked `#performance-metrics` while their own
   headings are `### Métricas de Desempenho` and `### 性能指标`. Both fixed to the
   anchor their own heading generates. ja-JP/ko-KR kept the English heading, so
   theirs already resolved.

4. A ja/ko sentence inverted its own meaning. Both rendered "which is the section
   designed to grow" with a bare demonstrative whose nearest referent read as
   Current Position - saying the opposite of the point. Rewritten so the clause
   attaches unambiguously to `## Performance Metrics`.

5. Cross-locale drift, flagged by the implementing agent rather than by me: after
   fixing EN, the four locales still stated the OLD false rule. Four documents
   asserting something measured to be wrong is worse than four saying nothing.
   All four now carry a faithful translation of the corrected paragraph, with
   code spans, each file's own anchor, and the ja/ko referent fix preserved.

Verified: all five claims executed against the built reader; every rewritten test
row proven to discriminate; gen-state-md-docs --check reports all 6 targets up to
date, so the edits stay outside the generated marker regions; locale heading
parity unaffected (prose only, no headings added); build:lib, lint and lint:ci all
exit 0.

Refs #3812

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3812): second false rule on the same page — scope the ranking to the section

A second isolated review found a second false falsifiable claim, and the failure
mode is the same one twice in a row:

  attempt 1: verified plain-vs-plain, wrote a claim about ALL FORMS
  attempt 2: verified bare stateExtractField, wrote a claim about THE DOCUMENT

Both times the claim covered a wider surface than what was actually executed. The
fix each time was not a better sentence, it was executing the surface the
sentence describes.

BLOCKER — "bold `**Phase:**` anywhere in the DOCUMENT wins" is false.

  ## Current Position / Phase: 1 of 5   +   ## Archive / **Phase:** 88

    bare stateExtractField(whole doc) -> "88 (other section)"
    PRODUCTION (slice then extract)   -> "1 of 5 (in section)"

#2956's section slice means production never hands another section to the
matcher; a bold line in `## Archive`, or in the YAML frontmatter, is simply not
seen. The ranking is real but scoped: it applies WITHIN `## Current Position`.
I verified against the bare function and wrote a claim about the system.

Every existing test placed its bold line inside the section, which is exactly why
nothing contradicted the claim. T5 now puts a bold `**Phase:**` in `## Archive`
and asserts production returns the in-section plain value, with the unscoped
reader asserted to DISAGREE so the row proves the scoping rather than assuming
it.

BLOCKER — the changeset still shipped the ORIGINAL retracted claim.

I corrected the page and left the release note saying "resolves to the first
occurrence ... a second entry added in good faith is silently ignored". The note
contradicted the page it announces, and the release note is what most people
actually read. Rewritten to the corrected rule.

MEDIUM — the concession was inverted. It read "wins even if it comes FIRST in the
file", which is the vacuous direction; the surprising case, and the one the very
next clause illustrates with an APPENDED bold line, is "even if it comes LAST".
All four locales reproduced the inversion faithfully, so it was an EN-source
defect rather than translation drift.

Two sharp edges now named, both measured: a bold `**Phase:**` followed only by
trailing spaces resolves to an EMPTY STRING and does not fall through to a valid
plain line below (T6 pins it); and `| **Phase:** | 3 of 4 |` short-circuits to the
bold form and returns the literal `"| 3 of 4 |"`. A page that teaches form
ranking has to say where the ranking bites.

Also fixed: all five files labelled the link `## Performance Metrics` while the
heading is `### Performance Metrics`. Anchors resolved correctly everywhere; only
the label's level was wrong.

Every clause in the final paragraph re-verified through the PRODUCTION chain
(stateCurrentPositionSlice -> stateExtractField), clause by clause, before being
written: bold in another section does not win; bold in frontmatter does not win;
bold appended last does win; first wins within one form; trailing-space bold
yields empty. All four locales carry the same corrected rule.

gen-state-md-docs --check reports all 6 targets up to date; heading counts
unchanged at 20/20 across all five files, so locale heading-parity is untouched;
build:lib, lint and lint:ci all exit 0.

Refs #3812

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3812): backfill changeset pr number

Refs #3812

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 20:50:15 -04:00
Tom Boucher
1051c6d8d4 fix(#3799): scope legacy cleanup to the install's resolved config dir and add --no-legacy-cleanup (#4013)
* fix(#3795): read the interrupted agent id before clearing the stale marker (#4006)

* test(#3795): the interrupted-agent read must precede the stale-id clear

* fix(#3795): read the interrupted agent id before clearing the stale marker

execute-plan's init_agent_tracking step ran `rm -f
.planning/current-agent-id.txt` BEFORE the existence check that read it,
so the interrupted-agent branch and the Task resume prompt it exists to
offer were unreachable (#3795) — a kill -9 mid-executor left the file,
and the next run deleted it before looking. The read now precedes the
clear; fresh-run semantics (no stale id leaking into the new spawn) are
preserved. A structural guard pins the order.

Emitted-Drift-Ack-Growth: execute-plan.md — #3795: +bytes from reordering the interrupted-agent read before the rm plus the explaining comment

* chore(#3795): changeset fragment (pr number backfilled after PR creation)

* chore(#3795): backfill changeset PR number (4006)

---------

Co-authored-by: sim <sim@local>

* fix(#3799): scope legacy cleanup to the install's resolved config dir and add --no-legacy-cleanup

* chore(#3799): changeset fragment (pr number backfilled after PR creation)

* chore(#3799): backfill changeset PR number (4013)

* test(#3799): mark the legacy-name fixtures with gsd-allow-legacy-name

---------

Co-authored-by: sim <sim@local>
2026-08-28 20:16:50 -04:00
sim
bf8905fcc9 fix(#4012): a hang never emits the events I was listening for
The init marker settled it. The diagnostic reported "THE REPORTER LOADED BUT
RECORDED NO TEST EVENTS — the events file contains only the reporter's own
reporter:init marker", which refutes the reporter-never-loaded hypothesis and
leaves exactly one explanation.

The runner spawns a child process per test file and surfaces a subtest's
test:start / test:pass / test:fail to the parent's reporter only once the child
REPORTS that test — which happens when it completes. The fixture hangs forever,
so it never completes, so it never reports. I had recorded exactly those three
event types: the precise set that a hang guarantees you will never see. The
feature could not have worked for the case it was built for.

test:enqueue and test:dequeue are emitted by the runner as it queues and begins
each file, independent of anything inside finishing. test:dequeue is what
actually means "in flight", and it is now the primary signal, with test:start
kept as a secondary one. A file is in flight when it has been dequeued and has
no terminal event.

The four branches now describe states that are all real: the events file absent
(reporter never loaded); the init marker alone (the runner dequeued nothing at
all — genuinely surprising now rather than the expected outcome); everything
dequeued and terminated (the files finished and the process hung afterwards, a
handle leak); and one or more dequeued-but-unterminated files, named, which is
the case this whole feature exists to report.

Verified against the exact shape the real hang produces, by executing
analyzeChunkEvents on a synthetic events file: init + enqueue + dequeue with no
terminal event reports hangs.test.cjs as in flight, and appending a test:pass
clears it. Four more unit tests cover the ordering and multi-file cases with no
subprocess, so this logic is now checkable without a runner round-trip — which
matters, because every defect in this feature so far was visible only remotely.

T1 is untouched and should now pass for the right reason.

Verification runs on the remote runner.

Refs #4012
2026-08-28 19:52:57 -04:00
sim
444137e63a fix(#4012): make the artifact say whether the reporter ever loaded
Down to 3 remote failures, all one chain. The explicit no-events reporting is
working — the diagnostic now states the events file does not exist, instead of
silently printing the generic message. But it then ASSERTED a cause: "the child
was killed before the reporter wrote even one event (process/spawn startup
stall, not a test hang)". That was a guess dressed as a finding, and the fixture
contradicts it: it starts a real test, so test:start should fire in milliseconds
against a 2000ms budget.

Two hypotheses remained and I could not separate them locally, because the local
test runner is hook-blocked here: either the custom reporter never LOADS in the
child, or it loads and no event reaches it before the SIGKILL.

Rather than guess a third time, the artifact now answers it. The reporter
appends a reporter:init line as its first action, before consuming anything, so
the file's contents discriminate: absent means the reporter never loaded;
init-only means it loaded and saw no test events; init plus events means it
works. The diagnostic has a branch for each and, where the cause is genuinely
unresolved, names both possibilities instead of picking one.

Also passes the reporter as a file:// URL via pathToFileURL. Node documents the
--test-reporter value as an import()-style specifier, and a bare absolute path
is not a portable one — notably on Windows. That is a correctness fix whichever
hypothesis holds, and it is a live candidate for the first. FIXED_OVERHEAD is
derived by reducing over the actual argv strings, so the longer URL is accounted
automatically.

Verified by execution, not assumption: composing the reporter against an EMPTY
event stream writes exactly one line, the init marker. That is the whole point
of the marker, so it is pinned by a test rather than left to inspection.

T1 stays red and untouched.

Verification runs on the remote runner.

Refs #4012
2026-08-28 19:41:32 -04:00
sim
5504724d7a fix(#4012): the reporter body must return nully, not an iterable
Fourth failure on this feature, and this one was caused by the previous fix's
lint workaround.

  TypeError [ERR_INVALID_RETURN_VALUE]: Expected nully to be returned from the
  "body" function but got an instance of Array.

When stream.compose is given an async FUNCTION as the body, that function must
return nully. The reporter ended with 'return []' under a comment asserting
Node "still requires the exported function to return an iterable" — exactly
backwards, and the direct cause of 41 failures across every run-tests.cjs
invocation.

That return existed only to dodge ESLint's require-yield after the previous
commit converted async function* to async function. A lint workaround became a
runtime crash, and the comment written to justify it stated the opposite of the
contract. Both are now corrected to what the runtime actually does.

Verified by EXECUTION rather than by reading: composing the real reporter
against a fake event stream completes with no error and leaves both handled
events durably on disk, with the ignored event type skipped. The failing form
was reproduced the same way first, so the diagnosis is not inferred from the
stack trace alone.

The new unit test requires the reporter directly and asserts the returned value
is nully — the assertion that would have caught this before it reached the
runner — plus the exact NDJSON written. No subprocess, so this half of the
feature is verifiable without a full runner pass, which matters because every
defect in this feature so far has only been observable remotely.

T1 remains untouched and red.

Verification runs on the remote runner.

Refs #4012
2026-08-28 19:29:51 -04:00
sim
7f2af28639 fix(#4012): the reporter sink must be a regular file, not devNull
Third failure on this feature, and this one broke everything rather than just
the diagnostic: 43 failures across every run-tests.cjs invocation.

  Error: EINVAL: invalid argument, fsync
  Emitted 'error' event on WriteStream instance

Node opens a WriteStream for a --test-reporter-destination and FSYNCS it on
close. fsync on /dev/null is EINVAL — it is a character device, not a regular
file. So os.devNull is not a usable reporter destination at all, and every
chunk crashed on exit.

The sink only ever needed to be a regular file that stays empty, since the
reporter writes its real output through appendFileSync to the path in
GSD_RUN_TESTS_EVENTS_FILE. It is now one fixed file inside the existing events
dir, pre-created rather than relying on the stream's create-on-open, and left
empty by design. One sink for the whole run, not per chunk, so its path length
stays constant — reporterOverhead feeds FIXED_OVERHEAD, which is computed once
before chunking, and a variable-length path would silently mis-account the
Windows argv ceiling.

The comment that named devNull now says why the destination must be a regular
file, so this does not get re-optimized back into the same crash.

Reaching for devNull was the mistake: it looks like the obviously correct way
to discard output, and it is, for a pipe or an fd — but not for something Node
is going to fsync. Each of the three failures on this feature was a different
edge of the same assumption, that a reporter destination behaves like ordinary
output.

The regression test asserts the closest externally observable consequence — a
normal run must not surface the EINVAL/fsync text. The argv construction lives
inside main() with no exported seam, and the sink is swept before a test could
stat it; adding a seam purely to assert that is left out rather than reshaping
production code for the test. Stated plainly rather than implied.

The failing T1 is untouched and still red.

Verification runs on the remote runner.

Refs #4012
2026-08-28 19:20:11 -04:00
sim
0abd137ec7 fix(#4012): the events reporter has to survive SIGKILL
The remote run proved the instrumentation did not work. Timing landed —
"chunk 1/1 was killed after 2006ms" — but the in-flight-file naming produced
nothing and fell through to the pre-existing generic message. The feature I
wrote to diagnose a kill was itself destroyed by the kill.

Root cause, confirmed rather than assumed. The reporter yielded strings, which
node pipes into the --test-reporter-destination WriteStream. That stream
BUFFERS. execFileSync's timeout sends SIGKILL, which is uncatchable and gives
nothing a chance to flush, so the events sat in a buffer that died with the
child. The parent's own timer reported correctly because it lives in the
parent — which is exactly why half the feature looked fine.

The reporter now writes each event with fs.appendFileSync, unbuffered and
durable at the moment it happens, to a path passed through
GSD_RUN_TESTS_EVENTS_FILE. Env vars do not count toward the Windows
32,767-char argv ceiling, so moving the path out of argv also REDUCES
FIXED_OVERHEAD; the accounting moved with it rather than being left stale. The
destination is now a fixed devNull sink that stays empty by design.

Silence was the reason this was invisible for a whole run. Failing to read the
events file now says so explicitly, and distinguishes a file that could not be
read at all from one that exists but is empty — the generic fallback firing
quietly is what let a broken feature look like a working one. A write is
unbuffered but not atomic, so a kill can still interleave a partial line; the
reader tolerates exactly one unparsable trailing line and reports the complete
ones before it.

The failing T1 was left red and untouched rather than weakened to pass. Three
new unit tests cover the reader directly, with no subprocess, so the parsing
half is verifiable without a full runner pass: missing file, existing-but-empty
file, and a truncated final line.

Also adds ndjson-reporter.cjs to GSD_SCRIPTS_LIB_FILES in bin/install.js —
scripts/lib/ ships, and omitting it meant the file would install everywhere and
orphan on uninstall. That single omission caused 4 of the 7 remote failures.

Verification runs on the remote runner.

Refs #4012
2026-08-28 19:10:54 -04:00
sim
11c24ce973 chore(#4012): regenerate the 19 install-tree goldens
scripts/ ships, so a new file under scripts/lib/ moves every install-tree
golden. The remote runner caught this as 26 failures across all 19 runtimes.

My fault in the dispatch: I limited the implementation's verification to
build:lib and lint:ci and left regen:derived out. Both of those passed, which
is precisely why a green local gate is not a substitute for the runner — the
goldens are only exercised there.

One line per golden: scripts/lib/ndjson-reporter.cjs joining the shipped tree.

Refs #4012
2026-08-28 18:58:05 -04:00
sim
8497833a15 fix(#4012): a killed chunk now names the file that was hanging
The per-chunk timeout fired correctly but reported almost nothing, so every
diagnosis cost a CI round-trip. run-tests.cjs logged chunk START only — no
timestamp, no duration, no end line — then on a kill printed all ~55 basenames
and asked the operator to work out whether output kept flowing (slow) or
stopped early (hang). It could not name the in-flight file because the child is
spawned with stdio inherit, deliberately, per #3597/#1051.

Three additions.

Per-chunk elapsed timing on every path, not just failures, so drift toward the
cap is visible before it becomes a kill. Every timing number in the
investigation behind this had to be reconstructed by hand from GitHub log
timestamps.

A second, machine-readable reporter running ALONGSIDE the human one, writing
NDJSON to its own file. On a kill that file is read back and the files with a
test:start and no matching completion are named, with the staleness of the last
event, so "stopped 480s ago at X" reads differently from "still emitting at
kill". stdio stays inherit and nothing is piped or tee'd — the maxBuffer and
live-output risks that shaped the original design are untouched.

Ranking of the killed chunk's files by known weight, flagging any absent from
tests/test-timings.json, since an unweighted file is an unknown quantity.

Two details that are correct rather than lucky. Passing --test-reporter at all
replaces node's implicit default, so the human reporter is now named explicitly
and reproduces node's own selection (spec on a TTY, tap otherwise) — visible
output is unchanged. And the destination path's chunk index is zero-padded to a
fixed width because FIXED_OVERHEAD is computed ONCE before chunking; a
variable-length path would have silently mis-accounted the Windows 32,767-char
argv ceiling and reintroduced #3597. The reporter flags are added to
FIXED_OVERHEAD exactly as --test-force-exit is.

The multi-reporter pairing and the stream.compose reporter contract were
confirmed against Node's v24 documentation, not recalled — the first draft
carried them as an unverified assumption and said so.

Also corrects a stale comment claiming the 600s cap sits "below the 20m job
cap". The lane is sharded 3x at timeout-minutes: 45; the windows shards were at
19m when chunk 1/5 was killed on b351c83e0 and c3e667df3. The per-chunk cap is
now the binding constraint, and the old silent-cancel model leads to the wrong
conclusion.

Verification runs on the remote runner.

Closes #4012
2026-08-28 18:50:55 -04:00