Commit Graph

927 Commits

Author SHA1 Message Date
Adnan
bdfc62889b fix(#3784): read the hybrid Current Plan: N of M shape, keep zero-padding, and name the accepted shapes on failure (#3791)
* fix(state): read the hybrid "Current Plan: N of M" shape

advancePlanCore derived the value FORMAT from the field NAME, so it
handled the legacy pair (`Current Plan` + `Total Plans in Phase`) and the
compound `Plan: N of M`, but not the hybrid of the two: the legacy field
name carrying a compound value with no Total Plans sibling. `legacyTotal`
is null so the legacy branch fell through, and the compound branch reads
the `Plan` field through a `^Plan:`-anchored pattern that never matches
`Current Plan:`. Both produced NaN against a file whose plan numbers are
plainly readable.

The shape is not exotic. An agent wrote it unprompted into a project's
STATE.md, believing it was the parseable form, and every subsequent run
in that project inherited the failure and worked around it by hand.

Track the field name and the value shape separately (`planSourceField`,
`planRawValue`) so write-back targets whichever field the value came
from. The legacy pair still takes precedence when both fields exist, so
a stray "of N" inside Current Plan cannot override an explicit Total
Plans — covered by a new test.

Also replace the caller's catch-all error. It reported "Cannot parse
Current Plan or Total Plans" for ANY transition failure, and named no
accepted shape, so a reader learned neither what failed nor what to
write. It now distinguishes "no result" from "unreadable plan position"
and lists all three shapes. The existing test asserted the literal
"cannot parse"; it now asserts the message names the shapes, which is
the property that makes it actionable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* fix(state): keep zero-padding when advancing a compound plan value

The compound write-back rewrote only the leading half of "N of M", so a
padded value drifted lopsided: "04 of 06" advanced to "5 of 06". Cosmetic
on its own, but a plan line that looks wrong is one the next writer tidies
by hand, and hand-tidying this particular line is what produced the hybrid
shape the previous commit had to teach the parser to read.

Pad the incremented number to the width it was written with. padStart never
truncates, so a value that outgrows its padding widens correctly: 09 of 12
advances to 10 of 12. Unpadded values are untouched — 2 of 6 still advances
to 3 of 6.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* fix(state): pass a literal field name to the compound write-back

The previous commit passed `planSourceField` — a variable — as the field-name
argument to `stateReplaceField`, which trips the state-write-path drift guard's
`unstripped_content_write` axis (ADR-3408 §8.3(b)). The guard is right to care:
a Title-Case literal cannot collide with a lowercase or snake_case frontmatter
key, so it is safe whatever the content argument is, while a variable could
hold anything and therefore requires its content to be demonstrably
frontmatter-stripped first.

The content argument here IS stripped — `body` is `stripFrontmatter(content)` —
but the guard does a narrow backward scan rather than dataflow tracking, by
design, and the nearest preceding assignment to `body` is another
`stateReplaceField` result. Rather than baseline a bypass or ask a future
reader to re-derive that the invariant holds, dispatch on the discriminator and
pass the literal.

Guard goes from 1 finding to 0; its own 32 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* chore(3784): add changeset fragment for #3785

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* test(#3784): cover the maintainer's AC1 write-back and AC6 reader-anchoring

Triage published six acceptance criteria; two were only half-covered.

AC1 asks that the hybrid write back to the SAME field with padding preserved.
The existing hybrid test used an unpadded value and asserted only `result.data`,
so it proved the parse but never the write. Now asserts the written content is
`05 of 06` on the original field, and that no separate `Plan:` field appears as
a side effect.

AC6 asks that the shared field reader not be loosened. Reading the hybrid is the
transition's job; `stateExtractField('Plan')` is line-anchored and has 13+
callers, so teaching it to match a name merely ENDING in "Plan" would be the
wrong fix and would silently change what those callers read. This holds by
construction here — the reader is untouched — but nothing locked it in. The new
test fails if anyone later reaches for that shortcut.

Also drops the changeset fragment written against the auto-closed PR number.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* chore(#3784): add changeset fragment for #3791

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* fix(#3784): write the advanced plan back to the field it was read from

Review findings 2-6 on #3791 were one defect seen from several angles: the
read path learned the hybrid `Current Plan: N of M` shape, the write path
did not follow it.

- `bumpLeadingNumber` now owns the increment for all three parse branches.
  Only the leading digits belong to this transition; the padding width and
  everything after it (` of M`, and the `\r` of a CRLF file) are the
  author's text and are preserved. The legacy branch wrote `String(newPlan)`,
  which turned `2 of 99` into `3` and `04` into `5`.
- `mutateCurrentPositionForAdvance` takes the plan field NAME. Its plan arm
  only ever looked for `Plan:`, so on a hybrid file the `## Current Position`
  section was never reached; combined with the body-level write being
  single-shot and bold-preferring, a file carrying the field at both sites
  advanced the header and left the section a plan behind. The parameter
  defaults to `Plan`, so the two callers that pass no plan are unchanged.
- Tests: both-sites-advance (fails without the section arm), legacy
  write-back content assertions (the previous test read only `data` and so
  could not see the lossy write), hybrid boundary at limit-1 and limit+1, a
  CRLF fixture, and an fc property pinning the padding-width contract.

Two characterization tests pinned `**Current Plan:** 02` advancing to `3`.
That dropped padding is the defect #3784 reports, so the expectation is
corrected to `03` rather than the fix being narrowed around it.

* fix(#3784): drop the unreachable advance-plan error branch, sync the doc

Findings 1 and 8 on #3791.

The `!resultData` arm could not fire: the transform callback assigns
`resultData` unconditionally, only runs once STATE.md is known to exist (the
missing-file case returns "STATE.md not found" upstream), and every
`advancePlanCore` return path sets `data`. It was a speculative second
failure mode with a message no caller could receive, and the comment beside
it claimed to distinguish two things that were never two. `!resultData`
stays in the condition as a type guard, which is all it ever was.

`docs/json-errors.md:142` quoted the old error literal verbatim and was the
sole occurrence in the tree; it now quotes the emitted one.

* chore(#3784): describe the write-back fix in the changeset

* fix(#3784): anchor the plan grammar and widen the schema row to match

Review round 3 on #3791: B1, B2, M1, M2, M3, M4 and the planSourceField nit.

B1 — `STATE_FIELD_SCHEMA.current_plan.acceptedShapes` widens to
`['N', 'N of M']`, which is what `src/state-md-schema.cts`'s own comment
instructed this PR to do on merge. `'N/M'` stays undeclared so row 23 keeps a
non-vacuous undeclared candidate to probe. Verified `gen-state-md-docs
--check` exits 0 and `--write` rewrites 0 of 6: the generated artifacts do not
surface this row, so there is nothing stale to regenerate.

B2/M4 — the discriminator was `/of\s+(\d+)/`, unanchored, so a total could be
read out of prose. `Current Plan: 4 — blocked on review of 2 PRs` parsed as
`4 of 2`, took the `currentPlan >= totalPlans` branch and WROTE
`Status: Phase complete — ready for verification` into the user's file. Both
shapes are now anchored at the start and every number comes from a capture
group via `planNumberFrom`, which rejects anything past `Number.MAX_SAFE_INTEGER`
rather than letting `data` and the persisted string disagree. Nothing on this
path calls `parseInt` on a raw field value any more.

The grammar keeps a trailing remainder after the total, because
`Plan: 2 of 5 in current phase` is a real tested shape. The refusal comes from
requiring `of <total>` to follow the leading number immediately, not from
forbidding a suffix.

M1 — `bumpLeadingNumber` is total. It returned its input unchanged when there
were no leading digits, so `+2` reported `advanced: true` while writing the
file untouched.

M2 — both section arms use replacer functions. File-derived text was being
spliced into a `String.replace` replacement string, where `$&` / `` $` `` /
`$'` expand: a value of `04 of 06 $&` spliced part of the document into itself.
`stateReplaceField` already used a function; these now agree with it.

M3 — the section arm targets the name the SECTION carries, and the body write
now writes both spellings, each with its own rendering. Keying off the header's
name left the other name stale in both directions: a legacy header beside a
`Current Plan:` section line, and a `**Plan:**` header beside one.

* fix(#3784): derive the shape error from the schema, widen the test coverage

Review round 3 on #3791: B3, m1, m2, m5 and the two test nits.

B3 — the accepted-shape set had two owners: the parser branches and an English
list hand-written beside them in `state.cts`. Nothing coupled them, so adding a
branch left the message stale and removing one left it advertising a shape that
errors, with no test able to see either. The message is now built from
`STATE_FIELD_SCHEMA.current_plan.acceptedShapes`, and the CLI test walks the
schema instead of restating the list. `Plan: N of M` is still spelled out
explicitly because no schema row owns the body-only `Plan` field —
`buildStateFrontmatter` never reads it into frontmatter, so it has no key to
hang a row on.

m1 — the property drove only the pre-existing `**Plan:**` branch, i.e. not the
branch under review. It now drives both compound spellings and ranges past 99
so the width transition is covered by the property rather than one example. A
second property covers the legacy pair's own preservation contract. Both were
mutation-checked: dropping the padStart turns 9 tests red.

m2 — degenerate boundary fixtures around the threshold (`0 of 0` is
phase-complete, not an error; `0 of 3` advances) plus the shapes the anchored
grammar must refuse, including Arabic-Indic digits.

m5 — `docs/json-errors.md` described rather than quoted the message, since it
is now schema-derived and a verbatim quote would be a third owner.

Nits — the CRLF assertion could not see a `\n` at index 0; the
`!/^Plan:/m` presence proxy is now an identity assertion on the whole
`## Current Position` body.

* fix(#3784): give the section plan write its own flag, and stop narrowing what parses

Review round 4 on #3791: Blockers 1-4, Majors 1-2, Minors 1-2.

B1 — the section fallback was guarded by `!mutated`, and `mutated` is
FUNCTION-wide, already set by the phase/status/lastActivity arms that
`advancePlanCore` always populates. A section spelling the field bold or as a
pipe-table row therefore skipped its fallback because an UNRELATED field had
been refreshed, and stayed a plan behind the header — the split-brain document
this arm exists to prevent. The arm now tracks its own `planWritten`.

Worth recording: the reviewer's fixture does not reproduce. The body-level
status write lands on the section's own `Status:` when the document has no
header `Status:`, so `mutated` is still false by the time the plan arm runs and
the fallback fires. The discriminating shape needs a header `Status:` to absorb
that write AND a bold section plan line. The mechanism was right; the example
was not, and the regression test uses the shape that actually fails.

B2 — `fallbackName` chose one name by ternary. In the legacy shape both values
are populated, so it always chose `Current Plan` and a `**Plan:**` section line
— which base did write — got nothing. Each name is now attempted independently
with its own fallback.

B3 — `PLAN_SHAPE_N` was anchored harder than `PLAN_SHAPE_N_OF_M`, so values base
parsed via `parseInt` began to hard-error: `Total Plans in Phase: 5 phases`,
`Current Plan: 3 (blocked)`. #3784's brief puts normalizing plan numbers beyond
this transition's read/write out of scope, so that narrowing was not licensed.
Both grammars now carry the same trailing tolerance. The prose defect stays
closed by the START anchor, not by forbidding suffixes.

Major 1 — the whole-body `Plan` write is scoped to documents that declare a
`Plan` field, instead of firing unconditionally where `stateReplaceField`'s
first match could be prose outside `## Current Position`.

Major 2 — the error message names both `Plan` spellings the parser accepts; it
previously omitted the sibling-paired form, which is the same
message-disagrees-with-parser drift the derivation exists to close.

B4 — the changeset claimed a guarantee B1 broke; it now describes what ships.

Minors — safe-integer boundary coverage at limit-1/limit/limit+1, and the CRLF
comment states the real mechanism (`stateExtractField`'s `(.+)` stops before the
CR; the trailing group is belt-and-braces, not the primary defence).

All three blocker regression tests verified red against the pre-fix source.

* test(#3784): pin the hybrid shape against #3807's ambiguity refusal

#4028 landed `advance-plan`'s multi-`Phase:` refusal on `next` after this
branch's last run, on the same function. The guard sits above the parse, so
a refused document is never parsed and the shape #3784 adds cannot reach the
mutation — but that is a property of source ordering, so assert it as
behaviour instead.

Fail-first proven, not assumed: with `phaseCandidates.length > 1` disabled,
the ambiguous hybrid document advances its FIRST entry's
`Current Plan: 04 of 06` to `05 of 06` and writes it — #3807's exact defect,
reached through #3784's shape. Both tests go red; both go green with the
guard restored.

The control pins the other direction: an unambiguous hybrid section still
advances, and its zero-padding still survives.

* fix(#3784): advance every spelling from its own text, refuse when they disagree

Round 6 review. B1 and M1 are one defect, so they are one fix.

`advancePlanCore` picked one field to parse from, computed `newPlan`, then
wrote BOTH spellings from that field's numbers. Two symptoms:

  B1  With `Plan` as the parse source, `Current Plan` was re-stamped with
      the number just derived from `Plan`. `Current Plan: 7` beside
      `Plan: 2 of 5` silently became `Current Plan: 3` — a value nothing
      derived for that field, no error, no diagnostic.
  M1  With the legacy pair winning, the `Plan:` line was re-rendered from
      a bare `${newPlan} of ${totalPlans}` built out of the sibling field.
      `Plan: 2 of 9` became `3 of 5`; `Plan: 03 of 05` became `4 of 5`.
      The changeset's claim that padding and everything after it survive
      was true only for whichever field happened to be the parse source.

Now: every spelling is advanced from its own raw text via
`bumpLeadingNumber`, so each keeps its own padding, its own total and its
own trailing annotation. Differing TOTALS are preserved, not reconciled —
`Plan: 2 of 9` beside a `Total Plans in Phase: 5` advances to `3 of 9`.

Differing CURRENT numbers are refused, with `reason:
"ambiguous_plan_position"` and both candidates named. Same posture as
#3807's multi-`Phase:` guard one field over: name the conflict, let the
caller resolve it, never pick. The guard sits immediately after the parse,
BEFORE the phase-complete branch — guarding only the normal advance would
let `Current Plan: 7` beside `Plan: 5 of 5` write a terminal "Phase
complete" into a document whose two spellings never agreed.

A field present but unreadable (`Plan: TBD`) is left exactly as authored.
Refusing the whole document because an unrelated line cannot be read would
be a narrowing #3784 does not license; writing a derived number over it is
the fabrication B1 was filed for.

The `planSourceField`/`planRawValue`/`useCompoundFormat` tracking is gone.
It existed only so the write path could ask which field the value came
from, and the write path no longer asks.

M2. The bare `Plan: N` + `Total Plans in Phase: M` shape is dropped. A
revision of this PR added it; base refused it. It cannot be given the
schema-row + forcing-test coupling the other shapes have, because `Plan`
is body-only and `buildStateFrontmatter` never reads it into frontmatter,
so there is no `current_*` key to hang a row on. Parser, the spelling in
`advancePlanShapeError`, and the lockstep test move together — the
invariant is the lockstep, not the length of the list.

N1. The whitespace narrowing (`5phases` no longer parses where `parseInt`
read 5) is documented in the changeset beside the other deliberate
narrowings, rather than loosened. Loosening restores the half-parse this
change exists to remove.

Tests: eight new cases plus a property that crosses the two spellings with
agreeing and disagreeing numbers — the review noted the existing
properties never did. Fail-first proven: restoring the old write path
reddens seven of the eight, both new property arms, and two pre-existing
padding tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

* fix(#3784): report Current Plan as updated only when it was written

The write became conditional in the previous commit — a `Current Plan:`
that is present but unreadable is left as authored — but the `updated`
push stayed unconditional, so `transitionCore` reported a field it had not
touched. `reconcileReportedFields` would have caught it against the
persisted bytes at the `state.cts` caller, but `transitionCore`'s own
`updated` is consumed directly (milestone-lock, the transition tests) and
has to be true on its own.

Covers the mirror of the unreadable-spelling case: `Current Plan: TBD`
beside a readable `Plan: 2 of 5`, where `Plan` is the parse source and the
legacy field is the one that cannot advance. Fail-first proven.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

* test(#3784): account for the new refusal in the output({error}) census

`tests/io.test.cjs`' A3 census asserts the exact population of
`output({error})` call sites in `src/`, per module. The
`ambiguous_plan_position` refusal added a 27th to `state.cts`, so the
census went red at 26/65.

Updated the way #3807 updated it when it added the ambiguous-POSITION
error one line above: bump the count and name the addition inline, so the
next person reads why the number is what it is. The alarm did its job —
it is the only gate that noticed a new user-visible error path had been
introduced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-01 15:03:20 -04:00
0xdhx
7c116b1c17 fix(#3697): warn when the phase-complete Requirements-line tokenizer under-selects REQ-IDs (#3744)
* fix(#3697): warn when the Requirements line under-selects REQ-IDs

`cmdPhaseComplete` tokenizes ROADMAP's `**Requirements**:` line by splitting
on `[,\s]+` and keeping tokens matching the anchored REQ-ID shape. That is
correct for the canonical comma list the template ships, and silently wrong
for every other form:

  `RANGE-01 … RANGE-05`  ->  the two ENDPOINTS only; the interior IDs are
                             never considered, yet `requirements_updated`
                             reports true with zero warnings
  `RANGE-01…05`          ->  ZERO IDs; the whole line is inert

The silence is structural: the only cross-check, `ghostReqIds`, is itself
`citedReqIds.filter(...)`, so an ID the tokenizer dropped is invisible to it
by construction — and to `traceabilityWriteMisses` and `requirements_updated`
with it.

Warn on both paths. This does not add range support: the selected set is
unchanged, so no existing ledger write changes. The trigger is ID-SHAPED
EVIDENCE only — an ID-shaped substring the tokenizer did not select, or a
range operator joining two IDs — with parenthetical citations and HTML
comments stripped before the scan, so the #2334/#2339 over-warning on
`None`, on the shipped `<!-- brackets optional -->` template comment, and on
annotated lines cannot return.

Regression tests extend the #2316/#2334 fixture family in tests/phase.test.cjs
(10 cases: 4 defect, 2 canonical controls, 4 negative-space controls).

Fixes #3697

* fix(#3697): rework under-selection detection onto tokens, not a free-text scan

Round 2, driven by the P4.6 cross-AI review (codex, gpt-5.6-sol) of b3ce71cb.
That review refuted 5 of 9 claims; three were false-positive classes in exactly
the category #2334/#2339 had to REMOVE:

  `RANGE-01, RANGE-02 - 3 points`        the bare-hyphen alternative read
                                         `RANGE-02 - 3` as a range
  `REQ-01, REQ-02 — locked per ADR-7.`   the trailing period kept `ADR-7.` out
                                         of the anchored filter, so the
                                         unanchored substring scan reported it
                                         as unparsed
  `REQ-01, REQ-02 (see (ADR-7), then ADR-8)`
                                         nested parens left `ADR-8)` behind

Replaces the free-text substring scan + loose range regex with three narrow,
token-based rules (R1 range-shaped token, R2 pure range operator flanked by two
selected IDs, R3 zero-selection with ID-shaped text). Also fixes the review's
CLAIM 9: the warning said IDs were "marked complete" when a ghost range marks
nothing — it now says "selected".

Side effect: the two false NEGATIVES the same review found are now covered —
`RANGE-01 through RANGE-05` and a parenthesised `(plus RANGE-02..RANGE-05)`.

NOT YET DONE (see the handoff prompt): regression tests for the four false
positives, the two new true positives, and the #3697-4 tightening the review's
CLAIM 8 asked for (it currently filters on the warning's phrasing rather than
asserting silence). Verified so far: tsc clean, the 10 existing #3697 tests
green, and a 20-case standalone harness covering every case above.

* test(#3697): pin the v2 token-detector boundary end-to-end

Six new cases + two hardenings for the review findings against v1:

- #3697-1 gains the worded spaced range (`RANGE-01 through RANGE-05`) —
  the operator set's `to|thru|through` arm was previously untested.
- #3697-5 (new): a tight range hidden inside balanced parentheses
  (`RANGE-01 (plus RANGE-02..RANGE-05)`) warns, names the range token,
  and ticks exactly RANGE-01 — the paren shave must not hide it.
- #3697-4 gains the four false-positive classes a free-text detector
  produced: numeric estimate (`- 3 points`), date annotation, em-dash
  citation with trailing period (`— locked per ADR-7.`), and nested
  parenthetical citations.
- #3697-3 and #3697-4 now assert the ENTIRE warnings channel is empty,
  not that one phrase is absent — a re-worded over-warning cannot pass.

Negative control: against the merge-base with its lib rebuilt, all 6
defect tests fail and all 10 controls pass.

* fix(#3697): close round-2 review findings — annotation false positives

Round 2 of the adversarial review (against 822a72a04) refuted five
claims; this closes the false-positive class and the cheap misses:

- R1's bare-hyphen arm now demands a full ID on BOTH sides
  (`REQ-01-REQ-05`): `LETTERS-\d+-\d+` is also a date-like annotation
  (`FY-2026-08`) and a sub-numbered ID, and warning on those is the
  expensive class. Tight hyphen shorthand with a live selection is the
  disclosed false negative; at zero selection R3 still catches it.
- R2 requires the endpoint pair to imply an INTERIOR (same prefix,
  gap > 1): `REQ-02 - REQ-03` selects both endpoints and can drop
  nothing, so an annotation hyphen between adjacent IDs stays silent.
- R3 skips placeholder-led lines: `None (per ADR-7)` is a declared-empty
  line citing its rationale, not unparsed residue.
- Token shave: quotes/backticks now shaved from alphanumeric tokens
  (`` `RANGE-02..RANGE-05` `` warns); punctuation-only tokens get a
  bracket-only shave so `(..)` surfaces its operator.
- 256-char token cap bounds the quadratic unanchored substring test.
- Warning text mentions range expansion only when a range rule fired.

Tests: 6 new cases (22 total). Negative control against the merge-base:
8 defect tests fail, 14 controls pass.

* fix(#3697): close round-3 review findings — half-spaced ranges, cross-prefix annotations, markdown wrappers

Round 3 of the adversarial review (against 2eb92dd0e) refuted four
claims; this closes them:

- Half-spaced ranges (`REQ-01 -REQ-05`, `REQ-01- REQ-05`) split at the
  tokenizer before R1's `\s*` can see them and under-selected silently.
  A glued-fragment rule warns when an operator is glued to a full ID
  with an ID-shaped neighbour on the open side and the endpoint pair
  implies an interior.
- Cross-prefix pairs around a separator no longer read as ranges:
  `REQ-02 - (ADR-7)` and `REQ-02 (...) (ADR-7)` are annotations, and
  real ranges are same-prefix by nature. `impliesInterior` now returns
  false on prefix mismatch and computes the gap with BigInt (parseInt
  lost precision past 2^53).
- The token shave now removes markdown emphasis markers and curly
  quotes, so `**None** (per ADR-7)` reaches the placeholder gate and
  `**RANGE-02..RANGE-05**` reaches R1.
- The unanchored-substring cap rises to 2048 (a markdown-link range
  with a long URL cleared 256); the anchored range regexes scan
  linearly and drop their cap.

Tests: 6 new cases (28 total; 377/377 file-wide). Negative control
against the merge-base: 11 defect tests fail, 17 controls pass.

* fix(#3697): round-4 review finding — word operators excluded from glued-fragment rule

`TOREQ-05` is a valid prefix-agnostic REQ-ID, and the glued-fragment
rule read it as `to` + `REQ-05`, warning on the canonical two-ID list
`REQ-01, TOREQ-05`. Glued fragments are now SYMBOL-operator-only
(`..`+, ellipsis, dashes): a word operator glued to an ID is an ID,
not a range spelling.

Tests: word-operator-prefixed ID control (misparse channel silent; the
fixture's ghost-ID warning legitimately fires, so the whole-channel
assertion stays with the registered controls) and an underscore-wrapped
tight-range defect case. 30 targeted cases; 379/379 file-wide; negative
control: 12 defect tests fail on the merge-base, 18 controls pass.

* fix(#3697): round-5 review findings — trailing word-op glue, dot shave, honest wording

- The glued-fragment TRAILING arm takes the word operators back: an ID
  must end in digits, so `REQ-01through` can never be an ID — the
  round-4 TOREQ collision was leading-arm-only, and symbol-only on both
  arms lost the `REQ-01through REQ-05` typo class.
- A trailing run of 2+ dots survives the punctuation shave: `REQ-01..`
  is a glued range operator, not sentence punctuation, and the shave
  was silently eating the `REQ-01.. REQ-05` form.
- The warning now says the line "could not be parsed as" a
  comma-separated REQ-ID list: `**REQ-01**, **REQ-05**` IS such a list
  — the selector just cannot parse decorated tokens — and a warning
  that misstates the input teaches readers to distrust it.

Tests: two new trailing-glue defect cases (32 targeted; 381/381
file-wide). Negative control: 14 defect tests fail on the merge-base,
18 controls pass.

* test(#3697): use t.after for cleanup per CONTRIBUTING test ruleset

CONTRIBUTING bans try/finally inside test bodies (it masks failures);
the approved shape is `t.after(() => cleanup(tmpDir))`. All seven
converted tests are this PR's own additions; the file's pre-existing
instances are untouched.

* chore(#3697): add changeset fragment for the Requirements-line under-selection warning

changeset-lint fails on this PR (fail_missing_fragment): src/phase.cts is a
user-facing surface and the branch carried no .changeset/*.md. Adds the Fixed
fragment via `npm run changeset -- --type Fixed --pr 3744`, symptom-led per
the house format, with the (#3697) backlink.

* refactor(#3697): extract the Requirements-line detector to a testable surface

Round-3 review Blocker 1 requires a fast-check property test over this
detector (`RULESET.TESTS.property-based-testing`: modules implementing
parsing contracts must include at least one), and Blocker 2 requires
limit-1/limit/limit+1 fixtures on its 2048-char token cap
(`RULESET.TESTS.boundary-coverage.fixtures`). Neither is expressible while
the logic is a closure inside `cmdPhaseComplete`: every existing #3697 test
reaches it by spawning the CLI, and a property test cannot pay a subprocess
per generated case.

So the selector and the three detection rules move to module scope as
`analyzeRequirementsLine` (pure, exported) plus
`formatRequirementsLineWarning`, and `cmdPhaseComplete` calls them. This
commit changes NO behaviour: `tests/phase.test.cjs` is untouched here, and
the pre-round suite passes against it unmodified (403/403).

Two things the move makes explicit rather than incidental. The selector and
the detector tokenize the SAME line DIFFERENTLY — the selector strips only
`[` and `]`, the detector also shaves quotes, emphasis and trailing sentence
punctuation — and that gap is deliberate: it is why `ADR-7)` is not selected
while `ADR-7` is still nameable in a warning. They now sit adjacent with the
reason written down, so they cannot drift apart silently.

And the stale citations in the moved comment are corrected. It pointed at
src/phase.cts:833,920,1078 for the `**Requirements**: TBD` seeds, which had
drifted to 1132/1237/1413, and at `templates/roadmap.md:32`, which is
`gsd-core/templates/roadmap.md:32`. Both are now anchored by content.

* fix(#3697): stop the warning claiming a misparse that did not happen

Round-3 review Major 3 and Minor 4. Both are the same defect: the warning
asserted more than the evidence supported.

MAJOR 3 — a correct comma list such as `RANGE-01, RANGE-02 — RANGE-05
deferred` warned "could not be parsed ... Range forms are not expanded;
rewrite the line". Reproduced: it selects RANGE-01, RANGE-02 AND RANGE-05,
i.e. every ID written on the line. Nothing was dropped, and the pinned
control only stayed silent because its pair was ADJACENT (gap == 1), so the
control was passing by accident of the fixture rather than by the rule.

The obvious fix — go silent — is not available. `RANGE-02 — RANGE-05` as a
range and as an annotation separator are textually identical, and no
token-level rule separates them; staying quiet re-opens the exact silent
under-selection #3697 is about. Deciding the ambiguity by assertion in
either direction is wrong. So it is DISCLOSED: the warning now has two
channels, chosen by whether any ID-shaped token was actually left unselected
(`droppedIdShaped`).

  * something was dropped (tight range, glued fragment, inert residue)
    -> "could not be parsed as a comma-separated REQ-ID list", as before.
  * nothing was dropped (only the spaced-operator rule fired)
    -> "contains what reads as a range between two cited REQ-IDs", stating
    both readings and saying explicitly that an annotation separator means
    the line is already correct.

This retires the "could not be parsed" wording for the four #3697-1 spaced
cases too, and that is a deliberate expectation change rather than a fix
counted twice: those lines never failed to parse either. They still warn,
still name the selected IDs, and still assert the endpoint-only marking is
unchanged; #3697-1 now also asserts the misparse channel stays SILENT.

MINOR 4 — `Deferred (see ADR-7)` reported `Unparsed text: ADR-7`, naming a
citation as requirement content it had failed to read. The trigger is
correct and stays: #3697's acceptance criterion asks for a warning "when it
selects zero IDs from a line that is non-empty and is not the `TBD`
placeholder", and inferring placeholder-ness from arbitrary prose is the
free-text heuristic this detector exists to avoid. What was wrong is the
wording, so the non-range arm now says "ID-shaped text that was not
selected" and names the escape the author actually has (`TBD` / `None`).

Tests: #3697-9 (three spaced forms — must warn, must NOT claim a misparse,
must offer both readings) and #3697-10 (`Deferred (see ADR-7)`, `N/A
(tracked in ADR-12)` — must warn, must not say "Unparsed text", must not
diagnose a range, must name the placeholder escape).

Reversion control: reverting the ambiguous channel fails #3697-9 (3 named
tests); reverting the R3 wording fails #3697-10 (2 named tests).

* fix(#3697): cap every token predicate, complete the dash set, cover the boundary

Round-3 review Blocker 2 and Nit 6, plus one self-found finding. All three
are about the detector's own predicates, so they land together.

BLOCKER 2 — the 2048-char budget had no boundary coverage.
`RULESET.TESTS.boundary-coverage.fixtures` requires limit-1 / limit /
limit+1 for any budget parameter. #3697-B1 and #3697-B2 now exercise 2047 /
2048 / 2049 against BOTH predicate families the cap guards, and each asserts
its fixture's exact length before asserting behaviour, so a mis-built
fixture fails loudly rather than passing at the wrong size. Clause (d) of
that rule — an input pushed within reserve-distance of the limit — has no
referent here: this is a hard cap with no reserve constant beside it, and
the test comment says so rather than leaving the omission to be re-derived.

NIT 6 — the cap guarded only the unanchored ID-substring regex. The
anchored range regexes were left uncapped, justified by a comment asserting
they scan linearly. The finding is right that this is informational (they
are anchored; the input is a local ROADMAP.md), but an asserted property is
cheaper to enforce than to defend, so all three predicates now share one
`short()` guard. #3697-B2 is what pins it: at 2049 the anchored scan must
now decline to classify.

SELF-FOUND (RV4 guard-shape census) — the range-operator set is a list this
code fixes at author time over a domain that grows without it, so the round
owes a census of what the enumeration reaches.

  reached:     `..`+, U+2026, U+2013, U+2014, ASCII `-`, to/thru/through
  NOT reached: U+2010 hyphen, U+2011 non-breaking hyphen, U+2012 figure
               dash, U+2015 horizontal bar, U+2212 minus sign
  consequence: a range spelled with any of those is SILENTLY under-selected
               — #3697's own defect, in the code that exists to fix it

Those five close. They are the same operator at a different codepoint and
carry none of the ASCII hyphen's collision risk, because they are not the
REQ-ID separator: `FY-2026-08` is date-shaped only with ASCII hyphens, so a
U+2010 never reaches the ID shape. They therefore join the NOHYPHEN arm
beside `—` and `–`; the strict full-ID-both-sides shape the bare hyphen is
held to is untouched, and #3697-12 pins that.

Still NOT reached, declined with reason rather than left unstated: `→`, `~`,
`..=`, `..<`, `until`, and `up to` (two tokens, so never one operator
token). Each is a symbol or word with an independent non-range use between
two REQ-IDs — the over-warning class #2334 cost three rounds.

Reversion control: reverting the uniform cap fails #3697-B2 (limit+1);
reverting the dash set fails #3697-11 (5 named tests).

* test(#3697): add the fast-check property coverage the parser rule requires

Round-3 review Blocker 1. `RULESET.TESTS.property-based-testing` (CONTEXT.md)
requires modules implementing parsing contracts to carry at least one
fast-check property test asserting a domain invariant, and the round-2 diff
had zero occurrences of `fc.` across its +311 test lines. Five properties,
1,900 generated cases:

  P1  soundness of silence (boundary containment) — for ANY canonical comma
      list of well-formed REQ-IDs, the selected set EQUALS the written set
      and nothing warns. This is the #2334 over-warning invariant and the
      #3697 under-warning invariant asserted as one statement, over
      generated IDs rather than hand-picked ones. It generalises #3697-4b:
      a prefix beginning with a word operator (`TORANGE-05`) is an ID, and
      P1 covers that class rather than the single example.
  P2  completeness — a same-prefix pair with an interior between them,
      separated by any of the nine spaced operators, ALWAYS warns.
  P3  the #2334 invariant — an ADJACENT pair around a separator can drop
      nothing, so it stays silent however it is annotated.
  P4  totality + idempotency — total over arbitrary strings, deterministic,
      and the formatter agrees with the analysis on whether there is
      anything to say (a warn with no text, or text with no warn, is a
      channel that can go silent or noisy on its own).
  P5  containment — every selected ID is ID-shaped and appears verbatim in
      the input.

Honest scoping, since a property test is easy to overclaim: P1, P3, P4 and
P5 hold against the round-2 code as well as this one — they are regression
guards, not bug-finders, and their value is that the invariants are now
stated and generatively checked rather than implied by examples. P2 is the
one that would have failed before the dash enumeration was completed.

fast-check v4 removed `fc.stringOf`, so the ID-prefix tail is built from
`fc.array(...).map(join)` with the alphabet pinned to the selector's own
`[A-Z0-9]` class.

These live in tests/phase.test.cjs rather than a new
`phase.property.test.cjs`: `lint-test-file-count` caps a production module
at 2 test files and phase.cts is already at its allowlisted entry, so a new
file would trade one gate for another.

* docs(#3697): document the ROADMAP Requirements-line grammar

Round-3 review Minor 5 — the change adds net-new user-visible warning output
for a grammar constraint documented nowhere under docs/. `type: Fixed` is
docs-exempt so this does not block, but a warning about a rule the reader
cannot look up is not actionable, and that is worth fixing whether or not a
gate demands it.

Added as a subsection of `phase complete` in docs/CLI-TOOLS.md, beside the
existing SUMMARY artifact-check advisory it is a sibling of: the supported
comma-list form, why ranges are deliberately not expanded, that `TBD` and
`None` are the entire placeholder vocabulary, and what each of the two
warning voices means — including that the range/annotation one may be
reporting a line that is already correct.

Existing file rather than a new one, deliberately: docs/ carries generated
indexes and zh-CN / ja-JP trees, and a new top-level page invites a parity
or index gate this change has no reason to touch.

* fix(#3697): rule-scope the warning-channel discriminator

Self-found at the round's pre-push review, against the Major 3 fix two
commits back. That fix chose the channel from a LINE-GLOBAL question — "was
any ID-shaped token left unselected?" — while the rules that produce the
warning are not line-global. The two disagree as soon as the line carries an
ID-shaped token no rule fired on:

  `RANGE-01, RANGE-02 — RANGE-05 deferred per (ADR-7)`

`(ADR-7)` survives the selector's bracket strip, so the global test called it
a drop and sent the line to the assertive channel — putting the false "could
not be parsed ... rewrite the line" claim back on a correct line. That is
review finding Major 3 returning through a side door, and it directly
contradicts #3697-4, which pins a parenthetical citation as NOT unparsed
residue.

The discriminator is now rule-scoped: R2 is the only ambiguous rule, so the
ambiguous channel requires that R2 fired, that no other rule did, and that
every endpoint R2 fired on was actually selected. The last conjunct is not
redundant — the detector shaves brackets and the selector does not, so R2 can
fire on a `(RANGE-02)` that was never selected, and that IS a drop:

  `RANGE-01 (RANGE-02) — RANGE-05`   -> assertive, correctly

`droppedIdShaped` is replaced by `spacedRangePairs` (R2's hits, so the
channel can ask about the endpoints the rule fired on) and the
`rangeReadingOnly` verdict.

Reversion control: against the line-global rule, #3697-9b fails. #3697-9c
passes under both rules — there the dropped token IS the R2 endpoint, so the
two agree; it is a regression guard, not a bug-finder, and is recorded as
such rather than counted as a second control.

* fix(#3697): hold every dash to the strict range shape, not just ASCII

Self-found at the round's pre-push review, and it CORRECTS a claim made two
commits back. That commit widened the range-operator set by five Unicode
dashes and asserted they "carry none of the ASCII hyphen's collision risk,
because they are not the REQ-ID separator". That reasoning was wrong. The
collision is a property of the SHAPE — `PREFIX-\d+ <dash> \d+` is also a date
(`FY-2026-08`) and a sub-numbered ID (`API-2-01`) — and the shape does not
care which dash sits in the operator slot, because the ID's own separator is
still ASCII either side of it. Measured:

  RANGE-01 (target FY-2026-08)   silent   <- pinned by #3697-4
  RANGE-01 (target FY-2026‐08)   WARNED   <- same line, U+2010

So the widening reintroduced the #2334 over-warning class on a date
annotation. It also exposed that the inconsistency PREDATES this PR: U+2013
and U+2014 were already in the loose arm at ce71dd399, so the en- and em-dash
forms of that same date annotation warned before round 3 ever ran.

One rule for every dash: a tight range spelled with any of the eight must
carry a FULL ID on both sides, exactly as the bare hyphen already had to.
`..`, `…` and the word operators stay loose — no date or sub-number reading
exists between two numbers, so the strict shape would cost them coverage for
nothing.

The cost is a false negative, and it is one the design already accepts:
`RANGE-01, RANGE-02-05` is silent today, deliberately, and now
`RANGE-01, RANGE-02–05` is too. That removes an inconsistency rather than
opening a gap, and a bare `RANGE-02–05` still warns — it selects nothing, so
R3 catches it.

Tests: #3697-13 (date annotation AND sub-numbered ID silent for all eight
dashes), #3697-13b (full-ID tight range still warns for all eight),
#3697-13c (loose operators keep their numeric endpoint), #3697-13d (the
accepted false negative is symmetric, and the bare zero-selection line still
warns).

* fix(#3697): close the round's own pre-push review findings

An adversarial cross-AI review of this round refuted 4 of its 10 claims. All
four were real. Every fix below is to code THIS round introduced.

1. THE SOFT VOICE CLAIMED TOO MUCH (refuted CLAIM 1).
   `REQ-01, (REQ-02), REQ-03 — REQ-05` took the range-reading voice and told
   the author "the line is already correct and nothing needs to change" — while
   `(REQ-02)` had been dropped by the selector, which does not strip
   parentheses.

   The channel choice is still right, and deliberately so: `(ADR-7)` and
   `(REQ-02)` are the SAME shape, so routing on "was anything unselected?" puts
   the false "could not be parsed" claim back on a line carrying a citation —
   the misroute fixed two commits ago. No rule can adjudicate this; the author
   can. So the voice stops asserting the line is correct (it now speaks about
   the SEPARATOR, which is all it has evidence about), and BOTH voices gained a
   factual clause naming ID-shaped text the selector skipped, with the reason
   (brackets are not stripped) and no verdict attached.

2. THE CAP SILENCED A LINE THAT USED TO WARN (refuted CLAIM 2).
   A 2049-char range token warned before this round and went silent after it:
   the "uniform cap" commit bounded the predicate and, with it, the warning.
   That is #3697's own defect, introduced by the fix for a nit.

   The cap bounds the WORK, not the warning. An over-cap token carrying `-` is
   now recorded as unclassified (a linear `includes`, never the unanchored
   regex the cap exists to keep off it) and gets its own voice: "could not be
   checked ... the REQ-ID selection on this line is unverified". Unclassified
   is reported, never treated as clean.

3. THE CAP WAS NOT UNIFORM (review MISSED finding).
   R2 capped the operator token but not its neighbours, so
   `<2049-char ID> .. <2049-char ID>` still ran REQ_ID_SHAPE_RE and BigInt over
   both endpoints unbounded. The glued rule had the same hole. Every
   participant is capped now.

4. PROPERTY P5 WAS VACUOUS (refuted CLAIM 5).
   It drew from a bare `fc.string()`, which over 500 samples produced max
   length 10 and ZERO inputs containing a REQ-ID — the loop body never executed
   an assertion. A containment property that never contains anything is a green
   test measuring nothing. The generator now interleaves real IDs with noise
   and the property ASSERTS it saw them (>50/500), so it can never silently go
   vacuous again. The free-form coverage it was actually providing survives,
   honestly labelled, as #3697-P6.

   The same finding refuted this round's claim that P2 distinguishes pre-round
   behaviour: every operator P2 uses was already in the pre-round operator set.
   P2 is a regression guard, and its comment now says so.

Also: docs/CLI-TOOLS.md repeated the broken channel claim verbatim (review
MISSED finding) and is corrected with the code.

Tests: #3697-9d (soft voice names the skipped ID, never claims the line is
correct), #3697-9e (over-cap token reported as unclassified, still warns),
#3697-9f (R2 and the glued rule cap their neighbours). #3697-B1/B2 now key the
boundary on the PREDICATE's verdict with `warn` asserted true at every length —
asserting `warn === false` at limit+1 was itself finding 2.

* docs(#3697): describe the third voice and the dash rule

Follow-on to the review-findings commit: that commit corrected the docs' claim
about the soft voice but left two things the code now does undescribed.

- There are THREE voices, not two. The over-cap voice ("could not be checked
  ... unverified") arrived with the fix for the review's CLAIM 2 and had no
  entry.
- Dash spellings require a full ID on both sides, and `..` / `…` / the word
  operators do not. That asymmetry is deliberate and load-bearing —
  `PREFIX-<digits><dash><digits>` is date- and sub-number-shaped — so a reader
  hitting `REQ-01-05` and getting silence has no way to find out why. The
  accepted cost (`REQ-01, REQ-02-05` unreported, bare `REQ-02-05` still
  reported) is stated rather than left to be discovered.

Documentation only; no behaviour change.

* fix(#3697): close the continuation review's findings

A continuation of the same adversarial reviewer, run against the reworked
round, refuted 6 of 7 claims. Four were real defects in this round's own work
and are fixed here; the other two are answered rather than changed, below.

1. THE SKIPPED-TEXT CLAUSE WAS ON ONE VOICE, NOT BOTH (refuted CLAIM A).
   The previous commit's message said both voices gained it. Only the soft
   return appended it. The assertive voice now carries it too — and, because
   that voice already names range tokens and inert residue under its own
   clauses, the note is filtered to what those did not already name. A warning
   that says the same token twice is one readers learn to skim.

2. THE CLAUSE'S WORDING WAS FALSE (also CLAIM A).
   It read "brackets and parentheses are not stripped". Square brackets ARE
   stripped by the selector — `[REQ-01, REQ-02]` is the documented form — so
   only parentheses qualify. Corrected in the message and in docs/CLI-TOOLS.md,
   which had inherited the same error.

3. THE OVER-CAP RULE STILL SILENCED A LINE (refuted CLAIM B).
   `oversizedTokens` filtered on `includes('-')`, which misses an over-cap
   OPERATOR: `REQ-01 <2049 dots> REQ-05` warned before this round, R2 declined
   to classify it once capped, and nothing reported it. That is the exact
   regression the field was added to close, one input over. Any token past the
   cap now counts — what it contains is irrelevant when we could not read it.

4. AND THEN OVER-REPORTED ONE (review MISSED finding).
   With (3) in place, a 2049-character CANONICAL REQ-ID was selected by the
   uncapped, fully-anchored selector AND flagged "REQ-ID selection on this line
   is unverified" — a contradiction inside one warning. A token the selector
   took was examined end to end, so it is excluded.

Two findings are answered, not changed:

  CLAIM C — the selector's own `REQ_ID_SHAPE_RE.test` is uncapped. True, and
  deliberate: this round does not touch what gets MARKED, and the pattern is
  anchored at both ends with no nested quantifier, so it is linear. The claim
  that "all predicate paths are capped" was too broad; the DETECTOR's are.

  CLAIM E — `REQ-01, REQ-02<dash>05` is silent for every dash. That is the
  documented, deliberate cost of holding dashes to the strict shape, and it is
  symmetric with ASCII, which behaved that way before this PR. The reviewer is
  right that "without losing a range spelling that should be detected" was too
  strong; a bare `REQ-02<dash>05` still warns.

Tests: #3697-9g (clause on the assertive voice, no repetition, bracket claim
true), #3697-9h (over-cap operator does not silence the line), #3697-9i (a
selected over-cap ID is never called unverified). #3697-9f is rebuilt — its
first version used the SAME id twice, so R2 could not have fired even uncapped
and it proved nothing; it now uses endpoints with a gap and fails when the
neighbour cap is removed.

Reversion control: all four fixes fail a named test when reverted in isolation
(#3697-9h, #3697-9i, #3697-9g, #3697-9f).

* fix(#3697): scope the over-cap exemption to what could actually pair

A second continuation of the same reviewer, against the reworked round,
confirmed the two claims that matter most and refuted three. This closes the
one real defect; the other two are answered below.

CLAIM J / CLAIM K (one defect, found from both directions). The previous
commit exempted EVERY selector-accepted token from `oversizedTokens`, on the
reasoning that the selector is uncapped and anchored so it examined the whole
token. True of that token's SELECTION — and not the same as "no rule was
suppressed by it". Two over-cap valid IDs either side of `..` are both
selected, so both were exempted, and R2 is capped: a line that warned before
this round went silent.

That is the third appearance of one class in this round — the cap suppresses a
check, and the suppression is not reported. Each fix for it over-corrected in
the opposite direction, which is why the rule is now stated in terms of what
was actually suppressed rather than in terms of the token: an over-cap token is
exempt only when it was selected AND nothing beside it could have paired with
it into a range (no range operator, no glued fragment, no second over-cap
token). Everything else is unexaminable and says so.

Two findings are answered, not changed:

  CLAIM M — the reviewer demonstrated, with driven evidence, a contextual rule
  that catches `REQ-01, REQ-02-05` while leaving `FY-2026-08` and `API-2-01`
  silent: recognise `PREFIX-a<dash>b` only when another SELECTED id on the line
  shares that prefix. That refutes this round's claim that the strict-dash
  trade was FORCED, and the claim is withdrawn — it is a design choice. The
  choice stands for this PR: the conservative rule is what ASCII already did
  before #3697, adopting a new contextual heuristic unreviewed at the end of a
  round is how the last three defects in this round were made, and #3697 asks
  for a warning rather than better range inference. Named here so the
  alternative is on the record rather than lost.

  Docs MISSED — CLI-TOOLS said every token over 2,048 characters "is not
  classified at all" and warns. Selection is not bounded; only range detection
  is. Corrected.

Confirmed by the same pass, and worth recording because they are the PR's
load-bearing promises: a 20,000-input comparison of the pre-extraction selector
against HEAD found `mismatches=0` (nothing about which REQ-IDs are MARKED has
changed), and the uncapped selector regex was measured linear from 100k to 800k
characters.

Tests: #3697-9f now asserts the range case is reported rather than silent, and
#3697-9j pins the exemption's scope in both directions. Reversion control:
restoring the blanket exemption fails both.

* fix(#3697): warn on zero selection, as the acceptance criterion asks

`Deferred`, `N/A`, `Pending`, `TBA` and `-` selected no REQ-IDs and stayed
SILENT, while three shipped artifacts said they warned: `docs/CLI-TOOLS.md`,
the `placeholderLed` census comment, and the advice string the command emits
to the user. The asymmetry was the tell — `Deferred (see ADR-7)` warned,
because the citation supplied the ID-shaped residue R3 required, while bare
`Deferred` did not. The claim was written into three places and never
executed once.

This is also #3697's AC-1b/AC-4 verbatim: "warn when `citedReqIds.length ===
0` while the raw capture is non-empty and not `TBD`".

R3b keys on the SELECTION being empty, never on what the prose means, so it
adds no free-text heuristic. It is deliberately not gated on ID-shaped
residue the way R3 is, and the negative space is what settles that: all
fifteen #2334/#2339 fixtures are held silent by non-zero selection or by
`placeholderLed`, and not one of them by the ID-shape gate — measured, not
argued. The gate was buying no negative space while costing the acceptance
criterion.

`tokens.length > 0` keeps an empty line and a comment-only line silent: the
tokenizer strips `<!-- ... -->` before splitting, so the shipped template's
own comment cannot reach the rule.

Selection behavior is unchanged. This warns; it never invents an ID.

Also extracts `warn` to a named const (round 3 review Minor 3) — this commit
adds a disjunct to exactly that predicate, and in the return literal a later
reordering would be a TDZ ReferenceError rather than a reader-visible error.

Tests: #3697-14 (six zero-selection lines warn and tick nothing, and the
warning names the TBD/None escape), #3697-14b (five placeholder spellings
stay whole-channel silent), #3697-14c (comment-only line stays silent).
Fail-first controls: all six #3697-14 cases fail against the pre-fix tree;
-14b and -14c pass at both ends, which is correct — they pin silence the
widening must preserve.

* fix(#3697): name the REQ-ID a glued delimiter dropped

`RANGE-01; RANGE-02` selects only RANGE-02 and marks only RANGE-02, with
`requirements_updated: true` — #3697's own half-success failure mode, reached
by one wrong delimiter, and silent before this rule. It is the issue's AC-1a
("a warning whenever the line contains ID-shaped content that the tokenizer
did NOT select") at the shape most likely to be typed by accident.

Round 4 review rated this Major rather than Blocker on the ground that the
case is indistinguishable from a parenthesised citation, since `(ADR-7)` also
shaves down to a bare ID. At the RAW token level it is distinguishable, and
that is what makes the rule shippable: `REQ-01;` is shaved of a trailing
DELIMITER, `ADR-7)` of a citation wrapper. R4 keys on that shave class and
requires the token to sit outside any parenthetical.

Measured before implementing: 0 false positives and 0 false negatives across
21 probes, including all fifteen #2334/#2339 negative-space fixtures. A first
cut without the parenthetical test scored 3 false positives — every one of
them a colon inside a citation (`(see ADR-7: section 3)`) — which is why that
test is the rule's boundary rather than an optimisation.

Adds the delimiter census the module did not have. The range-operator domain
was already censused; the comma-substitute domain was not. Swept 26
spellings: exactly two produce a silent under-selection, `; ` and `: `. Every
other spelling either selects both IDs or selects none and already warns. The
review hand-listed the semicolon; the colon is the sibling that sweep found,
and it fails identically.

`rangeReadingOnly` now excludes an R4 hit — the ambiguous voice claims nothing
was dropped, and must not speak for a line where something demonstrably was.

Tests: #3697-15 (four delimiter shapes warn, name EVERY dropped ID, and tick
exactly the unchanged selection), #3697-15b (three citation forms stay
whole-channel silent). Fail-first control: all four #3697-15 cases fail
against the previous commit's tree; -15b passes at both ends, pinning the
boundary the widening must not cross.

* fix(#3697): give the Requirements-line warning a stable machine kind

The warning's kind existed only in the prose of its message, so every consumer
and every test had to regex an English sentence — and rewording a message
silently un-asserted the tests that pinned it. Round 4 review Major 3.

The repo already had the settled seam for exactly these semantics.
`CONTEXT.md` records `diffLiveConfig` emitting `kind:'unverified'` for a
truncated scan, which is precisely this module's third voice; and
`WAVE_CLEANUP_WARNING` in `src/worktree-safety.cts` carries codes for the same
reason. ADR-3473 Decision 3 ("failure is a value") points the same way.

`formatRequirementsLineWarning` now returns `{ code, message }` instead of a
bare string, which also settles round 4 Nit 3 — `null` still means CLEAN, a
legitimate value, but the success arm is no longer a naked string one field
away from the shape the ADR standardises on.

The kind is carried ALONGSIDE the prose, never instead of it. `warnings[]` is
a documented `string[]` in `phase complete`'s JSON output, rendered by
execute-phase.md's "If has_warnings is true" step, so re-typing its elements
would be a breaking output-contract change for a shipped command. The code is
emitted as its own additive `requirements_line_warning` field, absent
entirely when the line is clean.

Vocabulary, exported so tests key on it rather than on string literals:
`req-line-misparse`, `req-line-range-reading`, `req-line-unverified`.

Tests: channel ROUTING in #3697-9/-9b/-9c/-9d/-9e/-9g/-10 now asserts the code;
message-content assertions stay where the user-visible wording is itself under
test. #3697-16 pins the code end-to-end through the CLI's JSON for four line
shapes and asserts warnings[] is still a string[]; #3697-16b pins that a clean
line emits no kind at all, because a field present on every run carries no
information. #3697-P4 holds kind-and-message-appear-together and
kind-is-in-the-declared-vocabulary over arbitrary input, so a channel added
later cannot ship without one.

* test(#3697): pin the divergence against the second parser of the same line

CLAUDE.md, KNOWN DEFECTS & ANTI-PATTERNS: "Generative Fix Divergence: when
sharing constants/arrays/parsers between parallel surfaces, add a parity
assertion test that fails if they diverge." Round 4 review Major 2.

`normalizePhaseReqIds` (src/gap-checker.cts) parses the SAME ROADMAP
`**Requirements:**` value — its own docblock says callers "may pass the
roadmap value through verbatim" — and diverges on four axes. Measured, not
inferred:

  line                                    phase complete      gap-checker
  RANGE-01..RANGE-05                      []                  5 IDs
  None (per ADR-7)                        []                  ["ADR-7"]
  (REQ-02)                                []                  ["REQ-02"]
  REQ-01a                                 []                  ["REQ-01a"]
  REQ-01, REQ-02                          both                both

This pins the divergence rather than removing it, which is the review's
second option and the correct one here: unifying the two would change what
`phase complete` MARKS, and "the ledger-writing set is byte-identical to base"
is the one invariant this PR holds fixed. Every axis is now asserted in BOTH
directions, so drift on either side fails here instead of widening silently.

The range axis is a DELIBERATE disagreement and is labelled as such — #3697
declines range expansion in terms ("I am not asking for range syntax to be
supported") while gap analysis adopted it under #1269.

The placeholder axis is the one worth reading twice: `None (per ADR-7)` is a
declared-empty line to `phase complete`, which reads the lead token, and a
one-requirement line to gap-checker, which strips parentheses first so the
citation survives its ID-shape filter. That is a citation being reported as a
requirement.

#3697-17b states the cost concretely: one line, five requirements in scope to
gap analysis and zero to phase complete. This PR is what makes that
contradiction visible, by finally giving the silent side a voice.

* fix(#3697): stop the skipped-text rider reporting a date, and close the 4b channel gap

Two round 4 review minors, both about a warning saying something it cannot
support.

MINOR 2 — false rider content. `REQ_ID_SUBSTRING_RE` is unanchored, so
`FY-2026-08` matches as `FY-2026` and lands in `unselectedIdShaped`.
`REQ_RANGE_TOKEN_RE`'s entire strict-dash arm exists to keep that shape
silent, and #3697-4 pins `RANGE-01 (target FY-2026-08)` as producing no
warning at all — but whenever some OTHER rule fired on a line that also
carried a date annotation, the rider told the author to "check whether any of
it is a requirement" about a date. Not a false warning, since the line was
warning anyway; false CONTENT, in the #2334 voice, through the side door.

Filtered at the MESSAGE rather than in the analysis: `unselectedIdShaped`
stays a faithful record of what the selector skipped — it is documented as a
fact that never routes — while the user-facing clause declines to assert
requirement-ness about a shape the design already ruled unadjudicable.
#3697-18b is the other half, so the filter cannot become a silencer: a
genuinely dropped REQ-ID is still named.

MINOR 1 — `#3697-4b` asserted only that the ASSERTIVE channel stayed silent,
so a regression routing `RANGE-01, TORANGE-05` into the AMBIGUOUS channel
would have passed. Whole-channel silence is not available on that fixture (the
pre-existing ghost-ID warning legitimately fires on the unregistered
`TORANGE-05`), so the precise assertion is that no Requirements-line warning
of ANY kind was emitted. The machine code added earlier in this round is what
makes that statable; before it, "both channels" could only have meant a second
prose regex.

* docs(#3697): record the Requirements-line seam in CONTEXT.md

CLAUDE.md names the CONTEXT.md glossary as a PR gate, and
`get_cochange_context(src/phase.cts, 45d)` ranks CONTEXT.md 4th at 25
co-changes — above src/init.cts and src/roadmap.cts. This PR introduced a
named seam, three warning kinds, a bound, a rule taxonomy and a deliberate
cross-parser divergence, and recorded none of it. Round 4 review Major 4.

The precedent is explicit rather than inferred: the directly analogous seam
is already there as `LIVE-CONFIG.GUARD.SEAM.truncation`, including its bound
and its boundary obligation — and that entry is the one this module's third
voice was modelled on.

Eight predicates, in the machine-oriented section beside it:

  .module                    the two exported functions and the code vocabulary
  .selector-identity         citedReqIds is byte-identical to base and is the
                             only thing reaching the ledger — a change to what
                             phase.complete MARKS is outside this contract
  .rules                     R1 / R2 / R2' / R3 / R3b / R4 / over-cap
  .kinds                     the three codes, and why they ride beside
                             warnings[] rather than inside it
  .cap                       2048, neighbours included, and the boundary rule
  .placeholder               the gate that actually holds the negative space
  .census-domains            both open domains with their NOT-reached members
  .gap-checker-divergence    the four axes, pinned not unified

The changeset type is `Fixed`, which exempts this PR from the docs/
co-change requirement — but the glossary gate is separate from that
exemption, and the 2048 cap in particular is a machine-canon-shaped fact
that until now existed only inside a source comment.

`docs/CONTEXT-INDEX.json` regenerated (269 predicates); lint:generated-sync
confirms all six targets in sync.

* docs(#3697): document what the command now does, in one changeset sentence

DOCS. The grammar section predated this round's two new rules, so it
under-described the behaviour it exists to make lookup-able:

- The placeholder paragraph enumerated three words; the rule is a DEFAULT.
  Any wording that selects no REQ-IDs warns, and the placeholders are matched
  as the LEAD token, so `None (per ADR-7)` and `**None**` are declared-empty
  too. The comment-only line is called out, because "any other wording" would
  otherwise read as covering the shipped template's own `<!-- ... -->`.
- The comma rule was implicit. `REQ-01; REQ-02` marks only REQ-02, and it is
  the quietest way to lose a requirement on this line — `requirements_updated`
  reads `true` either way — so it gets its own paragraph, with the
  parenthetical exemption stated beside it.
- The machine kind is documented where a consumer would look for it, with the
  instruction to key on the kind rather than the wording.
- The skipped-text note no longer implies it reports date shapes; it
  deliberately does not, and silently omitting that left the doc promising the
  behaviour this round removed.

CHANGESET (round 4 review Minor 4). CONTRIBUTING.md's format is
`**<Bold user-visible change>** — <symptom-led explanation>.` and both
canonical examples are one sentence; this fragment ran three. Now one, and
covering what the round actually delivers rather than only the range shape it
started from.

* test(#3697): keep phase.test.cjs off the docs-guard exemption fingerprint

A comment added earlier in this round named `docs/CLI-TOOLS.md` by path. The
docs-guard exemption ratchet (#3753 FIX 3) fingerprints literal `docs/`
references in exempt test files and fails when a new one appears, so that
comment turned four green gates red — `ci-docs-guard-registry` and the
registration lint — for a file that reads no documentation at all.

Caught by diffing the full suite's failing-name set against the same suite run
at `upstream/next` in a probe worktree: 33 of 37 failures reproduce at base
(install / config-home / shadowing tests under the sandbox HOME), and exactly
these 4 did not.

Rephrased rather than baselined. Adding the path to
DOCS_GUARD_EXEMPT_DOCS_PATHS is the sanctioned response when a test genuinely
starts READING a new docs path — the violation text asks the author to
re-confirm the exemption still holds. Nothing here reads documentation; the
guard matched prose. Baselining would have recorded a coupling that does not
exist and made the next reader wonder what phase.test.cjs does with
CLI-TOOLS.md. The comment still names where the contract is written, just
without planting a path string.

* fix(#3697): close four defects this round's own pre-push review drove

An adversarial cross-AI review of this round, run before the push, returned 6
CONFIRMED and 4 REFUTED. Every refutation was driven against the built tree,
and every one was a shape the author had not probed — the rules were correct
across the probe set and wrong just outside it.

(1) R4 FALSE POSITIVE, and it is the #2334 over-warning class arriving through
the rule added to close a different hole. `REQ-01, see ADR-7: section 3` fired:
`ADR-7:` is the same shave class as `REQ-01;`, and the parenthetical test does
not reach a BARE citation. The FP probe that scored this rule 0/0 only ever
tested the parenthesised form.

Fixed by requiring the dropped id's prefix to agree with a SELECTED id — the
module's own idiom, not a new heuristic: `reqEndpointsImplyInterior` already
demands an agreeing prefix for the same reason. Cost, stated in the census: a
dropped id whose prefix is on no selected id (`REQ-01, FOO-02: x`) stays
silent. Same trade the strict-dash rule takes — under-report a rare shape
rather than over-report a common one. Pinned as a declared blind spot by

(2) R4 FALSE NEGATIVE, on the DOCUMENTED form. `[REQ-01; REQ-02]` dropped
REQ-01 silently: the selector strips square brackets and R4's raw scanner did
not. The bracket spelling the shipped template recommends was the one shape the
rule could not see.

(3) The rider filter suppressed a REGISTERED requirement. `API-2-01` is a legal
requirement id — gap-checker's `parseRequirements` accepts it from
REQUIREMENTS.md — so a `\d+-\d+` filter hid a genuinely dropped requirement
behind a rule meant only to hide dates. Narrowed to a four-digit year segment.
The earlier #3697-18 case asserting `API-2-01` should be suppressed is REMOVED,
and the removal is recorded in place: its premise was refuted, it was not
inconvenient.

(4) An INVISIBLE line warned. A lone U+200B carried a token to the parser while
reading as empty to the author, so R3b fired with nothing on screen to explain
it. Zero-width and format characters are now stripped — stripped rather than
treated as delimiters, because splitting on one would fabricate two fragments
out of one ID.

Also corrects the documentation the same review found overstated: the line is
split on commas AND whitespace, and the ID shape is matched case-insensitively,
so `REQ-01 REQ-02` and `req-01, req-02` both select and neither warns. That was
pre-existing selector behaviour; this round is the one that asserted the docs
were true of it.

Tests: #3697-19 (four invisible-only shapes), -19b (embedded zero-width is
stripped, not split on), -19c (three citation forms), -19d (both bracket
spellings), -19e (both halves of the rider boundary), -19f (the declared blind
spot), -19g (the two documented tolerances). 500 tests in phase.test.cjs, 0
failures; lint:ci clean.

* fix(#3697): generalise the drop rule, and stop the invisible fix hiding a drop

The pre-push review's continuation refuted six of seven follow-up claims. The
first one is the one that mattered: the invisible-character fix committed in
c7dce173a INTRODUCED #3697's own defect. Stripping zero-width characters from
the detector wholesale made `REQ-01<ZWSP>, REQ-02` go SILENT — the selector
really does drop REQ-01, and the strip removed the only evidence of it. The
test written alongside asserted the tokens and the empty R4 result and never
asserted `warn`, so it DOCUMENTED the bug rather than catching it; that
omission was the reviewer's own MISSED finding.

An invisible is two different questions about one character, and the fix is to
stop conflating them: absence-of-content for the empty test, DECORATION on a
token for the drop rule. Neither is a reason to delete it from the line.

R4 is generalised accordingly, because the continuation drove four more shapes
a trailing-delimiter-only regex could not see — `REQ-01 ;REQ-02`,
`REQ-01 :REQ-02`, `**REQ-01;** REQ-02`, the backticked form — plus
`**REQ-01**, REQ-02`, where emphasis alone defeats the selector. These are one
class: decoration on a token the selector then cannot take. One rule, not four
patches; patching them individually is how a list stays short and wrong.

PARENTHESES ARE NOT DECORATION, and the suite caught me learning that: shaving
them made `REQ-01, (REQ-02), REQ-03 — REQ-05` report a glued delimiter that was
never there and broke #3697-9d's channel routing with it. A parenthesis is this
rule's citation marker.

The rider stops adjudicating an undecidable shape. `API-2-01` is a legal
requirement id and `API-2026-08` is too, while `FY-26-08` and `FY-2026-08-15`
are dates — no regex separates them, and both filters this round tried scored a
miss in each direction. It now NAMES the token and states the ambiguity, which
is the same thing the two warning voices already do about a range separator.
Filtering hides a real dropped requirement; reporting it bare asks the author
whether a date is a requirement; saying "this may equally be a date" does
neither.

The census and the docs are corrected to what the code does, including the part
that is NOT complete: the prefix gate does not stop a citation that SHARES a
selected prefix (`ADR-01, see ADR-7: sec 3` fires), and nothing at token level
separates that from a real drop. A prose heuristic on "see" is the free-text
detector this module exists to avoid, so the honest move is to say so.

CLAIM 17 — the invariant that actually matters — came back CONFIRMED on a
20,000-run fast-check property over arbitrary Unicode: `citedReqIds` is
identical to upstream/next's for every input, and marking is untouched.

513 tests in phase.test.cjs, 0 failures; lint:ci clean.

* fix(#3697): gate the drop rule on evidence, and stop an unmatched paren swallowing the line

Third pass of the round's own pre-push review, scoped to regression-hunting
rather than further polish. Three findings, all driven, all mine.

R4 OVER-WARNED on markdown styling. `REQ-01, see **REQ-7** for context` claimed
a dropped requirement: the previous cut treated any shaved decoration as
evidence, and emphasis is not evidence. Nothing separates that line from
`**REQ-01**, REQ-02` meaning to list one, so the rule now requires a positive
signal — a glued `;`/`:` (a list separator was INTENDED) or an invisible (the
token is CORRUPTED; nobody types one on purpose). Emphasis alone falls back to
the skipped-text rider, which names the id without asserting a drop, exactly as
`(REQ-02)` is handled. That is the #2334 class caught one cut before shipping.

R4 UNDER-WARNED on `**REQ-01**; REQ-02` — one shave pass cannot reach a wrapper
sitting behind a delimiter. Shaves to a stable point now.

The range OPERATOR lost its invisibles handling. `REQ-01 <ZWSP>..<ZWSP> REQ-05`
went silent, because the previous commit removed the invisible strip from BOTH
the tokenizer and R4 when only R4's was wrong. An invisible is two questions
about one character: for the classification rules it is noise and is stripped
from the token; for the drop rule it is the evidence and must survive on the
raw line. Stripping in both places hid a dropped id; stripping in neither hid a
range. The reviewer's MISSED finding named the missing control — regression
tests covered invisibles inside ids and not beside operators — and #3697-19i is
that control.

UNBALANCED PARENTHESES swallowed the line. `REQ-01, (note REQ-02; REQ-03`
reported nothing: a running-depth counter left the unclosed `(` open through
end-of-line, so every genuine drop after it inherited citation immunity. A
parenthesis confers that immunity only as part of a MATCHED span now — an
unmatched one is a typo, not a citation.

CLAIM 22 re-confirmed on a fresh 20,000-run property over arbitrary Unicode:
`citedReqIds` identical to upstream/next, marking untouched, warnings appended.

522 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite carries zero
head-only failures against a probe worktree at upstream/next.

* fix(#3697): make delimiter ADJACENCY the rule, and delete matched citations outright

Fourth and final pass of the round's own pre-push review. Three findings, and
they shared one root cause, so this is a narrower rule rather than a longer list
of shapes.

TOKEN-WIDE PAREN IMMUNITY LEAKED. `REQ-01, REQ-02;(note) REQ-03` is a single
whitespace token, so a matched parenthetical inside it conferred immunity on the
`REQ-02;` sitting OUTSIDE the parens, and the drop went silent. Matched spans
are now deleted from the line outright — which states what is actually meant,
that for this rule a citation is not on the line — and an UNMATCHED paren is a
typo that confers nothing. That also retires the running-depth counter whose
previous bug was the mirror image: an unclosed `(` swallowing the rest of the
line.

DECORATION WAS TESTED TOKEN-WIDE, so `REQ-01, see **REQ-7**; next topic` was
reported as a dropped requirement. It is a citation with sentence punctuation.
The rule is now ADJACENCY: styling is stripped, then the `;`/`:` must be
touching the id. `REQ-01;`, `;REQ-02` and `**REQ-01;**` qualify;
`**REQ-01**;` does not, because outside the styling that character is
punctuation. An invisible needs no adjacency test — nobody types one on
purpose, so anywhere in the token it is corruption rather than intent.

`**REQ-01**; REQ-02` therefore goes silent, and the test row asserting
otherwise is inverted rather than deleted quietly: it was added one commit ago
on the reasoning this pass refuted, and nothing distinguishes it from
`see **REQ-7**; next topic`.

Worth recording plainly: three successive cuts of this rule fired on a
citation, and each fix was a narrower definition of EVIDENCE, never a longer
list of shapes. The list-lengthening instinct is what produced the bug each
time.

The review's last MISSED finding named the missing control — the paren tests
all surrounded matched spans with whitespace, so none covered a span sharing a
token with an id outside it. #3697-19j carries both directions now.

526 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite zero
head-only failures against a probe at upstream/next; the invariant that
`citedReqIds` is identical to upstream re-confirmed on 20,000 arbitrary
Unicode inputs.

* fix(#3697): state R4's real boundary, and stop the over-cap voice masking a drop

Two defects, both found by this round's own pre-publication body claim-audit.

1. A DEMONSTRATED drop was discarded by the unverified voice. On
   `REQ-01, REQ-02: <2049 chars>` the analyzer names REQ-02 in
   delimiterDroppedIds and the formatter then reported `req-line-unverified`,
   whose message never mentions it — the one actionable finding masked by the
   token beside it. The over-cap channel now excludes a line carrying an R4
   hit, exactly as rangeReadingOnly already did and for the same reason: that
   voice's whole claim is that nothing could be checked, and R4 has already
   checked something. The assertive channel still carries the over-cap rider,
   so nothing about the cap is traded away. Pinned by #3697-19l, which fails
   against the pre-fix build and nothing else does.

2. Three shipped artifacts asserted behaviour the code does not have — the
   same class as this PR's round-4 blocker, re-committed. CONTEXT.md's rules
   predicate, the CLI tools reference, and the warning's own advice string all
   listed markdown emphasis as an R4 trigger. It is not: styling is shaved
   BEFORE the test and tolerated around an id, never a trigger on its own, so
   `**REQ-01**, REQ-02` and `**REQ-01**; REQ-02` are both silent. The trigger
   is exactly a glued `;`/`:` or an embedded invisible.

   The census predicate was wrong in a second way. Its 26-spelling separator
   sweep found only `;` and `:` because the sweep was SYMMETRIC-ONLY and
   therefore biased: one-sided attachment drops silently for every punctuation
   outside the set — `/ | & + . > \` and the full-width and non-ASCII forms
   `; , ؛` all measured silent. The domain is wide open and R4 covers two
   characters of it. Said plainly in all three places rather than widened
   here: every previous widening of this rule first fired on a citation, so it
   is not done blind at the end of a round.

Both blind spots are now PINNED as tests (#3697-19m styling-only, #3697-19n
one-sided separators) so the documents and the code cannot drift apart again —
which is what the round-4 blocker asked for.

tests/phase.test.cjs: 542 tests, 542 pass, 0 fail, 0 skip. lint:ci rc=0.
Both CONTEXT-INDEX consumers regenerated.

* fix(#3697): re-sweep the separator census properly, and say what it really found

The round-4 census in src/phase.cts concluded "exactly two — `; ` and `: `"
from a 26-spelling sweep. That conclusion was forced by how the sweep was
built, not by the code: it swept the ONE-SIDED form (`REQ-01; REQ-02`) for the
semicolon and colon, and only the BARE and SYMMETRIC forms (`|`, ` | `) for
every other separator. Different members of the domain were tested in
different shapes, so no other answer was reachable. Caught by this round's
pre-publication claim-audit of the response comment, reading the census
comment against its own swept list.

Re-swept fully crossed and driven through the built artifact: 21 separators x
{bare, trailing-space, leading-space, both-spaces} = 84 combinations. 26 select
both ids, 24 under-select and already warn, and 34 UNDER-SELECT SILENTLY. All
34 are one shape — a separator glued to exactly one of the two ids, e.g.
`REQ-01/ REQ-02` or `REQ-01 /REQ-02` — for every punctuation except `,` and
the `;`/`:` that R4 covers.

So R4 covers TWO CHARACTERS of a wide-open domain. That is now what the census
comment, the CONTEXT.md census-domains predicate and the CLI tools reference
all say. The set is deliberately not widened here: three successive cuts of
this rule fired on a citation, and a fourth at the end of a round with no
adversarial pass is how each of those got in.

Second false passage in the same block: styling-only decoration was described
as "left to the skipped-text rider, which names the id". A rider only exists
inside a message, and a message only exists once some rule sets `warn` — so on
a line where nothing else fires, `REQ-01, **REQ-02**` is wholly silent.
Describing it as handled reads as coverage. #3697-19m already pins the silence.

so the test matches the documented claim.

tests/phase.test.cjs: 552 tests, 552 pass, 0 fail, 0 skip. lint:ci rc=0.
Both CONTEXT-INDEX consumers regenerated.

* chore(#3697): regenerate both CONTEXT-INDEX.json after rebasing onto next

Rebased onto next @ f4fefb0be. The two generated indexes conflicted on the
replay and were resolved by regenerating, not by hand-merging:
`npm run gen:context-index` for docs/CONTEXT-INDEX.json and
`node examples/dynamic-context-management/gen-context-index.cjs --write` for
the example's copy. Against next, each now differs only in the eight
PHASE.REQ-LINE.SEAM.* predicates this PR adds (plus the PHASE class and the
count); every base-side change (the SEAM.* predicates, the ADR-3942 value
rewrites) is carried. `lint-example-parser-parity` and
`gen-context-index --check` both pass.

* fix(#3697): an over-cap token outranks the ambiguous range voice

Round 7 review, Minor 1. `rangeReadingOnly`'s guard conjunction checked
`hasGluedRangeFragment`, `inertIdShaped` and `delimiterDroppedIds` but not
`oversizedTokens`, so a line carrying a clean, fully-selected spaced range
*and* an unrelated token past the 2048-char scan cap was coded
`req-line-range-reading` — a code CONTEXT.md's PHASE.REQ-LINE.SEAM.kinds
predicate documents as "nothing was dropped" — over a token no rule (R1-R4
all skip over-cap tokens) had ever examined. Since the PR tells machine
consumers to branch on `.code` rather than parse prose, that is a false
"nothing to verify further" signal.

The review's suggested fix was to add `oversizedTokens.length === 0` to the
conjunction. That clause is right and is here, but on its own it routes the
line to the ASSERTIVE channel: `req-line-misparse`, whose message says the
line "could not be parsed as a comma-separated REQ-ID list" on a line where
every ID present was in fact selected. That is the #2334 over-warning class
this module's own channel-selection docblock exists to prevent — a false-clean
code traded for a false-assertion one. So the correct destination is
`req-line-unverified`: the line was not CHECKED, which is not the same as
clean, and nothing on it demonstrably failed to parse either.

Both non-assertive voices were carrying their own inline copy of the same
"nothing was demonstrably dropped" conjunction, and that duplication is what
let them drift: `rangeReadingOnly` omitted the cap, while the over-cap channel
excluded a spaced range wholesale via `!hasSpacedRange`. Extract it once as
`nothingDemonstrablyDropped` and have both read it. The new predicate is a
strict superset of the old `!hasSpacedRange` guard — `spacedRangePairs.every()`
is vacuously true when no spaced range fired — so the over-cap channel's
behaviour on every line without a spaced range is unchanged, and R2 firing on
an endpoint the selector did not take still routes to the assertive channel.

Exactly one input class changes routing: a clean fully-selected spaced range
beside an unexamined over-cap token, which moves from `req-line-range-reading`
to `req-line-unverified`. Pinned by `#3697-19n` (the twin of `#3697-19l` on
the other side of the boundary — there a demonstrated drop outranks the
unverified voice, here the cap outranks the ambiguous one), with `#3697-19o`
as the negative control asserting a clean range with nothing over the cap is
still a range reading.

* docs(#3697): state the warning-code precedence, and what the cap condition actually is

Two surfaces, one point. The round 7 finding cited CONTEXT.md's
PHASE.REQ-LINE.SEAM.kinds predicate as the documentation of what
`req-line-range-reading` claims, and it was right to: that predicate said "R2
alone fired on selected endpoints; nothing was dropped" with no mention of the
scan cap, describing a line the code could not distinguish from one carrying a
token it never examined.

The range reading is the weakest of the three claims and yields to the other
two: a demonstrated drop makes the line a misparse, and a token the cap left
unclassified makes it unverified. `docs/CLI-TOOLS.md` gains that sentence; the
CONTEXT.md predicate gains it plus two precision points that this round's own
pre-push adversarial review extracted over three passes, each with a driven
counterexample I reproduced before acting on it:

  * "nothing was dropped" overstates the rule. The discriminator is RULE-scoped
    by design (round 3, Major 3), so that a parenthesised ID-shaped token does
    not re-open the #2334 over-warning class — `(REQ-02)` is indistinguishable
    from `(ADR-7)` at token level and is carried by the skipped-text rider, never
    by this code. `REQ-01, (REQ-02), REQ-03 - REQ-05` selects three, names REQ-02
    as skipped, and is still a range reading. The predicate now says "no rule
    named a dropped ID", which is what the code tests.

  * The deferral condition is `oversizedTokens` being non-empty, NOT the presence
    of a token past the cap. Those differ: a long token the selector itself took
    can be exempt, because selection is uncapped and anchored and such a token
    was therefore examined. `RANGE-01 - RANGE-05, R-<2047 sevens>` yields
    `oversizedTokens=[]` and stays a range reading. Two earlier attempts to
    characterise WHEN the exemption applies were both refuted — the neighbour
    test admits any non-short neighbour, not just a range operator — so the
    predicate now states the condition and defers the exemption's own rule to
    SEAM.cap rather than paraphrasing it a third time.

Verified after the edit: deferral holds if and only if the cap left a token
unclassified, across all three counterexamples plus controls, and an exempt
selector-taken long token is exhibited. No behaviour change — predicate text,
one CLI-reference paragraph, and the two regenerated indexes. Predicate count
unchanged at 285 across 22 classes, 0 duplicate ids; lint-example-parser-parity
and both `--check` generators pass.

* test(#3697): rename the round 7 regression pair — 19n was already taken

Self-found immediately after the push, before anything was published to the
review thread. The two tests added this round were named `#3697-19n` and
`#3697-19o`, and `#3697-19n` was already in use: it is the declared-blind-spot
case for one-sided separators in both attachment directions, generated inside a
loop with a computed label, which is why a grep for a literal `test('#3697-19n'`
did not find it. The PR body's own *Declared blind spots* list already refers to
`#3697-19n` with that meaning.

Nothing failed. Duplicate test names do not error, and the suite stayed green at
567/567 — which is the argument for fixing it rather than against. Two concrete
costs: any TAP name-set differential collapses same-named tests under `sort -u`,
so one of the two becomes invisible to exactly the did-not-run and pass->fail
checks that name-set comparison exists to perform; and a reviewer reading
"#3697-19n" in the body now gets a different test than the one the body means.

Renamed to `#3697-19p` (over-cap beside a clean range) and `#3697-19q` (its
negative control), the next free ids in the series; the pre-existing `#3697-19n`
is untouched at its original 4 occurrences. Cross-references inside the renamed
block were updated with them. Negative control re-run under the new ids: 19p
still fails against pre-fix source and passes after, 19q passes at both ends.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-01 14:36:57 -04:00
0xdhx
900504f985 fix(#3776): decide nothing-to-commit from the staged diff, not from staging success (#3859)
* fix(#3776): decide nothing-to-commit from the staged diff, not from staging success

`cmdCommit`'s empty-diff guard tested `stagedPaths.length === 0`, but
`stagedPaths` records paths whose `git add` exited 0 — "did staging
succeed", not "is there anything to commit". Staging an already-committed,
unmodified file succeeds while contributing no diff, so the guard was
reachable only when every named path was missing from disk.

The ordinary empty-diff case therefore fell through to `git commit`, where
the only thing converting the failure back to `nothing_to_commit` was a
string match on git's output. Git runs the pre-commit hook before it
decides there is nothing to commit, so a rejecting hook pre-empted that
match and the caller was handed `commit_failed` carrying a gate message
about a commit that had nothing to gate.

Ask git whether the staged paths actually differ instead. Two conjuncts
are load-bearing: the `length === 0` short-circuit keeps the
all-missing-paths case exact (a pathspec-less `diff --cached` would test
the whole index, so unrelated staged work would suppress the guard), and
`!isMergeInProgress` keeps a merge from being abandoned — during a merge
git refuses a partial commit, so the pathspec describes nothing about what
would land.

Nine regression cases in tests/commands.test.cjs cover the brief's six
acceptance criteria plus the merge interaction. Against the pre-fix build
exactly one fails; the other eight pin behaviour that was already correct.

Residual sibling of #2608/#2693, which covered `git add` failing; this
covers `git add` succeeding and contributing nothing.

* fix(#3776): exempt a cherry-pick too, but never a revert

The empty-diff guard must not decide from a pathspec git will not honour.
That test was merge-only; git refuses a partial commit during a cherry-pick
for the same reason, so the guard would have fired there and reported a
silent `nothing_to_commit` where the pre-fix code surfaced git's refusal.

The three sequencer states do not agree, so this is driven rather than
reasoned by analogy (git 2.54):

  MERGE_HEAD        fatal: cannot do a partial commit during a merge.
  CHERRY_PICK_HEAD  fatal: cannot do a partial commit during a cherry-pick.
  REVERT_HEAD       permitted; behaves like an ordinary commit.

REVERT_HEAD is therefore deliberately excluded: enumerating it alongside the
other two — the obvious move — would suppress this fix during a revert and
reintroduce the very misreport it removes. Both new states are pinned by a
test, and the revert arm fails against the pre-fix build exactly as AC1 does.

`canScope` keeps its narrower merge-only test on purpose; widening it would
change pre-existing cherry-pick behaviour, which is outside this fix.

* fix(#3776): probe the working tree, not the index

`git commit -- <paths>` is a PARTIAL commit: it records the working-tree
content of those paths and ignores what is staged. The guard was probing
`git diff --cached` — the index — which answers a different question than
the commit asks.

Driven on the same path, in this order: `git add` an unmodified file, then
write to it, then probe.

  git diff --cached --quiet -- p   rc 0   ("nothing staged")
  git diff --quiet HEAD -- p       rc 1   ("the tree differs")
  git commit -m m -- p             committed the new content

So a working-tree write landing between the `git add` above and the probe —
another process in a shared checkout, which this project explicitly supports
— would let the guard report `nothing_to_commit` for a call that would have
recorded that content. Probing `HEAD` asks the question the commit answers.

Not reachable through this function single-threaded, because the staging
loop re-adds every named path immediately beforehand, so index and working
tree agree at the probe. The change is correctness by construction rather
than a fix for an observed miscommit.

An unborn HEAD makes `diff HEAD` fatal; that falls through to the commit as
any other probe error does, and is now pinned by a test — the first commit
in a repo must not be swallowed by an empty-diff guard.

Both added probes are now gated on the guard being able to fire at all, so
an unscoped commit and an `--amend` pay for neither.

Found by adversarial pre-filing review; the index/worktree distinction was
not something my own path-shape probes could have surfaced.

* test(#3776): pin the assume-unchanged boundary; correct a stale comment

`git update-index --assume-unchanged` makes `git add` stage nothing and makes
BOTH diff forms — `--cached` and `HEAD` — report no difference, so no
diff-based guard can see a change to such a path. `git commit -- <path>` is
the odd one out: it reads the working tree directly and records it.

So a modified assume-unchanged path now reports `nothing_to_commit` where it
previously committed. That is the answer consistent with this function's own
staging step, which honoured the flag one loop earlier — but it is a
behaviour change, and it belongs on the record as a decision rather than
surfacing later as a surprise.

Also corrects an AC3 comment still describing the `diff --cached` whole-index
form that the previous commit replaced.

* chore(#3776): set changeset fragment pr to 3859

The fragment carries the PR's own number, which is unknowable before the PR
exists. Backfilled post-create; the repo's changeset lint rejects the `pr: 0`
placeholder.

* fix(#3776): do not read an unanswered sequencer probe as "no merge"

`execGit` surfaces a spawn timeout as `exitCode: 1` (`_spawnResult`:
`result.status ?? 1`) — the same code `rev-parse --verify` returns for a ref
that does not exist. So the MERGE_HEAD and CHERRY_PICK_HEAD probes could not
tell "not in that state" from "never answered", and the empty-diff guard read
both as "not in that state". That is the one path in #3776 that did not fail
toward the previous behaviour: a timeout during a real merge decided
`nothing_to_commit` from a pathspec git will not honour and left the merge
unconcluded, where before it was a loud `commit_failed`.

Treat an unanswered probe as "assume the partial commit would be refused" —
which falls through to `git commit` and lets git speak for itself.

Routed into `partialCommitRefused` only, deliberately never into
`isMergeInProgress`. That flag also feeds the pre-existing `canScope`, and
widening it there is worse than the misreport it fixes: with `canScope` false
the commit runs bare, and a bare commit during a merge is PERMITTED — git
concludes the merge with the whole index under a message naming one file.
Driven: the whole-flag form reports `committed` where this form reports
`commit_failed`, and it drops the pathspec on the ordinary timeout, re-opening
the #2112 scope leak.

* fix(#3776): pin the empty-diff probe against diff-only configuration

`git diff` is porcelain and honours settings `git commit -- <paths>` does not,
so an unpinned probe let a caller's configuration decide whether the guard
fires. Driven against git 2.54, each with the paired `git commit -- <path>`
confirmed to record the change the unpinned probe reported as absent:

  diff.ignoreSubmodules=all      a gitlink bump is invisible to the probe
  .gitmodules  ignore = all      the same, and it needs NO local config at
                                 all — it is checked in, so it arrives with
                                 a clone
  diff=<driver> + textconv       two different blobs converge to one text,
                                 so the probe sees no change; no submodule
                                 involved

`--ignore-submodules=dirty` rather than `=none`, because `dirty` is what a
partial commit of a submodule path actually means: it records the gitlink,
which moves only when the submodule's HEAD does. Under `=none` a merely dirty
submodule work tree reports a difference the commit would not record, sending
an empty call back to `git commit` — the same misreport, re-entered from the
other side. `dirty` still overrides both `diff.ignoreSubmodules` and a
checked-in `.gitmodules` `ignore`, so the gitlink vectors stay closed.

`--no-ext-diff` is deliberately absent: `--quiet` short-circuits ahead of an
external diff driver, so `diff.<driver>.command` cannot invert the probe
(driven: rc 1 with and without the flag).

* docs(#3776): disclose the two outcome changes the changeset omitted

The body listed what stays unchanged and never named the arms whose
user-visible outcome moves, so neither would have reached the changelog:

  - a modified path under `git update-index --assume-unchanged` now reports
    `nothing_to_commit` where it was previously committed. `git add` already
    honoured the flag one loop earlier; the guard reports what staging did.
    Documented in a code comment and pinned by a test since the first round,
    but absent from the fragment.
  - naming a submodule whose work tree is dirty while its recorded commit has
    not moved now reports `nothing_to_commit` rather than `commit_failed`,
    because nothing would have landed. New in this round, from the
    `--ignore-submodules=dirty` pin.

* test(#3776): use helpers.cleanup() for the submodule fixture teardown

`local/no-raw-rmsync-in-tests` rejects a bare `fs.rmSync` in a test: the
helper carries the Windows-EBUSY retry budget (`maxRetries`/`retryDelay`)
that a raw call does not, and a submodule work tree is exactly the shape
that holds handles open on Windows.

Caught by CI, not locally — the round ran the two affected suites but not
`npm run lint:ci`, so the repo's own rule never fired until the push. The
chain now exits 0 locally against this tree.

* fix(#3776): never drop a named assume-unchanged path

`--assume-unchanged` is the one state where `git diff` and
`git commit -- <paths>` genuinely disagree: `git add` stages nothing,
both diff forms report no difference, and `git commit -- <path>` still
reads the working tree and records it. The empty-diff guard therefore
reported `nothing_to_commit` about content the caller named in `--files`
and git would have written.

commit is made — so suppressing its misreport must not be paid for by a
silent drop. Same rule the timeout routing already follows: a fix for a
misreport may not cost content.

The guard now asks `git commit --dry-run --porcelain` whether the commit
would record anything, and stands aside on rc 0. That is the same
decision the real commit makes, so there is no second implementation of
it to drift. It does not run the `pre-commit` hook (driven: a rejecting
one neither fires nor writes a marker), which is what matters — a firing
`pre-commit` is the whole of #3776. It is NOT hook-free in general: git
2.54 fires `post-index-change` here, so a repo using that hook sees it
once for the probe and once for the commit. Stated rather than claimed
away.

Asking git rather than reconstructing its answer was reached by
measurement. Comparing `git hash-object` against `HEAD:<path>` was tried
and is wrong three ways, each a silent drop of named content: it misses a
mode-only change (`chmod +x` leaves the blob identical while the commit
records `100755`); it cannot hash a submodule path at all (`fatal: Unable
to hash sub`, while the commit advances the gitlink); and the path it
needs must be parsed out of `ls-files` output, which `core.quotePath`
renders as `"caf\303\251.md"` by default. Each has its own arm, and the
non-ASCII arm pins `core.quotePath` so it cannot go vacuous.

Falling through on the `ls-files` tag alone — without asking whether
anything would land — is also wrong: an UNMODIFIED assume-unchanged path
would reach `git commit`, which with any unrelated modified file present
prints `no changes added to commit`, a string the fallback does not
match, and returns `commit_failed`. That is #3776 re-entered from the
other side, the same shape `--ignore-submodules=none` would have
re-entered it. Pinned by its own arm.

The `ls-files` read is an optimisation, not a gate: it keeps the dry run
off the hot path when no assume-unchanged entry is present, and when it
cannot answer the dry run simply runs, because the dry run needs nothing
from it. Failing closed there would drop content and failing open would
re-enter #3776 — both are wrong answers to a question that can be asked
directly.

`--skip-worktree` is not a second instance. A present, modified one exits
1 from `git add` and fails closed as `staging_failed` above the guard; an
absent one is skipped before `git add` runs (#2014) and is answered by
the `stagedPaths.length === 0` arm, exactly as it was pre-fix. Both
shapes pinned, because the shorter claim ("never reaches the guard") is
too strong.

* test(#3776): register fixture teardown so a failed assertion cannot leak

`bumpedSubmodule()` creates its sub-repo as a SIBLING of `tmpDir`, and
the unborn-HEAD arm creates `fresh` outside it too, so the describe's
`afterEach(() => cleanup(tmpDir))` reaches neither. Both were cleaned by
a trailing statement in the test body, which any failing assertion above
it skips — leaking a git repo into the temp root.

`bumpedSubmodule()` now records the path and a describe-scoped
`afterEach` drains it, which covers all three of its callers at once;
the unborn-HEAD arm takes `t.after`, the form already used elsewhere in
this file.

Negative-controlled both ways with a deliberate assertion failure
injected into the dirty-submodule arm, under an overridden TMPDIR:
before, one `*-sub` repo survives the run; after, none.

* fix(#3776): never read an unanswered dry-run probe as "nothing to record"

The `git commit --dry-run --porcelain` probe that decides the assume-unchanged
boundary is the one probe in the guard whose rc 0 is the reassuring answer, so
it inverts the diff probe's safety: `execGit` collapses a spawn timeout (or any
spawn error) to `exitCode: 1`, byte-identical to git's own "nothing to record",
and the guard then reported `nothing_to_commit` about content named in
`--files` that git was never asked to write. Same conflation the sequencer
probes already defend against.

Only a CONFIRMED rc 1 with no spawn error closes the path now; a timeout, a
spawn error, or rc 128 falls toward the real commit, where git speaks for
itself. Five arms in tests/commit-files-pathspec.test.cjs pin it (posix +
windows timeout shapes, rc 128, the ls-files optimisation's own timeout, and a
negative control on an unmodified path); the injection helper gains an optional
`matchArg` so the dry run can be targeted without intercepting the real commit.

Also corrects the comment that claimed both sequencer probes are gated on
`guardApplies` — the MERGE_HEAD probe predates this fix and is unconditional.

* fix(#3776): probe with --no-verify so a hook-firing git cannot close the guard

Round 4, review finding 3 (Minor). The `git commit --dry-run --porcelain`
probe's safety rested on an empirical claim about one git version: that
`--dry-run` does not run `pre-commit`. git 2.54 satisfies it, but the failure a
differing version would produce is silent and lands in exactly #3776's own
configuration.

A `pre-commit` that fires and rejects exits 1 — the same code git returns for
"nothing to record" — so the closure would read it as a CONFIRMED empty answer,
drop the content the caller named in `--files`, and report `nothing_to_commit`.
That is #3776 re-entered through the probe the fix added.

`--no-verify` forecloses it structurally rather than documenting the version
dependency. Driven on git 2.54: rc-identical in both directions (rc 0
would-record, rc 1 nothing) with and without the flag, so it is behaviour-
neutral where the version already agrees.

Two claims deliberately NOT widened: `--no-verify` does not suppress
`post-index-change`, which still fires on this call with or without it (driven
both ways); and the real `git commit` is untouched — #3776 is a bug about a
hook's message reaching the caller wrongly, never a licence to skip hooks.

The new arm pins the FLAG rather than an outcome, because the outcome it
protects is unobservable on a git that already declines to run the hook. It is
a seam assertion over the argv the guard actually issued, not a source grep.

* test(#3776): pin all-missing --files during a merge or cherry-pick

Round 4, review finding 1 (Major) and finding 7 (Nit, its coverage half). The
review asks for the `stagedPaths.length === 0` disjunct to be gated on
`!partialCommitRefused`, or for the combination to be documented and tested.
Documented and tested — the gating is refused, with cause.

The premise is confirmed: the state is reachable exactly as described, and
during a merge the `nothing_to_commit` report does not tell the caller the merge
is still open. The prescription is not. With every named path missing,
`stagedPaths` is empty, so `canScope` is false and the fall-through reaches a
BARE `git commit`, which git PERMITS during a merge and which then CONCLUDES it.

Driven, git 2.54, through cmdCommit with the prescription applied:

  cmdCommit(cwd, 'add the thing', ['.planning/never-produced.md'])
  -> { "committed": true, "hash": "8e6bf45", "reason": "committed" }
     MERGE_HEAD gone; HEAD is a 2-parent merge commit recording
     .planning/shared.md with the caller's resolution content.

So the gating trades a report that writes nothing for one that silently writes
the whole index under a message naming a path that does not exist, and reports
success. That is the same trade the timeout routing already refuses one block
up, which is why the sequencer states gate the DIFF branch only.

The behaviour is also pre-existing and unchanged by this PR: at 86452da7 the
identical short-circuit sat ABOVE the MERGE_HEAD probe, so it never consulted
the sequencer either. The residual — a merge held open behind a
`nothing_to_commit` report — is offered as a separate issue alongside the three
already deferred, not folded into this fix.

These are behaviour pins, not regression tests: they pass at base and red on the
gated implementation (both arms, verified).

* docs(#3776): state the git-version provenance once, not at three claims

Round 4, review finding 6 (Nit). The guard carries ~159 comment lines around 32
lines of executable logic, and the review's specific complaint is that "driven
against git 2.54" is repeated near-verbatim in three places, which makes the
decision tree harder to scan.

Hoists the provenance to a single block header and reduces the three repeats to
the observation each actually carries. One claim keeps its version explicitly
and now says why: the `--no-verify` reasoning is version-SENSITIVE rather than
merely version-observed, so it is the one place the version is load-bearing
instead of incidental.

The behavioural matrix stays inline rather than moving to an ADR or a doc block.
Every claim in it is a constraint on the four flags immediately below it, and the
value of having it here is that the next reader who wants to "simplify" one of
those flags meets the driven counter-example in the same screen. Splitting the
constraint from the code it constrains is how the flags get dropped.

Comments only. No behaviour change; suite and lint:ci unchanged.

* docs(#3776): correct two driven figures in the new guard commentary

Both found by this round's own pre-push adversarial review, and both re-driven
before adopting.

1. The empty-paths rationale said a bare commit during a merge produces a
   "three-parent commit". It produces a TWO-parent merge commit. Three was the
   token count of `git rev-list --parents -n1 HEAD` (commit + two parents) read
   as a parent count. `git cat-file -p HEAD | grep -c '^parent '` returns 2.
   This round's commit message for the pins already said two, so the tree
   contradicted itself.

2. The `post-index-change` disclosure said a repo using that hook "sees it once
   for the probe and once for the commit". Driven with a counting hook: git
   fires it TWICE per `git commit --dry-run`, and twice again for the real
   commit — 2/2/2 across the flagged probe, the unflagged probe and the real
   commit. The disclosure understated the cost by half in both halves.

Comments only. No behaviour change; suite 351/351 and lint:ci unchanged.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-01 13:59:11 -04:00
BeeHiggs
41466e8e88 fix(#4023): preserve decimal phase ids in init progress ordering and smart-entry output (#4110)
* test(#4023): reproduce decimal phase-id coercions

* fix(#4023): preserve decimal phase ids in progress signals

* test(#4023): align phase token contract expectations

* chore(#4023): point the changeset at PR #4110

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-01 13:58:20 -04:00
BeeHiggs
1a358ce0fd feat(#2761): bracket-tolerant read path — roadmap/validate/verify/state recognize bracket ids (epic #612 PR-2) (#2867)
* feat(#2761): gated heading-intro selection + one bracket identity grammar

Foundation. Two owner-level changes plus a federated convention resolver; no
reader consumes them yet.

1. GATED SELECTION, not an ungated widening.

   Widening every heading matcher requires the claim "no legacy ROADMAP contains
   a `[CODE.MM]` bracket followed by a digit", and that is false:
   `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and
   `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each
   as a phase — moving phase_count and total_phases and adding W006 on projects
   that never opted in. No narrowing rescues it: the premise is about documents
   we do not control.

   `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the
   pattern SOURCE at construction time. A project whose resolved
   `phase_id_convention` is not exactly 'bracket' compiles the same source string
   it compiled before. `baseline` is explicit because whether a site spells the
   any-bracket prefix or a bare `Phase\s+` is a fact about that site's history:
   handing the wider grammar to a bare site retro-grants tolerance it never had,
   in both directions — warnings appear, and a warning that fires today vanishes.

   Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through
   the base alternative, which captures nothing, so a reader saw no bracket, fell
   back to the legacy token rule, and counted a labeled icebox heading while
   excluding the label-less one beside it — two derivations of one ROADMAP
   disagreeing.

2. ONE bracket identity grammar, one width rule.

   The milestone width is reconciled with the emit validator: pad2 output, so
   two digits or 3+ with no leading zero. Earlier spellings diverged in both
   directions — admitting `002`, which the validator rejects, and a bare `0` pad2
   never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED
   a milestone no phase heading could then resolve into, recreating the
   on-disk-count fallback this epic removes. An unpadded bracket is now uniformly
   malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its
   directories is the surfacing signal.

   The milestone field is boundary-anchored, so a malformed run cannot match by
   its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays
   case-insensitive because readers compile `/i`, but identity helpers match
   `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:`
   failed every sentinel test. The qualified key shares the width, the `(?=-|$)`
   boundary and the single-sub-phase shape of the directory token, because
   phaseTokenMatches returns unconditionally on a qualified hit: a key matching a
   directory isPhaseDirName rejects would be a final wrong answer.

3. resolvePhaseIdConvention federates workstream -> root exactly as
   config-loader does — including that root is a fallback only when a WORKSTREAM
   is active, so a project-scoped directory stands alone. loadConfig cannot serve
   this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and
   this key is not among them. It governs the bracket-selection reads ONLY.

PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing
consumes it, and it is superseded rather than redefined.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#2761): roadmap.cts selects its heading grammar from the convention

Six matchers build their intro through the gated selector, and
cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each
resolve the convention ONCE per command and thread it down.

Three sites take the any-bracket baseline (they already tolerated
`[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`).
Handing the wider grammar to a label-only site retro-grants tolerance it never
had — and not only by adding matches: on a legacy repo an unchecked
`- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that
fires today.

Sentinel handling under bracket ADDS a rule rather than replacing one: a
bracketed heading is a sentinel when its bracket milestone is reserved
(`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog
convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a
mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly
the content this epic targets — add entries to the progress denominator. The
captured id is folded before the identity test, so a lowercase
`### [gsd.999] 07:` is excluded too.

The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention
once and hands it to all four of its heading/checklist patterns, but the single
`phaseTokenMatches` call that decides `disk_status`, `plan_count`,
`summary_count`, `has_context` and `has_research` was left two-argument — so
every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with
zero counts, on the PR's own headline verb, while the SAME build resolved those
same directories correctly in three other places on the same repo (W006/W007 via
phaseTokenFromDir, `state json` via the milestone filter, and the W021
milestone-complete read through this very helper's three-argument form). It
failed ONLY for the directory shape the convention exists to name: a
mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is
why nothing caught it. Measured, bracket vs its flat-legacy twin:
`[["01","no_directory",0,0],["02","no_directory",0,0]]` against
`[["01","complete",1,1],["02","planned",1,0]]`.

The oracle is the twin, computed in the same test run, plus exact literals —
`grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix
nor a future regression had any gate at all.

Disclosed: a ROADMAP written in bracket form before config.json is switched
reads as empty rather than mis-counted. Silent invisibility during the migration
window is the deliberate trade against claiming phases on projects that never
opted in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#2761): validate.cts selects its grammar; gated directory recognition

The W006/W007 feeders take the resolved convention as a threaded parameter.
These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them
where an ungated widening does the most damage: `### [RFC.2119] 5:` enters
roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on
disk" on a project that never opted in.

buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket
headings. Surfaced rather than filtered in place because roadmapPhases feeds both
a membership check and a missing-directory warning, and only the latter should
ignore an icebox item.

That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is
a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying
suppression on the token alone let an icebox heading silence a REAL phase that
happens to share its number — a false negative strictly worse than the warning it
removed. A token is suppressed only when no non-sentinel heading bears it.

Directory recognition is added as gated FUNCTIONS beside the exported RegExp
constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is
string-indistinguishable from the letter-prefixed-decimal family this repo
documents as ambiguous, and folding a branch in changes those constants' answers
on exactly that family. A RegExp constant has nowhere to attach a gate.

The recognizer mirrors the emit grammar and delegates the token to the canonical
owner, so recognizer and resolver agree on rejected input as well as accepted.
Both functions throw on a non-string, matching the call pattern they replace.

buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin
and like the sibling checklist scan in roadmap.cts, and for the reason that one
states: the bracket id has to ride along or the sentinel filter is blind to
`- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every
checklist token REAL, and the occurrence-aware un-suppression loop then deleted
the icebox token the HEADING scan had correctly marked sentinel — so `validate
consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE
ROADMAP shape where an icebox appears as both a bold bullet and a detail
heading. `validate health` stayed silent on that same repo, so the two verbs
disagreed — which is the disagreement `sentinelPhases` exists to close.

Both directions are pinned, because the failure mode of a careless fix here is
the opposite one: a real phase sharing a sentinel's token must still warn. It
does, in all four shapes that attack it (sentinel heading + real bullet,
lowercase sentinel, sentinel after the real heading, colon-less bullet).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2761): count bracket headings, and retire them, in both derivations

Both `total_phases` derivations select their grammar from the resolved
convention, in one commit — cmdStateSync already carries the comment that it
mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)",
so teaching one and not the other ships that divergence.

The #1514 retirement filter widens WITH the counter it protects. The canonical
gesture strikes the checklist BULLET and leaves the detail heading intact, so a
bracket-form retirement went undetected and the phase stayed in the denominator
forever. That is half a fix alone: the retired key is compared against
phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves
land here.

Under bracket the sentinel token rule composes as the full engine set {0, 999},
so this counter agrees with `roadmap analyze`, which has always excluded both —
otherwise the two derivations report different numbers for one ROADMAP and the
changeset's "excluded from every count" is false as written. The LEGACY path
keeps its pre-existing 999-only rule: widening it there would move legacy totals,
so the two stay split off the bracket path exactly as they are today.

The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not
the frontmatter total_phases. Sync's own counter never reaches that field — the
read derivation writes it — so asserting the frontmatter after a sync measures the
read path twice and lets a mutation to the write-path guard survive.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read

The shipped milestone-prefixed W021 gate keeps its ROOT-only config read,
verbatim base semantics. Federating it silently moved a legacy convention's
answer in BOTH directions on workstream repos — a W021 that fires at base
vanishing, and one that is silent at base firing. resolvePhaseIdConvention
governs the new bracket-selection reads only.

B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it
with an empty config) but selects its grammar from the convention. Inferring
'bracket' from the shape of a matched bracket ran a repo-failing check against a
legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution
widens with the heading read, so a bracket repo whose phases are on disk stays
silent, and a bracket sentinel is not reported as unstarted.

checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so
fenced examples cannot warn and heading level is structural. Its scope rules each
close a way it silently did nothing or fired wrongly: only a genuine MILESTONE
heading opens or closes a section (a `### Notes` used to reset scope and disable
both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed
phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a
phase; the full h2-h6 range is processed. Its section recognizer shares the one
milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the
id grammar and a section to the section grammar at once, silently re-scoping
every warning after it.

validate consistency suppresses bracket sentinels in its missing-directory
warning — the two verbs disagreed, health suppressing via notStartedPhases while
consistency did not. The legacy reading is untouched, including its pre-existing
wart that `### Phase 999:` still warns there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2761): scope the milestone by its bracket; select the disk-side filter

Two roadmap-parser reads, both of which made a bracket project's totals track
the disk instead of the ROADMAP.

The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name,
no version — but scoping matched STATE's `milestone: v2.0` STRING against a
heading, so the canonical form matched nothing and total_phases fell back to the
directory count. The rule was re-derived in THREE places: extractCurrentMilestone
plus two `milestoneBounded` guards; fixing one left the others falling back
regardless, so they are now one gated helper. It matches the CANONICAL padded
spelling only — accepting `0*N` bounded a milestone whose phases were invisible,
which un-suppressed a progress percent computed off an unscoped disk count.

getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a
bracket ROADMAP it collected nothing, so the filter degraded to pass-all and
buildStateFrontmatter counted every other milestone's directories — making the
bracket convention strictly worse than the M-NN one it supersedes on the property
that matters most: totals must track the ROADMAP, not the disk.

The DIRECTORY side of that same filter is selected with it. Teaching only the
heading scan was half a fix and a worse one: `milestonePhaseNums` became
non-empty, so the pass-all degrade stopped firing, but no bracket directory could
satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the
custom-id match captures the project code `GSD`, and stripProjectCodePrefix does
not strip a dotted prefix). Every bracket directory was rejected, and
completed_phases / total_plans / completed_plans / percent all collapsed to 0
while `state sync` went on writing a percent off the unfiltered disk — `state
json` reporting 0% on the same repo, in the same second, that STATE.md's body
called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and
total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)`
floors it at the ROADMAP count no matter how many directories are rejected.

The dir side matches on the milestone-QUALIFIED id, delegated to the owner's
gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B
puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one`
share the token `01` and only the qualified key separates them. The qualified ids
are kept in their own set — a hyphen in `milestonePhaseNums` would flip
`roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo
— and the branch is ADDITIVE: on a miss it falls through to the three legacy
checks, so a bracket project carrying legacy-shaped directories reads unchanged.

Both are resolved lazily and gated, so the legacy path pays neither a config read
nor a second scan and cannot change answer. The scoping call is also GUARDED:
resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a
GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only
planningDir call in extractCurrentMilestone sits inside the STATE-read try, so
the function returned normally on such an environment; an unguarded one here let
that escape and broke the never-throws invariant that getRoadmapPhaseInternal and
getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about.
Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the
workstream-name policy and GSD_PROJECT throws identically at base — but reachable
by any in-process embedder, which is precisely who that invariant is for. The
filter's own resolve call was already inside its try and is unaffected.

The milestone-qualified key is formed only for a token that is itself a bracket
phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration
heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to
`GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02:
the `-01` truncated, both such headings collapsing to one key, and the heading
claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting
`GSD.02-01-one`, the one it does. The guard drops those headings back to the
unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned
against the milestone-prefixed reading of the same ROADMAP, which is
base-identical on this shape.

Scoped precisely, because the fixture moves one number that the guard does not
touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket
heading COUNT this PR exists to add, not the splice — measured identical with and
without the guard, and identical to what the canonical `### [GSD.02] 01:`
spelling does on the same fixture (both read 2 with zero directories on disk,
where base reads 0). The claim is base-equivalent ACCEPTANCE, not a
base-equivalent reading.

One consequence is stated rather than fixed: a heading whose token carries a
hyphen still puts that hyphen into milestonePhaseNums and so still flips
`roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it
is what keeps the shape base-equivalent; excluding the token would have moved
answers versus base on malformed input. The comment at the qualified-set
declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out
of that flag's input, not hyphens in general.

The oracles ship with it, and they are the five numbers, not the one: the parity
gate now asserts total_phases, completed_phases, total_plans, completed_plans AND
percent, on both derivations, on two fixture shapes (one milestone; two
milestones with stale prior-milestone directories on disk). The oracle is the
flat-legacy twin, built in the same test run and compared number for number,
plus exact literals so a shared wrong answer cannot pass.

The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could
not serve, because buildStateFrontmatter's #2445 de-dup key captures only a
directory's leading integer and collapses `02-01-one` / `02-02-two` /
`02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's
[3,2,3,2,67], identically at base and before this fix, and structurally
unreachable from the bracket key space. That reasoning is only sound while it
stays true, so a characterization test holds the M-NN reading down on the two
numbers that do not depend on which directory wins the mtime race. Widen the
de-dup key and it fails, instead of quietly invalidating the changeset's
disclosure.

Also adds the call-site pin. The structural table pins transcription against the
selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping
verify.cts's milestone-complete site to the wider baseline grants a
fires-on-every-repo check tolerance it has never had, and every behavioural test
still passed. The pin reads the shipped sources and asserts the mode at each of
the 14 sites, count-exact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): pin the bracket read surfaces in the parity gate

This gate exists because #2043 fixed one bug across five hand-edited copies of a
rule and #2232 was the residual that survived, because a later reader could not
tell the copies were one rule. PR-2 adds two consumers, so they belong here.

Surface 7 — the heading read and the directory read must agree about WHICH phase
a `MM-<seg>` pair names, across the shared width corpus, and the bracket and
legacy spellings of one heading must yield the same token.

Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on
ACCEPTED input was already pinned; agreement on REJECTED input is where they
actually diverged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2761): changeset

Disclosures for the PR body (deliberate, not defects):

- phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and
  cannot serve as the convention resolver however the file is federated. This PR
  ships its own workstream->root resolver; adding the key and its value enum is
  later-slice work.
- Convention matching is strictly === 'bracket'. A misspelled value reads as
  not-configured and the project keeps legacy behaviour silently.
- An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing,
  bounds nothing, sections nothing, and is not a phase id. W005 on its
  directories is the surfacing signal.
- WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical
  inputs, none of which toDir can emit and none of which had a bracket caller at
  base:
    isSentinelPhaseId('GSD.0-01',    'bracket')  true  -> false
    isSentinelPhaseId('GSD.0999-01', 'bracket')  true  -> false
    getMilestoneFromPhaseId('GSD.2-01',   'bracket')  'v2.0' -> null
    getMilestoneFromPhaseId('GSD.002-01', 'bracket')  'v2.0' -> null
  The canonical pad2 sentinel spelling `[GSD.00]` still tests true.
- FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x ->
  milestone null) is preserved". After the unification that holds for the
  canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours;
  flagging the tension rather than editing it.
- The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is
  a sentinel when its bracket milestone OR its token is reserved. Under bracket
  the state-side token rule is the full {0, 999} set so both derivations agree;
  the LEGACY path keeps its pre-existing 999-only rule, unchanged.
- validate consistency's legacy reading is untouched, including the pre-existing
  wart that `### Phase 999:` warns there while validate health suppresses it.
- find-phase still cannot resolve a bracket phase directory. phase-locator.cts is
  outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls
  phaseTokenMatches without a convention, so whichever slice lands second must
  thread it through.
- Four of the five bracket readers scan raw ROADMAP content, so a bracket heading
  inside a fenced code block is read as a phase. Pre-existing for the legacy
  spelling; parity, not a new class.
- roadmapPhaseLookupSources gained no bracket source: nothing emits a
  milestone-qualified query into it yet.
- roadmap validate remains a separate, unfederated convention reader.
  Pre-existing and base-identical, but two verbs can disagree about the active
  convention on one project.
- _diskScanCache keys on cwd while the values it caches are now
  convention-dependent. Not reproducible through the CLI; pre-existing for the
  workstream dimension, widened here. Stated as inconclusive.
- A ROADMAP written in bracket form before config.json is switched reads as empty
  rather than mis-counted — the deliberate migration-window trade.
- THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that
  divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter
  applies the milestone filter; cmdStateSync does its own fs.readdirSync and never
  calls it, so on a repo carrying prior-milestone directories the read path
  reports the SCOPED percent and the sync body reports the WHOLE-DISK one.
  Measured on the true base build (d04592de), flat-legacy spelling, 3 in-scope
  phases with 1 complete plus 2 stale prior-milestone dirs: `state json`
  [3,1,3,1,33], sync body 60%. The bracket twin of that repo now reads the same
  two numbers — 33 and 60. Scoping the sync counter would move every legacy
  repo's percent, which a bracket read-path PR must not do. The gate pins both
  sides, so the mirror cannot silently become a one-sided fix.
- THE PARITY ORACLE IS THE FLAT-LEGACY TWIN, NOT THE M-NN ONE, and that is a
  measurement finding rather than a preference. buildStateFrontmatter's #2445
  de-dup key captures only a directory's LEADING integer, so the M-NN dirs
  `02-01-one` / `02-02-two` / `02-03-three` all key to `2` and two of the three
  are dropped before they are ever counted: base reads [3,0,1,0,0] where the
  flat-legacy twin of the same repo reads [3,2,3,2,67]. Present identically at
  base and at HEAD, untouched here, and structurally unreachable from the bracket
  key space — `GSD.02-01-one` does not match that pattern at all, so every bracket
  directory keys to its own name. The source line already carries a
  `phase-id-owner:` sanction recording the divergence. Mirroring it under bracket
  would mean manufacturing a collision that cannot occur, so the gate compares
  against the flat-legacy spelling, which is uncontaminated. This paragraph is
  itself pinned: a characterization test holds the M-NN reading on the two
  numbers that do not depend on which directory wins the mtime race, so widening
  the de-dup key in a later slice fails the suite rather than silently making
  this disclosure false.
- THE `phaseTokenMatches` CALL-SITE CENSUS, stated so the remaining gaps are
  auditable rather than implied. 13 call sites outside the owner (phase-id.cts).
  THREE are three-argument: verify.cts:2229 (the W021 milestone-complete read,
  already was), roadmap.cts:436 (`roadmap analyze`'s directory lookup, threaded
  by this PR) and roadmap-parser.cts:792 (the disk-side milestone filter, added
  by this PR). The other TEN are two-argument and stay that way — phase.cts ×5
  (220, 277, 444, 585, 1547), phase-locator.cts:62, smart-entry.cts:243,
  init.cts:1414, milestone.cts:551 and verify.cts:2467. All ten are untouched by
  this PR and base-identical.
  One of them sits in a file this PR DOES edit, so it is named rather than left
  to a reader's grep: verify.cts:2467, `verify schema-drift <phase>`. Measured on
  a bracket repo across base / pre-fix branch / this HEAD, all three agree on all
  three argument forms — `verify schema-drift GSD.02-01` and `… 01` both report
  "Phase directory not found" on every build, and `… GSD.02-01-one` resolves on
  every build through the exact-directory-name fallback. So the user-visible
  shape of what stays broken is: a bracket phase is addressable there by full
  directory name only, exactly as at base. Threading the convention into a
  function this PR never touched, in the last round before ship, is the wrong
  trade; it is where the same one-argument fix goes next, alongside
  milestone.cts:551 and init.cts:1414.
- A BRACKET HEADING WHOSE TOKEN CARRIES A HYPHEN (`### [GSD.02] Phase 02-01:`,
  a mid-migration spelling) forms NO milestone-qualified key, and therefore
  scopes through the unqualified legacy path — base-equivalent ACCEPTANCE, which
  is the claim, and not a base-equivalent reading: `total_phases` on that shape
  moves 1 -> 2 for the same reason it moves on the canonical `### [GSD.02] 01:`
  spelling, because counting bracket headings is what this PR does. Such a token
  still flips `roadmapUsesHyphenedIds`, as it also does at base. The comment at
  the qualified-set declaration now claims only that narrower, true thing.

The `pr:` field carries the sub-issue number as a placeholder — it must be
updated to the real PR number when the PR is opened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2761): point the changeset at PR #2867

* test(#2761): fast-check properties for the convention-selection layer

CONTRIBUTING.md mandates a generative property test for parser/bijective-contract
changes; PR-2 shipped six example-based files and none. This adds the missing
layer, scoped to what PR-2 actually contracts — WHICH pattern each reader
compiles, decided by the resolved `phase_id_convention` — rather than restating
PR-1's grammar round-trip properties, which already live in
tests/adr-612-bracket-grammar.test.cjs.

Four properties: P1 an opted-in repo reads the ADR-canonical label-less bracket
heading/dir and a non-opted-in repo is byte-blind to the identical input; P2
every non-bracket convention agrees with the hand-transcribed BASE source over
generated content, including bracket-DOTTED legacy prose (`[RFC.2119] 5:`) that
must never be claimed as a phase; P3 nine per-field mutations are rejected and
the one case variation folds instead; P4 both sides of a phase comparison derive
the same key under the same convention.

Generators template every input from raw primitives — nothing is seeded through
renderPhaseId/toDir, the p2() tautology that made #2258 round 1's property test
structurally unable to find B1. Domain reaches past 99 into the 3+-digit branch
(round 2's numArb-capped-at-99 miss), forces sub-phases in at weight, and pins
both sentinel milestones.

Falsified against the COMPILED lib, not the source: five deliberate mutants
(gate never fires; gate always fires; milestone width widened to \d+; the #612
convention forwarding dropped from phaseKeyFromDir; extractPhaseToken's bracket
branch ungated) each fail the specific property that should catch them —
16/2, 16/2, 15/3, 17/1, 17/1 pass/fail — and the lib restores byte-identical.

An earlier draft of P2 held vacuously: its base regex omitted the markdown
furniture the selected one carried, so every realistic `### Phase NN:` line
matched neither side. The gate-always-fires mutant did not kill it. Both are now
compiled through one function, and that mutant kills P2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): adversarial malformed bracket tokens across the tolerant readers

The existing boundary coverage stopped at shapes the emit grammar rejects
(unpadded `[GSD.3]`, wrong-case, `12A`). It never exercised a STRUCTURALLY
broken token — a non-numeric milestone, a bracket that never closes, a bracket
nested in another — which is the input a tolerant reader is most likely to
half-read, and the one the PR's own regex commentary is explicit about.

read-tolerance (roadmap heading scan + validate's dir and variant builders):
ten malformed headings, each asserted to be read as a phase by NO convention and
to give the opted-in repo the same answer as the legacy one; the corpus driven
through `roadmap analyze` end to end; malformed DIRECTORY names asserted
unrecognized and non-throwing on all four conventions; and the two variant
builders asserted to agree, since a widening that reaches only one splits
`validate consistency` from `validate health` (the #3242 Bug B shape).

coherence (verify.cts W021): the same six broken shapes asserted to raise no
W021 of their own AND not to re-scope the W021 that follows them — the G2
failure mode reached from a different shape, where a heading that is not a phase
but IS read as a section silently moves later warnings onto the wrong milestone.

Both files gain a pathological-input time bound. Nested quantifiers over a long
unclosed bracket are the classic ReDoS shape and two commits on next (#2828,
#2944) were CodeQL-flagged for exactly that, so the bound is asserted rather
than argued from reading the pattern. The probes themselves parse no regex.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2761): document "bracket" as a phase_id_convention value

The row listed only `"milestone-prefixed"` and `null`, so after two shipped
slices (#2258 grammar, this PR's read path) the convention had no documented
enum value. CONFIGURATION.md is also a top-10 historical co-changer of both
src/verify.cts and src/state.cts and was absent from this PR.

The row states the boundary rather than the ambition: `"bracket"` changes the
READ path only, there is no migrator and no emit yet, and a project on any other
value compiles the patterns it compiled before. That keeps the docs honest for
the two releases before PR-3 and PR-4 land, instead of describing a convention a
user cannot yet migrate to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2761): retype the changeset Added, drop the docs-exempt marker

`Fixed` was wrong by CONTRIBUTING.md's own definition — a fix restores
documented behavior, and bracket read tolerance is the second slice of a
capability that did not exist before #2258. The type also carried a
`docs-exempt` marker, and `Fixed`/`Security` are exempt from the docs-required
lint, so the typing had the effect of routing around a gate this change should
pass. It now passes it: `lint-docs-required` returns ok_docs_updated on the
CONFIGURATION.md row added in the previous commit.

Body gains one sentence pointing at that row and restating that `"bracket"` is a
read-path opt-in until the migrator and write path land.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): pin the version-less bracket milestone scoping gap; narrow the claim

The changeset asserted that milestone scoping "recognises the ADR-canonical
`## [GSD.02] Foundation` heading and applies to the phase DIRECTORIES too." A
CLI probe on that exact heading form falsifies the second half: with no `vN.N`
in the milestone heading the directory side does not scope, and directories from
BOTH the prior and the later milestone are admitted. Measured 4 dirs counted
where the milestone declares 2.

Every bracket fixture in the suite writes `## [GSD.02] v2.0: …`, so nothing
covered the form the ADR actually specifies — and the state.cts doc comment
calls that version-less form canonical.

Mechanism, in extractCurrentMilestone: the bracket scope branch selects the
right currentSection, but `preambleCutoff` keys off a pattern requiring a
version or status emoji, so a version-less roadmap falls back to the current
milestone's own offset and every PRIOR milestone lands in the preamble — whose
phase-stripping regex only strips `Phase N:`-labelled headings, so bracket phase
headings survive it. Independently, `computeSectionEnd` accepts a boundary only
on a version/emoji heading, so the section runs to EOF and every LATER milestone
is swept in. Two sites, bidirectional.

Not fixed here: it changes milestone scoping, which is shared with the legacy
path. Five characterization tests pin today's reading plus a versioned CONTROL
proving the version string is the only difference, and the changeset sentence is
narrowed to what the code does. The DEFECT assertions are written to be
INVERTED by the fix, not deleted — that inversion is its regression proof.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2761): re-anchor the branch's own cross-file line citations after the rebase

Three of this branch's code comments cite sibling call sites by line number, and
the rebase onto 178ec000 moved two of the three targets:

  roadmap-parser.cts  validate.cts:210 -> :218   (const g = capturing ? 1 : 0)
                      state.cts:1715   -> :1752  (const bg = … 'bracket' ? 1 : 0)
  roadmap.cts         verify.cts:2229  -> :2355  (phaseTokenMatches 3-arg form)

`state.cts:1715` had drifted 37 lines and now lands on the retirement skip, not
the capture-offset idiom the sentence is about — the citation read as evidence
for a claim the cited line does not support.

planning-workspace.cts's `config-loader.cts:618/:649` was checked and is still
correct; left alone.

Comment-only. Build, drift guard and the bracket suites re-run unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2761): scope the version-less bracket milestone heading too (B1)

computeSectionEnd and the preambleCutoff scan in extractCurrentMilestone
(roadmap-parser.cts) only recognized a milestone boundary heading that
carried a vN.N token or a status emoji. The ADR-canonical bracket
heading (## [GSD.02] Foundation) carries neither, so on that shape
computeSectionEnd fell through to content.length (sweeping every LATER
milestone into scope) and preambleCutoff fell back to the current
milestone's own offset (leaking every PRIOR milestone's bracket phases
into the preamble, whose Phase-N: strip regex never matches them).

Under the bracket scope branch, both sites now also accept a
`#{1,2}\s+\[CODE.MM\]` boundary, built from phase-id.cts's BRACKET_ID_SRC
(single owner of the bracket-id grammar) rather than a re-typed literal.
`#{1,2}` is the deliberate discriminator: a bracket PHASE heading is
level 3 and shares the same `[CODE.MM]` prefix, so a `#{1,3}` boundary
would swallow it too. Reachable only when bracketScopeConvention ===
'bracket' was already resolved (i.e. the bracket scope branch actually
fired), so version-bearing/emoji headings and non-bracket conventions
take the exact pre-existing code path byte-identically — confirmed by
the full adr-612 suite staying green.

Inverts the four DEFECT assertions in the
"#612 PR-2 CHARACTERIZATION: a version-less bracket milestone does not
scope" describe block (tests/adr-612-bracket-phase-counting.test.cjs)
into their regression-proof form, per the block's own doc comment, and
reframes the describe title/comments accordingly. Corrects the
.changeset/2761-bracket-read-tolerance.md fragment, which described the
directory-side version-less gap as an open, un-closed bound.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): thread sentinelPhases into validate health's W006 loop (B2)

cmdValidateHealth's W006 loop (src/verify.cts) destructured only
roadmapPhases from buildRoadmapPhaseVariants, not sentinelPhases —
unlike cmdValidateConsistency, which already skips sentinelPhases with
the identical guard a few hundred lines up. A heading-only bracket
icebox/pre-milestone entry ([GSD.999] / [GSD.00]) therefore gained a
false W006 "no directory on disk" from validate health while validate
consistency correctly stayed silent on the very same ROADMAP — the two
validators contradicting each other.

Threads sentinelPhases through and skips it before the existsOnDisk
check, mirroring the consistency guard exactly. Gated the same way
sentinelPhases already is (empty unless phase_id_convention is
'bracket'), so a legacy repo's W006 reading — including its own
pre-existing wart where a legacy `### Phase 999:` still warns on both
verbs — is untouched; confirmed by the existing "INHERITED WART,
unchanged" test staying green.

Adds the paired-agreement regression test (#612 PR-2 B2 describe block
in tests/adr-612-bracket-read-tolerance.test.cjs): a sentinel-only
bracket roadmap must produce no missing-directory warning from EITHER
validator, plus a CONTROL proving a real phase with no directory still
warns on both. Confirmed red (health false-W006) against the pre-fix
code before applying the fix.

Corrects the .changeset/2761-bracket-read-tolerance.md fragment, which
described the asymmetry as already closed and in the wrong direction
(it credited validate health with already staying silent, when health
was the one falsely warning).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#2761): pin mixed-shape preamble cutoff and boundary heading levels

Closes two self-flagged coverage gaps in the B1 fix (commit 08d5b0c4)
ahead of adversarial review. No src change — all three new tests are
green against the code as committed.

1. earliest-of-either preambleCutoff comparison: only exercised where
   the version/emoji match and the bracket match happen to land on the
   same heading. Adds the mid-migration mixed shape (version-bearing
   PRIOR + version-less CURRENT) and asserts scoping outcomes (accepts
   booleans + total_phases), not internals.

2. `h.level <= 2` conjunct in computeSectionEnd: provably redundant
   whenever the selected milestone heading is level 2 (every existing
   fixture), since `h.level > level` alone already implies it there —
   a mutant deleting the conjunct would have survived every prior test
   in this file. Adds a level-3 CURRENT-heading fixture (with a real
   PRIOR milestone so the preamble side-channel can't independently
   rescue the truncated phases) that makes the conjunct's deletion
   test-visible, confirmed by hand-mutating a throwaway copy of the
   compiled output (never touching tracked src or the real build) and
   observing the assertion flip. Also pins a level-1 companion case
   (#{1,2} tolerance, not just level 2).

NOT included here: the other mixed-shape direction (version-less PRIOR
+ version-bearing CURRENT) turned out to be a genuine, currently-unfixed
gap — reported separately rather than silently patched or weakened, per
instruction not to touch src while a probe run is in flight.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): engage bracket boundaries when the current milestone heading is version-bearing (B3, self-caught)

Found during round-2 self-verification of B1 (commit 08d5b0c4), while
closing the mixed-heading-shape coverage gaps flagged in my own review
notes. The B1 fix resolved `bracketScopeConvention` only inside the
`if (headingMatches.length === 0)` gate that also drives SELECTION's
own bracket fallback (which heading counts as "current"). That gate is
correct for selection, but `bracketScopeConvention` also feeds
computeSectionEnd's and preambleCutoff's boundary detection further
down — which accidentally inherited selection's gate instead of having
its own.

Trigger shape: the CURRENT milestone heading is itself version-bearing
(`## [GSD.02] v2.0: Current Milestone`), so the primary version-string
match succeeds immediately — headingMatches.length !== 0 from the very
first check — and the entire bracket-resolution branch was skipped. A
sibling milestone (PRIOR or LATER) that is version-less then got
neither the version/emoji boundary rule (it has none) nor the bracket
boundary rule (never resolved), reproducing the original #612 defect
(total_phases falling back to the whole-disk count) through a
structural shape B1's own fixtures never exercised — every one of them
is uniformly version-bearing or uniformly version-less across all
three milestones, never mixed with CURRENT specifically being the
version-bearing one.

Fix: resolve `bracketScopeConvention` unconditionally, decoupled from
`headingMatches.length`. SELECTION is deliberately left untouched — the
`if (headingMatches.length === 0 && bracketScopeConvention === 'bracket')`
fallback that picks which heading is "current" keeps its original gate
byte-for-byte (confirmed by diff: that line is unmodified). Only the
convention *resolution* moved out from behind it, so boundary detection
can consult it regardless of which branch selected the heading. The
extra `resolvePhaseIdConvention` call this now costs on every
invocation (previously paid only when the version match found nothing)
is the accepted cost: a non-bracket repo still resolves to something
other than 'bracket' (or null on a poisoned env, caught exactly as
before), so `bracketMilestoneHeadingRe` stays null and every downstream
branch is byte-identical to today — confirmed by the full adr-612 +
roadmap-parser + state + verify + health-validation suite staying green
(1260/1260) and the all-version-bearing/legacy fixtures showing no
behavior change.

TDD: tests/adr-612-bracket-phase-counting.test.cjs describe block
"#612 PR-2 B3: bracket boundaries engage even when CURRENT is
version-bearing but a sibling is not" — 4 tests, confirmed red against
pre-fix code (leak-in booleans true/true, total_phases 4) before this
change, green after.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): reject same-milestone continuation headings as boundaries (B1)

Gate-2 adversarial review Blocker 1: the B1/B3 boundary fired on ANY
as the one currently selected — a version-less checklist/detail split
(`## [GSD.02] Foundation (Phase Details)`, or an ad-hoc continuation
heading) truncated the current milestone's own section instead of
being recognised as a continuation of it. The `(Phase Details)`
re-append only searches VERSION-STRING matches, so a version-less
continuation heading was cut out and never re-appended — a confidently
wrong, non-degraded phase count for a still-incomplete milestone
(repro8 case 1: 1/1/100 instead of 2/1/50; repro5: same, on a fully
version-less roadmap with no sibling milestones at all).

Introduces one shared helper, isBracketMilestoneBoundary(headingText,
level, selectedBracketId), used by both computeSectionEnd and the
preambleCutoff bracket scan, replacing the ungated `h.level <= 2 &&
bracketMilestoneHeadingRe.test(...)` inline check. `selectedBracketId`
(case-folded via phase-id.cts's foldBracketId, matching the branch's
own fold-before-identity convention) is derived from `selected[0]`,
which is the full matched heading line on BOTH selection paths
(version-string and bracket-fallback), so one extraction covers both.

Level cap stays at `level > 2` for now (temporary — ADR-612's content
discriminator replaces it in the next commit); same-milestone rejection
is the change this commit is scoped to.

DEVIATION from the reviewed plan, caught empirically: applying the
same-milestone rejection at the preambleCutoff site (as literally
specified) regressed an existing pin ("boundary heading level: a
level-1 CURRENT milestone heading also scopes correctly") and a
fenced-heading case (repro10 A3) — because preambleCutoff's job is
"where does the earliest milestone-shaped heading sit, scanning from
the TOP of the document," and the selected heading's own occurrence is
always a correct answer to that question regardless of same-id-ness;
rejecting it let the earliest-of-either comparison fall through to a
stray LATER heading instead. `selectedBracketId` is threaded through as
`null` at the preambleCutoff call site for this reason — bracket-shaped
(and, from the next commit, phase-tail) discrimination still applies
uniformly at both sites; only the same-milestone component is
call-site-specific, since it encodes a "keep scanning past this
heading" instruction with no counterpart in a top-of-document search.

Tests: new describe block "#612 PR-2 B1 round-2: a same-milestone
continuation heading is not a boundary" — RED-turned-GREEN fixtures for
repro8 case 1 and repro5, plus PINs for repro8 case 3 (trailing
different-id icebox still terminates) and repro10 A1 (all-version-
bearing + icebox + Phase Details stays exactly 2/1/50 — no double-count
from the same-milestone exclusion interacting with the pre-existing
detailsMatch re-append). syncedTotal()/syncedPercent() assertions
omitted from the repro10 A1 pin: that fixture carries dirs outside the
current milestone, which exposes the SEPARATE Major 1 defect
(cmdStateSync's body percent from an unfiltered disk scan) — asserted
once Major 1 is fixed, not here.

Full suite green (796/796 across the targeted adr-612 + roadmap-parser
+ state files); node scripts/lint-phase-id-drift.cjs clean; eslint
clean on both changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): bracket boundary discriminates by content, not heading level (B2)

Gate-2 adversarial review Blocker 2: three sites disagreed about which
heading levels are a bracket milestone. The selector
(roadmap-parser.cts's bracket-fallback SELECTION branch,
`^#{1,3}\s+\[CODE.MM\]`) and `isMilestoneBounded` (state.cts) both
admit level 1-3, but isBracketMilestoneBoundary's level cap only
admitted level 1-2 (`h.level <= 2`, from the B1 commit). A `###`-level
bracket milestone heading was therefore SELECTED and BOUNDED but never
TERMINATED: computeSectionEnd ran with level=3, a level-3 SIBLING
milestone survived the pre-existing `h.level > level` (not-deeper)
filter, failed the version/emoji test (version-less), then failed
`h.level <= 2` — falling through to `return content.length` and
sweeping the sibling milestone's own phases into the current one.
Reproduces trek-e's original #612 defect verbatim ("a safe degrade
became a confidently-wrong persisted number") on a heading level the
selector and bounding predicate both already admit (repro2 case C:
4/75% instead of 2/100%; mechanism confirmed directly via repro7 —
extractCurrentMilestone returned the whole 214-byte document).

ADR-612 Decision 1 (docs/adr/612-bracket-phase-id-convention.md:56)
specifies the discriminator as CONTENT, not level: "a phase heading is
a bracket followed by a digit-then-colon ([GSD.02] 05:); a milestone
heading is a bracket followed by a name." Replaces the `level > 2`
rejection with BRACKET_PHASE_TAIL_RE — built by interpolating
phase-id.cts's single-owner phaseHeadingPrefixSrcFor(ANY_BRACKET,
'bracket', false) plus the digit + optional-tag + colon tail every
phase-heading counter in this file already spells, not a re-typed
grammar — and widens the level check to a depth-sanity cap of 3
(mirroring the selector's own `#{1,3}` ceiling; NOT itself a
phase/milestone discriminator). Covers the dotted sub-phase heading
form (`[GSD.02] 05.03:`) via the same `[\w][\w.-]*` token, pinned by a
new fixture — the shape where a regex slip in the tail grammar would
hide.

preambleCutoff's own raw-scan regex is widened from `^(#{1,2})` to
`^(#{1,3})` in lockstep: the outer pattern's level ceiling must track
the helper's cap, or a level-3 PRIOR milestone heading is invisible to
that scan and its own phase heading leaks into the preamble
un-stripped (a real double-count this widening closes, verified
against repro2 case C directly).

The existing "boundary heading level: a level-3 CURRENT milestone
heading still scopes correctly" pin (3e562f12) now passes via a
DIFFERENT mechanism than before — its own neighbours are version-
bearing, so it previously passed via the version/emoji rule (the level
cap was never actually exercised by that fixture, per the round-2
review's own finding); with the content discriminator, the SAME
fixture's level-3 phase headings are now correctly excluded because
they are phase-tail-shaped, not because they are too deep. A
deliberate mechanism change, confirmed by re-running that test green
after this commit.

Also updates the "every selector call site declares the right
baseline" governance pin (adr-612-bracket-heading-selection.test.cjs):
BRACKET_PHASE_TAIL_RE is a new, legitimate ANY_BRACKET call site in
roadmap-parser.cts (always passing the literal 'bracket' convention,
since its only caller is already gated on bracketBoundaryActive) —
EXPECTED count bumped 1->2, with a matching BASE_SITES transcription
entry (identical src to every other ANY_BRACKET site, since the
function is pure).

Tests: new describe block "#612 PR-2 B2 round-2: the bracket boundary
is a CONTENT discriminator, not a level cap" — RED-turned-GREEN for
repro2 case C (exact total AND truthful percent, since
isMilestoneBounded already returns true at #{1,3}) and repro7's
mechanism, a PIN for the dotted sub-phase form, and a re-pin of repro8
case 3 (icebox) under the new mechanism.

Full suite green (907/907 across the targeted adr-612 + roadmap-parser
+ state + phase-id files); node scripts/lint-phase-id-drift.cjs clean;
eslint clean on all changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): fence-aware preamble cutoff on the bracket branch (Blocker 3)

Gate-2 adversarial review Blocker 3: preambleCutoff's bracket scan used
a raw content.match/matchAll — blind to fenced code blocks — while its
sibling computeSectionEnd (a few lines above it) already consumed
tokenizeHeadings(content), which strips fences. The two halves of one
boundary semantic disagreed about what a heading is.

A fenced markdown example in the preamble containing a bracket heading
(ADR-612's own docs do exactly this) was textually the earliest
`#{1,3} [CODE.MM]` match: preambleCutoff landed INSIDE the fence,
`preamble = content.slice(0, preambleCutoff)` ended with an unclosed
opener, and the unbalanced fence then blinded
getMilestonePhaseFilter's own tokenizeHeadings(scope) call — every
heading in the returned scope vanished, phaseCount degraded to 0, and
the pass-all filter admitted every directory on disk (repro11's
mechanism, confirmed directly: fence count 1/odd, tokenizeHeadings(scope)
-> only "Roadmap"). Regression vs round-1, which had no bracket pattern
to blind and so fell back to the correct heading (repro12 bracket row:
2/1/50 at round-1, 4/3/75 at HEAD).

Fixed by hoisting one tokenizeHeadings(content) call
(currentMilestoneHeadings) shared by computeSectionEnd and the
preambleCutoff scan, which now iterates that same fence-aware token
list instead of a raw regex. HeadingToken.text is already hash-stripped
and trimmed, so isBracketMilestoneBoundary needs no `^#{1,3}\s+`
re-derivation at this site (that spelling would not match h.text — a
note the round-2 review called out explicitly, confirmed while
porting). selectedBracketId stays `null` here, unchanged from the B1
commit's same-milestone-exclusion reasoning.

DISCLOSED, not fixed (explicitly out of scope per the round-2 review's
own minimal-fix note): the LEGACY (non-bracket) anyMilestonePattern
raw-match path shares the identical fence-blindness hazard and stays
byte-identical — a bracket repo whose preamble has a fenced
VERSION-BEARING heading still has the legacy raw-match win the
earliest-of-either min() (repro12's LEGACY control: 4/3/75, unchanged
across base/round-1/HEAD/this commit). Pinned here so a future reviewer
files this as a known, pre-existing gap rather than a new regression.

Tests: new describe block "#612 PR-2 Blocker 3 round-2: preambleCutoff
is fence-aware (bracket branch only)" — RED-turned-GREEN for repro12's
bracket row and repro11's mechanism (fence balance + non-degraded
phaseCount + correct per-directory admission), a PIN for repro12's
LEGACY control (the disclosed gap, explicitly unchanged), and a PIN for
repro10 A3 (a fenced heading INSIDE the current section must still not
terminate it).

Full suite green (1072/1072 across the targeted adr-612 + roadmap-
parser + state + phase-id + markdown-sectionizer files); node
scripts/lint-phase-id-drift.cjs clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): scope cmdStateSync's disk scan by milestone under bracket (Major 1)

Gate-2 adversarial review Major 1: `state sync` wrote a Progress
PERCENT computed from an UNFILTERED whole-disk scan, beside the
milestone-scoped total_phases/completed_phases it writes into the same
STATE.md via the refreshed frontmatter (syncStateFrontmatter ->
buildStateFrontmatter, which has always applied getMilestonePhaseFilter
for the READ path). cmdStateSync's own `fs.readdirSync` chain (the
WRITE-path scan) never called the milestone filter at all, unlike
buildStateFrontmatter's identical-purpose scan. One command therefore
wrote two contradictory numbers into one file: on the ADR-canonical
version-less bracket fixture (4 dirs, 3 complete; asserted milestone =
2 phases, both complete), base wrote total_phases:2/completed_phases:2
(correct, from the READ derivation) alongside body Progress 75% (wrong
— from the unfiltered WRITE derivation; repro3).

Fixed by threading `getMilestonePhaseFilter(cwd)` through the same
`.filter()` chain buildStateFrontmatter already applies, gated on
`syncConvention === 'bracket'` (falling back to a pass-all predicate
otherwise) — so totalDiskPlans/totalDiskSummaries/diskCompletedPhases/
syncTotalPhases become milestone-scoped under bracket, byte-identical
under legacy.

DEVIATION (approved, stated plainly): an earlier phrasing of this fix
called for mirroring buildStateFrontmatter's filter UNCONDITIONALLY.
Implemented GATED instead — an unconditional filter would ALSO move
every LEGACY repo's persisted percent, since the milestone-scoping-vs-
whole-disk divergence this closes is engine-wide, not bracket-specific.
The gate keeps legacy byte-identical, which is the binding constraint:
this is a bracket read-path PR, not a legacy behavior change.

Nit 2 (informational, no code change): 10 calls to
extractCurrentMilestone on a legacy repo cost 10 config.json
existsSync + 10 readFileSync (0 before B3); accepted, unmemoized cost,
unaffected by this commit.

Also folds in two minors from the round-2 review:
- Corrects .changeset/2761-bracket-read-tolerance.md: the sibling-
  exclusion sentence now states it holds at any heading level 1-3 and
  across a milestone split over two headings (true again now that
  Blockers 1 and 2 are fixed); the percent sentence states plainly that
  `state sync`'s body percent is now milestone-scoped under bracket,
  and unaffected under legacy.
- Records the read/write scoping divergence at currentMilestoneRawRanges
  (src/roadmap-parser.cts) in a comment: it did not receive B1/B2's
  bracket boundary fixes, currently harmless (its only consumer falls
  back to whole-content mutation, and every mutation there is still
  Phase-labelled-only, not bracket-widened), but live the moment the
  write path is bracket-widened — flagged so a future PR closes it in
  lockstep with that work, not after.

Tests: 6 pre-existing tests in tests/adr-612-bracket-phase-counting.test.cjs
needed fixture updates, not logic changes — they used the default
single directory (`GSD.02-01-setup`, phase "01"), which the SENTINEL/
retirement/mixed-heading fixtures in those tests never declare as a
real phase (only 04/05/06/999/etc are declared). Before this fix,
cmdStateSync's unfiltered scan counted that off-roadmap directory
anyway; after this fix the milestone filter correctly excludes it,
which for several of these fixtures made `state sync` a no-op (the
computed 0% coincided with STATE.md's initial template default) and
broke `syncedTotal()`/`syncedPercent()`'s ability to observe anything.
Updated each to pass an EXPLICIT directory naming one of the fixture's
REAL declared phases, preserving each test's original numerator/
denominator intent. One test — "shape 2 WRITE" — was substantively
rewritten: it was a CHARACTERIZATION of the Major 1 bug itself ("the
DISCLOSED legacy gap, mirrored — not closed"), and now correctly pins
bracket closing to 33% (agrees with the read path) while legacy stays
at the disclosed 60% (unchanged, deliberately, per the gating decision
above).

Full suite green: `npm test` 1449/1449 (0 fail, 0 skipped, 0 todo,
single-shard "all" run — includes issue-2765-brace-expansion-lockfile
passing); `npm run lint:ci` clean (0 errors; 2 pre-existing timing-
assertion warnings in files this PR does not touch); node
scripts/lint-phase-id-drift.cjs clean; node scripts/changeset/lint.cjs
ok.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): preambleCutoff identity is offset- and child-aware (round-3 Blocker 1)

Gate-2 round-3 re-verify Blocker 1 (NEW): the round-2 B1 deviation
(39c42a89) threaded `selectedBracketId` as the real value at
computeSectionEnd but as bare `null` at the preambleCutoff scan. The
deviation's rationale — "the selected heading's own occurrence is
always a correct earliest answer" — was right, but `null` disables the
same-milestone check for EVERY candidate, not just the selected one.
Any bracket-shaped heading earlier than the selected milestone was
accepted as a boundary regardless of identity: a same-id checklist/
overview heading preceding the version-bearing selected heading (cases
A, B — the version lands on the LATER half of a split, or a plain
overview heading with no "(Phase Details)" spelling), or a DIFFERENT-id
bracket-shaped PROSE heading with no children of its own sitting above
the current milestone's content (case D — `## [ADR.612] Heading
convention used by this roadmap`). In every case the region between
that false boundary and the real sectionStart was silently dropped —
a completed phase vanished and `state sync` persisted a confident 0%
where base and round-1 both correctly wrote 50%. Regression vs base
AND round-1 (not merely "under-fixed", per the round-3 review's own
severity note).

Fixed with two changes, both scoped to the preambleCutoff scan only
(computeSectionEnd already threads the real `selectedBracketId` and is
untouched):

(a) `h.offset === sectionStart` now bypasses BOTH the same-milestone
    check inside isBracketMilestoneBoundary (passing the REAL
    `selectedBracketId` for every other candidate) and the new child
    rule below — the selected heading's own position is definitionally
    the correct answer, so neither discriminator should run against it
    (rejecting it would mean rejecting the heading against ITSELF).
    Closes cases A and B — verified by the reviewer's own one-liner,
    reproduced here.

(b) New `bracketHeadingHasMatchingChild`: an otherwise-accepted
    candidate (bracket-shaped, not phase-tail-shaped, not the same id
    as the selected milestone) must ALSO have a next-strictly-deeper
    heading carrying its OWN bracket id to count as a boundary. This is
    what a genuine sibling milestone has (its own phase children share
    its bracket id — `## [GSD.01] Setup` / `### [GSD.01] 01: …`) and an
    unrelated bracket-shaped prose heading does not. A candidate with
    no such child at all (childless — e.g. an empty prior milestone, or
    one immediately followed by a same-or-shallower heading) degrades
    to NOT a boundary — over-inclusive, the safe direction: its own
    heading text stays in the preamble, contributing nothing to any
    phase count (not phase-shaped). Closes case D, which (a) alone does
    not — verified: without this rule, `[ADR.612]`'s prose heading is
    indistinguishable from a genuine prior sibling at this site.

As a side effect, also neutralizes Nit 2 (a colon-less `[GSD.02] 05`
heading spuriously terminating the preamble): a colon-less bracket
heading is not phase-tail-shaped so isBracketMilestoneBoundary alone
would accept it, but it is — precisely because it is malformed/
incomplete rather than a real milestone — childless, so the child rule
rejects it too. Pinned.

Known interaction with the fence-blind SELECTION path (disclosed by
the reviewer, not introduced here, tracked for the next commit): when
`sectionPattern` selects a FENCED version-bearing heading (an
extremely pathological shape — a fenced example whose text happens to
match STATE's asserted version), no token exists at `sectionStart`, so
the `h.offset === sectionStart` bypass never fires and the loop falls
through to the ordinary same-id / child-rule checks. This composes
with the round-3 Major 1 fix (next commit) rather than introducing a
new defect — SELECTION itself is untouched by any of this — but is
worth stating plainly rather than rediscovering.

Tests: new describe block "#612 PR-2 Blocker 1 round-3: preambleCutoff
identity is offset- and child-aware" — RED-turned-GREEN for cases A, B
(rv-attack1) and D (rv-attack1b) with syncedTotal()/syncedPercent()
assertions (the persisted 0% is the point), a PIN for a genuine prior
sibling with real children (still excluded), a PIN for a childless
prior sibling (degrades to not-cutting, over-inclusive/safe), a PIN
for the colon-less Nit 2 shape, and the reviewer's rv-mech1 mechanism
re-run as a proper test (scope now equals the full input document,
phaseCount 2, both dirs accepted).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 937/937 pass (930 baseline + 7 new). node
scripts/lint-phase-id-drift.cjs clean; eslint clean on both changed
files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): fence-aware version/emoji half of preambleCutoff on the bracket branch (round-3 Major 1)

Gate-2 round-3 re-verify Major 1 (NEW): ff6bf0a8 (round-2 Blocker 3)
made the BRACKET half of preambleCutoff's "earliest milestone-shaped
heading" search fence-aware, but left the VERSION/emoji half a raw
`content.match` even on the bracket branch. A fenced VERSION-BEARING
example heading in a bracket repo's preamble (ADR-612's own docs
illustrate the LEGACY heading shape exactly this way, inside a fenced
authoring-guide block) was still textually the earliest match for that
raw regex, winning the min() and un-suppressing a wrong persisted 75%
that base correctly suppressed (rv-attack3c fixture C1: base
suppressed the percent entirely — `isMilestoneBounded` false — HEAD
wrote 75% where truth is 50%).

Fixed by deriving the version/emoji half from the SAME fence-aware
`currentMilestoneHeadings` token list as the bracket half, on the
bracket branch only — the exact `/^Phase\s+\S/i` / `/v\d+\.\d+|✅|📋|🚧/i`
pair `computeSectionEnd` already uses against `h.text`. The non-bracket
(legacy) path is untouched: it keeps the raw `content.match`, byte-
identical to before, including its own fence-blindness (repro12's
LEGACY control, pinned unchanged in the round-2 Blocker-3 test block —
not re-pinned here to avoid duplicating an already-covered assertion).

Not rated Blocker (per the review) because it is not a regression vs
round-1 and the fixture (a version-BEARING fenced example in a bracket
repo) is rarer than the already-fixed bracket-heading case; still
fixed now rather than disclosed, per this arc's own precedent (every
prior "disclose instead of fix" call in this PR has been overturned on
re-review).

Tests: new describe block "#612 PR-2 Major 1 round-3: preambleCutoff's
version/emoji half is fence-aware on the bracket branch" — RED-turned-
GREEN for case C1 (readTotal + syncedPercent, so the persisted 75% is
directly observed, not just the read-path total), PIN for case C2 (the
already-fixed fenced-bracket-heading shape, unchanged).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 939/939 pass (937 + 2 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean on both changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): correct changeset claims + mark runtime-gated BASE_SITES row (round-3 minors)

Gate-2 round-3 re-verify Minor 2: three changeset sentences in
.changeset/2761-bracket-read-tolerance.md were overstated in a new
direction after round-2:

- The split-milestone claim ("across a milestone split over two
  headings") was true only when the version-bearing heading came
  FIRST (repro8 case 1); false when it came LATER (round-3 Blocker 1
  case A). Now restated to say plainly "with the version-bearing
  heading in EITHER position" — true again now that round-3's Blocker
  1 fix lands earlier in this range.
- The "counted from the phases... rather than from every directory on
  disk" claim was false on cases A/B/D (a strict subset of the
  milestone's own phases). Restated as "ALL of the phases... not a
  subset", and extended to state that an unrelated bracket-shaped
  heading with no phase children of its own (case D's `[ADR.612]`
  shape) does not truncate the milestone either — true now, not before.
- "Each widened read is SELECTED by the project's phase_id_convention"
  was literally false for BRACKET_PHASE_TAIL_RE, which is RUNTIME-gated
  (via its only caller, isBracketMilestoneBoundary, itself only
  consulted when bracketBoundaryActive) rather than selector-gated.
  Restated behaviourally: "every widened read ENGAGES only when the
  project's resolved phase_id_convention is bracket" — true for both
  gating mechanisms, so it no longer implies a selector call this site
  does not make.

Minor 1: the STRUCTURAL IDENTITY test's BASE_SITES row for
BRACKET_PHASE_TAIL_RE (added in the B2 commit) asserts a property of
`phaseHeadingPrefixSrcFor` — the function — not of the call site; it
would pass unchanged even if the site were deleted. Safety at that
specific site rests entirely on a runtime gate the test cannot see.
Added `runtimeGated: true` to the row and threaded it into the
generated test's own title (`… [runtime-gated, not selector-covered]`),
so the gating mechanism is visible in test OUTPUT, not only in a source
comment that could drift silently.

No production code changed. Targeted suite green (48/48 in the
affected file); `node scripts/changeset/lint.cjs` ok; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): harden round-3 preamble cutoff — subtree child scan + level cap

Team-lead review of f87bba0e found two edges in the round-3 preamble-
cutoff code, both hardened here.

AMENDMENT 1 — bracketHeadingHasMatchingChild (2e06aef5) checked only
`headings[index + 1]`, the IMMEDIATE next heading, not the candidate's
whole subtree. A genuine prior sibling milestone whose section opens
with a non-bracket subsection before its first phase heading
(`## [GSD.01] Setup` / `### Notes` / `### [GSD.01] 01: Old`) was
therefore wrongly rejected as a boundary — its real phase heading sits
TWO headings deep, not one — leaking its entire section into the
preamble unstripped.

CONFIRMED RED, not merely theoretical (built and ran the fixture
against f87bba0e before touching the fix, per instruction): scope
membership DOES drive the disk-side filter on this shape.
`GSD.01-01-old`'s directory was wrongly admitted into the CURRENT
milestone's filter via the leaked heading's qualified key
(`GSD.01-01`) — 3/2/67% where truth is 2/1/50%.

Fixed by scanning the candidate's full SUBTREE: continue past a
non-matching deeper heading instead of returning false on the first
one; only a same-or-shallower heading actually closes the subtree and
yields "no match found". A candidate whose entire subtree closes with
no same-id hit (including a genuinely childless one) still degrades to
`false` — over-inclusive, safe, unchanged from before.

AMENDMENT 2 — c483552a ported the version/emoji half of preambleCutoff
to the token-based scan with no level cap; the raw
`content.match(anyMilestonePattern)` it replaced was anchored
`^#{1,3}\s+`. A level-4+ version-bearing heading in the preamble
(`#### v2.0 notes`) therefore won the scan on the bracket branch where
the raw pattern — and the legacy path, unaffected — ignores it
outright. Fixed with `if (h.level > 3) continue;`, mirroring the
depth-sanity cap isBracketMilestoneBoundary already applies to the
bracket half of this same scan.

Tests: new describe block "#612 PR-2 round-3 hardening: subtree child
scan + level cap on preambleCutoff" —
- RED-turned-GREEN for the Notes-intervening fixture: exact 2/1/50 (was
  3/2/67), plus the disk-filter observable (`GSD.01-01-old` now
  correctly excluded).
- PIN for the level-4 preamble heading: the scope now PRESERVES the
  heading's text (was silently dropped before this fix — harmless in
  this minimal fixture's total_phases specifically, since the dropped
  text carries no phase-shaped content, but a real correctness gap
  against the raw pattern's own ceiling) — asserted via scope content,
  not total_phases, since that number is invariant here either way.
- PIN for the LEGACY control on the same level-4 shape — unchanged,
  confirming the raw content.match path is untouched.

Re-verified the existing genuine-prior-sibling and childless-sibling
pins (round-3 Blocker 1 commit) still pass under the subtree scan —
both fixtures' outcomes are unchanged since their same-id hit (or its
absence) was already at the first deeper heading.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 942/942 pass (939 + 3 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean on both changed files. Full `npm test` + `npm run
lint:ci` deferred to the team lead's own run per instruction.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): bracketHeadingHasMatchingChild requires a same-id PHASE child (round-4 Blocker 1)

Gate-2 round-4 re-verify Blocker 1 (NEW): the round-3 hardening's
subtree scan (fbfd0fca) proved SAME-ID-NESS but never asked whether
the matching child was PHASE-shaped. Case F1 re-opens round-3's case D
one heading later: `## [ADR.612] Heading convention` is followed by
its OWN sub-heading `### [ADR.612] Examples` — same bracket id as the
candidate, but MILESTONE-shaped (a name, no digit-then-colon), not a
phase. Same-id-ness alone satisfied the subtree scan and re-cut the
preamble at exactly the shape the round-3 hardening was written to
close.

Failing input: `## [ADR.612] Heading convention` / `### [ADR.612]
Examples` (prose) / `### [GSD.02] 01: One` (the current milestone's
own first phase, now unreachable) / `## [GSD.02] v2.0: Foundation` /
`### [GSD.02] 02: Two`. Truth 2/1/50. HEAD read 1/0/0, and `state sync`
reported "nothing to do" (exit 0, `{synced:true,changes:[]}`) because
its wrong 0% happened to equal the STATE.md seed — a half-done
milestone read as untouched with no write-path signal at all.

Fixed with the reviewer's one-conjunct addition: a same-id child only
counts if it is ALSO phase-tail-shaped (`BRACKET_PHASE_TAIL_RE`) — the
same single-owner discriminator `isBracketMilestoneBoundary` already
uses one level up for the identical distinction (phase vs milestone),
reused here rather than re-derived. This is exactly what the
changeset's own wording already claimed ("no phase children of its
own") — the code now matches the sentence rather than the other way
around.

Docstring updated at the function itself: the rule is "same-id PHASE
child", not "same-id child".

Tests: new describe block "#612 PR-2 round-4 Blocker 1: the same-id
child must be PHASE-shaped" — RED-turned-GREEN for F1 with
syncedTotal()/syncedPercent() (the persisted 0% — and the
report-nothing-to-do write-path silence — is the point), PINs for F11
(colon-less same-id child) and F11b (bullet-only phase list): both
correctly stay excluded either way, and the leak the phase-shape
requirement newly creates for these two shapes is INERT — a colon-less
heading forms no qualified key (getMilestonePhaseFilter's own
phase-heading pattern requires the colon too) and a bracket bullet
never matches the legacy-only BULLET_PHASE_LINE_PATTERN — confirmed
directly via getMilestonePhaseFilter, not merely inferred. Re-verified
the four existing child-rule pins (F2 subtree-closure, F3 deep-nested
same-id, F4 level-4 same-id, F5 childless-at-EOF) are unaffected by the
phase-shape requirement, since every one of them already used a
colon-bearing same-id child.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 946/946 pass (942 + 4 new). Zero drift measured across the full F1-F12
corpus except F1 itself (F9/F10/F12 remain red, deferred to the
separate Major 1 fix). node scripts/lint-phase-id-drift.cjs clean;
eslint clean on both changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): fence-aware phase counting, milestone bounding, and bracket-fallback selection (round-4 Major 1)

Gate-2 round-4 re-verify Major 1 (NEW): roadmapPhaseCount is a
fence-blind raw `.exec()` over the scope string, duplicated in TWO
independent copies (buildStateFrontmatter's read path, cmdStateSync's
write path). With the bracket alternative now compiled into it (#612),
a fenced EXAMPLE phase heading in the preamble inflates total_phases
and persists a wrong percent that base got right. Two further
fence-blind sites participate: isMilestoneBounded (a raw
`.test(roadmapRaw)`) and the bracket-fallback SELECTOR inside
extractCurrentMilestone (a raw `content.matchAll`, only reachable when
version-string selection finds nothing).

Failing inputs:
- F10 (clean isolate, version-bearing selection): a fenced
  `### [GSD.02] 05: Example phase` in the preamble inflates
  total_phases 2->3, persisting 33% where truth is 50%. LEGACY control
  on the same shape is correct on every build — not a pre-existing
  hazard being inherited, bracket-only.
- F9 (version-less selection): a fenced example carrying the project's
  OWN milestone id additionally confuses the bracket-fallback selector
  (the fenced heading gets SELECTED), compounding with the same
  fence-blind counter. base suppressed the percent; round-1 and HEAD
  both wrote 33%.
- F12 (isMilestoneBounded isolate): the ONLY `[GSD.02]` heading in the
  document is inside a fence, and the asserted milestone genuinely has
  no section at all — HEAD persisted 67% where base correctly
  suppressed the percent (the milestone is absent from the roadmap).

Fixed at the CONSUMER level, not the producer — extractCurrentMilestone's
returned scope string is deliberately UNCHANGED, since every other
consumer of that string needs its full content fidelity and legacy
identity forbids touching the shared string (this branch's own
precedent, ff6bf0a8/c483552a, was producer-level; here the ruling is
consumer-level because the string is shared far more broadly than the
two round-3 fixes' narrower producer edits):

(a) New `countRoadmapPhaseHeadings` (src/state.cts, immediately above
    extractRetiredPhaseNumbers) — ONE shared implementation for both
    call sites, replacing two independently-maintained copies. BRACKET
    convention counts via `tokenizeHeadings(scope)` at levels 2-4,
    testing each heading's hash-stripped text directly — fence-aware by
    construction, since tokenizeHeadings never produces a token for a
    fenced line. LEGACY convention keeps the exact pre-existing raw
    `.exec()` loop, byte-for-byte. A pre-existing, deliberately
    PRESERVED asymmetry between the two original call sites — the read
    path always excluded a bare `/^999\b/` token, the write path never
    did — is threaded through as an explicit
    `includeUnconditional999Check` parameter per call site, so sharing
    the implementation does not silently unify (and thereby move)
    either total.
(b) isMilestoneBounded's bracket branch now scans
    `tokenizeHeadings(roadmapRaw)` for a matching heading (level <= 3)
    instead of a raw regex test. Legacy version-string branch untouched.
(c) The bracket-fallback SELECTOR now builds its candidate set from
    `tokenizeHeadings(content)` instead of `content.matchAll`,
    reconstructing a match-shaped array so every downstream consumer of
    `headingMatches` sees the identical shape the raw-regex path always
    produced. This is the ONE site in this entire arc where SELECTION
    itself changes — selection SEMANTICS are otherwise unchanged (same
    pattern, same first-match-wins by document order); only the
    candidate set is now fence-aware. Pinned that unfenced selection is
    byte-identical.

Zero drift measured across the full historical corpus (repro2-13,
rv-attack1/1b/3c, rv-mech1, rv2-amend1/2, and F1-F11b) except the three
target fixtures.

Tests: new describe block "#612 PR-2 round-4 Major 1: four fence-blind
sites on the bracket path" — RED-turned-GREEN for F10 (with
syncedTotal()/syncedPercent()) and F9 (both layers), PINs for F10's
LEGACY control and F10c (non-phase-shaped fence, unaffected either
way), RED-turned-GREEN for F12 (asserts the percent KEY is absent from
`state json`'s output and that `state sync`'s body stays at its
unmodified seed — the persisted-suppression signal, not merely a
total_phases number), and an explicit PIN that unfenced bracket-fallback
selection (first real milestone-shaped heading wins, no fences
involved) is unaffected.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 952/952 pass (946 + 6 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean on all three changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): correct docstring overstatement + stale consumer-count sentence (round-4 minors)

Gate-2 round-4 re-verify Minor 1 + Nit 1. No production code changed.

Minor 1: bracketHeadingHasMatchingChild's own docstring said a
rejected (no-same-id-PHASE-child) candidate's degrade "contributes
nothing to any phase count" — true of the candidate's OWN heading
text, but not of its SUBTREE, which is what actually stays in the
preamble. F7 (`## [GSD.01] Setup` / `### [GSD.07] 01: Foreign`) shows
a DIFFERENT-id bracket PHASE heading inside a rejected candidate's
subtree DOES form a qualified key and CAN admit a foreign directory —
3/2/67%, stable across base, round-1 and HEAD (base via its own
pass-all degrade). Not a regression, still the declared over-inclusive
/ never-under-inclusive safe direction — the comment now says that,
with F7's numbers cited, at the call site that actually decides
`isBoundary` (roadmap-parser.cts's preambleCutoff loop) rather than
only at the helper's own definition.

Minor 2 (changeset) — VERIFIED, no wording change needed: re-ran F1,
F9, F10, F12 at this HEAD. The "no phase children of its own... does
not truncate it either" sentence (naming the `[ADR.612]` shape
directly) is now literally true — F1 reads 2/1/50. The "counted from
ALL of the phases... not a subset" sentence is now true on every
measured shape — F1/F9/F10 all read 2/1/50, F12 correctly suppresses
the percent. No carve-out for F9/F10 is needed since round-4 Major 1
(3be5c412) closes both; per the fix-round instruction to "only carve
out anything genuinely left," nothing is.

Nit 1: tests/adr-612-bracket-heading-selection.test.cjs's runtimeGated
row claimed `BRACKET_HEADING_INTRO_RE` has "no other consumers" — true
when round-3's f87bba0e wrote it, stale since 2e06aef5 (round-3's own
earlier commit) had already added two more uses inside
bracketHeadingHasMatchingChild. Corrected to state the true count
(three consumers) and re-confirm the conclusion is unaffected: all
three are still nested inside the same bracketBoundaryActive runtime
gate, and BRACKET_HEADING_INTRO_RE is built from BRACKET_ID_SRC, not
phaseHeadingPrefixSrcFor, so it was never a selector site regardless of
consumer count.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 952/952 pass (comment-only changes, no count movement). eslint
clean; node scripts/changeset/lint.cjs ok.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): restore the bracketId guard in countRoadmapPhaseHeadings (round-5 Blocker 1)

Gate-2 round-5 re-verify Blocker 1 (NEW, introduced by 3be5c412): the
merge that created the shared countRoadmapPhaseHeadings helper dropped
the `bracketId && ` guard both original inline loops carried before
calling isSentinelPhaseId. Every other isSentinelPhaseId call site in
src/ (roadmap-parser.cts, roadmap.cts, validate.cts x2, verify.cts)
keeps the guard; state.cts's shared counter was the only one of seven
without it.

When the phase-heading-intro grammar's LEGACY alternative matches (a
`### Phase 00:` heading in a `phase_id_convention: "bracket"` repo —
the mid-migration shape this PR exists for), the bracket capture group
is `undefined`, so the unguarded call became
`isSentinelPhaseId("undefined-00", 'bracket')` — measured TRUE, so
phase 00 (and 000, 0a, 0.5, 999.1 — any token whose splice with the
literal string "undefined" happens to fall in a sentinel range) was
silently dropped from the denominator. `getMilestonePhaseFilter` (which
still carries its own guard) counts the phase and admits its directory
regardless, so the filter and the counter disagree — a half-done
milestone reads as 100% complete, persisted.

One-line fix, restoring the guard every sibling call site already has:

    if (bracketId && isSentinelPhaseId(`${bracketId}-${token}`, 'bracket')) continue;

Line count: `git diff --stat src/state.cts` -> 1 file changed, 1
insertion(+), 1 deletion(-).

Tests: new describe block "#612 PR-2 round-5 Blocker 1:
countRoadmapPhaseHeadings restores the bracketId guard" — RED-turned-
GREEN for G3 (3/2/67, was 2/2/100) and G3d (the mixed bracket+legacy
mid-migration shape, same numbers) with syncedTotal()/syncedPercent(),
PIN for G3's legacy control (unaffected), PIN for G3b (isolates the
counter with no directory to admit), PIN for G3c (legacy 01/02 only,
no sentinel-shaped token present).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 957/957 pass (952 + 5 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): extractRetiredPhaseNumbers is fence-aware on the bracket path (round-5 Major 1)

Gate-2 round-5 re-verify Major 1 (NEW) — a FIFTH fence-blind site on
the bracket path, missed by 3be5c412's own enumeration of "four".
extractRetiredPhaseNumbers' line scan (`scope.split(/\r?\n/)`) has no
fence awareness. This PR compiles the bracket alternative into
`introSrc` ("the retirement filter has to widen with the counter it
protects" — the function's own pre-existing comment), so a FENCED
authoring EXAMPLE showing the #1514 retirement gesture in bracket
spelling is now indistinguishable from a real one: it retires a
genuine phase, shrinking the denominator and persisting a confident
100% where base correctly read 50%.

Fixed with the SAME consumer-level ruling this arc has used at every
other fence-blind site, reusing markdown-sectionizer's existing
exported `stripFencedCode` rather than hand-rolling a second fence
parser (single-owner rule) — retirement lines are BULLETS, not
headings, so `tokenizeHeadings` doesn't serve here; `stripFencedCode`
is the general-purpose fence stripper the tokenizer itself is built on.
Gated on `convention === 'bracket'`; the LEGACY line scan stays the raw
`scope` string, byte-identical — its own fenced-example hazard is
pre-existing (wrong at base too) and out of scope.

Line count: `git diff --stat src/state.cts` -> 1 file changed, 9
insertions(+), 2 deletions(-) — one import added, four lines inside
the function (a comment + the `scanScope` computation + the changed
`.split()` call).

Tests: new describe block "#612 PR-2 round-5 Major 1:
extractRetiredPhaseNumbers is fence-aware on the bracket path" —
RED-turned-GREEN for G2 (fenced example in the preamble) and G2b (the
same example placed INSIDE the milestone section, ruling out a
preamble-scoping artifact — the site itself was fence-blind wherever
the fence sits), PIN for the LEGACY control (unchanged, pre-existing,
out of scope — base is wrong on this shape too).

Also folds in the "four fence-blind sites" correction: 3be5c412's
commit message and any restatement of it should read FIVE going
forward; this commit's own message states the count correctly.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 960/960 pass (957 + 3 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): state sync's counter excludes the bracket 999 icebox token (round-5 Major 2)

Gate-2 round-5 re-verify Major 2 (NEW) — `includeUnconditional999Check`
left `state json` and `state sync` reporting different totals for one
bracket repo. Under bracket, READING-B puts the sentinel in the
bracket, so `isSentinelPhaseId("GSD.02-999", 'bracket')` is false —
the `/^999\b/` TOKEN rule is the only thing excluding a
`### [GSD.02] 999:` icebox heading, and it ran on the read path
(buildStateFrontmatter, `true`) and on getMilestonePhaseFilter
(unconditional), but not on cmdStateSync's own counter (`false`). One
`state sync` call could leave a single STATE.md with its own
frontmatter (percent 50, from the read-path re-sync inside
writeStateMd) and body (percent 33, from the write-path counter that
alone still counted the icebox heading) disagreeing — falsifying this
PR's own stated invariant that sharing countRoadmapPhaseHeadings made
"the two counters must see the same phases" structural.

Functional change is one argument, exactly as specified: the write
call site now passes `syncConvention === 'bracket'` instead of the
literal `false`. `syncConvention === 'bracket'` is `false` for every
non-bracket value, so the LEGACY path resolves to the exact same
`false` it always did — this file's own pre-existing, deliberately-
unchanged read/write divergence on that path is untouched. The READ
site (`:1860`) is NOT touched — its historical behaviour applied
`/^999\b/` to legacy and bracket alike, so changing it would move
legacy READ totals, exactly the class of mistake this arc's own
Blocker 1 (this round) was.

Line count: the functional change is ONE argument
(`false` -> `syncConvention === 'bracket'`); the surrounding comment
was rewritten because the previous one asserted the now-superseded
behaviour ("preserving this file's pre-existing... divergence... a
bare 999 token is not excluded here") and leaving it would mislead the
next reader — not a structural change.

Tests: new describe block "#612 PR-2 round-5 Major 2: state sync
excludes the bracket 999 icebox token like the read path" —
RED-turned-GREEN for G1, asserting `state sync`'s body percent equals
`state json`'s own percent (both 50, not 33 vs 50), PIN for G1's
legacy control (33 vs 50 unchanged, deliberately).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 962/962 pass (960 + 2 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): restore indent parity at the two line-start-anchored bracket-heading reconstructions (round-5 Minor 1)

`HeadingToken.offset` (tokenizeHeadings) is the LINE-START character offset,
not necessarily the `#` character's own offset — a ≤3-space-indented ATX
heading has both. Two round-4 reconstructions built on tokenizeHeadings
inherited this gap and accepted indented headings their raw, line-start-
anchored predecessors (`^#{1,3}\s+\[...`) never matched:

  1. roadmap-parser.cts's bracket-fallback SELECTOR (extractCurrentMilestone,
     ~line 377) — an indented, version-less `[GSD.02]` milestone heading
     could be reconstructed into headingMatches, then mis-parsed downstream
     (selectedBracketId null, level fallback to 1), leaking a SIBLING
     milestone's phases into the counted scope (G6: HEAD read 3/2/67 instead
     of 2/1/50 — a real phase heading's own directory belonging to the NEXT
     milestone got counted).

  2. state.cts's isMilestoneBounded (~line 1571) — the same gap let an
     indented-only `[GSD.02]`-shaped heading wrongly bound a milestone
     absent from the roadmap, un-suppressing a percent that should stay
     suppressed (mirrors round-4's F12 fenced-only case, but via indentation
     instead of a fence).

Fix: one added conjunct per site — `content[h.offset] === '#'` (source-named
`roadmapRaw` in state.cts) — filtering to tokens whose LINE-START offset IS
the `#` character, i.e. exactly the set the raw line-start-anchored regex
would ever have matched. Restores byte-for-byte raw parity; no other logic
in either function changes. computeSectionEnd and the preamble version/
emoji-token scan are untouched, as instructed — they consumed tokenizeHeadings
output before this arc and are out of scope here.

Also corrects roadmap-parser.cts's now-provably-false docstring claim that
`h.offset` is unconditionally "the same `#`-character coordinate space
`content.match().index` used" — true only for the survivors of the new
filter, not for every token tokenizeHeadings produces.

Line count: the FUNCTIONAL change is exactly 2 lines (one added `&&` conjunct
per call site — `git diff --stat` on the two source files shows 20
insertions/7 deletions, but only those 2 lines change behavior; the rest is
docstring/comment rationale, per this round's "state the line count" ask).

TDD: both fixtures verified RED at HEAD before this commit, GREEN after,
via /tmp/pr612rev/rv5-attack.cjs G6 and a locally-authored isMilestoneBounded-
isolating probe (G6 alone doesn't distinguish the two sites — its unindented
phase headings already satisfy isMilestoneBounded's loose prefix regex
either way, so a second, indentation-only fixture was needed to prove that
site's fix is not a no-op; verified by temporarily reverting just that one
conjunct, confirming 100%-wrongly-bounded RED, then restoring it, confirming
suppressed-percent GREEN).

New tests (adr-612-bracket-phase-counting.test.cjs):
  - RED (G6): indented version-less bracket milestone heading — 2/1/50, not
    the pinned-before-fix 3/2/67.
  - PIN (G6c unindented control): identical document, no indent — 2/1/50
    unaffected on every build.
  - RED (isMilestoneBounded site, indented-ONLY): mirrors round-4's F12
    shape (fenced-ONLY → indented-ONLY) — percent stays suppressed instead
    of the pinned-before-fix wrongly-bounded 100%.

Full G1-G10 (rv5-attack.cjs) + G2b/G3c/G3d/G6c (rv5b.cjs) re-verified
zero-drift against TRUTH after this change. Targeted suite (adr-612-*,
roadmap-parser, state, verify): 962 -> 965 (+3), 0 fail.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): changeset discloses counting-set narrowing + correct stale consumer-count sentence (round-5 minors)

FIX 5 (Minor 2): .changeset/2761-bracket-read-tolerance.md did not disclose
that bracket phase-heading counting is narrower than the raw-regex
predecessor in two ways the review's G4/G5 fixtures surfaced: a level-5
heading (`##### [GSD.02] 05: ...`) is no longer counted (the counter's
tokenizeHeadings scan caps at level 4, matching the selector/isMilestoneBounded
ceiling), and a space-less heading (`###[GSD.02] 05: ...`) is no longer
counted (CommonMark requires ≥1 space/tab after the hashes, which
tokenizeHeadings correctly enforces and the old raw regex did not). Both
G4 and G5 moved from round-1's wrong (inflated) values back to base's
original values as an incidental side effect of routing through
tokenizeHeadings — never a deliberate feature of this PR, and previously
undocumented. One clause added to the existing run-on paragraph; no other
wording in the changeset touched.

FIX 6 (Nit 1): tests/adr-612-bracket-heading-selection.test.cjs:76 —
`BRACKET_PHASE_TAIL_RE` has TWO consumers as of round-4's 65d257ce
(isBracketMilestoneBoundary's own use, plus bracketHeadingHasMatchingChild's
same-id-PHASE-child conjunct), not the "no other consumers" the comment
claimed. Same correction pattern round-4 already applied to this row's
BRACKET_HEADING_INTRO_RE neighbor (that sentence's own staleness was fixed
in 4d7184b8): note the true consumer count, confirm both stay nested inside
the same bracketBoundaryActive runtime gate (verified at
src/roadmap-parser.cts:178 and :250, both reached only through the
`if (bracketBoundaryActive)` block starting at :529), and record which
commit and which fix introduced the drift. Comment-only; no assertion
logic changed.

Line count: 2 files, 8 insertions / 2 deletions total — one added clause
in the changeset (1 line changed) and one comment block replacing the
single stale line in the test file (6 comment lines replacing 1).

Verified: targeted suite (adr-612-*, roadmap-parser, state, verify)
unchanged at 965/965 pass (comment/prose-only diffs, no test count change).
`node scripts/changeset/lint.cjs` and `npx eslint` on the touched files both
clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): restore the bracketId guard on the bracket-only /^0\b/ sibling rule (round-6 Blocker 1)

Round-5's Blocker 1 was `isSentinelPhaseId("undefined-00")` — the merge
that introduced the shared counter dropped the `bracketId &&` guard on the
sentinel check. This is the identical failure one line further down, in the
sibling rule this PR itself added: when the phase-heading grammar's LEGACY
alternative matches (`### Phase 0:` in a `phase_id_convention: "bracket"`
repo — the mid-migration shape this PR exists for), `bracketId` is
`undefined`, and the unguarded `/^0\b/` fires on the bare token anyway.
Neither the LEGACY branch of this same function nor
`getMilestonePhaseFilter` has a `/^0\b/` rule at all, so the filter counts
the phase and admits its completed directory while the counter refuses to
count its heading — a milestone with an unstarted phase 02 persists as a
confident 100%.

`/^0\b/` matches `0` and `0.5` (word boundary before the `.`) but not `00`
(no boundary between the two zeros), which is exactly why round-5's G3/G3d
fixtures (`### Phase 00:`) never tripped this one — same defect class,
different token spelling.

Fix: one word, mirroring the guard round 5 restored two lines above —
`if (/^0\b/.test(token)) continue;` -> `if (bracketId && /^0\b/.test(token)) continue;`.

Expected and intentional side effect: `roadmap analyze` and `state json`
now disagree again on this bracket-repo shape (analyze phase_count=2, json
total_phases=3) — exactly as they already do under the legacy convention
today (verified via /tmp/pr612rev/rv6c.cjs on both conventions). That is
the counter regaining agreement with `getMilestonePhaseFilter` (the tighter
constraint — it is what actually decides `completed_phases`), not a new
break; the counter/filter disagreement is what was wrong.

Out of scope, deliberately NOT fixed here (Minor 1, disclosed via a PIN
test only): the bracket-SPELLED `### [GSD.02] 0:` shape has the same
counter/filter disagreement, but reads 2/2/100 on base too — never closed
by any build in this arc, so it is a pre-existing gap rather than a
regression this commit could introduce.

Line count: src/state.cts is exactly 1 insertion / 1 deletion (one word,
`bracketId && ` prepended to the existing condition).

TDD: T0 (`Phase 0:`) and T05 (`Phase 0.5:`) verified RED at HEAD before this
commit (json 2/2/100, sync body 100%) via /tmp/pr612rev/rv6b.cjs, GREEN
after (3/2/67 on both derivations, matching TRUTH). Zero-drift verified by
diffing the FULL corpus (rv5-attack, rv5b, rv4-attack, rv-attack1,
rv-attack1b, rv-attack3c, rv-mech1, rv2-amend1, rv2-amend2, rv6-attack
H1-H19, rv6b) between a pre-fix and post-fix build: the only differing
lines in the entire diff are T0, T05, and H14 — the three target fixtures.

New tests (adr-612-bracket-phase-counting.test.cjs):
  - RED (T0) + PIN (T0L legacy control)
  - RED (T05) + PIN (T05L legacy control)
  - PIN (B0, bracket-spelled `0` — base-parity characterization, Minor 1,
    not fixed this round)
  Round-5's own G3 test (`### Phase 00:`) already serves as the T00 pin —
  `/^0\b/` never matched `00`, so it is unaffected and untouched.

Targeted suite (adr-612-*, roadmap-parser, state, verify): 965 -> 970 (+5),
0 fail. `lint-phase-id-drift.cjs` and `npx eslint` on touched files clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): changeset re-attributes the phase-count level cap + discloses the bare-0 carve-out; correct two stale test comments (round-6 Minor 1 + Nits)

FIX 2 (Minor 1 — disclosure only, NOT a code fix): the bracket-SPELLED
`### [GSD.02] 0:` / `0.5:` shape has the same counter/filter disagreement
round-6's Blocker fixed for the legacy spelling, but it reads 2/2/100 on
BASE too — never closed by any build in this arc, so it is a pre-existing
gap rather than a regression this round could introduce. Two changes:

  - Pinned as a base-parity characterization: new PIN test (B0) in
    04b5a95f's describe block asserting the exact unchanged value with a
    comment stating why it's deliberately not touched.
  - .changeset/2761-bracket-read-tolerance.md: one qualifying clause added
    to the "counted from ALL of the phases … not a subset" sentence — a
    bare `0`/`0.x` phase token (however spelled) keeps `roadmap analyze`'s
    own pre-existing sentinel reading under bracket too, so it stays
    excluded from these counts. Carried over, not newly introduced by this
    PR — `roadmap analyze` has always read it this way.

FIX 3(a) — same changeset paragraph misattributed the phase counter's
2-4 level cap to "the milestone-boundary machinery in this PR" (the
selector, `isMilestoneBounded`, the preamble scan) — that machinery caps
at `###` (level ≤3), not 2-4. The 2-4 cap belongs to
`getMilestonePhaseFilter`'s own phase scan. Re-pointed the attribution;
the paragraph's earlier, correct ≤3 claim (selector/isMilestoneBounded)
is untouched.

FIX 3(b) — tests/adr-612-bracket-heading-selection.test.cjs:76-82 (added by
1395bd89, round-5's own Nit-1 correction) claimed both
`BRACKET_PHASE_TAIL_RE` consumers are reached "only through the
`if (bracketBoundaryActive)` block starting at :529". Verified at HEAD:
`bracketHeadingHasMatchingChild` is (only caller :593, inside that block).
`isBracketMilestoneBoundary` is not — it has two callers, an inline
`bracketBoundaryActive &&` conjunct at `:473` (inside `computeSectionEnd`)
and the `:529` block at `:567` — the very fact this row's own earlier
"exactly two callers" paragraph already stated correctly. Corrected the
mechanism claim; the CONCLUSION (every consumer is still gated on the same
flag) is unchanged, exactly as round-5's own Nit-1 fix left round-4's
conclusion unchanged when it corrected the consumer count.

FIX 3(c) — tests/adr-612-bracket-phase-counting.test.cjs:2311,2340 still
said "four fence-blind sites" after round-5 (357ba671) added a fifth
(the retirement scan) in its own block below. Reworded both the section
comment and the `describe()` label to read as historical scoping of
round-4's own fix ("the four sites known at round 4 … a fifth was found at
round 5, see its own block below") rather than a live exhaustive claim.

Line count: 3 files, changeset 1/1, heading-selection test 17/8,
phase-counting test 6/2 (comment/prose-only; B0's own PIN test landed in
04b5a95f alongside T0/T05 since it was investigated as part of that same
Blocker's defect class, not in this commit — noted here for the record).

Verified: targeted suite (adr-612-*, roadmap-parser, state, verify)
unchanged at 970/970 (comment/prose-only diffs, no test count change).
`npx eslint` and `node scripts/changeset/lint.cjs` both clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): W021's bracket remediation hint stops pointing at a command that hard-errors (round-7 Minor 2)

`checkBracketCoherence`'s W021 (added by 94abf5df) attached the fix hint
`Run \`gsd-tools roadmap upgrade --convention bracket\` to migrate
(dry-run by default)` to every bracket-convention W021 — both sub-checks
(`missing-bracket` and `mismatch`) share the single `addIssue` call at
src/verify.cts:2255-2256. But `roadmap-command-router.cts:204` throws
unconditionally for any `--convention` value other than
`milestone-prefixed`:

  $ gsd-tools roadmap upgrade --convention bracket
  Error: Only --convention milestone-prefixed is supported

This contradicts the PR's own two disclosures: the changeset ("`bracket`
is a READ-path opt-in until the migrator and write path land") and
docs/CONFIGURATION.md:188 ("There is no bracket migrator and no bracket
emit yet"). Reachability is the exact mid-migration repo this PR targets —
any bracket project with one un-migrated heading gets an unfollowable
instruction on every `validate health`.

Fix: one string. Replaced the hint with what a user can actually do today —
manually align the heading's bracket milestone to its section — and named
the tracked future landing (#612 PR-3) instead of a command that errors.
The milestone-prefixed sibling hint at src/verify.cts:2240 (a different,
already-functional convention/command pair — verified against
roadmap-command-router.cts:65) is untouched.

Line count: src/verify.cts is exactly 1 insertion / 1 deletion (one string
literal).

No test in the suite previously asserted this string's content
(`grep -rn "upgrade --convention bracket" tests/` was empty), so the
unfollowable hint shipped unpinned — `lint-fix-has-regression-test.cjs`
would not have caught a string-only change without a new test. Added one:
asserts the fix string both (a) does not match `--convention bracket`
(the specific pinned-before-this-fix hazard) and (b) equals the new string
exactly, covering the invariant a future edit must not re-break: no
unsupported `--convention` value named in remediation text users are
expected to run verbatim.

Verified end-to-end (not just unit-level) via /tmp/pr612rev/rv7f.cjs: the
W021 issue's `fix` field now reads the new string; `gsd-tools roadmap
upgrade --convention bracket` (and its --dry-run variant) still correctly
hard-error — that command remains unsupported, only the hint text changed.

Targeted suite (adr-612-*, roadmap-parser, state, verify): 970 -> 971 (+1),
0 fail. `lint-phase-id-drift.cjs` clean; `npx eslint` on touched files:
0 errors (1 pre-existing unrelated no-elapsed-assertion warning, same file,
same line this arc's prior rounds already disclosed).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): correct the bare-0 changeset disclosure + disclose W021's second sub-check (round-7 Minor 1 + Nit 1)

FIX 1 (Minor 1) — round-6's bare-0 changeset clause (1395bd89) was wrong in
three falsifiable ways:

  (a) Its own illustration (`### Phase 0:`) is exactly the LEGACY spelling
      04b5a95f made COUNTED. The residual exclusion after that fix applies
      only to the BRACKET-spelled token (`### [GSD.02] 0:`) — the clause
      named the wrong shape as its example.
  (b) "excluded from these counts" over-scoped the carve-out. The phase-0
      directory is admitted into `completed_phases`/`total_plans` in BOTH
      spellings (measured: B0 at HEAD has `completed_phases=2`,
      `accepts {"GSD.02-0-bootstrap":true}`). The exclusion that survives
      lives in `total_phases`/`phase_count` (heading counting) only.
  (c) The pre-existing "so `roadmap analyze` and `state json` report the
      same number" sentence (present since before this arc's bracket work)
      is now false for the bare-0 LEGACY-spelled shape: 04b5a95f's own
      commit message discloses this exact re-divergence as expected —
      `roadmap analyze` phase_count=2 vs `state json` total_phases=3 on T0
      — matching the disagreement legacy already carries today. The
      changeset still asserted unconditional agreement.

Rewrote both sentences: dropped the `### Phase 0:` example, scoped the
carve-out to `total_phases`/`phase_count`, named the bracket-spelled token
as the one that keeps analyze's sentinel reading, stated plainly that the
legacy-spelled form in a bracket repo IS counted (the mid-migration guard),
and qualified the "report the same number" claim to the `999` token, with
the bare-0 legacy-spelled shape named as the one exception and why.

FIX 3 (Nit 1) — `checkBracketCoherence` has always had two sub-checks
(its own docstring: "Two sub-checks, both surfaced as W021") but both the
changeset and docs/CONFIGURATION.md:188 described only the `mismatch`
sub-check. The `missing-bracket` sub-check — which fires on every
legacy-spelled heading in a bracket repo, the noisier of the two on a
mid-migration project — was undisclosed in both places. One clause added
to each.

Line count: 2 files, 1 line changed each (both are single-paragraph/
single-row files; git diff --stat reports 1/1 per file though several
distinct clauses were edited within that one line each).

Verified: `node scripts/changeset/lint.cjs` clean. Targeted suite
(adr-612-*, roadmap-parser, state, verify) unchanged at 971/971
(prose-only diffs, no test count change, no code touched).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): resolvePhaseIdConvention's docstring no longer claims loadConfig drops the key

Upstream #2997 (aa7697fe, landed in `next` during this PR's final
verification) added `phase_id_convention: get('phase_id_convention') ?? null`
to `_baseConfig` in config-loader.cts — `loadConfig` now surfaces
`phase_id_convention` in its resolved config. This function's docstring
gave "loadConfig merges against CONFIG_DEFAULTS and drops keys it does not
know, and `phase_id_convention` is not among them" as the reason for
reading config.json directly instead of calling `loadConfig(cwd)`. That
rationale is now stale against live `next`.

Corrected the comment to state the two reasons that actually survive #2997:

  1. The workstream->root federation this function performs is a standalone
     resolution run against a GIVEN cwd, not necessarily the same base a
     `loadConfig(cwd)` call elsewhere in the codebase would federate from.
  2. Convention-ENUM validation is still #612 PR-4 work — this function
     returns the raw string unvalidated, exactly as the now-surfaced
     resolved key would.

Noted that #2997 surfacing the key makes consuming it from resolved config
(instead of re-reading config.json here) a natural PR-4 consolidation —
not this PR's scope. The "cycles were never the obstacle" close and every
other paragraph in the docstring (federation rationale, SCOPE note) are
untouched; they still hold.

Comment-only — no code behavior changed. This worktree's own history does
not contain aa7697fe (git merge-base --is-ancestor confirms neither branch
is an ancestor of the other; not rebasing per instruction), so this is a
textual correction against a documented external fact, not a functional
sync with upstream.

git diff --stat:
 src/planning-workspace.cts | 17 ++++++++++++-----
 1 file changed, 12 insertions(+), 5 deletions(-)

Verified: `npm run build:lib` clean, `node scripts/lint-phase-id-drift.cjs`
clean, `npm run lint:ci` exit 0 (same 2 pre-existing unrelated eslint
warnings as every prior round this arc, 0 errors; lint-fix-has-regression-test
PASS). Full suite not re-run per instruction (comment-only diff).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): resolve phase_id_convention against the caller's workstream

`resolvePhaseIdConvention` took no workstream, so it resolved from
`planningDir(cwd)` — which falls back to `GSD_WORKSTREAM` only when its `ws`
argument is `undefined`. Every caller that passes a workstream by ARGUMENT
(they cannot set the env var per iteration) therefore read the convention from
the ROOT config while reading that workstream's ROADMAP.

Two reproduced consequences:

- A workstream that explicitly declares its own `phase_id_convention` had it
  ignored. Flipping ONLY the root config between bracket and milestone-prefixed
  changed which milestone that workstream extracted.
- `--workstream foo` and `GSD_WORKSTREAM=foo` disagreed on the same repo: the
  arg form fell through to the root config, the env form did not.

`resolvePhaseIdConvention(cwd, ws?)` now forwards `ws` to `planningDir`, and
both roadmap-parser call sites pass theirs — `extractCurrentMilestoneScoped`
(which already reads STATE from `planningDir(cwd, ws)`) and
`getMilestonePhaseFilter`'s lazy branch. `undefined` keeps the env fallback, so
convention-less call sites are byte-identical; `null` still means "explicitly no
workstream". The `undefined` vs `null` discriminator on `phaseIdConvention` is
untouched — only the base the `undefined` branch resolves FROM moves.

workstream-inventory's two sites (`countRoadmapPhases`, `inspectWorkstream`)
passed a literal `null` where they meant `undefined`, pinning every workstream
to the legacy grammar. On a bracket workstream whose milestone declares 3
phases that returned phaseCount 0 and fell back to the on-disk directory count;
it now returns 3. A non-bracket workstream is unchanged (legacy control pinned).

The workstream -> root federation is preserved: a workstream that declares no
convention still inherits the root, as config-loader does. That inheritance is
pinned so the fix cannot be over-applied into isolation.

* fix(#2761): classify missing phase details per occurrence, not per token

Under READING-B a phase's sentinel status lives in the BRACKET, so
`[GSD.999] 01` and `[GSD.02] 01` share a token and are not the same phase.
Both sides of the missing-detail check were keyed by the bare token anyway, and
both produced false negatives in `missing_phase_details`:

- The checklist scan built a token -> bracket-id Map, FIRST-WINS. Of two entries
  sharing a token, whichever the author wrote first classified both. With
  `- [ ] **[GSD.999] 01: Icebox**` above `- [ ] **[GSD.02] 01: ...**` the real
  phase inherited the icebox's sentinel verdict and vanished from the report;
  swapping the two bullets — same document, same phases — reported it. Pinned
  with a test asserting BOTH orders.

- The detail set was `new Set(phases.map(p => p.number))`, also token-keyed, so
  `[GSD.02] 01`'s heading marked token `01` present and satisfied
  `[GSD.03] 01`, which has no heading anywhere. Order-independent, same class.

Both sides now key on an occurrence key: the owner's `bracketQualifiedKey`
(fold- and padding-insensitive, so `[gsd.2] 01` and `[GSD.02] 01` are one
phase), falling back to a fold-normalized composite for the two shapes that
grammar refuses — a token carrying its own hyphen, which splices to an id whose
trailing segment the qualified-key grammar truncates (the hazard
`getMilestonePhaseFilter` guards with its own `!token.includes('-')`), and any
id it does not accept. With no bracket id the key IS the bare token, so the
legacy path keeps its exact keys and dedupe order.

The emitted value is unchanged — `missing_phase_details` stays an array of bare
tokens, matching `phases[].number`. Only the classification moved to the
qualified key, so two brackets' `01` both missing report `01` once instead of
one silently covering for the other.

* test(#2761): replace the two wall-clock ReDoS assertions with algorithmic bounds

Both guards asserted elapsed wall-clock time — `Date.now()` against a 20s
ceiling in the coherence suite, `process.hrtime.bigint()` against 1s in the
read-tolerance suite. Those measure the host machine rather than the SUT and
flake on a loaded CI runner (RULESET.TESTS.no-timing-assertion). Per
RULESET.TESTS.delete-bad-tests they are REPLACED, not skipped, and the
behavioral property each one guarded is preserved.

The property is "the widened bracket patterns do not backtrack
catastrophically", which is a claim about growth, so it is now stated by
scaling the input instead of by timing it:

- validate health runs the pathological unclosed bracket at 4,000 and 16,000
  characters and must return the same correct result (no W021) at both.
- The four reader entry points run five ReDoS shapes at 5,000 and 20,000 and
  must name exactly the phases each shape should name at each size, with the
  variant cardinality unchanged across the two.

Catastrophic backtracking is superlinear, so a regression cannot complete the
4x leg under any ceiling, while a bounded matcher is indifferent to the
scaling. The one attack that is well-formed-but-oversized now pins its reading
precisely (`['1'.repeat(n)]`) rather than being lumped in with the malformed
ones. A positive control asserts the readers still extract a well-formed
bracket heading, so "names no phase" cannot pass by the readers being inert.

`{ timeout }` is a hang backstop, not an assertion: it turns a runaway into a
deterministic failure instead of a suite that never returns.

* fix(#2761): give the bracket grammar one owner and teach the drift guard to see it

The bracket milestone-intro grammar was re-typed verbatim in three readers —
roadmap-parser's bracket-fallback selector, state's `isMilestoneBounded` and
verify's `checkBracketCoherence` — which is exactly what #2761's own gate
forbids ("no token literal outside src/phase-id.cts"). `check:phase-id-drift`
reported clean the whole time: its detector only ever knew the phase-NUMBER
token grammar, so the bracket class `[A-Z][A-Z0-9_]*` was invisible to it.

Ownership. `src/phase-id.cts` now exports the class as
`BRACKET_PROJECT_CODE_SRC` and the intro in the two shapes its readers need:
`bracketMilestoneIntroSrcFor(milestone)` (pinned to one milestone) and
`BRACKET_MILESTONE_INTRO_CAPTURING_SRC` (milestone captured). The pinned
builder owns the pad2 spelling rule too — "canonical spelling only, not `0*N`"
was previously restated in prose beside each copy, a convention two files had
to keep agreeing on by hand. `BRACKET_ID_SRC`, `BRACKET_ID_PREFIX_RE` and
`BRACKET_QUALIFIED_KEY_RE` now compose the class rather than re-spelling it.
All three call sites consume the owner; the regex sources are byte-identical to
what they spelled, asserted against hand transcriptions of the pre-fix lines.

Guard. `scripts/lint-phase-id-drift.cjs` gains a bracket rule
(`findBracketGrammarDrift`), wired into `scanRepo` and tagged `kind`. It
deliberately does NOT copy the token rule's `line.includes(CANON_REF)` escape:
that escape is line-level, and verify's copy referenced the owner for the
MILESTONE field on the same line as the re-typed PROJECT-CODE class — so a
bracket rule with that escape would have kept passing on the very site under
review. Partial ownership is the drift; only a dedicated `// phase-id-owner:`
comment suppresses it.

Proof, end to end: planting the shipped verify.cts literal back into src/ makes
`npm run check:phase-id-drift` exit 1 naming `[bracket] src/verify.cts:1475`;
restoring it returns the gate to ok.

Tests. phase-id-drift-guard carries all three shipped literals as negative
fixtures, the same-line-owner-reference case, the case-widened evasion variant,
the sanction rules, and a temp-tree scan proving the rule is wired into
scanRepo rather than merely exported. continuation-grammar-parity drives the
pinned and capturing shapes over a 12-entry corpus and requires the same
verdict from both plus the same captured milestone — widening either alone
fails there.

* test(#2761): drop the out-of-scope source-grep exemption from the selector pin

The baseline-selector pin in tests/adr-612-bracket-heading-selection.test.cjs
read `src/*.cts` with readFileSync and regex, claiming the no-source-grep
escape with a source-text-is-the-product reason. CONTEXT.md's documented scope
(RULESET.TESTS.no-source-grep.exemption) reserves that escape for tests whose
subject is a runtime CONTRACT FILE — STATE.md, config.toml, hooks.json, agent
.md — and `src/*.cts` is none of those.

It was also broader than it looked: eslint-rules/no-source-grep.cjs matches the
marker in ANY comment in the file, so one block's claim disarmed the rule for
the whole ~700-line suite.

The pin itself is worth keeping — the BASELINE ARGUMENT at each call site is a
fact no behavioural test can recover (flipping verify's milestone-complete site
from LABEL_ONLY to ANY_BRACKET grants a tolerance it has never had, and every
behavioural test still passes), so pinning it does require reading the authored
source. That reading moved to `scripts/lint-phase-id-drift.cjs` — the seam's
own guard, where source scanning is sanctioned (`warn` scope) and already
happens for the grammar rules — as `countSelectorBaselines` /
`scanSelectorBaselines`, which return a structured census. The test asserts on
the returned data and touches no file text.

CONTEXT.md is unchanged: the exemption scope was not widened to fit the test.
No marker remains in the suite, so the rule is live across all of it again, and
eslint passes with the escape removed rather than relocated. A floor assertion
pins that the census actually found the five consumers, so the "no other file
consumes the selector unpinned" check cannot pass on an empty scan.

* docs(#2761): rewrite the changeset as a lead + deltas instead of one paragraph

The fragment was a single ~6,400-character paragraph that opened on internal
mechanics, buried the user-visible change, and named an internal test path
(tests/adr-612-bracket-phase-counting.test.cjs) that means nothing to a reader
of the CHANGELOG.

It now leads with what a user sees — bracket-style phase IDs are recognized on
the read path across roadmap, validate, verify and state — followed by compact
bullets for the behavioral deltas, the opt-in caveat, and the upstream
consequences. Both merge-added disclosures are kept: the #3185 legacy-sentinel
Phase-0 delta with the reason the bracket counter keeps the narrower rule, and
the enumerator's convention-argument default flip with the four newly-scoped
read surfaces and the note that archival and milestone-completion paths are
unchanged. Two deltas from this review round are stated as their own bullets:
per-occurrence classification in `missing_phase_details`, and workstream-scoped
convention resolution including the workstream-rollup count change.

The internal test path is gone and the body is down to ~3,850 characters. The
bullets render as sub-bullets under the CHANGELOG entry and the `(#NNNN)`
suffix still lands on a trailing paragraph rather than mid-list.
`npm run lint:changeset` passes.

* test(#2761): use helpers.cleanup in the drift-guard temp-tree test

`local/no-raw-rmsync-in-tests` rejects a bare `fs.rmSync` in tests — the shared
helper carries the Windows-EBUSY retry budget. The temp-tree scan added with
M3's guard coverage test used the raw call.

* docs(#2761): correct the occurrenceKey comment on qualified-key normalization

The comment claimed `bracketQualifiedKey` is "fold- and padding-insensitive, so
`[gsd.2] 01` and `[GSD.02] 01` are one phase". The fold half is right; the
padding half is not. `BRACKET_MILESTONE_NUMERIC_SRC` is `(?:[1-9]\d{2,}|\d{2})`,
so `bracketQualifiedKey('GSD.2-01', 'bracket')` returns null and that input
takes the composite fallback instead.

No behavior change — the two key spaces use different separators and cannot
collide — but the claim was checkable and wrong. Restated: the owner case-folds,
and padding-tolerance is not a property it has or needs, because each accepted
milestone has exactly one canonical spelling and `[GSD.2]` is malformed rather
than an alternate spelling of `[GSD.02]`.

* ci(#2761): give the coverage-gate job the heap floor the shard jobs already have

The gate's report steps parse ~2.5GB of merged raw V8 dumps in one
process at the runner's implicit ~4GB ceiling, so pass/fail comes down
to GC timing (run 31338081337 OOM'd; next passes the same volume at
28s). Deterministic locally: crashes at --max-old-space-size=4096,
completes in 11s at 8192. PR shard data is within 0.15% of green next
runs — the load is pre-existing, only the ceiling was missing. #2952
set 6144 on the shard-collection step; the gate job was missed.

* Revert "ci(#2761): give the coverage-gate job the heap floor the shard jobs already have"

This reverts commit 03c74a97c911e9ddc53675aded9c9bbb89f14cd1.

* fix(#2761): thread the resolved convention into cmdStateValidate's phase-directory lookup

#3208 replaced cmdStateValidate's `startsWith` prefix test with the canonical
`phaseKeyFromDir(...) === selectedPhaseKey` comparison. That is the right
surface, and it is why the lookup now needs the resolved `phase_id_convention`,
which the rewrite does not pass.

`phaseKeyFromDir` deliberately refuses to read a bracket directory without an
explicit signal (ADR-2121: a bracket dir is string-indistinguishable from the
legacy letter-prefixed-decimal family), so un-threaded it returns the whole dir
name as the key — `GSD.02-05-real-work` -> `GSD.02-5-REAL-WORK` — while the
STATE side is the bare `05` that `parsePhaseFromProse` yields. Both sides of one
comparison derived under different conventions is #2562's defect class, and this
file's other three `phaseKeyFromDir` call sites already thread against it.

Observable: a bracket repo whose phase directory plainly exists reported
`valid: false` and "no phase directory matches phase 05", and the drift scan
(plan-count mismatch, verification status) never ran. The pre-#3208 `startsWith`
missed the same directory but skipped silently, so this is a visible-failure
regression on bracket repos, not a new miss.

Non-bracket conventions are byte-identical by construction: `extractPhaseToken`
branches only on `=== 'bracket'`, so null / 'milestone-prefixed' / unresolvable
compile the same path as the un-threaded call. The flat-legacy twin assertion
pins that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2761): note the cmdStateValidate convention threading in the changeset fragment

The fragment described the bracket read path as of the pre-merge branch. 504c64ff
added a shipped behaviour change — `state validate` now resolves bracket phase
directories — that the fragment did not mention, so the rendered changelog would
have under-described what ships.

Body text only; `type` and `pr` are untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): invert the inherited-wart characterization — #3225 fixed it upstream

Required by the merge of next @ 86101ee6; test-only, no src delta.

This branch disclosed rather than fixed a pre-existing upstream wart: on a
LEGACY repo, `### Phase 999:` warned from `validate consistency` while
`validate health` suppressed it — the two verbs contradicting each other.
The case was pinned as a characterization test ("INHERITED WART, unchanged")
precisely so it would INVERT if upstream ever closed it, rather than rot
silently.

ae7dc529 (#3225) closed it, by adding the `isSentinelPhaseId` guard to this
very loop. So the assertion inverted on the merge — as designed. Flipped to
assert the FIXED behaviour rather than deleted: it is the negative-space
proof that this branch's `sentinelPhases` guard never had to grow a legacy
reading of its own, and it reds if a future conflict resolution keeps our
guard while dropping upstream's.

Added a scope control alongside it (a legacy NON-sentinel `### Phase 09:`
with no directory still warns), so deleting the loop outright cannot pass.

Union proven load-bearing in BOTH directions against the merged tree —
neither guard subsumes the other:
  - drop `isSentinelPhaseId(p)` (ours only) -> 1 red, exactly this case.
    Note upstream's own #3225 tests stay GREEN there: they cover the
    disk-side loops (sentinel dir on disk) and the gap-numbering filter,
    not the ROADMAP-side loop. This case is that site's only coverage,
    which is the second reason to keep it rather than delete it.
  - drop `sentinelPhases.has(p)` (upstream only) -> 4 red bracket suites
    (icebox-not-missing, health/consistency agreement, occurrence-aware
    suppression, checklist-index suppression).

tests/adr-612-bracket-read-tolerance.test.cjs 79 -> 80, all green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): pin the three re-homed W006/W007 reads the merge left unfalsifiable

#3309/#3310 deleted the helpers this PR threaded (collectDiskPhases,
collectDiskPhaseEntries, collectArchivedPhaseDirNames,
forEachArchivedPhaseToken) and rebuilt their reads inside
src/planning-snapshot.cts + src/health-diagnostic-rules/*.cts. The convention
threading moved with them in the merge commit — but a mutation sweep over the
re-homed sites found three where reverting the convention argument changed real
CLI output and NOT ONE existing test went red. The pre-migration sites were
covered indirectly, through helpers that no longer exist, so the coverage did
not survive the relocation even though the behaviour did.

Each case is pinned at the CLI with its flat-legacy twin as the byte-identity
control, plus a non-vacuity control in the opposite direction:

  archivedPhaseTokens      revert -> "Phase 05 in ROADMAP.md but no directory on
  (planning-snapshot.cts)            disk" for a phase whose only directory is
                                     under .planning/milestones/v1.0-phases/
  roadmapPhaseCheckboxes   revert -> "Phase 09 in ROADMAP.md but no directory on
  (planning-snapshot.cts)            disk" for an unstarted `- [ ]` phase
  W007's extractPhaseToken revert -> "Phase GSD.02-77-orphan exists on disk"
  (roadmap-disk-consistency)         instead of "Phase 77 exists on disk"

Red/green: each single-line revert reds exactly its own case and nothing else;
every legacy control stays green in all three runs.

NOT pinned, and disclosed as droppable rather than given an unfalsifiable test:
buildValidPhaseSet's extractPhaseToken(dir, convention) in W002. That rule
unions disk tokens with roadmapDeclaredPhases and archivedPhaseTokens, and any
STATE.md reference a bracket disk token would rescue is already rescued by the
ROADMAP half — probed directly, the argument makes no observable difference. It
is threaded because it restores collectDiskPhases(planBase, convention)'s
derivation exactly, not because a test needs it.

Test-only; revertable independently of the merge commit.

* test(#2761): pin the two reads the #3165 extraction left unfalsifiable

DROPPABLE, offered as such. Test-only; the merge commit is correct without it.

#3428's extraction gave `collectAnalyzePhases` TWO call sites — the scoped
milestone window and the truncated-window recovery path. The merge threads
`convention` into both and swaps `detailKeys` with `phases` across the recovery.
Neither of those was falsifiable by the suite as it stood:

  - passing a NULL convention at the FALLBACK site, with the scoped site still
    threaded, is green across all 518 tests of the bracket, roadmap and
    milestone-window files. Every pre-existing bracket assertion reaches the
    enrichment through the scoped window, so none of them can observe the
    fallback's reading at all;
  - dropping `detailKeys = fallbackScan.detailKeys;` — the line this merge
    authored — is likewise green, because no fixture that reaches the recovery
    path carries a checklist bullet, so `missing_phase_details` is `null`
    either way.

That is the same hole class lap 3 closed in 9914359c: an argument whose revert
changes real CLI output with zero reds.

Pinned at the CLI on the shape that reaches the fallback on a bracket repo — a
MID-MIGRATION ROADMAP: bracketed ACTIVE milestone, its phase-detail sections
and checklist bullets sitting after an intervening CLOSED legacy milestone (so
the window closes over prose only), plus one legacy `### Phase N:` section of
its own. Six tests:

  - a NON-VACUITY control that removes the phase directories, so #3428's own
    precondition fails and the result is the empty one the recovery exists to
    replace — this is what proves the rest read the FALLBACK's output rather
    than the scoped scan's;
  - the directory assertion (`disk_status`/`plan_count`/`summary_count` for
    canonical `{CODE}.{MM}-{PP}-slug` dirs) + a flat-legacy twin asserting
    byte-equal shape, and that the twin is the right answer rather than a
    shared wrong one;
  - `missing_phase_details: null` for phases the recovery just found, and its
    companion direction — a bullet with no heading anywhere is STILL reported,
    so a fix that merely suppressed the report does not pass;
  - a non-bracket repo taking the same path unaffected.

Red/green, each mutation reverted afterwards:
  - fallback call site -> `null` convention: exactly 2 reds, both directory
    assertions in this block; 307 other tests green, including upstream's own
    #3428 tests in tests/milestone-window-single-owner.test.cjs.
  - `detailKeys` swap deleted: exactly 2 reds, both `missing_phase_details`
    assertions; 414 other tests green.
  - `matchPhaseDirs` 3rd argument dropped (the lap-2 site): 4 reds — the 2
    already on record plus this block's 2, which is the point: one seam, now
    reached by two paths.

DISCLOSED IN THE TEST BODY WITH MEASURED OUTPUT, not fixed: `hasPhaseEntries`
(src/roadmap-parser.cts) is convention-blind, so on a PURE-bracket ROADMAP the
window classifies COMPLETE and #3428's recovery is gated off entirely. That
document returns `{"scope":"complete","phase_count":0,"next_phase":null,
"phases":[]}` — the "genuinely empty milestone" answer, the indistinguishability
#3184 introduced `scope` to remove — where the flat-legacy twin returns
`{"scope":"truncated","phase_count":2,...}`. Widening it changes the value of an
upstream-owned output field on bracket repos, so it is a behaviour slice (the
PR-2.5/PR-4 convention-less-readers question), not a merge-round change. The
mid-migration shape pinned here is the reachable half.

* fix(#2761): mirror the bracket terminator in the milestone-scope write guard

parser's terminator vocabulary — "a level 1-3 heading that is not a Phase
heading and carries a milestone signal". On this branch that vocabulary is
convention-SELECTED: `computeBracketSectionEnd` adds `isBracketMilestoneBoundary`
as a terminator arm, and the ADR-canonical `## [GSD.09] Hidden` carries NO
vN.N token, NO ✅/📋/🚧/🔄 marker, and not the word "Milestone". Left
unmirrored, the guard accepted exactly the description it exists to reject.

NOT a defect on clean next — a bracket heading terminates nothing there. The
branch widens the terminator set, so the branch owns the mirror.

MEASURED at the CLI seam before the fix (bracket fixture, one milestone,
one phase):

  $ gsd-tools phase add $'Sneaky\n## [GSD.09] Hidden'   -> exit 0, written
  $ gsd-tools phase add 'Innocent follow up'            -> exit 0, written
  $ gsd-tools roadmap milestone-scope
    { "scope": "complete", "phases": ["01","1"], "phase_count": 2 }
  $ grep '^### Phase' .planning/ROADMAP.md
    ### Phase 1: Sneaky
    ### Phase 2: Innocent follow up      <- in the document, out of the window

After the fix the first add exits 1, ROADMAP.md is byte-unchanged and no
phase directory is created. The legacy twin (same text, no
`phase_id_convention`) still exits 0 — opt-in only, base behaviour preserved.

SHAPE
- `findMilestoneScopeHeadingLines(text, convention)` — REQUIRED, the same
  tripwire `scanMilestonePhaseIds` carries in the merge commit and for the
  sharper reason: a blind call here fails OPEN (the guard quietly ACCEPTS a
  window-narrowing description), so a future call site must fail to COMPILE.
  Census is one caller, which pays nothing for it. A non-bracket value takes
  the pre-existing path byte-identically. The bracket arm routes through
  `isBracketMilestoneBoundary`, the same single-owner phase-vs-milestone
  discriminator `computeBracketSectionEnd` consults; no second bracket-heading
  grammar is spelled here.
- `selectedBracketId` is deliberately `null`, so the same-milestone
  CONTINUATION exemption never fires and a value naming the ACTIVE milestone
  is flagged too. That is this predicate's third stated conservatism and
  rests on its own existing argument: which milestone is active is a property
  of the document at write time, not of the text being validated, and
  over-rejecting is one-directional.
- `assertDescriptionPreservesMilestoneScope` takes `cwd` and resolves through
  the branch's tolerant try/catch — this guard runs BEFORE `loadConfig` and
  before the ROADMAP existence check, so an unresolvable convention must
  degrade to the legacy vocabulary, never turn a rejection into a crash.
- The error's marker list gains the bracket form only when the project is on
  the bracket convention.

RED/GREEN
- Revert the bracket arm -> 2 reds (phase add; insert + add-batch), the
  non-bracket control and both no-false-positive cases stay green.
- Revert the probe threading in the merge commit -> 1 red (the probe case).
- 6 new cases in the #3262 file, each with a non-bracket or
  no-false-positive control: fenced bracket milestone heading and bracket
  PHASE heading are both non-violations.

Droppable: revert this commit and the merge stands on its own; the branch
then ships the gap as a disclosure instead of a fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changeset): narrow the enumerator claim to the directory set — per-entry rendering in progress/stats/init-manager is display-slice scope

Round-7 review confirmed 3 of the 4 surfaces named by the closing claim
parse each directory or heading with legacy-only patterns that live in
files this PR does not touch (commands.cts:1766, commands.cts:2116,
init.cts:2241). The claim now states exactly what this slice delivers:
the scoped directory set. Their per-entry conversion is display work,
deferred to the epic's display PR with the statusline/progress-card
gates it belongs beside.

* fix(#2761): de-accident the unmatched-milestone bracket fixture, re-pin to #3480's withhold contract

tests/adr-612-bracket-phase-counting.test.cjs:515 ("a milestone that
does NOT match STATE is not scoped in") asserted total_phases === 0.
Since today's merge brought in 70b5c1a1 (#3354/#3480, already in
`next`), buildStateFrontmatter withholds total_phases (omits the key,
read back as null) for a milestone that is genuinely sectioned but
matches no ROADMAP heading, instead of substituting a computed number
— the fixture's `state json` read now returns null, failing the
strictEqual(0) assertion.

The original single-section fixture's 0 only ever survived by
accident: hasMilestoneSectioning requires >=2 milestone-vocabulary
headings to call a ROADMAP sectioned, and its isPhaseHeading helper
recognizes only the legacy `Phase N:` text form — so the bracket
phase heading `### [GSD.03] 09: Not this milestone` (title containing
the word "milestone") miscounted as a second milestone heading,
tipping hasMilestoneSectioning true and routing to the disk-count
branch. With that miscount removed, the same one-section fixture
reads 1 (the non-matching milestone's phase count leaking into v2.0's
total) — proving the pinned 0 was never validating the scoping this
test claims to exercise.

Replaced the fixture with two genuine milestone sections (GSD.03,
GSD.04), neither matching STATE's v2.0, with phase titles carrying no
incidental vocabulary — the real #3354 shape. Re-pinned the assertion
to null, matching upstream's own tested contract (tests/state-
document.test.cjs, "#3354 with nothing stored, the key is omitted
rather than written from the dir count").

Verified: file green 3/3 runs (104/104), full state-suite regression
771/771 green. The isPhaseHeading gap that made the old fixture
accidental is a real, pre-existing, upstream-owned limitation
(hasMilestoneSectioning only guards >=2-section conflation, not a
single non-matching section leaking through) — not introduced by

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): exempt two self-authored-fixture regexes from #3441's unbounded-quantifier rule

Both sites parse the STATE.md the test itself just wrote to its own
tmpdir — fixed-size fixture output, not adversarial or document-scale
input. Exempted with the justification-comment pattern the tree's own
tests use for this exact case (settings-integrations, copilot-install,
research-agent-profiles). The sites predate the rule; #3441 landed on
next this morning and this branch picked it up in the catch-up merge.

* fix(#2761): fix forward two next-movement test regressions from the round-11 rebase

origin/next's #3573 newly threads STATE's stored `milestone:` value into
`state json`'s buildStateFrontmatter call (previously always `undefined` on
that read surface). Two pre-existing PR-2 fixtures reach code paths that
call never exercised before that merge:

- RED (repro2 case C): a version-less, all-bracket-id ROADMAP now hits
  getMilestonePhaseFilter's pre-existing (unaffected by this branch) row-5
  `versionResolved && !headingFound => SCOPE.UNSCOPED` rule, which withholds
  `progress.percent` even though the scoped total_phases/completed_phases
  are still correct. That row-5 rule is load-bearing for six other
  version-less-document pins in this same file; narrowing it broke seven of
  them in testing, so production code is untouched here. Reassert the
  test's real claim (scoping, via total_phases/completed_phases) and
  disclose the now-withheld percent instead of silently dropping it.

- PIN (repro12 LEGACY control): a fenced-example-vs-real version heading
  fixture. #3573 routes this read through sliceMilestoneWindow (fence-aware)
  instead of the legacy anyMilestonePattern raw scan (fence-blind) this pin
  was disclosing as out of scope, so the fixture no longer exercises the
  blind path — total_phases moves from the whole-doc 4 to the correctly
  scoped 2. Re-pinned with an upstream-attributed comment, matching this
  file's existing house style for prior origin/next movements (see the
  sibling "PIN (G3 LEGACY control)" comment).

Both are test-only fix-forwards: no production code changed. The remaining
12 failures in this file at 27363e9d0 (dropped seam commits' VERIFICATION.md
naming + fixture updates) are pre-existing and deferred to the stacked
follow-up per the task's own scope.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): thread phaseIdConvention explicitly at the two write-adjacent enumerator call sites (round-11 BLOCKER)

listMilestonePhaseDirs' phaseIdConvention param lost its `= null` default
(phase-locator.cts), so an omitted convention now means "resolve from
config" instead of "explicitly not bracket" — a deliberate flip, but
milestone.cts and state.cts were not in the PR-2 diff and both call the
enumerator without threading it:

- milestone.cts cmdMilestoneComplete (the single #3597 shared derivation
  feeding the stats loop, --dry-run preview, and the real archive/rename
  pass) now resolves phase_id_convention once and threads it explicitly,
  so a bracket project's `milestone complete` archives its real
  bracket-declared phase directories instead of silently inheriting
  whatever the enumerator's lazy default resolves to.
- state.cts cmdStateUpdateProgress's own enumerator call threads the same
  resolved convention. Empirically this does not change the #3217 withhold
  gate (scope is assigned before headingConvention resolves in
  getMilestonePhaseFilter, so it's convention-independent either way) or
  the reported percent (already correctly threaded via
  computeUpdateProgressPreview -> buildStateFrontmatter); it closes a
  second, silently-resolved answer to the same question this file's own
  ONCE-and-THREAD rule (~:2300) already states as policy.
- state.cts's other listMilestonePhaseDirs call site (the state-sync
  scope-only read, ~:4791) is left unthreaded on purpose, with an inline
  note explaining why: only `.scope` is consumed, and `.scope` is set
  before convention resolution in getMilestonePhaseFilter, so it cannot
  disagree with a threaded convention.

Both enumerated sets are pinned by new tests (round-11 BLOCKER block in
tests/adr-612-bracket-phase-counting.test.cjs), including a mutation-style
regression check on the milestone.cts fix (forcing the un-threaded default
back on makes the pinned test fail, catching a regression to the
pass-all-degrade legacy reading).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): correct the changeset's false archival claim and amend ADR-612 for the round-11 M2(2) split scope

The changeset (.changeset/2761-bracket-read-tolerance.md) asserted "The
archival and milestone-completion paths are unchanged" — false: both
paths reach the widened enumerator, and this PR now threads their
convention explicitly (previous commit). Replaced the closing paragraph
with an accurate description of what changes for a bracket project at
those two call sites, and notes both enumerated sets are pinned by tests.

ADR-612 (docs/adr/612-bracket-phase-id-convention.md) amended per M2(2):
the 2026-08-03 PR-2/PR-4 boundary proposal is added in-body (PROPOSED,
not stamped — mechanics per docs/contributor-standards.md:143 reserve ADR
ratification to maintainers), rescoped to what actually ships in PR-2
(#2761) now that the round-11 M1 split moved the completion-seam
threading (isPhaseArtifact/scopeToPhase, phase-id.cts:964-1090) out into
a separate, stacked follow-up PR referenced generically via epic #612:

- state.cts read-tolerance (both total_phases derivations, the #1514
  retirement filter) moves into PR-2's row, alongside milestone.cts,
  reflecting the round-11 fix above.
- the write-observability point is restated in terms of what actually
  ships: explicit threading at two named call sites, pinned by tests,
  rather than silent inheritance.
- the completion-seam threading is named and explicitly excluded from
  PR-2's scope, with a new ratify item (5) and a Negative consequence
  bullet covering the sequencing cost.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): repair round-11 response gaps — disk-side sentinel bug, test-fixture bug, scanMilestonePhaseIds caller drift

Four independently-verified gaps in the round-11 repair response:

1. isSentinelPhaseId (src/phase-id.cts) treated a bare, untagged phase
   directory under phase_id_convention: "bracket" (`0-bootstrap`, no
   `{CODE}.{MM}-` prefix) as sentinel milestone 0 by falling through to the
   legacy leading-int rule when the bracket-tag match failed. This silently
   dropped a real, on-disk, milestone-declared phase directory from
   `listMilestonePhaseDirs` (phase-locator.cts:424, the only unguarded call
   site) and undercounted completed_phases/percent. Mirrors the carve-out
   already present on both heading-side counters (state.cts's
   countRoadmapPhaseHeadings guards its bare-0 exclusion with `bracketId &&`;
   roadmap-parser.cts's scanMilestonePhaseIds composes the bare-token rule as
   999-only) — under bracket convention, milestone 0 is expressed only via an
   explicit bracket tag, so an untagged leading 0 is a real phase token.

2. tests/adr-612-bracket-phase-counting.test.cjs's writeProject fixture
   hardcoded `01-VERIFICATION.md` for every "complete" phase dir regardless
   of the dir's real phase token. Under legacy convention this file already
   correctly failed #3511's strict isPhaseArtifact match for any dir other
   than phase 01 (the fixture never modeled what it claimed to); under
   bracket convention the pre-existing ambiguity fail-safe admitted it
   regardless. The two readings' disagreement was mistaken for a missing
   completion-seam threading (isPhaseArtifact/scopeToPhase convention
   awareness, correctly split out to the stacked #3644 per round-11's M1).
   Replaced the hardcoded name with verificationNameFor(dir), deriving the
   real per-phase filename from the production normalizePhaseName the same
   way cmdScaffold does — the 7 failing "flat-legacy-twin" / CHARACTERIZATION
   assertions pass on #2867's own code with no seam threading required, and
   the genuine seam-dependent cross-phase-stray-exclusion test the fixture
   fix would otherwise have hidden lives on the stacked branch instead.

3. tests/roadmap-parser.test.cjs's two #3577 tests still called
   scanMilestonePhaseIds with the old single-Set return shape; this PR
   changed it to `{ ids, qualifiedIds }` for every other caller but missed
   these two, which don't touch #2761/bracket code at all. TypeErrors at
   runtime, not silently-wrong assertions. Updated both call sites.

4. .changeset/2761-bracket-read-tolerance.md gains a paragraph disclosing
   fix 1 above, so the changeset stays accurate to what actually ships (the
   round-11 BLOCKER was exactly this changeset going stale once).

Verified: tests/adr-612-bracket-phase-counting.test.cjs +
tests/continuation-grammar-parity.test.cjs + tests/roadmap-parser.test.cjs =
354/354. adr-612-{coherence,grammar,heading-selection,read-tolerance,
selection.property}.test.cjs + collision-characterization = 287/287
(CONFIRMED-CLEAN set, unaffected). Full unfiltered suite run separately.

PR #3643 (round-11's M1 split, opened draft per that round's explicit
requirement) was auto-closed by this repo's own draft-PR policy 11 seconds
after opening — draft PRs are unconditionally closed here. Reopened as
non-draft #3644 (same branch, same commits, DO NOT MERGE / stacked-on-#2867
marker in the body, enhancement template) since the repo's own bot confirms
non-draft contributor PRs are tolerated even off-template.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): narrow the state-update-progress pin claim to what mutation testing actually proved, add the test that covers the rest

Round-11 BLOCKER response gap (verified, not a guess): the changeset and
ADR-612 amendment point 2 both claimed the round-11 tests pin
`cmdStateUpdateProgress` so "a future change to the enumerator's default
cannot silently move ... what state update-progress reports without failing
a test." Mutation testing disproves this. Reverting BOTH the state.cts
explicit `phaseIdConvention` thread AND simulating the phase-locator.cts
pre-#612 hardcoded-null default (`phaseIdConvention: null` at that one call
site) leaves the existing PIN test ("writes a real percent for a
bracket-scoped milestone") green, because that percent comes from
`computeUpdateProgressPreview` -> `buildStateFrontmatter`, a separately and
already-correctly-threaded derivation the state.cts inline comment at ~:965
already candidly documents.

What the reverted thread DOES change, empirically, is `phaseDirs`/`totalPlans`
— the enumerated `.value` this call site feeds into the #3233 zero-plans
no-op gate a few lines below. That is the one place a regression at this call
site is observable in the command's output. Added a test that pins exactly
that: a bracket milestone whose declared phases carry no plans on disk,
alongside a decoy directory that plainly does not belong to the milestone
(no bracket tag, no phase token) but does have a plan. Correctly scoped, the
decoy is excluded and the #3233 no-op fires (`updated: false`). Degraded to
the pass-all legacy reading, the decoy is swept in, `totalPlans` flips
nonzero, and the no-op never fires (`updated: true`).

Mutation-tested against both scenarios:
- Reverting ONLY the state.cts explicit thread (falls back to `undefined`,
  which `getMilestonePhaseFilter` resolves via the identical
  `resolvePhaseIdConvention(cwd, undefined)` call the explicit thread also
  makes): new test stays green — confirms the single-hunk thread really is
  pure single-derivation hygiene, exactly as the existing inline comment
  claims, for this test too.
- Forcing `phaseIdConvention: null` at that call site (the combined-revert
  scenario the round-11 mutation testing actually exercised): new test FAILS
  (`updated: true, percent: 0` instead of the expected `updated: false`).
  Restoring the real code makes it pass again.

Narrowed the changeset and ADR-612 point 2 prose to match: `cmdMilestoneComplete`'s
enumerated set is pinned against a pass-all-degrade regression (genuine,
already verified in round 11); `cmdStateUpdateProgress`'s own call site is now
pinned against the #3233 zero-plans no-op specifically, not against the
reported percent, which the prose no longer claims. Added a one-line pointer
to the new test in the state.cts inline comment. No production behavior
changes — test and documentation only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(#2761): reconcile ADR PR-2 module map

* fix(#2761): restore #3639's dir-aware W007 sentinel exclusion lost in the rebase replay

The rebase replayed this file's pre-#3639 patch over next, reverting the
isSentinelPhaseId(token) -> isSentinelPhaseDir(dirName) fix: the extracted
token is milestone-stripped, so a bracket sentinel (GSD.999-07-icebox,
GSD.00-01-backlog) was invisible to the id predicate and false-fired W007.
Restores upstream's call and comment verbatim; the branch's convention-aware
token remains for the diagnostic message only.

Caught by next's own #3639 regression tests (CI shard 2). Local: the file's
18/18, health-diagnostic-rules 165/165, full npm test exit 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MeAsZNfhiUQioFPA4ygEGS

* docs(#2761): ratify scope and correct surface claims

* test(#2761): strengthen legacy selection and drift claims

* docs(#2761): correct the enumerator claim, scope list, and ratification receipt

Three documentation corrections, none touching production code.

The changeset claimed the shared phase-directory enumerator "now defaults its
convention argument to 'not yet resolved' rather than 'resolved, and not
bracket'". That is false: `phase-locator.cts:375` still destructures
`phaseIdConvention = null`, and the lazy resolve-from-config fires only on
`undefined` (`roadmap-parser.cts:941`, `:1928` — whose own comment records that
"explicit null still means 'resolved and non-bracket'"). Measured: only 4 of 17
`listMilestonePhaseDirs` call sites thread a resolved convention (`milestone
complete` and `state`'s three). The changeset went on to name `progress`,
`stats`, `phase list` and the init manager view as now receiving a correctly
scoped set — those are precisely callers that omit it. Replaced with what the
code does, and the deferral stated plainly.

The ADR's PR-2 row omitted `scripts/lint-phase-id-drift.cjs` (+141/-14) and
`scripts/lint-phase-enumeration-drift.cjs`, both changed by this PR. An
under-claim rather than an over-claim, but the row is the epic's
scope-of-record.

Ratify item 5 asserted maintainer ratification while citing only the review that
raised the question. It now cites the review that granted it (#2867 review
`5012940978`, 2026-08-24) and quotes its terms, so the claim carries its receipt.

Found by an adversarial review pass over the round-12 diff. (#2761)

* fix(#2761): drop the no-op --json that strict argv now rejects

Surfaced by the rebase onto next, not by a change in this PR's subject.

#3884 ("failure is a value — strict argv", e20744eac) made an unrecognized
flag a hard error: `state validate --json` now exits 1 with
`unknown flag "--json"; accepted: --strict` on stderr and EMPTY stdout,
where the token was previously accepted and ignored. `state validate` never
had a `--json` flag — JSON is its only output shape — so the argument was a
silent no-op from the start.

The helper parsed that empty stdout, so all five subtests in the
"state validate resolves bracket phase DIRECTORIES" suite failed identically
with `SyntaxError: Unexpected end of JSON input` at the JSON.parse, masking
what they actually assert.

Dropping the token restores the same envelope the helper already parsed. No
assertion changes. Verified against the fixture the suite builds: exit 0, and
the output carries the S005 plan-count warning the drift assertions require
with no S004 phase-directory warning — i.e. the bracket directory resolves,
which is the behaviour these tests exist to pin. 94/94 in the file.

* fix(#2761): exempt the phase-counting suite from the docs-guard lane

Surfaced by the rebase onto next. #3753 (107eb8c1d) added
lint-docs-guard-registration, which requires every test file that reads a
docs/ path to be either registered in the docs-guard lane or carry an
explicit marker. It flags adr-612-bracket-phase-counting.test.cjs, which
reads no docs/ path at all.

The file's only docs/ occurrence is the ADR-612 Decision 1 citation in a line
comment; every read call it makes targets a tmpdir .planning fixture. It trips
Detector 3, whose DOCS_TEMPLATE_LITERAL_RE sees an odd prose backtick in the
comment block above that citation as opening a template literal and reads the
span between them — citation included — as a docs/ path expression. That
detector documents this trade in its own header: it favours recall, and says a
false positive costs one docs-guard-exempt marker with a reason. The baseline
records the same class ("comment-only mentions") for 46 of its entries.

Registration was the wrong side of the trade here: it would run this suite in
the docs-guard lane on docs/ changes whose content it never reads.

Three pieces, matching what the gate requires and what its 54 existing entries
already do:
  - the marker in the file's header window, written without backticks so it
    cannot itself disturb the parity tracking findExemption does;
  - the basename in DOCS_GUARD_EXEMPT_BASELINE, since the ratchet fails a new
    marker until the baseline is deliberately updated — the reviewable diff is
    the point;
  - the FIX 3 per-file fingerprint, derived with the lint's own
    extractDocsPathReferences rather than retyped, so the exemption fails loudly
    if the set of docs/ paths this file mentions ever changes.

Gates: lint-docs-guard-registration 0 violations, ci-docs-guard-registry 51/51.

* docs(#2761): correct the bare-0 comment and state what opting into bracket costs

Round-13 review items Minor 2 and Minor 3, both still live on the previous head.

Minor 2 — src/phase-id.cts. The comment block claimed "Bare `0` is admitted
alongside it because a 0.x sentinel is a legitimate identity that predates
padding", which contradicted both the shipped constant and its own next
paragraph. BRACKET_CANONICAL_NUMERIC_SOURCE is `(?:[1-9]\d{2,}|\d{2})`;
measured against it, `0` is rejected while `00`, `05`, `99`, `100` and `999`
are admitted. The paragraph four lines below already recorded the removal
("the earlier `(?:\d{2,}|0)` ... admitted ... a bare `0` that pad2 never
produces"), so the block asserted and denied the same fact. The stale sentence
is replaced with what ships: `00` is the backlog sentinel's canonical padded
identity, `\d{2}` already covers it, and nothing needs the unpadded spelling.
No behaviour change — the constant is untouched.

Minor 3 — docs/CONFIGURATION.md. The phase_id_convention row described the
read-path widening but never said what a repo GIVES UP by opting in. Added:
on a bracket repo a heading whose bracket is followed directly by a digit is
read as a phase heading, so `### [RFC.2119] 5:`, `### [v1.0] 2024:` and
`### [ADR.612] 3:` — legal prose headings under any other convention — are
claimed as phases and move phase_count, total_phases and W006. Those three
shapes are the ones phase-id.cts's own selector comment names as the reason
the widened read is selected at construction time from this value rather than
applied everywhere; the row now carries that trade instead of only its
upside.

Gates: tsc --noEmit exit 0, npm run lint:ci exit 0, npm run test:unit
33890/33891 with the single failure being emitted-attribution's base drift
against a next that moved after the rebase (133/133 against the rebase base;
this push re-bases onto the current tip, which resolves it).

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-31 14:03:48 -04:00
0xdhx
ef9ce3e598 fix(#3702): count asterisk, plus and ordered markers as deferred-items entries (#3739)
* fix(#3702): deferred-items counts `*`, `+` and ordered markers as list items

`deferred-items.md` has no template and no mandated shape, but its parser
recognised only the `- ` hyphen marker. Asterisk bullets, plus bullets and
dot-terminated ordered lists — all lists in CommonMark and GFM — contributed
ZERO entries on both the headless and the heading-delimited path, and a mixed
file dropped its non-hyphen entries while keeping their hyphenated siblings,
under-reporting without ever looking empty.

The restriction was a regex literal inherited from the Gaps seam, where the
template genuinely mandates the hyphen YAML-lite form; nothing in the module's
stated rationale distinguishes `*` from `-`.

Widened on the deferred path only:
- `splitGapsEntriesCore`'s entry opener, `extractGapEntryFields`' line-0 strip
  and `rawGapEntryText`'s line-0 strip take a `BulletMarkers` parameter that
  DEFAULTS to the hyphen-only set, so `## Gaps` keeps its template-mandated
  grammar byte-for-byte and the module still has exactly one grouping pass.
- `splitDeferredHeadingEntries`' body-bullet test, `stripLeadingBulletMarker`
  and `acknowledgeDeferredItem`'s status-field regexes move in lockstep —
  widening what OPENS an entry without widening what is STRIPPED before field
  extraction would surface an entry that can never resolve.

Unchanged, and pinned by tests: prose-only and bare headings still contribute
nothing ("prose is not an item"); a table under a leaf heading still yields
exactly its rows, since table lines are skipped before the body-marker flag can
be set and a `|` row is not a list marker; the paren-terminated ordered form
`1)` is out of this fix's scope.

* docs(#3702): changeset fragment (pr: 0 placeholder pre-create)

* fix(#3702): widen the forensic-audit prose entry rule to match the parser

Sibling site of the same defect class, found by a defect-class sweep of the
deferred-items consumers. `/gsd-progress` check 7 does NOT go through
`gsd-tools query` — it globs `deferred-items.md` and has the model read entries
by a prose rule that mandated "one entry per top-level `- ` line". Left as-is,
the marker widening would hold on the CLI path while the one consumer that
bypasses the parser kept reporting "No unresolved deferred items" for a file
written with `*`, `+` or an ordered marker: the same false negative, surviving
in the only place the fix could not reach by code.

Also pass DEFERRED_BULLET_MARKERS explicitly where the heading path extracts
fields. It was already correct — stripLeadingBulletMarker pre-strips the widened
set from every line, so the default hyphen strip is a no-op there — but relying
on that leaves a detection site and a strip site nominally on different marker
sets, which is exactly the asymmetry the BulletMarkers doc comment warns about.
Explicit is local; inferred is a trap for whoever edits the strip next.

Out of scope, noted rather than fixed: forensic-audit.md globs only
`.planning/phases/*/` and so misses archived milestone phases that
`scanDeferredItems` covers. Pre-existing, a different defect, and not this
issue's ruling.

* docs(#3702): note the milestone-close halt for heading-shape non-hyphen files in the changeset

A heading-delimited deferred-items.md written with */+/ordered markers
previously parsed to zero and closed silently; it now yields entries whose
heading shape acknowledgeDeferredItem refuses, halting complete-milestone
until hand-edited. User-visible, so the fragment states it.

* chore(#3702): set changeset fragment pr to 3739

* fix(#3702): CR-normalise the heading path and the acknowledge writer (review B1, M4, m2)

B1 — `splitDeferredHeadingEntries` stored RAW lines; on a CRLF file every
body line but the last still carried its `\r`, the `$`-anchored marker
strip failed on it, the marker survived into field extraction and the
field was lost — a `**Status:** resolved` that was not the file's final
line resurfaced its entry as open. The heading path now stores CR-stripped
lines like the headless path already did, and the strip regex tolerates a
trailing CR on its own. Round 1's CRLF test put `**Status:**` on the last
line, the one position `collectSection`'s `.trimEnd()` had already
de-CR'd; the new tests put it first and mid-body.

M4 (pre-existing on `next`) — `acknowledgeDeferredItem` found the status
line on a CR-stripped copy but rewrote the raw line with a `$`-anchored
`.*`, which cannot consume `\r`; `replace` returned its input, and the
writer reported `ok` over byte-identical content. The rewrite now runs on
a CR-stripped line. The comment that claimed `.*$` consumed the `\r` is
corrected — it was the bug, stated as the design.

m2 — the indent probe for an inserted `status:` line ran on the raw line
and fell back to indent 0 on CRLF; it is CR-stripped too.

* fix(#3702): derive every deferred-items marker regex from one source (review M3, N1, N2)

M3 — round 1 carried the marker alternation in FOUR places: the
`BulletMarkers` pair and two inline literals inside
`acknowledgeDeferredItem`, under a doc comment saying the interface
existed so a detection site and its strip site could not drift. All four
now derive from `DEFERRED_MARKER_ALT`; drift is impossible rather than
discouraged. A parity test pins the vocabulary against
`markdown-sectionizer`'s `iterateBullets` on everything the two grammars
are meant to agree on, and names the two points they deliberately differ.

N1 — the ordered marker is `\d{1,9}\.` (CommonMark §5.2), not `\d+\.`.

N2 — the marker is followed by `[ \t]`, not `\s`, which also accepted
`\r`; the tab remains accepted (CommonMark-legal) and the divergence from
`iterateBullets`' literal space is pinned rather than papered over.

The four regexes are exported for the parity test only.

* fix(#3702): an ordered marker opens an entry only from `1.` or inside a run (review B2, m1)

B2 — `\d+\.` alone read ordinary prose as a list: "2026. was a bad year
for this module" and, under `### Notes`, "3. is the number of retries we
settled on." both opened an entry on round 1, the second straight through
the "prose is not an item" contract that round's AC4 claimed to preserve.
CommonMark §5.3 faces the same ambiguity when an ordered list would
interrupt a paragraph and resolves it by requiring the list to start with
1; `matchListOpener` applies that rule wherever an ordered marker is seen,
with the run carried per list (headless) or per leaf-heading body. Numbers
after the first are ignored, as CommonMark ignores them. Stated cost,
pinned: a hand-numbered list starting at 2 reads as prose — every ordered
record in the #3702 scan starts at 1.

Both reviewer cases are pinned as prose; the ruling's `1. alpha / 2. beta`
shape still counts.

m1 — the 9-digit boundary is pinned at both sides (`999999999.` opens,
ten digits is not a marker), and the 3-vs-4-space indentation cliff is
pinned as deliberately NOT applied: the parser is indent-lenient because
surfacing a questionable hand-written entry beats dropping a real one.

* fix(#3702): thematic breaks close the list and fenced code never opens an entry (review M1, M2)

M1 — `- - -` was a phantom `"- -"` entry on base; widening the marker set
added `* * *` and `+ + +` to the class, and `* * *` is the separator an
author writing in the `*` style is most likely to use. A CommonMark §4.1
thematic break (plus the `+ + +` gesture, which is the same garbage as an
entry name) now closes the open entry on the headless path and is dropped
from the body on the heading path — neither an item nor a continuation.

M2 — neither splitter was fence-aware, so `+ `-prefixed diff lines and
`1.`-numbered repro steps inside a code block counted as entries; #3702's
wild records carry exactly those blocks. Both splitters now classify lines
by the sectionizer's own `scanFencedBlocks` (so `~~~`, indented and
unterminated fences behave as `stripFencedCode` would): fence content
never opens an entry, is continuation inside an open one — keeping the
span invariant `acknowledgeDeferredItem` re-verifies — and is discarded
before the first.

* test(#3702): range the #2287 deferred-items property over marker × shape × line ending (review B3)

The `#2287` property hard-coded `- ` and filtered `\r\n` out of its
arbitraries, so the widened marker set — an enumerated domain, exactly
what a property is for — was never under it. It now ranges over
`{-, *, +, ordered}` × `{headless, heading}` × `{LF, CRLF}`, with the
heading shape placing `**Status:**` first or last: the review's
prescription (markers × line endings) would not have reached B1, which
lives on the heading path only, so the shape axis is the load-bearing
addition. Ordered entries are numbered from 1, so the B2 run rule is
under the property too.

A second property drives `acknowledgeDeferredItem` over every unresolved
headless entry across the same marker × line-ending grid — the one that
reaches M4 (a CRLF rewrite that reported `ok` and wrote nothing) and m2.

* test(#3702): pin the milestone-close halt on a heading-delimited `*`/`+`/`1.` file (review m3)

A heading-delimited `deferred-items.md` written with a non-hyphen marker
previously parsed to zero entries and let `complete-milestone` close
silently; it now yields entries whose heading shape `acknowledgeDeferredItem`
refuses, which the milestone loop turns into `record_ack_failure` → exit 1.
The loop is prose in a workflow, so the test drives the two CLI calls it
makes: `audit-open --json` must list the entry, and
`audit-open acknowledge --text <the audit's own text>` must refuse with the
heading-delimited message and write nothing.

* docs(#3702): changeset and forensic-audit prose carry the round-2 grammar

The changeset names the CRLF fixes, the ordered start-at-1 rule, thematic
breaks and fences. The `/gsd-progress` forensic-audit step is the one
prose parser of this file and must state the same grammar the code has.

* fix(#3702): round-review refinements — run ends at a paragraph, rejected ordinals unstripped, breaks at any indent, fenced fields, `## Gaps` scope

Findings from the pre-push adversarial review of round 2, each pinned:

- An ordered run ENDS at a paragraph that follows a blank line (CommonMark
  §5.3); a non-indented line with no blank before it is lazy continuation
  and keeps the run open. `1. a` / blank / `paragraph` / blank / `5. x` is
  one entry, not two.
- The heading path strips the marker off every body line before field
  extraction (#3457); a line whose ordinal `matchListOpener` REJECTED must
  not be stripped, or "3. status: resolved" as prose loses its `3. ` and
  reads as a resolved field. `splitDeferredHeadingEntriesDetailed` now
  carries a per-line opener flag and only accepted openers are stripped —
  in headless regions of a heading-shaped file too.
- A thematic break is recognised at any indent, matching the parser's
  indent-lenient reading of items; `    * * *` was a phantom `* *`.
- Fenced lines carry no FIELDS either: a `status: resolved` quoted inside a
  code block no longer resolves its entry on either path.
- Block structure (breaks, fences) is a property of the GRAMMAR, carried as
  `BulletMarkers.blockStructure`: the deferred set opts in, the Gaps set
  does not, so `## Gaps` is byte-for-byte on its `next` behaviour — the
  round-2 M1/M2 change had reached it through the shared splitter.

* test(#3702): the property exercises the rejected-ordinal branch; the N2 control is independent

Round review: the widened #2287 property numbered every ordered run from 1
and so never generated an ordinal the start-at-1 rule rejects — it could
not tell round 1 from round 2 on B2. Each entry may now carry a decoy prose
line beginning with a non-1 ordinal, placed where it cannot end a run
(before the first headless entry; first in a heading body), followed by a
`status: resolved` that must never become a field; and a decoy-only
heading body must yield no entry.

The N2 assertion accepted a tab, which round 1's `\s` accepted too, so a
`[ \t]` → `\s` revert alone stayed green. NBSP, form-feed and vertical-tab
are now asserted refused — the assertion that fails on that revert on its
own, and the disclosure that `[ \t]` narrows what round 1 accepted.

* fix(#3702): the splitter records its own opener flags; an opener clears the blank-line memory

Round-review continuation, two state defects in the ordered-run logic:

- `blankSeen` survived the headless splitter's opener branch, so an opener
  followed by a lazy continuation line read as "paragraph after a blank" and
  ended the run — `1. a` / blank / `2. b` / lazy / `3. c` folded `c` into `b`.
  The opener branch now clears it.
- The heading path re-derived per-line opener flags for headless regions
  without the paragraph reset, re-accepting a rejected `3. status: resolved`
  under a stale run and stripping it into a field. `GapsEntrySpan` now
  carries the flags the splitter itself computed, and the heading path reads
  them; the re-derivation is deleted.

* fix(#3702): ordered-run memory is per indent — nested runs resolve, nested ordinals never inherit the top-level run

Round-review continuation 2: nested openers consulted the TOP-LEVEL run
flag and never wrote their own, so a nested `1. / 2.` run under a hyphen
entry rejected its `2. status: resolved` (round 1 resolved it), while a
nested `3. status: resolved` under a nested `- ` bullet inherited an open
top-level run and was stripped into a false field.

`OrderedRuns` keys the memory by indent: a new opener at indent d resets
every deeper level, a paragraph after a blank at indent d ends the runs at
d and deeper, a thematic break or a heading clears all. Both splitters use
it; the top level still decides entry boundaries, nested levels decide
only which continuation lines are accepted openers for field stripping.
Pinned for LF and CRLF.

* fix(#3702): run levels — one top level at or above the base, CommonMark column indents, a fence ends its level's runs

Round-review continuation 3:

- A dedenting top-level list (`    1.` / `  2.` / `3.`) lost its entry
  boundaries: the exact-indent run lookup rejected the shallower ordinals
  before the boundary check ran. Every indent at or shallower than the
  list's base is now ONE level, in both splitters.
- `indentOf` counted characters, so a tab and a space aliased to one level
  and `\t1. nested` / ` 2. status: resolved` resolved falsely. Indent is now
  measured in CommonMark columns (§2.2: a tab advances to the next multiple
  of 4), for the run level and the entry-boundary check alike.
- A nested run survived a fenced block. A fence is a non-list block: its
  opening delimiter ends the runs at its level and deeper, exactly as a
  paragraph after a blank does.

* fix(#3702): the indent measure is grammar-scoped — Gaps keeps next's character count

`blockStructure: false` promised the Gaps grammar byte-for-byte parity with
`next`, but the CommonMark-column indent measure added for the deferred
grammar was shared by the whole splitter core, so tab-indented Gaps input
changed entry boundaries in BOTH directions:

  `\t- a` / `  - b`  — next folded into one entry, HEAD split into two
  `  - a` / `\t- b`  — next split into two,      HEAD folded into one

`indentWidth` now keys the measure on the grammar: columns for the deferred
set, raw character count for Gaps. The opt-out covers indent semantics, not
only fences and thematic breaks.

Four cases pin both halves — the two flipped Gaps pairs, the two Gaps pairs
that never moved, and the same tab/space pairs on the deferred path returning
the opposite (column-measured) verdict by design.

* fix(#3702): the acknowledge path reads and writes through one classifier

Round 3, Blockers 1 and 3, and Minors 7 and 8 — one mechanism, so one commit.
Every consumer of an entry's lines now reads the splitter's own per-line
verdict instead of a re-derivation of it.

B1. Round 2 widened the WRITER's status-line finder to the deferred marker set
while `extractGapEntryFields` still de-bulleted line 0 only. A nested
`  * status: pending` was therefore selectable by the writer and invisible to
the reader: acknowledge rewrote it in place, returned `ok`, and the item stayed
outstanding on every later audit. Measured against a `next` build, `*`, `+` and
`1.` each resolved on base and stopped resolving at round 2's head — a
regression, not a gap in new behaviour. The hyphen form of the same shape was
already broken on `next` and is fixed here too: one classifier cannot be right
for three markers and wrong for the fourth.

`parseGapEntryFieldLine` is now the single place a line is classified as a
field, and it reports the offset at which the VALUE begins. The rewrite happens
at that offset rather than through a second regex, so a line the classifier can
select is one whose rewrite it has already located — the selection and the
rewrite cannot disagree. Both `DEFERRED_STATUS_FIELD_RE` and
`DEFERRED_STATUS_REWRITE_RE` are deleted rather than widened. A read-back guard
returns `rewrite_not_readable` rather than `ok`; it is unreachable by
construction today and is the fail-loud floor under the next divergence.

B3. This is the end state the round-3 review prescribed on both #3739 and
#3773: #3773's shared classifier, parameterised by this PR's marker set, with
this PR's two status regexes deleted. #3773 lands first. Its hyphen-only strip
is consistent with `next`'s hyphen-only splitter today, so the writer/reader
divergence is created by THIS merge, which is why widening every consumer
belongs to the PR that widens the domain.

m7. The heading path marker-stripped its lines before calling the reader, so
the reader's fence scan ran over text the splitter never saw: `- ```sh` is an
ordinary bullet to the splitter but strips to a fence opener, and a
`**Status:** resolved` after it was suppressed as fence content — a resolved
entry resurfaced as open. Stripping now happens inside the reader, after the
fence scan.

m8. `rawGapEntryText` stripped a marker off line 0 unconditionally, but on the
heading shape line 0 is the heading TEXT: `### 1. Race in the writer` was
silently renamed to `Race in the writer`, and the name is the key acknowledge
matches on. Line 0 is stripped only when the splitter accepted it as an opener.

Also removed: `splitDeferredHeadingEntries`, whose sole caller only null-checked
it (round 3, M4 — the claim was zero callers, which was wrong; the wrapper's
`.map` was waste at the one call site), and `stripLeadingBulletMarker`, which
this change leaves with no callers at all. The export surface narrows to the two
splitter regexes the behavioural parity test reads (M6).

[PEER-ASK pr-order-12d5]
q: Reviewer blocked both on merge order. I'm declaring #3773 lands first and
   building the end-state shape into #3739 now (both my status regexes
   deleted). Does that match your plan?
reply: CONFIRMED - same order, derived independently. #3773 cannot carry the
   fold: `DEFERRED_BULLET_MARKERS`/`BulletMarkers` have zero occurrences at
   `next` (verified), so the prescribed end state is not executable inside
   #3773 without absorbing this PR's work.
deadline: 03:55 UTC (answered before it)
fallback: declare #3773 first, adopt end-state shape in #3739, push+comment
decision: proceeded as stated; #3773 lands first, this PR carries the widening
   of every consumer.

Refs #3740

* test(#3702): pin the detect/strip symmetry, and drop a white-box test that could not reach it

Round 3, Blocker 2 and Minors 6 and 9.

B2. The regression shipped green because no fixture put a marker on a nested
status line. Four markers x {nested status line}, each asserting the entry
READS BACK as acknowledged rather than that acknowledge merely reported `ok` —
reporting `ok` over a line the reader skips is the whole defect. Plus the bare
capitalised `Status:` case (the reader stores it case-sensitively, so the
writer must not select it), and an idempotence test, which is the failure the
defect actually produced: the item resurfaces, is acknowledged again, and never
settles.

Each of these was run against the pre-fix build first: all five fail there and
pass here. Two further assertions in the block are labelled CONTROL because
they held pre-fix — they guard the new offset-based rewrite and the opener-flag
threading against regressing, and calling them regression tests for a reported
defect would overclaim.

M6. The round-2 parity test asserted that four writer-side regexes embedded the
same source string. That is true of a detect/read asymmetry too, so it could
not have caught B1 — and two of the four regexes were widened into `export =`
purely to let it read them. Replaced with a behavioural test that drives the
real seam: every marker that opens an entry must also resolve it through
acknowledge. The structural assertion is kept for the two splitter regexes,
which really are two copies of one alternation.

m9. `expectedResolved` was computed and immediately voided; the loop beneath it
already asserts both polarities.

m7/m8 coverage lands here too: a bullet whose content is a fence opener must
not suppress the entry's fields, and a heading beginning with a list marker
must keep it in the entry name.

* docs(#3702): document the deferred-items entry shape where the file is written

Round 3, Major 5, and #3702's own item 2. The widened grammar was documented in
the reader (`forensic-audit.md`) but not at the write site, where
`executor-examples.md` still said only "log to deferred-items.md" — so the
question the issue actually raised, which shapes count, remained unanswered
anywhere a human writes the file.

States what opens an entry (`-`, `*`, `+`, and `1.` when the list starts at
`1.`), that `1)` is not a marker here, that a separator closes the list and
fenced content is never an entry or a field, and that an entry without an
explicit `status: resolved` stays open by design.

* chore(#3702): regenerate the changeset through the generator

Round 3, Minor 10. The fragment was hand-named against 64 generated names on
`next`, and its body ran ~250 words against CONTRIBUTING's one-sentence form.
Regenerated via `npm run changeset`, which is also what the random three-word
name is for: concurrent PRs never collide.

* fix(#3702): the fence gate lives on the seam both sides call, not just the reader

Found by the pre-push adversarial review of this round, and it is a regression
this round introduced rather than a pre-existing one.

`extractGapEntryFields` applied `fencedLineSet` before classifying; the
acknowledge writer's status-line search did not. So a `status:` line inside a
fenced block was SELECTED by the writer and SKIPPED by the reader — the write
produced a line nothing reads, the read-back guard refused it, and the entry
became impossible to acknowledge at all: `audit acknowledge` raised an internal
error and `complete-milestone` halted on it.

Measured, `- alpha` / fence / `  status: pending` / fence:

  next          ack=ok                   -> reads back "acknowledged"
  round-2 head  ack=ok                   -> reads back ""      (the B1 defect)
  before this   ack=rewrite_not_readable -> refuses entirely   (worse than next)

`entryFieldLines` is now the seam — per line of an entry, the field it declares
or `null`, fences included — and the reader and the writer both go through it.
That makes "the writer cannot select a line the reader will not read back"
structural rather than asserted, which is what the previous commit's message
claimed while a second read-side filter still lived outside the classifier.

Two comments corrected with it. The read-back guard is NOT "unreachable by
construction": this round shipped a reachable path to it, which is precisely
what an invariant asserted in a comment is worth. And the M6 replacement test
put its marker only on the entry opener, so it passed against the defective
build — the exact weakness it was introduced to fix in round 2's test. It now
marks the nested status line too, and fails pre-fix like the rest.

Round-3 tests against the pre-fix build: 10 of 12 fail there, and the 2 that
hold are labelled CONTROL because they guard this round's new code rather than
pin a reported defect.

* fix(#3702): one end-of-file CRLF algorithm, adopting #3773's with its B4 closed

Round-4 M1. Two open PRs shipped two different answers to "what line ending
does an entry that ENDS THE FILE get?", and the review's ruling was that the
disagreement needs one answer, not two. Neither shipped answer was that one.
Measured on builds of both heads:

  case                                     #3739 r3   #3773   here
  undelimited single entry, CRLF preamble    pass      FAIL    pass
  LF-dominant list, one stray CRLF at EOF    FAIL      pass    pass
  (the other five)                           pass      pass    pass

This PR's content.endsWith('\r\n', matchIndexInContent) reads the terminator of
the PREVIOUS line, so it propagated an isolated CRLF into an LF-dominant list --
refuted by #3773's own LF-dominant fixture, ported here. Withdrawn.

#3773's crlfAtEof asks the right question -- does anything before the entry,
within scope, contradict CRLF -- and fails closed. But its scope goes EMPTY for
an undelimited single-entry list, because the entry-list region runs from the
first entry's start to the insertion point and those coincide; crlfAtEof('') is
false by its own before.length > 0 guard, so 'preamble\r\n\r\n- alpha' gained a
bare \n in a CRLF document. That is #3773's B4, verified by driving its head.

Adopted here with the scope widened to everything preceding the insertion point
where the preferred region is empty, rather than asserting LF from no evidence.
That only ever loosens a scope carrying zero information, and the predicate
stays fail-closed over the wider one. An entry at offset 0 of an undelimited
document has no evidence under either scope and stays LF.

Tests: 10 added. Negative control, driven -- 1 of the 10 fails against this
branch's own pre-fix head (the stray-CRLF fixture); B4 fails against #3773's
head; the remaining 8 are the scope counterexamples ported with the function,
which were regression pins in #3773 and are guards here. Each still kills a
simpler algorithm: drop any one and a refuted scope passes again.

Four deferred-items suites 450/450, 0 skipped. npm run lint:ci exit 0.

* fix(#3702): drop the unreachable rewrite_not_readable guard (B3)

Round-4 B3: the status had zero test coverage in either file. The review
offered two branches -- drive it from a test, or delete it and stop carrying an
untested terminal status. Taking the second, with the reason stated rather than
assumed.

Why it cannot be driven. Round 3 added the guard after a fenced `status:` line
proved the writer could select a line the reader would not read back. Round 3
then closed that divergence STRUCTURALLY, by routing the writer's line selection
and the reader's field extraction through one entryFieldLines seam. The guard
now detects a state construction prevents: 21 document shapes were driven
against it -- fence openers on the bullet line for every marker in the widened
set, duplicate and triplicate status lines, bolded and nested variants, fences
between duplicates -- and none reached it. The only seam that would is routing
the internal call through the module's exports so a test could stub it, which
reshapes production surface for a test.

Why leaving it undriven is not free. RULESET.TESTS.mutation-score runs Stryker
incrementally over changed files at an 80% threshold and says to treat a
surviving mutant as a failing test specification. An undriven `if` on a changed
file is exactly that, on both the condition and the .toLowerCase() comparison.

What this gives up, stated rather than hidden: if a future change re-splits the
writer's selection from the reader's extraction, acknowledgeDeferredItem returns
ok over an item that stays outstanding -- the original #3702 defect class. One
correction to the review's framing: match_verification_failed does NOT backfill
it. That check runs BEFORE the write and compares the matched span to the
target, so it cannot see a post-write read-back failure. The protection against
re-splitting is the shared seam and the round-3 tests that pin it, not a runtime
assertion. A comment at the removal site records all of this.

Removing it also drops the union member from both files, which resolves the PR
body's internal contradiction (it claimed no type-signature changes while adding
one) and the duplicate-status surface #3773 collides on.

No test changed behaviour: 450/450 across the four deferred-items suites, 149/149
across the audit suites, npm run lint:ci exit 0 -- the same figures as before the
removal, which is itself the evidence that nothing exercised the branch.

* fix(#3702): the deferred fence gate is indent-unbounded, like the rest of the grammar (M2)

Round-4 M2. scanFencedBlocks is CommonMark, which caps a fence delimiter's
indent at three spaces -- a fourth makes it an indented code block instead. This
grammar had already opted out of that cliff for entry openers ([ \t]*) and for
THEMATIC_BREAK_RE (^[ \t]*), but not for fences. So a fence at four spaces was
not a fence to the gate, and a `status: resolved` line inside it RESOLVED the
entry containing it.

That is not an exotic shape. A fenced block written under a NESTED bullet sits
at four spaces, so ordinary hand-written deferred-items.md files reach it.
Driven before the fix at indents 4, 5, 8 and a leading tab: all four silently
resolved. It is the #3702 silent-resolution defect class in a new place.

gsd-core/references/executor-examples.md, added by this PR, states flatly that
"nothing inside a fenced code block is an entry or a field". The review offered
fixing the parser or bounding that claim in three places. Fixing it -- the claim
is the one users will rely on, and the grammar had already chosen unbounded
indent everywhere else.

NO second fence dialect (the rule blankIndentedFenceDelimiters states). The
classification is still done by scanFencedBlocks, the one exported CommonMark
state machine, over a de-indented VIEW of the same lines. Run lengths, backtick
vs tilde, closer-must-match-and-not-trail, info-string rules and the
unterminated-at-EOF case remain that engine's answers. Indent is the only
dimension hidden from it, and it is exactly the dimension this grammar has
already declared it does not measure. Index alignment is 1:1 -- map preserves
length -- so every returned line index still addresses the original line.

Scope is the deferred grammar only. Both marker-parameterised call sites gate on
markers.blockStructure, which the Gaps set does not set, so Gaps reaches an empty
set. Verified, not asserted: the 47-fixture Gaps differential (marker x
line-ending x separator x fence x break x key-shape x list-shape) is
BYTE-IDENTICAL across this change, 8033 bytes both sides.

Tests: 14 added, of which 8 fail against the pre-fix source and pass here; the
other 6 are the deliberate controls -- indents 0 through 3, which must NOT move,
and the Gaps opt-out guard.

Four deferred-items suites green; the 58 suites touching uat/deferred/sectionizer
run 6045 tests with an IDENTICAL failing set before and after this change (17
pre-existing environment failures -- installs and an unpinned GSD_EMITTED_BASE;
emitted-attribution passes 259/259 in isolation with its base pinned). lint:ci
exit 0.

* fix(#3702): changeset, both prose parsers, and the minors (M3, M4, m1-m3, m5, n1-n2)

M3 -- the changeset omitted a user-BREAKING change. Measured against next: a
heading-delimited deferred-items.md written with `*`, `+` or `1.` went from
"0 entries, so complete-milestone has nothing to acknowledge and closes" to
"1 entry, the CLI writer refuses the heading shape, ACK_FAILURES accumulates,
exit 1". The `-` form already halted and is unchanged. That is release-note
material: a close that used to succeed now fails, and the correct response is to
fix the file, not revert. Also names the fence-indent fix below, and adds #3740
so #3773's issue is attributed here as it is absorbed.

M4 -- gsd-core/workflows/progress/steps/forensic-audit.md is a SECOND,
model-executed parser of the same grammar, and prose cannot carry a parity test.
Its widened text stated the start-at-1 rule, fences and separators but not the
`1)` exclusion nor the nine-digit ordinal cap, both enforced in code with pinned
tests. Both stated now, along with the round-4 fence-indent rule. (No ack
fragment: the size ratchet's currentSizes does a NON-recursive readdirSync of
gsd-core/workflows and agents, so a file under workflows/progress/steps/ is
outside its scope -- verified by reading the helper, not by the green.)

n1 -- executor-examples.md documented that the BOLDED status key is matched
case-insensitively and left the bare key's rule to inference. Driven: bare
`Status: resolved` is NOT read, so the entry stays open with no warning, while
`**Status:**` is. Stated explicitly, with the digit cap and the any-indent fence
rule (n2).

m1 -- boundary coverage was 2/3. limit (999999999.) and limit+1 (1234567890.)
were pinned; limit-1 (12345678.) added, per RULESET.TESTS.boundary-coverage.

m2 -- THEMATIC_BREAK_RE and the tab-expanding indent counter are hand-rolled
CommonMark rules with no in-repo peer to compare against, so the parity
assertion is against the SPEC: eight positive and five negative fixtures, plus
the two DELIBERATE divergences pinned as deliberate (`+` is a separator here but
not in CommonMark, because `+` is a list marker in this grammar and `+ + +`
would otherwise be a phantom entry; indent is unbounded). One fixture was
initially wrong -- `-- -` IS a CommonMark break, since the spec allows free
spacing between the three characters -- and the parser was right.

m3 -- the result union is hand-duplicated in audit.cts as part of a deliberate
structural view of uat.cjs, so the fix is not to delete a copy but to make drift
observable. Every REACHABLE status is now driven from a fixture; four of the six
(ambiguous, unsupported_heading_shape, already_resolved, match_verification_failed)
had no assertion anywhere in the suite before this. match_verification_failed is
still undriven and the test says so rather than omitting it.

m5 -- DECLINED, with the measurement. The review is right that `(\s*)` in the
opener and `/^[ \t]*/` in the reader disagree about \f, \v and NBSP, but its
prescribed narrowing was implemented, driven and REVERTED: as shipped, an entry
indented with any of those surfaces, parses its status field, acknowledges, and
reads back acknowledged -- a complete round-trip. Narrowing turns all three into
SILENTLY DROPPED entries, which is the #3702 defect class itself and the opposite
of this file's stated fail-safe rule. A latent inconsistency in the safe
direction is not worth a live regression in the unsafe one. Pinned by three
round-trip tests so the prescription cannot be re-applied silently; if it is ever
closed, the direction is to make the readers agree with the opener, not to make
the opener reject lines it accepts today.

Four deferred-items suites 475/475, 0 skipped. lint:ci and lint:changeset exit 0.
The 47-fixture Gaps differential is byte-identical at 8033 bytes.

* fix(#3702): the pinned `## Gaps` phantom now cites its issue (m4)

Round-4 m4. The second assertion in the Gaps byte-for-byte test pins a real
defect as expected output: a spaced hyphen thematic break in `## Gaps` is read
as an ITEM, so `- - -` surfaces a phantom open gap named `- -`. Reproduced on
pristine next at 389bc86e0 across nine separator shapes -- every spaced hyphen
form is affected, `---`/`----`/`* * *`/`___` are not, and the dividing line is a
space after the first hyphen (the Gaps opener is /^(\s*)(-)\s/ with no
thematic-break concept at all).

Filed as open-gsd/gsd-core#3898. The pin stays: scope-limiting Gaps is the point
of the blockStructure opt-out, and this assertion is the only thing that would
notice the Gaps path moving. What was missing was the tracking -- a pinned defect
with no issue behind it reads as intended behaviour to the next reader. The
comment now says which it is and what the expectation becomes when #3898 lands.

* fix(#3702): an unterminated fence runs to the end of its entry, never past it (B1, B2)

Round 4 de-indented every line before `scanFencedBlocks`, so a fence
opened at any indent — and `scanFencedBlocks` runs an unterminated
fence to end-of-document — so one stray delimiter swallowed every entry
after it into the entry before it. `- a` / blank / four-space ``` /
blank / `- b` yielded ONE entry where `next` yields two: a widening
that made an already-counted item vanish, on the mixed-file shape #3702
exists to close. Reproduces at indent 0 as well.

The bound is the entry. CommonMark closes a fence with its container
and a container at the next item at its level; this parser extends
that to a document-level stray delimiter, where CommonMark would
swallow to EOF and the fail-safe rule (surface, don't drop) will not.
`scanFencesFrom` reports the unterminated opener and the walk supplies
the bound — the next line shaped like a top-level item — then RESCANS
from it, so a later delimiter is read on its own terms. Still one
fence dialect: every block boundary is `scanFencedBlocks`' answer.
Entry-scoped `fencedLineSet` (the field reader) already ran an
unterminated fence to the end of its lines, so reader and walk agree
by construction.

Tests: the M2 pin that asserted `[]` for a stray fence before an item
flips (the item counts); the round-4 "runs to end-of-file, exactly as
CommonMark says" test is retitled — its assertion stands because the
bound is the entry — and extended with the next entry; a new block pins
the review reproduction at both indents, a terminated deep fence still
gating, the gated status inside the bounded fence, the rescan case, and
the heading-tokenizer caveat (at indent 0 the tokenizer applies
CommonMark's own fence rule, so a heading after a stray delimiter is
body text there, exactly as on `next`).

Reverted in isolation against the final tree: 3 named tests fail.

* fix(#3702): `0.` starts an ordered list (M1)

The start-at-1 rule applied unconditionally dropped ONLY the first item
of a `0.`-numbered list — the run then started at `1.` — which is the
mixed under-report that looks like a clean parse. CommonMark §5.2
permits any 1-9-digit start and a `0.` list is ordinary; a sentence
opening with "0." is not a shape anyone writes. The threshold is now
`> 1`. The cost is restated accurately in the doc comment and pinned:
a list starting at 2 or more, at a paragraph position, reads as prose
until its first `0.`/`1.` line — the prefix, not the whole list.

Boundary tests at the threshold itself: `0.`, `1.`, `2.` starts, `00.`/
`01.`, and the prefix-loss case. Reverted in isolation: 1 named test
fails.

* fix(#3702): a non-1 ordinal is an item wherever a list is already open at its level (M2)

The per-indent run memory recorded whether the previous opener was
ORDERED, so a bullet item closed the run and `1. a` / `- b` / `5. c`
folded `5. c` into `b` — another mixed-file under-report. In CommonMark
`5. c` there opens a fresh ordered list (start=5): a non-1 start is
refused only where it would interrupt a PARAGRAPH (§5.3), and after a
list item it interrupts nothing. `ListRuns` now records "a list is
open here"; the start rule applies where no list is open at the line's
level — the positions a sentence can occupy — so the round-2 B2 pins
(doc start, after a heading, after a paragraph) hold unchanged.

Two round-2 pins move with it, both CommonMark-backed: `1. alpha` /
`- beta` / `2. gamma` is three items, and a nested `3. status:` after
a nested bullet is a nested item (a field line, as `- status:` would
be); the "rejected ordinal is not stripped" pin is re-anchored at a
paragraph position, where it still holds. Reverted in isolation: 4
named tests fail.

* docs(#3702): the two runtime-loaded docs state the grammar the parser ships (B3, M4 parity)

`executor-examples.md` (the write-site doc) and `forensic-audit.md`
check 7 (the model-executed parser) both asserted "never silently drops
a possibly-open item" over a grammar that dropped three measured shapes.
Both now carry the round-5 grammar — `0.`/`1.` starts, a non-1 ordinal
inside an open list, an unclosed fence ending with its own entry — and
the fail-safe sentence is kept with what it does NOT cover named
beside it: a fenced line, a separator, and an ordered list numbered
from `2.` upward at a paragraph position, and nothing else.

* docs(#3702): changeset reflects the merged contract

The "Breaking, and deliberate: … HALTS complete-milestone" paragraph
described a refusal that #3781 removed from `next`; a heading-shaped
file written with a newly recognised marker now surfaces its entries
and `complete-milestone` acknowledges them in place. The fragment cites
#3702 alone — #3740 and #3775 closed on `next` through #3940 and #3989;
this PR's shared reader/writer classifier subsumes both fixes rather
than closing either issue. The round-5 grammar (ordered start, unclosed
fence bound) is stated in the user-facing sentence.

* fix(#3702): the heading-shape insert lands on a line the reader reads, and keeps a closing `#` sequence

Found by the round's pre-push adversarial review. An entry whose body ends
in a fenced block — closed, or unclosed and therefore running to the
entry's end — received `status: acknowledged` AFTER its last non-blank
line, i.e. as fence content: the writer returned `ok` and the reader
never saw the marker, the item stayed outstanding. That is the #3702
class itself (a write nothing reads), on the shape #3781 just opened.
The insert now walks back over blank AND fenced lines, classified by the
reader's own `fencedLineSet`, so the marker lands on a line the reader
reads; pinned as round-trips for an unclosed fence, a closed fence, and
a pending entry ending in an unclosed fence before a heading. Reverted
in isolation: the round-trip test fails.

Separately, the leaf line-0 rewrite (`### status: open ###`) dropped the
closing `#` sequence; it is kept now. Cosmetic, pinned.

* docs(#3702): the prose parser states the bare-key case rule; both docs say what an unclosed fence does, no more

`forensic-audit.md` check 7 called `status: resolved` case-insensitive
where the code reads a bare key lower-case only (the bolded form in any
case; the value case-insensitively) — `executor-examples.md` already said
so, the model-executed parser did not. And both docs claimed "a stray
delimiter cannot hide the entries after it", which overstates B1: an
UNCLOSED fence ends with its entry; a closed pair of delimiters is a
fence, whatever sits between them, as CommonMark reads it. Found by the
round's pre-push review.

* fix(#3702): a heading whose text is a fence delimiter is a heading, not a fence

Second finding of the round's pre-push review, one door over from the
first: for a leaf headed `### ```` (or `~~~`) the entry-level fence scan
read line 0 — the heading TEXT, not a Markdown line — as a fence opener,
so every body line was fenced: the reader read no field under it, and
the writer's marker (placed by the same scan) landed on a line nothing
reads — `ok`, item outstanding. `entryFencedLines` now owns the entry's
fence view for reader and writer alike, and a leaf's line 0 never opens
a fence (the leaf tell is `openerFlags[0] === false`; a pending or
headless entry's line 0 is a marker line, never a delimiter). Pinned for
both delimiters, read and write; reverted in isolation the pin fails.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-31 13:44:47 -04:00
Tom Boucher
62b0d939b6 feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey (#4083)
* feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey

Add an optional `timeoutConfigKey` field to the reviewer lane descriptor,
resolved in `resolveLanePlan` at invocation time and falling back to the
frozen `timeoutFloorMs` when unset or invalid, in the same spirit as the
existing `promptBudgetKey`/`modelConfigKey` fields. All 12 shipped lanes
declare `review.timeouts.<slug>` on both surfaces (the descriptor and their
capability.json manifest), validated by capability-validator.cjs.

For the antigravity lane, the native `agy --print-timeout` flag — previously
a second hardcoded literal (`540s`) independent of the outer cap — is now
derived from the same resolved outer timeout in `antigravityArgv`, preserving
the existing 60-second buffer relationship (ADR-2782 D6: the outer bound is
declared data, the inner one is handler-owned).

The antigravity default timeoutFloorMs stays at 600s per the maintainer's
disposition; users raise it through the new config key instead.

* docs(#3274): document review.timeouts.* and extract resolveTimeoutMs helper

Address code-review findings on the timeoutConfigKey change: extract the
inline timeout-resolution logic into a named, exported, directly-tested
resolveTimeoutMs helper (matching the file's existing configString/
normalizeHost convention); document the new review.timeouts.* federated
config keys in docs/CONFIGURATION.md, docs/reference/capability-manifest.md,
and docs/how-to/ship-a-reviewer-lane.md; add the changeset fragment.

* fix(#3274): resolve native antigravity timeout in resolveLanePlan, not the runner

gsd-test caught two design mistakes in the prior commits:

1. SpawnPlan.argv is documented and tested as fully resolved by
   resolveLanePlan (model/effort/output/prompt already folded in) — leaving
   the antigravity '{{nativeTimeout}}' marker unresolved until the runner's
   antigravityArgv violated that contract and broke tests that read
   plan.argv directly (tests/antigravity-reviewer.test.cjs,
   tests/review-default-reviewers-workflow.test.cjs). Fix: '{{nativeTimeout}}'
   is now a fifth ARGV_PLACEHOLDER member, resolved by resolveLanePlan itself
   via the new nativeTimeoutToken() helper, exactly like the other four.
   antigravityArgv reverts to its pre-#3274 four-argument form. Also missed
   updating capabilities/antigravity/capability.json's invoke.args to match
   the descriptor, which broke the manifest/descriptor parity test.

2. tests/reviewer-config-federation.test.cjs enforces a deliberate, narrow
   invariant (#3691 narrows #2797): qwen, cursor, and coderabbit — the three
   lanes with neither a model flag nor a host — may own no config key beyond
   their own prompt-budget key. Adding review.timeouts.<slug> to all 12 lanes
   violated it. Fix: those three keep timeoutConfigKey: null and own no
   review.timeouts.* key, matching their existing modelConfigKey: null. The
   other 9 lanes are unaffected.

* chore(#3274): backfill changeset PR number (pr:0 -> 4083)

---------

Co-authored-by: sim <sim@local>
2026-08-30 14:26:02 -04:00
Dennis Kim
8487f0ed42 enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage

- pin configured, absent, and malformed branch-list behavior
- require opposite CLI and execute warning outcomes

* feat(01-01): warn on configured protected branches

- resolve the base branch union configured protected branch names
- expose exact boolean CLI comparison output for workflow callers
- keep execute-phase warning advisory and within its byte budget

* test(01-01): add failing protected branch config coverage

- cover valid list persistence and null unset
- reject hostile shapes while preserving the prior value

* feat(01-01): validate protected branch configuration

- register git.protected_branches as a canonical config key
- require a non-empty array of non-blank branch names

* test(01-02): add failing ship protected-branch controls

- Execute both workflow warning blocks with exact predicate arguments
- Require true and false results to produce opposite warning outcomes
- Preserve the none-strategy feature-branch offer contract

* feat(01-02): warn at ship on protected branches

- Reuse the typed protected-branch predicate in ship preflight
- Keep raw base resolution for PR targeting and advisory branch creation
- Prove execute and ship warning blocks with opposite-result controls

* test(01-02): add failing protected-branch docs parity

- Require the canonical schema key in both English config references
- Pin the non-empty string-array type and absent default
- Require synchronized multi-branch examples and advisory semantics

* feat(01-02): publish protected branch configuration contract

- Document the optional non-empty string-array field in both references
- Explain resolved-base union and absent-field compatibility
- Keep execute and ship warnings advisory under branching_strategy none

* fix(01): CR-01 honor active workstream branch policy

* fix(01): WR-01 assert protected config path selection

* docs: add changeset fragment for #3648

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx

* fix(#3648): resolve base_branch precedence inversion and round-1 findings

Blocker 1/2: production config resolution was flat-first, so a project
that migrated to git.base_branch but still carried a stale flat
base_branch got the old value back. Add base_branch to
normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos
pattern: canonical nested wins) and route readEffectiveGitConfig's
test seam through the same normalization so it can't silently diverge
from production again. Adds a regression test with both keys set that
fails without the fix.

Blocker 3/4/5: restore the handle_branching case-selector prose and
"none" contract sentence that #3389's tests anchor on, and revert the
unrelated prose/comment compaction in the same step — both were
drive-by edits outside #3552's scope.

Also addresses review majors/minors: delete readConfigBaseBranch and
readConfigProtectedBranches (dead in production, only self-tested);
--is-protected now fails closed (reports protected) instead of
silently answering false when the base branch can't be verified;
trim configured protected-branch names; fix HOME-without-USERPROFILE
vacuous isolation on Windows; correct the drift-ack's byte accounting.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N

* test(#3648): add failing legacy-key hoist safety coverage

Round-2 review found normalizeLegacyKeys block 5 records a normalization
carrying the DISCARDED flat value on the canonical-wins branch. Probing
that turned up a second, unreported defect in the same helper shape:
blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no
object guard, so a config whose section key holds a string is spread into
index keys —

  {"git":"main","base_branch":"release"}
    -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}}

The resolved value is accidentally still correct, so nothing fails and no
diagnostic fires. But normalizations.length > 0 sets configDirty, and
config-loader then serializes that shape back into the user's
config.json — a read that silently corrupts config.

The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"}
explicitly; this is the input it would have caught.

Covers both defects across blocks 1 and 5, with object/array/null
negative controls that must stay green in both phases, and a fast-check
property over arbitrary `git` values.

* test(#3648): pin fail-closed handling of malformed protected_branches

Replaces the test that pinned the fail-OPEN behaviour. The old
assertion — ['develop', 42] yields isProtected === false for 'develop' —
locked in the exact failure #3552 exists to close: config-set validation
is bypassable by a direct edit of .planning/config.json, so a user who
believes 'develop' is protected got a silent false and no warning.

It was also inconsistent with the fail-CLOSED direction twelve lines
away, where an unverified base reports protected and writes a
diagnostic. A protection predicate must not have two opposite failure
directions depending on which input is bad (#3648 review Blocker 3).

New coverage: a bad element drops only itself, a non-array contributes
no names, an empty list is well-formed rather than malformed, and
--is-protected surfaces the rejection. Both negative controls — a clean
list reports nothing rejected and writes no diagnostic — must stay green
in either phase, so the reject channel cannot fire unconditionally.

* fix(#3648): drop only invalid protected_branches and report them

Partition git.protected_branches instead of discarding the whole list on
one bad element, and carry the rejections out through
ProtectedBranchStatus so --is-protected can name them on stderr. Valid
names keep protecting; the user finds out the rest were ignored.

A non-array value still contributes no names — a bare string is not a
list of branch names — but is now reported rather than swallowed. An
empty array stays silent: declaring no extra protected branches is a
valid choice, not a misconfiguration.

writeDiagnostic is hoisted out of the unverified-base branch since both
arms now use it.

* test(#3648): prove the predicate diagnostic survives both call sites

The workflow bash stub now emits a stderr diagnostic the way the real
command does, which is what makes a swallowed `2>/dev/null` visible to a
test — previously the stub was silent on stderr, so discarding it changed
no observable behaviour and the call sites could drop the explanation
undetected.

Adds the Minor 2 binding check as well: ship must expose the predicate
result as IS_PROTECTED rather than only echoing a warning, asserted by
running the extracted bash and reading the bound value, not by grepping
the workflow source.

Both tests carry opposite-outcome controls — an empty diagnostic must
leave the text absent, and a false predicate must bind false.

* fix(#3648): surface the predicate diagnostic and bind ship's result

Drop `2>/dev/null` from the --is-protected call at both call sites. The
fail-closed explanation and the new rejected-entry warning both go to
stderr, so discarding it left the user with a bare "protected branch"
warning on a branch that is not protected and no way to tell a real
match from a degraded-git guess. `git branch --show-current` keeps its
own redirect — that one is genuine noise.

ship.md binds IS_PROTECTED and its prose now branches on the variable,
so the following steps have evaluable state instead of having to infer
it from warning text in tool output.

execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth
319 bytes (was 331 before the redirect came out). Baseline re-verified
against the current rebase base by blob id; the ceiling check passes
with 755 bytes of margin.

* test(#3648): restore negative space for the readFile config seam

The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm
it pinned survives verbatim in readEffectiveGitConfig's readFile branch —
the JSON.parse catch, the non-object guard, the git-section object guard,
.trim() and blank-string rejection — and the four surviving readFile
injections were positive-path only. protected_branches was never driven
through this seam at all.

Restores nine cases against the seam, including protected_branches
partitioning, plus a control proving loadConfig still wins when both
seams are supplied.

Records honestly what the suite pins. Mutating the built lib shows
.trim() is KILLED, while the non-object guard and the blank-string
rejection SURVIVE — both are unreachable through this entry point for
the same reasons the deleted suite documented against its own
equivalents: a JSON-parsed non-object carries no relevant own-property
either way, and a blank value is rejected a second time downstream by
the resolver's truthiness check. They stay as defence-in-depth and are
labelled known-unkillable rather than left looking like coverage this
suite does not provide.

* test(#3648): distinguish detached HEAD from a missing branch argument

`args[1] ?? ''` collapsed two different situations into one: a detached
HEAD, where `git branch --show-current` legitimately prints nothing, and
the flag being called with no argument at all. Both answered false, so
the right outcome arrived by an unintentional path and a caller bug was
indistinguishable from normal operation.

Asserts the detached case stays silent and the missing-argument case
reports, with a control that the two diagnostics differ.

* fix(#3648): report a missing --is-protected branch argument

Answer false either way, but say so when the flag arrives with no
argument. A detached HEAD passes an explicit empty string and stays
silent, since that is a normal state rather than a misconfiguration.

* docs(#3648): state exact-name matching and per-entry rejection

isProtected is exact string equality, so a git-flow project must
enumerate every release/* and hotfix/* by name. #3552 only asked for an
integration-branch field, so the implementation satisfies the letter of
the issue while leaving its git-flow motivation partly unserved — say so
where users will meet it rather than leaving them to discover it.

Also documents the Blocker 3 behaviour change: an invalid entry is
ignored with a warning naming it and the remaining names still apply.

Both statements land in docs/CONFIGURATION.md and
gsd-core/references/planning-config.md, and the config-field-docs parity
test asserts each in both so the two cannot drift.

* refactor(#3648): extract isValidProtectedBranches for cross-surface pinning

The `git.protected_branches` check inside `cmdConfigSet` and the resolver's
per-entry filter in `git-base-branch.cts` are deliberately different shapes —
all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot
fail the guard open. Nothing structural keeps their two definitions of "usable
branch name" in step.

Lifting the write-side check into a named, exported predicate lets a property
test ask both surfaces about the same value and assert they agree, which is the
fast-check gap the round-2 review flagged. No behaviour change: the predicate is
the same expression, called from the same place.

* fix(#3648): stop --is-protected rewriting the config it is asking about

`gsd_run query git.base-branch --is-protected` runs on every execute-phase and
every ship. It resolved config through `loadConfig`, whose normalize-then-write
path rewrites `.planning/config.json` whenever any legacy key normalizes — so a
boolean question was silently editing the user's checked-in config. This PR had
widened the trigger by adding a fifth normalization block (top-level
`base_branch` -> `git.base_branch`), making it fire for exactly the projects the
feature targets.

`loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution
is unchanged, only the two write-back side effects are suppressed. The predicate
passes `persist: false`; the ~30 other callers are untouched, so a legacy config
is still migrated by ordinary use.

Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and
reflows whitespace even when the values are equivalent. Three tests, each with
its own control: the end-to-end CLI leaves the file byte-identical while still
answering `true` from the legacy key (proving the config WAS read); an ordinary
persisting load of the same fixture DOES change the bytes (proving the fixture
is live rather than inert); and `persist:false` vs default over one directory
returns deep-equal config while differing on the write. Reverting the one-line
`persist: false` fails the first of those and only that one.

Also from the review:

- `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through
  the same precedence authority production uses". It does not, and cannot — it
  reproduces two of production's steps over a single file. The comment now names
  what the seam covers and what it does NOT (root/workstream deep merge, builtin
  and global defaults, federated merge), and the seam now applies production's
  flat-then-nested lookup so it stops disagreeing about a surviving flat key.

- The missing-argument diagnostic promised "answering false", which the
  fail-closed guard on the same call can contradict by printing `true`. It now
  states what it did with the argument and leaves the answer to stdout.

* test(#3648): re-pin block 5 on #3760's refusal contract

#3767 landed on next while this PR was in review and fixed the non-object
config-section defect properly: a present-but-non-object section now BLOCKS its
own migration — value preserved, no Normalization pushed, refusal reported via
`skipped[]` — rather than being rebuilt from a plain-object view. That supersedes
this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread
but still dropped the section value silently, and which the round-3 review
correctly called out as destruction in place of corruption. The rebase drops that
commit and routes block 5 through the upstream helper.

This file's tests asserted the superseded design, so they are rewritten to pin
block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's
suite was written — against the contract that now governs it: ordinary hoist into
an absent/null/object section, canonical-nested-wins, and refusal for each of
string/number/boolean/array sections with the exact `skipped` entry.

Two controls keep it from passing vacuously: the refusal must be scoped to block
5 (an unrelated block still normalizes in the same call), and a property over
arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually
exclusive per key, that a refusal leaves both the section and the legacy key
untouched, and that a hoist manufactures no index key the input did not carry.

* docs(#3648): correct the Git Query and Config Loader module contracts

CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct
`.planning/config.json` read. Since this PR it is the EFFECTIVE configuration
resolved by the Config Loader — a materially different authority, carrying the
root/workstream deep merge, flat-then-nested lookup and builtin/federated
defaults. The `--is-protected` predicate, `git.protected_branches`, and the two
invariants that distinguish the predicate from the plain query (fails closed on
an unverified base; must not write) were undocumented entirely.

The Config Loader entry now states that loading is not side-effect-free by
default and documents `options.persist`.

docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and
no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write`
was run and produced no diff: the manifest indexes roster NAMES, not row prose,
so a description edit cannot move it.

Also closes the global-defaults minor: `git.protected_branches` is inert in
`~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key
appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is
section-wide and predates this PR, so the fix is to state the scope where users
meet it rather than to quietly extend the resolution set for two new keys.

* fix(#3648): close four defects found by the round-4 external review

Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially
against this branch. Four findings reproduced against source; each is fixed with a
failing-first test and a control, and each fix was verified by reverting it and
watching exactly the intended test fail.

1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was
   only half closed. `loadConfigResolved` re-enters itself with a bare
   `{ workstream: null }` when a workstream has no config.json of its own, and
   that literal discarded every other option — so the recursive pass ran at the
   DEFAULT persistence and rewrote the ROOT config. Reproduced: with
   GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected`
   rewrote `.planning/config.json` despite `persist:false`. Both recursions now
   forward `options` and override only `workstream`; the explicit override still
   wins the hasOwnProperty check, so spreading cannot let `workstreamContext`
   reintroduce a workstream.

2. Both workflow call sites failed OPEN, and aborted under `set -e` (both
   reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty
   string when the query fails, so `[ "$X" = true ]` was simply false: no
   warning, no trace — a silent hole in the guard whose only job is to warn. The
   bare assignment also aborted the step under `set -e`. Both sites now degrade
   VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the
   check did not run. Deliberately not fail-closed — claiming "protected" on no
   evidence would warn on every branch whenever gsd-tools is unavailable.

3. `isValidProtectedBranches` and the resolver disagreed on a sparse array
   (antigravity). `.every()` skips holes; the resolver's `for...of` yields
   `undefined` for them, so `["main", , "develop"]` was accepted by config-set
   and rejected by the resolver. The cross-surface property passed only because
   `fc.array` cannot generate a hole. The predicate now indexes, and the
   generator punches holes so that axis is actually falsifiable. JSON cannot
   express a hole, so this is unreachable in production — but two definitions of
   one predicate must not contradict each other.

4. A top-level `protected_branches` silently outranked `git.protected_branches`
   (antigravity). Routing the key through `get(key, {section, field})` gave it
   flat-then-nested precedence, which is back-compat for keys
   `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has
   no legacy form, so that invented an undocumented alias. It now resolves
   nested-only through a new `getNested`, in production and in the test seam.
   `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's
   refusal path can leave behind — and a control pins that distinction.

Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed
only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`);
a git command that runs and exits non-zero counts as a clean negative, so a cwd
that is not a repository answers `false`, not `true`. Verified pre-existing on
next @ 738f42f4, so the documentation was over-claiming rather than the code
regressing — but an over-broad contract is exactly what the module docs must not
carry.

Both workflow byte figures re-derived after the call-site change:
execute-phase.md 92356 -> 92865 (+509), ship.md 36784 -> 37227 (+443).

* test(#3648): pin git config read parity

* docs(#3648): document git query contracts

* fix(#3648): expose protected branch default

* test(#3648): snapshot planning tree for read-only query

* test(#3648): pin planning snapshot stray-write detection

* fix(#3648): resolve merge conflict from #3078's ack-fragment sweep

next swept the fully-spent 2818/3003 ack fragments this branch had
appended to (#3078, a84f7563). Rebased onto upstream/next and took
the deletions on both, then moved the #3552 append into a new
fragment of its own.

Rebasing onto the current base also left execute-phase.md only 34
bytes under the frozen ADR-857 Phase 6 margin ceiling (93400 bytes) —
intervening next PRs consumed the rest while this PR was in review.
Extracted the "none" arm's protected-branch-warning bash block into
gsd-core/workflows/execute-phase/steps/protected-branch.md (content
unchanged, matching the existing steps/ extraction pattern used
elsewhere in this file) so the inline growth is a one-line pointer
instead of the full block. 93366 -> 93385 bytes (+19), 15 bytes
inside the ceiling.

* fix(#3648): drop stale ack entry for the new step file

The extracted execute-phase/steps/protected-branch.md needed no
acknowledgment of its own — the differential-attribution check flagged
the entry as stale once the build ran, so removed it and kept the two
growth entries (execute-phase.md, ship.md) that actually needed one.

* fix(#3648): follow the step-file reference in the bash-extraction test helper

extractProtectedBranchWarningBash() read the "none" arm's bash block
directly out of execute-phase.md. That block now lives in
execute-phase/steps/protected-branch.md (byte-ceiling extraction);
the helper follows the step-file reference and extracts from there
when no inline block is found, so the three execute-phase tests that
execute this bash for real keep exercising the actual behavior.

* fix(#3648): regenerate INVENTORY-MANIFEST.json and satisfy the CRLF-fragile lint rule

- gen-inventory-manifest.cjs --write to pick up the new
  execute-phase/steps/protected-branch.md entry (already covered by
  docs/INVENTORY.md's generic workflow_steps wildcard row, so no
  INVENTORY.md edit is needed).
- Reworked the step-file-reference lookup in
  extractProtectedBranchWarningBash() to avoid a bare-\n regex split
  on file content (local/no-crlf-fragile-split), using the same
  line-array scan the function already uses elsewhere.

* fix(#3648): regenerate golden install-tree fixtures for the new step file

npm run gen:install-tree, adding gsd-core/workflows/execute-phase/
steps/protected-branch.md to all 19 runtime install-tree fixtures.
CI's tests/golden-install-tree.test.cjs caught this on push — I'd
verified the differential-attribution and INVENTORY-MANIFEST checks
but missed this separate golden-fixture check for the new file.

* fix(#3648): add the canonical gsd_run preamble to the new step file

CI's runtime-launcher-parity suite requires exactly one canonical
resolver preamble in every workflow .md that calls gsd_run. The
inline "none"-arm block never needed one (execute-phase.md already
carried a preamble elsewhere in the same file), but the extracted
execute-phase/steps/protected-branch.md is now its own file with no
preamble of its own. Ran node scripts/sync-runtime-launcher.cjs to
insert it (execute-phase.md itself is untouched — still 93385 bytes,
inside the ADR-857 ceiling).

That preamble defines its own gsd_run(), which shadows the mock
tests/git-base-branch.test.cjs injects for the three #3648 tests that
execute this bash for real — without stripping it, those tests reached
the real gsd-tools.cjs on the machine running them instead of the
test's fixture. Preamble correctness is already covered by
tests/runtime-launcher-parity.test.cjs, so extractProtectedBranchWarningBash()
now strips the preamble line before handing the block to the harness;
it only needs to exercise the #3552 warning logic.

* fix(#3552): address PR 3648 review feedback on protected branch warnings

- Fix execute-phase handle_branching branching_strategy=none instruction
  to "Read and execute execute-phase/steps/protected-branch.md"
- Use io.error(..., ERROR_REASON.USAGE) for cmdGitBaseBranch usage errors
- Align git.protected_branches schema default to (none) without fallback []
- Relocate CONTEXT.md forward-referencing sentence into module body
- Sanitize control and ANSI characters in renderRejected diagnostics
- Clean up out-of-scope whitespace hunks in gsd-tools.cjs

Emitted-Drift-Ack-Growth: execute-phase.md — #3552: execute-phase handle_branching adds a pointer to execute-phase/steps/protected-branch.md for branching_strategy=none so the protected-branch check executes while keeping execute-phase.md within the ADR-857 Phase 6 margin ceiling (93400 bytes). 93392 bytes, 8 bytes inside the ceiling.
Emitted-Drift-Ack-Growth: ship.md — #3552: ship preflight step 3 now asks the same typed git.base-branch --is-protected predicate as execute-phase, binding IS_PROTECTED and warning without refusing execution or blocking the branching_strategy=none feature-branch offer; it degrades visibly (rather than silently reading an empty result as "not protected") when the query itself fails to run. 36841 bytes, well inside the XL cap (98304, tests/workflow-size-budget.test.cjs).

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-29 17:00:52 -04:00
0xdhx
472f585f7c fix(#3726)!: require --confirm before milestone complete mutates (#3774)
* fix(#3726): require --confirm before milestone complete mutates

`milestone complete <version>` is a one-way door — ROADMAP.md and
REQUIREMENTS.md archived, every phase directory in the milestone MOVED,
STATE.md rewritten — and ran unconditionally on first invocation through
every invocation path, including `query milestone.complete <version>`,
whose `query` meta-prefix reads as a read-only namespace but performs no
filtering (#167's invocation-compatibility shim + #3243's dotted-form
normalization).

The gate lives on the destructive command itself, not on the `query`
prefix (the prefix is an intentional invocation mechanism, not a
permission boundary — restricting it would break dozens of shipped
workflow callers). Without --confirm and without --dry-run the command
now refuses via error() before reading anything beyond its arg checks,
so an unconfirmed invocation is a guaranteed no-op on disk. --dry-run
still previews with no confirmation needed and is now documented in the
usage block (it was only documented for the sibling archive-quick).
--force keeps its narrow meaning — bypassing the TRUNCATED-scope and
unstarted-phase guards — and does not double as the mutation opt-in.
--confirm follows the existing `phases clear --confirm` idiom in the
same module.

complete-milestone.md's two invocations pass --confirm (the workflow has
gathered explicit user intent by that step). Existing tests get
--confirm appended — pre-change behavior is exactly confirmed behavior —
and a #3726 regression block covers: refusal + full-tree byte-identity
on both invocation forms, --force not satisfying the gate, --dry-run
still passing without confirmation, and --confirm proceeding. The
refusal tests fail against pre-fix code (negative control run).

Fixes #3726

* docs(#3726): document the --confirm requirement in CLI-TOOLS and COMMANDS

Cross-AI review of the fix diff (codex, pre-create) caught three shipped
doc sites still instructing the now-refused bare invocation: the
CLI-TOOLS.md milestone-complete synopsis + flag table, and COMMANDS.md's
two guard-override instructions (`--force` alone now refuses without
--confirm). Localized CLI-TOOLS copies already lag the English synopsis
(no --force/--dry-run either) and follow the translation pipeline, not
this fix.

* chore(#3726): set changeset fragment pr to 3774

* test(#3726): confirm-gate CI repairs — QA scenario caller + growth ack

Two CI reds from the --confirm gate, both this branch's own misses:

- tests/qa/scenarios/milestone-rollover.json invoked `milestone complete
  1.0 --force` as a JSON arg-array fixture — a caller shape the test
  sweep (which grepped runGsdTools/runSdkQuery in tests/*.cjs) never
  enumerated. Adds --confirm; the scenario's boundary-crossing contract
  is otherwise untouched.
- complete-milestone.md's +420-byte --confirm note trips the
  emitted-attribution growth ratchet. Acknowledged as a #3726 append to
  the existing complete-milestone.md entry in
  3409-unreachable-guard-arms.json (two ack sources may never name the
  same path, per that fragment's own precedent).

Local: lint-emitted-drift-ack ok; loop-walk.qa 115/115 green sandboxed.

* docs(#3726): CLI-TOOLS.md guard-override sentences say --force --confirm

Review Major 1: the truncated-window and unstarted-phase guard paragraphs
still told the reader to "Pass `--force` to override", which now refuses
(--force alone does not satisfy the confirmation gate), while the flag
table 470 lines later said the opposite. Mirror the docs/COMMANDS.md pair
so the file no longer contradicts itself.

* docs(#3726): synopsis renders --confirm and --dry-run as alternatives

Review Nit 1: `milestone complete <version> --confirm [--dry-run]` read as
"a dry run still needs --confirm", the opposite of AC 3. Render the pair
as `(--confirm | --dry-run)` in the CLI-TOOLS.md synopsis and the usage
docblock, and let the flag rows carry the rule.

* test(#3726): pass --confirm in base-added milestone fixtures; re-file the growth ack

Rebase onto next (26 commits) surfaced three tests the gate now refuses:
the #3685 write-flag contract pair in tests/milestone.test.cjs and the
`milestone complete` boundary fixture in tests/state-contract.test.cjs
all invoke the command bare. Each now passes --confirm (a mutating run is
exactly what they assert on).

The +420 byte complete-milestone.md growth ack rode on
3409-unreachable-guard-arms.json, which #3078 swept from next as fully
spent — hence the modify/delete conflict. Re-filed under a fresh fragment
named for this issue, never resurrecting the swept one.

* test(#3726): pin the present-but-falsy arm of the confirmation gate

Review Minor 1: the boundary triple covered absent and present but not
present-but-falsy. The gate is an exact-token match, so --confirm=false
and --confirm=0 refuse today — pinned (canonical + query forms, whole
.planning/ tree byte-identical) so a future `=`-aware or prefix-matching
parser cannot silently turn --confirm=false into a confirmed run of an
irreversible command.

* test(#3726): drop --confirm from dry-run-only invocations

Review Nit 2: --confirm was mass-appended to 14 pre-existing --dry-run
invocations that never needed it, so each stopped standing as incidental
proof that a preview needs no confirmation. Reverted to the pre-PR form;
the dedicated AC-3 test carries the explicit assertion.

* docs(#3726): sync the localized CLI-TOOLS synopsis with the confirm gate

REQ-I18N-02 (docs/features/internationalized-documentation.md) requires
translations to stay synchronized with the English source. The four
localized CLI-TOOLS.md guides still advertised a bare
`milestone complete <version>`, which now exits 1. Render the English
synopsis verbatim — `(--confirm | --dry-run)` plus the `[--force]` and
`[--archive-quick]` flags the translations had also fallen behind on.

* test(#3726): drop --confirm from the remaining preview-only invocations

Round 2 reverted the --confirm appends on --dry-run-only invocations in
tests/milestone.test.cjs, but four more sat in two files the sweep missed:
tests/milestone-archive.test.cjs (three) and
tests/milestone-window-single-owner.test.cjs (one).

Each is a preview run whose whole purpose is to document that a preview
mutates nothing, so `--dry-run ... --confirm` contradicted the semantics
the test exists to pin. Dropping the token restores each as incidental
proof that a preview needs no confirmation; the dedicated AC-3 test keeps
the explicit assertion.

No assertion added, relaxed, or removed — the change is four tokens.

* chore(#3726): migrate the emitted-drift ack from a fragment to a commit trailer

#3954 (ADR-3942) moved emitted-drift acknowledgments out of
tests/emitted-drift-acks/ and into git commit trailers, and the fragment
directory no longer exists on next. The reason this PR's fragment carried
moves verbatim into the Emitted-Drift-Ack-Growth trailer on this commit;
the fragment file is removed rather than resurrected.

Emitted-Drift-Ack-Growth: complete-milestone.md — #3726: +420 bytes (40186 -> 40606). The archive_milestone step's two `milestone complete` invocations now pass the required --confirm flag (the command refuses to mutate without it — the archive is irreversible), with a note explaining the flag and pointing at --dry-run for previews. Deliberate runtime-loaded workflow text for the new gate, not converter drift.

* fix(#3726): name --confirm in the version-required refusal

The documented arg-discovery path (gsd-tools.cjs top-level usage: invoke
the command without args and the error lists what is required) stopped at
`version required for milestone complete (e.g., v1.0)` — one required
argument short. Discovering --confirm took a second round trip through the
gate. The refusal now reads `… — and --confirm to mutate`, pinned by a test
that also asserts the version-less invocation leaves .planning/ untouched.

* test(#3726): pin the milestone complete docs against a silent regression

The changeset is `type: Fixed`, which the docs-required lint exempts, so
nothing in CI would notice a later edit that reinstated the bare-`--force`
override prose or dropped `--confirm` from the synopsis. Four tests in
tests/milestone.test.cjs now pin: the synopsis line in docs/CLI-TOOLS.md
and its four localized mirrors; the `--confirm` flag row; both
guard-override instructions in docs/CLI-TOOLS.md and docs/COMMANDS.md,
by guard name (a substring match on each instruction's `--force
--confirm` text); and — as an identity ratchet over the
milestone-complete sections — every `--force` sentence or clause that
lacks `--confirm`, so a new bare instruction in its own sentence or
clause fails whatever its wording. Named residual: a bare instruction
spliced into the same clause as a compliant one coalesces with it and
passes the ratchet; the by-name pins are what keep the four known
instructions from losing the pairing that way. The file is registered
in scripts/docs-guard-registry.cjs so the pin runs on the PR that
changes those docs, not only after merge.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-29 17:00:45 -04:00
Tom Boucher
b431ae9f0d fix(#3898): a separator-shaped line in ## Gaps is skipped, not made an entry (#4057)
* test(#3898): a spaced-hyphen thematic break in ## Gaps is not an entry (failing first)

* fix(#3898): a separator-shaped line in ## Gaps is skipped, not made an entry

splitGapsEntriesCore's opener regex (/^(\s*)-\s/) matched a spaced
hyphen thematic break, fabricating a gap named '- -' with result
'unknown' — surfaced by audit-uat as an outstanding finding that
cannot be cleared by editing any entry, because there is no entry,
only the separator the author wrote for readability. Five shapes were
affected (- - -, -, and wider/indented variants); the unaffected ones
(---, ----, * * *, ___) were safe only by accident — the path never
matched them, not because it understood breaks.

A line whose content after the opening marker is solely hyphens and
whitespace (with at least one further hyphen) is now skipped entirely —
neither an opener nor a continuation. Deliberately option 2 from the
issue, not a full thematic-break concept: a break does not close the
Gaps list, entries after it keep parsing, and the deliberately-frozen
byte-for-byte Gaps behavior changes ONLY for documents carrying such a
separator (which previously produced a phantom). A real entry whose
truth begins with a hyphen (- truth: "-5 error budget...") is untouched
— its remainder contains non-hyphen characters.

* fix(#3898): review fold-ins — span-contiguous skip, property coverage

The skip is narrowed to where the phantom came from: a separator-shaped
line BETWEEN entries (nothing open, or it would open a top-level entry).
One landing strictly inside a live entry (indent > baseIndent) folds
back as a continuation line, so entry lines and the GapsEntrySpan agree
byte-for-byte — the span invariant and the #3805 ack writer's identity
re-verification both hold (the review traced the unconditional skip to
a match_verification_failed refusal in that corner). Adds the parser-
convention property test (arbitrary hyphen counts/indents/spacings) and
a span-contiguity pin.

* chore(#3898): changeset fragment (pr number backfilled after PR creation)

* chore(#3898): backfill changeset PR number (4057)

---------

Co-authored-by: sim <sim@local>
2026-08-29 15:51:58 -04:00
Tom Boucher
44ddc6dc46 fix(#3865): init.* phase queries accept --phase <N> as the positional alias (#4054)
* test(#3865): init.* phase queries accept --phase as positional alias (failing first)

* fix(#3865): init.* phase queries accept --phase <N> as the positional alias

The phase-taking init.* queries read their phase token at args[2]
blindly: '--phase 60' made the literal '--phase' the phase, which the
locator can never resolve — the reported incident answered a well-formed
phase_found:false, plan_count:0 with exit 0 for a phase holding seven
committed plans (ADR-3473 §8.4's later strict validation turned the
paired form into a usage error instead, still not the alias).

normalizePhaseAlias (shared by all eight phase-taking handlers — the
issue's four plus phase-op, review, discuss-phase-assumptions, todos,
the same class) rewrites '--phase N'/'--phase=N' into the caller-owned
positional slot before flag parsing, so handlers and strict validation
see exactly the argv the positional form produces. A valueless --phase
is a usage error naming the flag. Any other flag-shaped args[2] now
resolves to undefined (the commands' designed use-the-current-phase
input) instead of passing flag text down as a phase name.

isFlagToken exported from command-arg-projection (single owner of the
flag-shape predicate).

* fix(#3865): review fold-ins — honest no-position-given comment, fail-closed throws

The helper's comment claimed flag-shaped args[2] resolves to 'the
commands' designed use-the-current-phase input' — no such cross-module
behavior exists: execute-phase/plan-phase/verify-work usage-error
'phase required', the find-based queries answer phase_found:false, and
todos drops its area filter. Reworded to what actually happens, so the
contract-grade comment cannot mislead a future edit. Adds the
parseNamedArgsOrExit-style throw after each error() call as a
fail-closed backstop against a returning fail().

* chore(#3865): changeset fragment (pr number backfilled after PR creation)

* chore(#3865): backfill changeset PR number (4054)

---------

Co-authored-by: sim <sim@local>
2026-08-29 15:22:34 -04:00
Tom Boucher
d9e906a744 fix(#3864): smart-entry classify() matches the verif* status stem (#4052)
* test(#3864): smart-entry classify must match the verif* status stem (failing first)

* fix(#3864): classify() matches the verif* status stem, aligning with normalizeStateStatus

"verified"/"verification" contain no "verify" substring, so the
exact-word branch never matched: a STATE.md declaring status: verified
fell through to situation "unknown" (or idle-stranded on a clean tree
with unpushed commits — differently wrong, which made it look
intermittent). state-document's normalizeStateStatus already matches
any verif* stem; the classifier now uses the same stem, so the two
owners agree on the invariant. verify_failed is tested earlier in
classify(), so failed verification still wins. Negative controls
(completed/executing/planning/paused) unchanged and pinned.

* test(#3864): review fold-in — the real handler-written verification status pinned both ways

'Phase complete — ready for verification' (state-document.cts's own
Status default) carries the verif stem but names completion: at 5/5 it
must stay complete (isComplete beats the stem), mid-project it must
route to verify-pending (pre-fix it fell through to unknown/idle-
stranded — a real behavior change beyond the literal verified repro,
sanctioned by the issue's own verification→verify-pending table).

* chore(#3864): changeset fragment (pr number backfilled after PR creation)

* chore(#3864): backfill changeset PR number (4052)

---------

Co-authored-by: sim <sim@local>
2026-08-29 15:07:14 -04:00
Tom Boucher
4048ba2e80 fix(#3860): Quick Tasks lookup accepts milestone-suffixed headings; schema-aware section selection (#4050)
* test(#3860): quick-tasks heading tolerance for milestone suffixes (failing first)

* fix(#3860): prefix-match Quick Tasks heading + schema-aware section selection

The exact-anchored predicate (^quick tasks completed$) never matched a
milestone-suffixed heading ('Quick Tasks Completed (v1.1+)'), so every
/gsd-fast append AND every reset failed with QUICK_TASKS_SECTION_ABSENT
before the columns were ever looked at — the reported table was already
canonical. A heading is a section label, not data: the predicate is now
prefix-anchored with a word boundary (Completedness still excluded),
hoisted to one shared isQuickTasksHeading beside QUICK_TASKS_SECTION_ABSENT
so the two call sites cannot drift.

Section selection is now schema-aware (the issue's deliberate-choice
ask): among ALL matching sections, the first whose table parses with a
recognized Quick Tasks schema wins — a legacy (v1.0) table first in
document order no longer shadows a usable (v1.1+) one below it. When no
section is usable, the FIRST is returned so the error names the real
problem (unrecognized schema) instead of a false 'no section'.

Fixes one over-assertion in the new tests (v1.1+'s own next ordinal IS 2;
pins v1.0 rows byte-identical instead).

* fix(#3860): review fold-ins — level-bounded bodies, splice guard, convention pin

Adversarial review caught that collectSections ends a candidate's body
only at the next MATCHING heading: an intervening '## Deferred Items'
table (canonical STATE.md layout — templates/state.md puts it after the
Quick Tasks section) would be swallowed into the body, and the append's
last-table-line scan would splice the quick-task row into that WRONG
table — silent corruption on the primary /gsd-fast path. Each candidate
is now re-collected through collectSection with an offset-precise
predicate, restoring the pre-#3860 level-bounded stop (next same-or-
higher heading) for both probing and splicing.

Also pins the both-schema-valid tie-break (document order, newest-on-top
— the layout the issue itself demonstrates; this layer has no
active-milestone signal) with a test, guards the Deferred Items layout
with a dedicated splice-target test, and updates resetQuickTaskRows's
doc comment to name the new pipeline.

* chore(#3860): changeset fragment (pr number backfilled after PR creation)

* chore(#3860): backfill changeset PR number (4050)

---------

Co-authored-by: sim <sim@local>
2026-08-29 14:52:24 -04:00
Tom Boucher
fd63889d1f fix(#3854): write normalization preserves tight multi-line lists (#4049)
* test(#3854): write normalization must preserve tight multi-line lists (failing first)

* fix(#3854): no blank before a bullet whose previous line is an indented continuation

_normalizeMd's 'separate a list from a preceding paragraph' rule inserted
a blank before any bullet whose previous line wasn't a bullet — but an
INDENTED CONTINUATION of the previous multi-line item also isn't a
bullet. Every .md write (phase.complete in the report, but any write
through platformWriteSync) therefore converted tight lists to loose
ones: +61 blank lines on the reporter's 1015-line ROADMAP, one before
each bullet following a wrapped item. Tight and loose lists render
differently, so this was a rendering change plus misleading diff noise;
one-shot (idempotent afterwards), which is why integrity checks on
headings/content passed.

The guard is the mirror image of the after-a-bullet rule two lines
below, which already excludes indented next lines. Paragraph→list and
heading→list separations — the rule's purpose — are pinned unchanged by
the new suite.

* fix(#3854): review fold-ins — ceiling tracks next's 281 + this branch's marker (282), header/require nits

The ceiling is not ratcheted but must track the tree: origin/next raised
it to 281 (sibling branch's marker file); this tree adds one more
(shell-command-projection-md-normalize), so 282/282. Also fixes the test
header's stale pre-rename filename and hoists the inline require to the
file's single import.

* chore(#3854): changeset fragment (pr number backfilled after PR creation)

* chore(#3854): backfill changeset PR number (4049)

---------

Co-authored-by: sim <sim@local>
2026-08-29 14:37:21 -04:00
Tom Boucher
400db94e02 fix(#3894): quick path honors workflow.research_before_questions; key resolves from global defaults (#4047)
* test(#3894): research_before_questions must resolve globally and order quick.md (failing first)

* fix(#3894): quick path honors workflow.research_before_questions; key resolves from global defaults

Two layers, one key. The quick workflow ran its discussion phase
before its research phase unconditionally — neither quick.md nor its
steps ever read workflow.research_before_questions, though the key is
documented, schema-registered, /gsd-settings-writable, and honored by
/gsd-discuss-phase and /gsd-new-project. A gray-area answer given
without research is then written to <quick_id>-CONTEXT.md as a locked
decision downstream agents are told not to revisit — an evidence-free
choice made unfalsifiable (the reporter's #3714 misresolution).

- quick.md Step 4 now carries the same research-before-questions check
  the two honoring paths make: when enabled, research-phase executes
  before discussion-phase; false/unset keeps the written order. Both
  sections stay section-manifest gated.
- src/config-loader.cts forwarded workflow.post_planning_gaps from
  ~/.gsd/defaults.json but silently dropped this key — same file, same
  nesting, one resolved and one didn't. Now forwarded with the same
  flat + nested-alias fallback shape, added to the resolution-keys
  lockstep canary and the #3532 shadowed-warning set (nested alias
  reporting generalized over both keys).

Emitted-Drift-Ack-Growth: quick.md — #3894: +Step 4 ordering rule (the research-before-questions check the discuss-phase and new-project paths already make); a real behavioral gate, not incidental bloat.

* fix(#3894): review fold-ins — gate the CONTEXT.md reference, colon slash-forms

- quick/steps/research-phase.md directed the researcher subagent to read
  <quick_id>-CONTEXT.md under DISCUSS_MODE with no existence hedge — but
  under the new ordering (research BEFORE discussion) the file cannot
  exist yet when the researcher is dispatched. The reference now says
  read-only-if-present with the #3894 reason; the alignment purpose
  still applies on the default ordering.
- quick.md's new rule used the hyphen slash forms (/gsd-discuss-phase,
  /gsd-new-project); source artifacts under gsd-core/workflows must
  author the colon form the install-time converters key on — the same
  file already uses /gsd:new-project and /gsd:quick elsewhere.

* docs(#3894): planning-config row names the flat CONFIG_DEFAULTS alias

config-field-docs requires every CONFIG_DEFAULTS key to appear in the
doc; the row documented the canonical namespaced form only. Adds the
same alias sentence post_planning_gaps's row carries, plus the #3894
quick-path note.

* chore(#3894): changeset fragment (pr number backfilled after PR creation)

* chore(#3894): backfill changeset PR number (4047)

---------

Co-authored-by: sim <sim@local>
2026-08-29 13:51:42 -04:00
Tom Boucher
192eb1dfbd fix(#3886): git commit timeout reported as commit_timeout; stale lock surfaced; 30s band (#4046)
* test(#3886): a timed-out git commit reports commit_timeout, not commit_failed (failing first)

* fix(#3886): git commit timeout reported as commit_timeout; 30s band; stale-lock surfaced

cmdCommit's git commit invocation did not distinguish a spawnSync
timeout from a real non-zero exit (#2608 fixed this for the staging
loop only): a slow pre-commit hook crossing the 10s cap was
SIGTERM'd mid-hook and reported as reason commit_failed with whatever
partial stderr git had flushed (in the reporter's case an incidental
CRLF warning), while the kill left a stale .git/index.lock blocking
the next attempt.

All three commit sites now check isSpawnTimeout before the
nothing-to-commit/ordinary-failure branches: cmdCommit reports
reason commit_timeout + timed_out:true and names the stale lock's
path (surfaced, not auto-deleted — deleting a lock a live git holds
is destructive; the caller recovers deliberately); the subrepo
counterparts do the same within their per-repo result / rollback
error. The commit calls also move to the 30s band the push call
already uses — husky+lint-staged alone idles ~4s on Windows before
any task runs.

* fix(#3886): review fold-ins — git-path lock resolution, shared band constant, executor contract row, precedence pin

- The stale-lock path is resolved via git rev-parse --git-path
  index.lock, never a literal .git/index.lock join (#3588 row 8's
  class: a linked worktree's .git is a FILE, so the literal path cannot
  exist there while the real lock — under <gitdir>/worktrees/<name>/ —
  blocks the next commit; this repo leans on linked worktrees).
- COMMIT_TIMEOUT_MS hoisted; all three sites and their messages build
  from it (the subrepo variant also regains the stdout fallback the
  primary site had).
- agents/gsd-executor.md's commit-result contract gains the
  commit_timeout row with the OPPOSITE retry advice from
  staging_timeout (remove the stale lock, then retry once) — an
  executor matching the doc previously had no handling for the new
  reason.
- Precedence pin: a timeout whose partial output contains 'nothing to
  commit' must still read as a timeout (branch-reorder mutant).

Emitted-Drift-Ack-Growth: gsd-executor.md — #3886: +commit_timeout row to the commit-result contract with the retry guidance OPPOSITE staging_timeout's (remove the stale lock, then retry once); the executor previously had no handling for the new reason.

* chore(#3886): changeset fragment (pr number backfilled after PR creation)

* chore(#3886): backfill changeset PR number (4046)

---------

Co-authored-by: sim <sim@local>
2026-08-29 13:27:14 -04:00
Tom Boucher
0bf778c352 fix(#3849): phase allocation counts numbers held by sibling git worktrees (#4042)
* test(#3849): phase allocation must skip numbers held by sibling worktrees (failing first)

* fix(#3849): phase allocation counts numbers held by sibling git worktrees

Both allocators (cmdPhaseAdd, cmdPhaseAddBatch) chose max+1 over numbers
gathered from ONE checkout — headers, bullets (add only), on-disk dirs.
Every sibling git worktree carries its own .planning/ on its own branch,
so a phase minted there was invisible and the same number was allocated
twice (the reported incident: two Phase 441s, one with six written
plans, surfaced a day late by human memory).

New shared horizon collectSiblingWorktreePhaseNums: one
git worktree list --porcelain, then per sibling — phase-dir names (the
cheap scan that would have caught the incident) and the WHOLE sibling
ROADMAP.md headers (a row can predate its dir; milestone-scoping would
be wrong — a number used under any milestone on another branch is
taken). Widen, never refuse: unreadable sibling / no .planning / not a
git repo / git unavailable each contribute nothing and allocation is
unchanged. Reuses isSentinelPhaseId and the allocators' own patterns.

Secondary (#1229 never reached batch): cmdPhaseAddBatch now also scans
roadmap bullets — a bullet-only 'Phase N' row was invisible to batch
allocation, exactly the condition #1229 was filed for.

Also fixes the two new tests' result-key access (output.phases, not
output.results).

* fix(#3849): review fold-ins — subprocess band, bounded test git, linked-worktree fixture

- execFileSync options now match the repo's git band (10s window,
  windowsHide, 4MiB maxBuffer) — a spurious 4s timeout silently reverted
  to the pre-fix collision.
- test git helper bounded (15s) per local/no-unbounded-spawn.
- new fixture: allocation FROM a linked worktree counts the main
  checkout — the incident's actual topology direction.
- fixture-setup rmSync carries the sanctioned lint-disable (setup, not
  teardown; cleanup() still owns directory removal).

* fix(#3849): exempt the sibling-worktree scan from the enumeration-drift guard

collectSiblingWorktreePhaseNums reads a SIBLING checkout's phases dir —
a different question from the cwd-scoped listMilestonePhaseDirs the
guard routes everything to (which cannot see another worktree's
.planning at all). Function-scoped, per ADR-3180 Decision 4(a): any
other re-derivation in phase.cts is still caught. The GREEN bench
caught the omission.

* chore(#3849): changeset fragment (pr number backfilled after PR creation)

* chore(#3849): backfill changeset PR number (4042)

---------

Co-authored-by: sim <sim@local>
2026-08-29 11:35:19 -04:00
Tom Boucher
331747ea99 fix(#3817): count the truncation remainder — display truncates, counting must not (#4034)
* test(#3817): audit-open counts must include the truncation remainder

* fix(#3817): count the truncation remainder — display truncates, counting must not

* chore(#3817): changeset fragment (pr number backfilled after PR creation)

* chore(#3817): backfill changeset PR number (4034)

---------

Co-authored-by: sim <sim@local>
2026-08-29 08:53:21 -04:00
Tom Boucher
213a2fff63 chore(#3813): delete the caller-less listMilestoneArchiveDirs seam; #1883 contract now pins the live path (#4029)
* test(#3813): pin the #1883 unreadable-milestones contract on the live planning-snapshot path

* fix(#3813): delete the caller-less listMilestoneArchiveDirs seam; #1883 contract now pins the live path

* chore(#3813): changeset fragment (pr number backfilled after PR creation)

* chore(#3813): backfill changeset PR number (4029)

* chore(#3813): docs-exempt marker — internal dead-code removal

---------

Co-authored-by: sim <sim@local>
2026-08-29 07:44:22 -04:00
Tom Boucher
b811ea16fc fix(#3807): advance-plan refuses an ambiguous multi-Phase Current Position (#4028)
* test(#3807): advance-plan must refuse an ambiguous multi-entry Current Position

* fix(#3807): refuse an ambiguous multi-Phase Current Position before advancing

* chore(#3807): changeset fragment (pr number backfilled after PR creation)

* chore(#3807): backfill changeset PR number (4028)

---------

Co-authored-by: sim <sim@local>
2026-08-29 03:50:03 -04:00
Tom Boucher
3a4c3cb83e fix(#3805): audit-uat honours the audit_acknowledged marker via the shared predicate (#4025)
* test(#3805): audit-uat must honour the audit_acknowledged marker

* fix(#3805): route audit-uat's UAT and VERIFICATION scans through the shared acknowledged predicate

* chore(#3805): changeset fragment (pr number backfilled after PR creation)

* chore(#3805): backfill changeset PR number (4025)

---------

Co-authored-by: sim <sim@local>
2026-08-29 02:51:54 -04:00
Tom Boucher
f4fefb0bef fix(#3804): audit-uat enumerates all three phase-archive layouts (#4022)
* test(#3804): audit-uat must see all three phase-archive layouts

* fix(#3804): enumerate all three phase-archive layouts (flat, workstream-archived, workstream-active)

* chore(#3804): changeset fragment (pr number backfilled after PR creation)

* chore(#3804): backfill changeset PR number (4022)

---------

Co-authored-by: sim <sim@local>
2026-08-29 01:14:58 -04:00
Tom Boucher
51ca9f39ba fix(#3801): register inline_plan_threshold in the defaults manifest and correct the docs (#4019)
* fix(#3801): register inline_plan_threshold in the defaults manifest (default 2) and correct settings-advanced

* chore(#3801): changeset fragment (pr number backfilled after PR creation)

* chore(#3801): backfill changeset PR number (4019)

* test(#3801): parse the defaults table with the shared markdown-table parser

---------

Co-authored-by: sim <sim@local>
2026-08-28 23:48:42 -04:00
Tom Boucher
83273f9642 fix(#3798): the profile closure follows command references into workflow spawn surfaces (#4009)
* test(#3798): tiered profiles must install the agents their workflows spawn

* fix(#3798): the profile closure follows command references into workflow spawn surfaces

* chore(#3798): changeset fragment (pr number backfilled after PR creation)

* chore(#3798): backfill changeset PR number (4009)

---------

Co-authored-by: sim <sim@local>
2026-08-28 18:02:44 -04:00
Tom Boucher
ab69b9ce56 enhance(#3987): guard slug re-derivation and the swallowed-precondition shape — §8.5 was guardable after all (#3999)
* feat(#3987): guard slug re-derivation, and record why the swallow shape cannot be guarded

Epic #3473's Decision 1 requires the wrong call site be UNREPRESENTABLE. #3984
measured that two of the nine §8 rules had no guard at all and recorded both as
"Shipped - test-covered". This closes one of them, proves the other cannot be
closed the same way, and corrects two false claims I merged yesterday.

1. §8.3 - scripts/lint-slug-derivation-drift.cjs.

   generateSlugInternal (src/core-utils.cts) is the canonical owner; #3883 removed
   11 inline copies. Nothing prevented a twelfth: no slug guard existed in
   scripts/ or eslint-rules/.

   The detector is STATEMENT-scoped and matches the shape the real copies took -
   one statement carrying BOTH .replace(<negated class>, '-') and
   .replace(/^-+|-+$/, ''). Statement scoping is what buys the precision: the
   loose LINE-level form yields 18 hits with 7 unrelated, a material
   false-positive rate. Measured on the tree: 5 flags, 2 TRUE, 3 SANCTIONED,
   0 FALSE.

   The three sanctioned sites are allowlisted with a reason each, following
   lint-phase-enumeration-drift's form rather than a bare denylist. The owner
   itself is listed explicitly even though it escapes by construction - an
   implicit escape is a latent bug, and the next person to touch line 192 would
   not know the guard depended on it.

2. Both TRUE positives were live defects, not style.

   scripts/qa-smell-ratchet.cjs reproduced the canonical formula including the
   60-cap but trimmed BEFORE truncating - the #2849 bug - and never
   transliterated. The divergence is total, not cosmetic:

     canonical  "privet-mir-privet-mir-privet-mir-privet-mir-privet-mir-prive"
     inline     "tail"

   Cyrillic collapsed to nothing and only the ASCII remainder survived, so the
   ratchet was keying on wrong identifiers for any non-ASCII input.

   tests/planning-inspect.test.cjs carried a helper whose comment claimed parity
   with getPhaseDirFromPhaseId. That function now transliterates; the helper did
   not, so the test asserted against a stale formula while looking correct. Both
   now route through the seam.

3. §8.5 - measured, and deliberately NOT shipped.

   A candidate detector (swallowing catch + errno-retry-set test in the same
   function) gives 26 flags across 11 functions: 0 TRUE, 26 FALSE. Every one is
   best-effort unlink/rm/close cleanup, lost-rename-race backoff, or a deliberate
   null fallback. The file-scoped variant is worse at 71.

   Worse than the noise: the only known true instance was removed by #3885, so
   there is NO POSITIVE CONTROL - the guard cannot be shown capable of failing,
   which this repo requires of every drift guard. Shipping it would add a guard
   nobody can trust and nobody can test.

   The ADR now records the measurement and the reason, keeps §8.5 at
   "Shipped - test-covered", and points at the #1884 regression test as what
   actually enforces it. An honest "not detectable at acceptable precision" beats
   a guard that only ever passes.

4. Two claims I merged into the ADR yesterday were wrong.

   §8.9 said 17 of 19 subsumed children have a test citing their issue number,
   and that #3364 and #3812 have none. Both halves are false, and the claim came
   from a NUMBER-GREP - inside an amendment whose own subject is that a text match
   is not a fact.

     #3364 IS cited: tests/runtime-marker-resolution.test.cjs:107,
       T3 installMarkerResolvesWhenEnvAndConfigAbsent_3897 (#3364), asserting at
       :115-119.
     #3812 IS covered: tests/gen-state-md-docs.test.cjs:374, asserting at :382.

   Corrected to 19 of 19.

   #3812 does carry a real finding, though a different one: it is PARTIALLY
   DELIVERED on a CLOSED issue. The shipped fix declares cardinality for
   frontmatter keys, but #3812's stated acceptance was about the
   ## Current Position BODY section, and docs/reference/state-md.md:196-208 still
   has no normative single-valued/overwrite sentence and no pointer to
   ## Performance Metrics for history. Recorded in the ADR and left for #3812 to
   re-open - fixing it here would bury a scope question inside an unrelated PR.

Note on B6: this ADDS a guard, and B6 said the net count must fall. #3951 already
amended that clause - a guard ledger is a claim about COVERAGE, not count - which
is what makes adding this one honest rather than contradictory.

Verified: the guard flags 0 on the fixed tree, and PROVES IT CAN FAIL - a fresh
inline copy planted in src/ makes it exit 1 naming the exact statement. All three
sanctioned sites were confirmed exempt BY the allowlist, not by accident of the
pattern, by re-attributing each to a non-exempt path and watching it flag.
build:lib, lint and lint:ci all exit 0.

Refs #3987

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3987): add the changeset fragment

Doc-only, so it carries forward from the verified sha rather than costing a
second matrix run.

Refs #3987

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3987): §8.5 IS guardable — I was wrong, and the guard found a live defect

Two orthogonal reviews. The correctness review overturned my central judgment,
and it was right.

1. I concluded §8.5 was "not detectable at acceptable precision" and recorded
   that in the ADR. False.

   My evidence was 26 flags / 0 TRUE / 26 FALSE. The reviewer pointed out what I
   had not: all 26 false positives are CLEANUP verbs - rmSync 54, unlinkSync 43,
   closeSync 17, chmodSync 12 - and the obvious narrower predicate was never
   tried. A swallowed cleanup is legitimate best-effort. A swallowed CREATION is
   a precondition silently lost, which is exactly the #1884 shape.

   Measured properly, in three stages:
     swallowing catch                                     911
     + try-block calls a CREATION verb                     24
     + enclosing function references a *_ERRNOS set         0

   0 flags, 0 false positives. The `*_ERRNOS` naming key is empirically total -
   all 10 retry/tolerate sets in src/ follow it.

   My second claim was worse. I wrote that no positive control exists because
   #3885 removed the only true instance, so the guard "cannot be shown capable of
   failing". That is self-refuting: this very PR's slug guard proves-it-can-fail
   on a synthetic tree, and the pre-#3885 blob is available as exactly such a
   fixture. It is now the control, and it works in both directions - the rule
   flags 0c43d853e^:src/planning-workspace.cts at line 210, the line the fix
   commit's own message cites, and reports zero on the post-fix code.

   I stopped at the first negative result on the option that meant less work.

   Shipped as eslint-rules/no-swallowed-precondition.cjs, wired into the existing
   src/**/*.cts ESLint block rather than a scripts/lint-*-drift.cjs: no script in
   scripts/ requires typescript/espree/acorn, and scripts/ ships to consumers, so
   a .cts-parsing standalone guard would add a devDep at consumer runtime. The
   ESLint block already parses .cts for free.

2. The guard immediately found a live defect of the same class.

   src/capability-lock.cts swallowed a mkdirSync on the lock directory, then
   acquireLock classified the follow-on failure as `code !== 'EEXIST' → return
   null`. A real EACCES/EROFS makes openSync(lockPath,'wx') fail ENOENT, which is
   not EEXIST - so a fatal filesystem error was laundered into "lock
   unavailable". Same defect as #1884, different laundering target.

   Fixed the way #3885 fixed #1884: the creation failure propagates. Regression
   test proven fail-first by hand - with the fix stashed, EACCES was laundered to
   null; restored, it throws.

   The strict rule does NOT catch this shape (its errno classification is an
   inline literal, not a named set). The rule is deliberately left strict: the
   broadened form had 2 false positives - capability-lock.cts:408, the deliberate
   EEXIST steal protocol, and commonjs-marker.cts:131, which returns a distinct
   documented outcome. The gap is noted in code rather than papered over with a
   noisy predicate.

3. The security review found the slug guard's exemption FAILED OPEN.

   currentFunction was never reset, and only a column-0 `function` declaration
   updated it, so exemption bled from an allowlisted declaration to the next one.
   generateSlugInternal exempted 50 lines for an 11-line function. A
   re-derivation planted anywhere in that window was silently exempt - the same
   fail-open shape that produced a blocker in #3897, and an allowlist is a
   SUBTRACTION so a mismatch fails open by construction.

   Extent is now tracked by real brace depth, and a test plants a violation after
   each allowlisted function's real closing brace and asserts it IS flagged.

4. Also from the security review: the guard was a CI-DoS and narrower than I
   claimed.

   Its unbounded [^\]]* was re-scanned from every `.replace(/[^` start: 54.3s on
   a 1.28MB line. It imported MAX_REGEX_LITERAL_LEN and never called
   readRegexLiteralAt - the bounded tokenizer that exists for exactly this. Now
   routed through it with a 2MB file cap: ~200ms.

   15 of 25 genuine re-derivations evaded. Widened to catch replaceAll, {1,},
   \s*-wrapped classes, escaped ], literal new RegExp(...), five trim spellings,
   .split().join(), and multi-line .replace( args - still 0 false positives.
   Two forms still evade and are documented as deliberate gaps with negative
   tests: the two-statement/temp-var form and new RegExp built from a variable.
   Both need data flow, and guessing at it is how a guard becomes noisy.

   Also fixed: // inside a string truncated the line, a ; inside the collapse
   regex split the statement (a one-character bypass), and SCAN_EXT omitted
   .mjs/.tsx/.jsx.

5. A regression I introduced, caught by the same review.

   qa-smell-ratchet.cjs top-level-required a build output that is not
   git-tracked, so the script hard-failed MODULE_NOT_FOUND before build:lib -
   including for --help, which previously had no build dependency. The require is
   now lazy at the point of use.

6. Four of my own tests were vacuous or weak.

   T9's input yielded an identical string under the buggy formula, so it passed
   on the implementation it was meant to catch. T12 compared maxLen null vs 60 on
   an 18-char name, where they agree trivially. T9-T12 all asserted
   generateSlugInternal directly, so they would pass unchanged if both call-site
   fixes were reverted. And prove-it-can-fail was scoped to scanRepo, never the
   CLI - dropping main()'s exit-code line would have kept every row green.

   All rewritten with discriminating inputs, per-call-site rows that red when the
   fix is reverted, 59/60/61 boundaries, an entirely-non-alphanumeric row, and a
   CLI row asserting the real subprocess exit code and both sanitizeForReport
   sites.

Verified: both guards flag 0 on the tree and both prove they can fail. The
swallow rule's control is confirmed in both directions - pre-#1884 shape flagged,
post-#3885 shape clean. build:lib, lint and lint:ci all exit 0.

Refs #3987

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3987): record that §8.5 IS guardable, and correct a correction that made a ledger worse

Three ADR corrections, two of them to text this branch wrote hours ago.

§8.5 advances to Enforced. Its previous entry said the rule was not detectable
at acceptable precision. That was wrong twice: the 26 false positives were
uniformly CLEANUP verbs, which is a reason to narrow the predicate rather than
abandon it, and the claim that no positive control exists was self-refuting - the
pre-#3885 blob is available as a fixture and this repo's own guards prove-it-can-
fail on synthetic trees. Narrowed to creation verbs plus a *_ERRNOS reference:
911 -> 24 -> 0 flags, 0 false positives, control confirmed in both directions.
The entry keeps the wrong reasoning visible, because a high false-positive count
being evidence the predicate is wrong - not evidence the rule is unguardable - is
the transferable part, and the first negative result is most seductive when it is
also the answer that means less work.

§8.9's correction is itself corrected. The original 17-of-19 claim was CORRECT
for the predicate it stated; this branch silently swapped cited -> covered and
declared 19 of 19. #3812 appears in zero test files. Changing what a word means
to make a ledger read better is a worse failure than the miscount it claimed to
repair. Both predicates are now reported separately - 18 of 19 cited, 19 of 19
covered - because §8.9 asks for a test NAMING each child, so 18 is the number
that answers it. #3812 is also re-opened for real, rather than the first draft's
promise that it could be.

§8.3 stays Shipped - test-covered rather than advancing. The slug guard catches
the copy-paste class and a dozen variants, but two forms still evade by decision
(temp-var split, new RegExp from a variable) because both need data flow. Naming
them keeps the status honest: the wrong call site is much harder to write, not
unrepresentable.

Closes #3987

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3987): backfill changeset pr number

Refs #3987

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3987): replace my own wall-clock assertion, and close the guard that let me write it

CI went red on ubuntu shard 2/3. The failing test was mine, and the failure was
the test, not the code.

  a 1.28MB line ... scans in well under a second (was 54.3s pre-fix)  7368ms

It asserted ELAPSED TIME. ~200ms locally, 7.4s on a shared CI runner. The bound
introduced for the MAJOR-2 DoS fix works - 7.4s against a 54.3s pre-fix baseline
is the fix doing its job - but an absolute wall-clock threshold on shared
hardware is a race, not an assertion. CLAUDE.md says so directly: "Clock Seams:
Do not assert on wall-clock time." I wrote the anti-pattern the project bans, in
a PR about guards.

Raising the threshold would only move the flake. The row now asserts a
DETERMINISTIC bound instead: an instrumentation seam on drift-scan.cjs counts
readRegexLiteralAt calls and characters examined, and the test asserts
charsExamined stays under an absolute ceiling. Measured on the same 1.28MB
fixture: 120,000 calls, 48,000,000 chars - two orders under the ceiling. The
pathological fixture is kept; only the thing being asserted changed.

Proven to still discriminate: with MAX_REGEX_LITERAL_LEN raised to simulate the
unbounded pre-fix behavior, the same fixture does not complete in 120 seconds,
versus ~0.3s bounded. It is a real regression test, not a tautology.

Then the second half, which is the same defect class as the rest of this PR.

  eslint-rules/no-elapsed-assertion.cjs matched only the EXACT identifiers
  ^(elapsed|duration|took|ms)$.

I used `elapsedMs`. It evaded the rule entirely. tookMs, durationMs,
elapsedTime and msElapsed evade the same way. A guard that cannot see the
violation it exists to catch is exactly what this PR is about - it just happened
to be an existing rule rather than one of the two I came here for, and it was
found because I committed the violation it should have blocked.

Widened to /^(?:elapsed|duration|took|ms)(?:[A-Z]\w*)?$/ plus a narrow
start/endMs delta pair. Deliberately NOT a blanket *Ms suffix: a first draft did
that and produced 2 false positives on `timeoutMs` in
plan-phase-stall-detection, which is a configured timeout and not a measurement.
Verified negative on params, items, forms, terms, dirnames, timeoutMs,
cacheTtlMs and staleAfterMs.

Measured over the five files carrying camelCase timing identifiers: 0 true
positives beyond my own, so nothing else needed rewriting. The rule's own test
file gains a row asserting `elapsedMs` flags, proven to fail against the
pre-widening rule - the same prove-it-can-fail standard both new guards in this
PR are held to.

Refs #3987

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3987): a comment I added leaked a Claude reference into every runtime install

The runner went red with 4 failures in tests/install.test.cjs:

  Leaking: .hermes/scripts/lib/drift-scan.cjs
  Leaking: .qwen/scripts/lib/drift-scan.cjs

The instrumentation seam added for the deterministic bound carried a comment
naming CLAUDE.md as the source of the no-wall-clock-assertions rule. scripts/
SHIPS to consumers, so that comment was installed verbatim into hermes and qwen
trees, and the install suite scans for exactly this - a Claude-specific reference
reaching a non-Claude runtime.

The rule is real and worth citing; the filename is not portable. The comment now
says "this repo's test rules" and states the rule inline, which is what a reader
of an installed tree actually needs anyway.

Worth noting what caught it: not lint, and not the two guards this PR adds - the
install suite's full-tree scan, which exists precisely because a shipped file is
read by runtimes that have never heard of CLAUDE.md. Same lesson as the rest of
this PR from the other direction: the check that matters is the one that can see
the surface where the defect actually lands.

Refs #3987

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3987): a test fixture swallowed 46 git exit codes and produced a silent false negative

CI red on ubuntu shard 3/3:

  tests/health-validation.test.cjs:2029
  expected exactly one W024, got [{"code":"W006", ...}]

Not caused by this branch, and the evidence is decisive rather than a hunch: the
SIBLING test at :2039 builds the IDENTICAL fixture with the identical
commitsAhead and asserts the same thing, and it PASSED in the same process, same
file, same run. Same input, both outcomes - which rules out logic, ordering,
sharding and environment, and leaves a per-invocation nondeterministic failure
inside one fixture build.

The mechanism is an unchecked exit code, 46 times over. The W024 fixture performs
~46 runGit spawns and never checks a single one. runGit returns failures as DATA
and never throws, so one silently-failed `git commit` yields 19 commits instead
of 20, or a silently-empty `git rev-parse HEAD` yields a blank state_head. Either
drops readStateHeadFreshness below the advisory threshold, W024 never fires, and
only W006 remains.

Reproduced exactly: 20 commits -> ["W006","W024"]; 19 -> ["W006"]; blank
state_head -> ["W006"] - byte-identical to the CI assertion dump.

The arithmetic is what hid it. At threshold-1 and threshold+1 a lost commit still
produces the asserted answer; only the exactly-at-threshold cases sit one commit
from a false negative. Two of the seven tests are in that position, and CI hit
one. That is why it had never been seen before, and why it surfaced now: this
branch adds three test files, which reshuffles the cost-weighted shard partition
and moved this file into a chunk where the latent flake fired.

My files were checked as suspects first and cleared: all fixtures mkdtemp-unique,
no process.chdir, no .planning/ writes, no git spawns, and node --test gives
per-file process isolation regardless.

Fixed at the cause, not the symptom. A mustGit wrapper throws on a non-zero exit
with the command, exit code and stderr, and all nine call sites route through it.
The fixture now asserts its OWN preconditions before the assertion under test
runs - the seed head is non-empty, and `git rev-list --count <seed>..HEAD` equals
the requested commitsAhead - so a fixture that did not build what it claims fails
loudly as a FIXTURE ERROR naming got-versus-asked, instead of quietly handing a
weaker input to the assertion.

Proven: dropping one commit now raises
  FIXTURE ERROR: requested commitsAhead=19 but git rev-list --count reports 18
where it previously produced a silent ["W006"] pass-for-the-wrong-reason. 64/64
tests in that block pass unperturbed.

Deliberately NOT done: no threshold change, no retry, no loosened assertion, no
skip. The assertion was correct; the input was silently wrong.

Worth naming, because it is the same shape from the other side: this PR ships
eslint-rules/no-swallowed-precondition.cjs, whose entire subject is a swallowed
precondition failure being laundered into a plausible downstream outcome. This
fixture is that defect in test code - the swallowed git failure was laundered
into a legitimate-looking "W024 did not fire". The rule does not cover test
fixtures, so the connection is noted at the fix site rather than enforced.

Refs #3987

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3987): two tests wrote to committed files; the shard packing decided when that mattered

CI red on windows-latest shard 1/3 only:

  "gen-exit-code-registry: CLI" > "a --write run redirected to a tmpdir leaves
  every committed artifact untouched"
  AssertionError: hooks artifact must be untouched

The Linux runner passed the same sha at 40425/40425. It is Linux-only, so a
Windows-scheduling defect is structurally invisible to it.

Root cause, established by measurement rather than inference.

tests/cli-exit.test.cjs appended a corruption marker to the REAL COMMITTED
hooks/lib/exit-code-registry.js, held it corrupted across a full subprocess, and
restored it in a finally. tests/exit-code-registry.test.cjs reads that same real
file before and after its own subprocess and asserts byte equality. If it samples
while the other test holds the file corrupted, it fails. The landmine is
pre-existing, from 2ea5efc15 (#3911).

What this branch changed is WHEN the two run together. scripts/run-tests.cjs
shards by cost-weighted LPT over the sorted unit list, so adding three test files
repacks the bins:

  merge-base c3e667df3 (838 files): cli-exit -> shard 1, exit-code-registry -> shard 3
  HEAD       03b342601 (841 files): BOTH in shard 1, same argv chunk, one
                                    node --test process, concurrent

Co-location is necessary but not sufficient - Linux shard 1/3 also had both and
passed. Windows loses because TEST_CONCURRENCY defaults to 2 there against 4
elsewhere, spawn cost is ~10x, and the sibling corruptor holds one of only two
slots through a ~90s tsc compile. That turns a sub-second overlap into seconds.

Not a path-separator or case-sensitivity issue, and not CRLF - .gitattributes
pins * text=auto eol=lf. Redirection was not at fault either: ensureScriptsOut
derives all five --out flags correctly and gen-exit-code-registry.cjs honours
them with no __dirname escape.

Fixed at the cause: no test writes to a committed file any more. Both corruptors
now copy to a mkdtempSync tmpdir, corrupt the COPY, and point the generator at
it. Repinning or reordering the shards would have turned CI green while leaving
the landmine armed for the next reshuffle.

That required closing an inconsistency between two sibling generators.
gen-exit-code-registry.cjs already accepts
--out/--scripts-out/--hooks-out/--dts-out/--sh-out and honours them under
--check; gen-hooks-cli-exit.cjs hardcoded OUTPUT_PATH and had no flag surface at
all, so its corruptor could not be redirected anywhere. It now takes --out in the
same style, honoured by both --write and --check, and is a no-op when absent -
verified: a bare --check on the default path still exits 0.

ensureScriptsOut moved to tests/helpers/exit-code-artifact-flags.cjs and both
test files import it. Hand-rolling a second copy of the flag derivation would
have been a re-derivation of exactly the kind this PR ships a guard against.

Verified: both tests still detect corruption (proven by defeating the check and
watching them red, with a positive control showing an uncorrupted copy exits 0);
SHA-256 of hooks/lib/exit-code-registry.js and hooks/lib/cli-exit.js identical
before and after running both rewritten bodies, and git reports nothing under
hooks/ modified - that is the property that was violated. A repo-wide search for
the corrupt-then-restore-in-finally shape against hooks/ found no other
instances.

One detail worth recording: the tmpdir test keeps --declaration pointed at the
real committed declaration rather than copying it, because the generated banner
embeds path.relative(REPO_ROOT, declarationPath) - copying it would produce a
false drift unrelated to the injected corruption. The declaration is read-only on
that path and never written.

Refs #3987

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 14:40:02 -04:00
Tom Boucher
dd4f179672 feat(#3970): per-task external-tracker content-resolution seam (#4000)
* feat(#3970): per-task external-tracker content-resolution seam

Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute
plus a new optional `taskContentResolver` capability-manifest field let a
capability resolve a task's action/verify/acceptance-criteria/read_first/done
content from an external issue tracker instead of PLAN.md's inline body.

- src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId`
- src/task-content-resolution.cts: new leaf module — split/find/build/resolve,
  with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed
  resolution, never a silent fallback to possibly-stale inline text
- src/task-command-router.cts: new `task resolve-content --plan --task-id --raw`
  CLI verb wiring the module into a real process exit code
- gsd-core/bin/lib/capability-validator.cjs: validates the new
  `taskContentResolver` manifest field (feature-role only, cross-capability
  trackerPrefix uniqueness)
- gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md,
  docs/reference/capability-manifest.md: wire the seam into the per-task loop
  and document it as a new `execute:task` point outside the existing
  contribution/step/gate vocabulary (unconditional in autonomous mode)

Closes #3970

* fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard

Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646
Phase 1) found three defects:

1. execute-plan.md's task-content-resolution bullet fired on any
   tracker-id-bearing task with no check that it wasn't type="checkpoint:*",
   contradicting ADR-3646 Decision 1 (a checkpoint task must never enter
   resolve-content). plan-document.cts already parses trackerId: null
   unconditionally for checkpoint tasks; only the workflow prose needed the
   fix, so the bullet now explicitly excludes checkpoint tasks.

2. task-content-resolution.cts's parseResolverDeclaration accepted any
   non-empty trackerPrefix with no grammar check, while capability-
   validator.cjs's KEBAB_RE enforces kebab-case at install time — a
   Generative Fix Divergence gap. Added the same grammar (as a literal
   regex, documented as intentionally not shared across the .cts/.cjs build
   boundary) plus a parity test asserting the two surfaces agree across a
   valid/invalid trackerPrefix table.

3. task-command-router.cts's routeResolveContent path-traversal guard on
   --plan had zero test coverage. Added a test exercising a
   ../../../etc/passwit-shaped path and asserting the USAGE rejection names
   the offending path.

* fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs

Two findings caught by an isolated security-review pass on the task
content resolution seam:

- ResolverFailedError/ResolverMalformedOutputError embedded raw,
  unsanitized subprocess stderr/stdout (attacker/model-influenced via
  the tracker-id argv token) into .message. A hostile or buggy resolver
  could smuggle a newline plus a forged "Error: " line, or terminal
  escape sequences, into a diagnostic io.cjs's error() writes verbatim
  to stderr. Fixed at the constructor (task-content-resolution.cts) via
  io.cjs's existing formatDiagnosticToken(), so every caller of
  resolveTaskContent gets a safe .message by construction.

- capability-validator.cjs's validateTaskContentResolverFields had no
  upper bound on taskContentResolver.invoke.timeoutMs, letting a
  manifest declare an effectively unbounded value and defeat the
  "bounded subprocess" design intent. Added a 120000ms ceiling specific
  to this field, without touching the shared isPositiveIntegerMs()
  helper (still used unbounded by the reviewer lane's timeoutFloorMs
  and probe timeoutMs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion

gsd-test (remote dockerized matrix) came back red with 5 failures on this
PR; all five are real defects, fixed here.

- tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's
  execute-plan.md entry pointed at line 415, which ffc190df4's
  checkpoint-exclusion caveat (added near line 221) shifted down by one
  line. The actual "validated downstream by gsd-tools uat
  classify-coverage" descriptive mention now sits at line 416. Updated
  the allowlist entry's line number to match.

- tests/task-command-router-resolve-content.test.cjs: the path-traversal
  test asserted the outside-project-scope diagnostic against the thrown
  ExitError's own .message. io.cts's error() (ADR-3889) writes its
  human-readable message to fd 2 via writeAllSync and then throws a bare
  `new ExitError(1)` with no message argument — by design, so the
  exception carries no duplicate text and the thrown ExitError's message
  defaults to "process exit 1" (cli-exit.cts's ExitError constructor).
  Root cause was the test, not the source: task-command-router.cjs's
  outside-project-scope rejection already calls error() correctly and the
  diagnostic text is genuinely emitted, just on fd 2, not on the
  exception. Fixed the test to capture fd-2 writes (mirroring
  tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this
  same file's own captureStdout for fd 1) and assert against the captured
  stderr text instead of err.message. This was masked locally because a
  manual `node -e` sanity check that only inspects the caught exception's
  .message cannot see what the real node:test run actually failed on.

Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3970): backfill changeset PR number (pr:0 -> pr:4000)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 13:17:04 -04:00
Tom Boucher
2012e8cc7f fix(#3781): span-carrying heading walk unblocks acknowledge on heading-shaped deferred items (#3998)
* test(#3781): heading-shaped deferred entries must be acknowledgeable

* fix(#3781): span-carrying heading walk unblocks acknowledge on heading-shaped deferred items

* test(#3781): table fixture counts the row; BLOCKER 1 updated to the supported contract

* chore(#3781): changeset fragment (pr number backfilled after PR creation)

* chore(#3781): backfill changeset PR number (3998)

---------

Co-authored-by: sim <sim@local>
2026-08-28 10:21:57 -04:00
Tom Boucher
9f1996b8f9 fix(#3775): ack matches exactly the status-line case shapes the reader reads back (#3989)
* test(#3775): bare Title-case status lines must ack through the reader-visible path

* fix(#3775): match exactly the status-line case shapes the reader reads back

* chore(#3775): changeset fragment (pr number backfilled after PR creation)

* chore(#3775): backfill changeset PR number (3989)

---------

Co-authored-by: sim <sim@local>
2026-08-28 08:59:55 -04:00
Tom Boucher
3a6c0412a9 enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4) (#3976)
* enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4)

Extends ADR-1703's portability rule catalog with a second production-runtime
rule: it flags an exact-case read of a Windows case-varying environment
variable (PATH, PATHEXT, ComSpec, USERPROFILE, TEMP, TMP, APPDATA) off any
receiver that is not process.env itself, matched via an env-shaped-receiver
check to avoid colliding with ordinary `.path`-named properties elsewhere in
the tree.

Exports the seam's private `_envGet` as `envGet` so the rule's remediation
message names a real helper, and fixes the one pre-existing violation the
tightened rule found (`src/runtime-hooks-surface.cts`'s `env.APPDATA` read).

Closes #3624

* fix(#3624): extractStaticName recognizes non-computed Literal destructuring keys; add missing accessor-call test case

Review findings from the code-review + isolated-adversarial passes:
- extractStaticName only matched non-computed Identifier keys, so a
  destructuring like `const { 'PATH': v } = opts.env;` (the issue's own I8
  acceptance case) silently evaded the rule. Widened to accept a Literal key
  regardless of computed, which is safe for MemberExpression too (its
  non-computed property is always an Identifier by grammar).
- Added the missing RuleTester valid case for "a case-insensitive accessor
  call" (envGet(env, 'PATH')) from the issue's Done-when checklist.

* docs: backfill changeset PR number for #3624 (PR #3976)

---------

Co-authored-by: sim <sim@local>
2026-08-28 08:49:03 -04:00
Tom Boucher
4f32209f78 enhance(#3267): reduce handleEvaluate complexity below refactor-trigger's own threshold (#3978)
* fix(#3267): reduce handleEvaluate complexity below refactor-trigger's own threshold

handleEvaluate scored 26 (then 21 after later #3261 commits) against the
complexity-triggered-refactor feature's own default threshold of 15. Extracts
the read-and-analyze loop (analyzeTouchedFiles) and the artifact/baseline/
ledger write path (finalizeEvaluation) into named helpers, per the issue's
suggested direction. Behavior-preserving: every existing test in
tests/refactor-trigger-cli.test.cjs is unchanged, and every degrade-path
reason code (REFACTOR_INVALID_PHASE, REFACTOR_GIT_UNAVAILABLE,
REFACTOR_NO_TOUCHED_FILES, REFACTOR_FILE_UNREADABLE,
REFACTOR_ANALYZER_UNSUPPORTED, REFACTOR_ANALYZER_UNPARSEABLE,
REFACTOR_BASELINE_WRITE_FAILED, REFACTOR_STRICT_NOT_ENFORCING) keeps its
current value and emission path.

The four complexity-trigger.cts lexer functions (scanFunctions,
stripLiterals, skipTypeExpr, skipGenericParamList) are deliberately left
untouched, per ADR-1953 D5 — they are a hand-rolled lexer state machine,
densely branchy by construction, and refactoring them to lower the metric
would be exactly the "split a coherent function to satisfy a metric"
behavior D5 exists to prevent.

Closes #3267

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs: backfill changeset PR number for #3978

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 08:12:27 -04:00
Tom Boucher
d24e22b156 enhance(#3912): gsd-tools declares outcomes, pinned at v1 (#3983)
* enhance(#3912): gsd-tools declares outcomes, pinned at v1

ADR-3889 §4. Phase 6 already moved error()'s terminator onto the seam, so what
remained was the declaration — and the pin that makes it invisible today.

The census corrected two documented figures before any code changed.
ERROR_REASON has exactly 25 members (the ADR and epic were right; an earlier
note of mine claiming 23 was wrong and is corrected). And output({error}) is
**64 sites across 9 files, not the 60 ADR-2980 ratified** — the module shape
holds but the total drifted +4: frontmatter 7 not 6, phase 4 not 2, roadmap 3
not 2. That matters because this phase's criterion demands the pin be asserted
over the enumerated population rather than sampled; asserting over a stale 60
would leave four sites unpinned while claiming full coverage, which is the
shape of failure this epic exists to remove.

The issue does not state the fact that shapes the design: output() never
touches the exit code. Confirmed by reading it — it writes fd 1 and returns.
So a declared outcome for those 64 sites had nowhere to be READ. The mapping
was never the work; wiring somewhere for the declaration to land was.

The seam already existed twice over. cli-exit.cts holds two globalThis-Symbol
cells, each because the module is emitted to three locations and a module-level
`let` would let instances disagree, and runMain already maps a code returned by
main(). A third cell inherits that solution. output() records DEGRADED for any
{error} payload — key-order agnostic, which is exactly why the "42 sites"
figure undercounts — and runMain projects the cell only when main() returns
nothing, so an explicit return still wins.

error() maps its reason through a table over the closed 25-member enum, leaving
all 278 call sites untouched; 226 of them pass no reason at all. The version
gate lives in error(), NOT in projectOutcome: registered names are
version-invariant there, so mapping a reason straight through would make USAGE
project to 64 under v1 and break the pin on its first line. projectOutcome is
left exactly as Phase 2 shipped it, DEGRADED's 0/80 asymmetry included.

Proven rather than asserted. v1 is byte-identical across three real CLI paths —
config-get plain, config-get --json-errors, and an output({error}) path —
matching exit code and exact bytes against the pre-change build. Under
GSD_EXIT_CONTRACT=v2 the same commands now exit 66 (CONFIG_KEY_NOT_FOUND ->
NO_INPUT) and 80 (DEGRADED), both looked up through the registry. An
anti-vacuity test pins that v1 and v2 genuinely differ for at least one reason,
because without it a mapping where everything projects to 1 under both versions
would satisfy every other assertion and the declaration would be theatre.

A1 iterates all 25 enum members and A3 asserts over the measured 64-site
population, so a 26th reason or a 65th site fails until it is given a mapping —
the drift guard this phase needs, given ADR-2980's own count had drifted +4
unnoticed.

Verification runs on the remote runner.

Refs #3912

* fix(#3912): the outcome cell must never lower an exit code

The remote run caught a fail-open that this phase introduced, in the phase
whose entire purpose is removing fail-opens.

`state validate --strict` on a missing STATE.md exited **0** where it must exit
1. Mechanism: `runMain` projected the pending outcome whenever `main()` returned
void, and under v1 DEGRADED projects to 0 — so a `process.exitCode` already set
non-zero by the command was clobbered down to success. Confirmed live against a
fixture, before and after.

This refutes a review conclusion recorded earlier in this phase, that the cell
was "fail-closed and can never mask a failure as success". It could, and did.
Recording that plainly so the assumption is not repeated: the cell's danger was
never only that it might add a failure — it was that projecting it
unconditionally overwrites whatever decision came before.

Projection is now guarded: it may set a code only when none is set, and an
already-non-zero exit code always wins. The full precedence — explicit `main()`
return, then an existing non-zero exitCode, then the declared outcome — is
written at the projection site. A regression test drives a void return with a
pre-set non-zero code and a pending DEGRADED, and fails against the pre-fix
build.

The second failure was my test encoding the wrong contract, not a code defect.
It asserted `output({found:false, error: undefined})` records DEGRADED because
the KEY is present. `JSON.stringify` drops undefined, so the payload the user
receives is `{"found":false}` — carrying no error at all, and calling that
degraded would hand back exit 80 under v2 for output that reads as clean. The
discriminator is a serializable error VALUE, not key presence. The test now
pins `{error: undefined}` as explicitly NOT degraded, and the design doc's
wording is tightened to match.

Verification runs on the remote runner.

Refs #3912

* docs(#3912): the versioned exit contract, and a flag defect the docs found

Diataxis pass for Phase 8, plus a real fix that only surfaced because writing
the how-to meant running its own examples.

The docs. ADR-2980's "Revisit if" clause asked for exactly the versioned
projection this phase provides, so it gets an amendment naming #3912 /
ADR-3889 section 4 as that boundary: v1 stays 0 byte-for-byte, v2 projects
DEGRADED to 80. The amendment also records the count drift rather than
restating a stale figure — the ADR ratified 60 output({error}) sites in 9
modules; the AST-measured population is 64 across the same 9 (frontmatter 7
not 6, phase 4 not 2, roadmap 3 not 2). The pin is asserted over the
enumerated 64. json-errors.md gains the outcome-declaration reference,
including the precedence order a review pass got wrong and the suite refuted:
an explicit main() return, then an already-set non-zero process.exitCode, then
the declared outcome. Projection may only ever set a code, never lower one.

A how-to is owed here and is written, not skipped. Under v1 nothing changes,
so the audience is an operator opting into v2 and needing to know what the
codes mean for a CI gate — a migration, which is how-to shaped. It covers
turning v2 on, the code table, why 80 is "ran and reported a condition" rather
than a crash, and how to split a gate that treats any non-zero as fatal. No
tutorial: there is no new entry point to learn, and under the default contract
a reader would be walked through observing nothing.

The defect. Running the how-to's own Step 1 example returned

    $ gsd-tools --exit-contract=v2 state validate --strict
    Error: Unknown command: --exit-contract=v2          (exit 64)

while the same flag trailing the subcommand worked and exited 80. The flag
half-worked, by argv position. resolveContractVersion scans argv
non-destructively, so the token survived into the dispatcher, which treats
argv[2] as the command name. --json-errors had already solved precisely this
at gsd-tools.cjs:4455, under a comment naming the hazard verbatim: "The argv
splice must happen here too, otherwise the dispatcher below sees
--json-errors as an unknown command." The later flag never got the same
treatment.

Fixed rather than documented around: the version is resolved first — which
memoizes the cell and makes an invalid value throw early — and then every
--exit-contract= token is spliced out of the dispatcher's argv copy.
--exit-contract is now listed in TOP_LEVEL_USAGE, where it never was. The
regression test pins leading position, trailing position, agreement between
the two, and a loud failure on v3 rather than a silent fall back to v1.

Neither review engine would have caught this: the defect is invisible in the
diff, because the diff does not touch argv handling. It surfaced only from
running the documentation's own example. Writing a how-to is an execution pass.

Verification runs on the remote runner.

Refs #3912

* fix(#3912): the flag splice has to run before the run-with-timeout return

An isolated review of the previous commit found that the fix did not deliver
what it claimed, and that two of its own tests were weak. All three findings
reproduced by execution before any change was made.

The fix was placed below a return. main() intercepts `run-with-timeout` at
gsd-tools.cjs:4436 and returns from there — above both the --json-errors block
and the --exit-contract splice added in the previous commit. So the flag still
died in leading position for that one command:

    $ gsd-tools --exit-contract=v2 run-with-timeout 5 -- node -e "..."
    Error: Unknown command: run-with-timeout        (exit 64, child never ran)

The previous commit message and the test's describe-block both claimed
position-independence unconditionally. That was an overclaim, not a gap left
open, and it is the part worth naming: the fix was verified by hand on the
commands I happened to think of, and `run-with-timeout` returns before the
code I was verifying.

Both global-flag blocks now run above the interception, with a comment naming
it so a later edit cannot slide them back down. Moving --json-errors up fixes
the identical pre-existing bug for that flag, verified failing beforehand
(exit 1, sdk_unknown_command). Fixing the sibling is deliberate: same defect,
same block, and a known-broken twin next to a fixed one is not a resting state.

Two tests were not pulling their weight. The invalid-value test was vacuous —
it passed against the pre-fix build, because `--exit-contract=v3` already
exited 1 there and already printed the resolve error lazily through
error() -> getContractVersion. Both its assertions held before the fix, so it
pinned nothing. The real discriminator is that the pre-fix build emits BOTH
"Unknown command: --exit-contract=v3" and the resolve error, while the fixed
build emits only the latter; the test now asserts that absence.

The leading-position and leading==trailing tests asserted proxies — "not 64",
"no Unknown command", "the two agree" — none of which pin a value, and all of
which would survive both positions being identically broken. With a .planning
directory and no STATE.md, state-snapshot exits exactly 80 under v2 and 0
under v1 in both positions. Those numbers are pinned now. The multi-token case
the descending splice loop exists for is covered too, and run-with-timeout has
regression tests for both flags.

The lesson is narrower than "test more". Hand-verifying the production
behavior does not verify that the test would have caught its absence. The
pre-fix binary has to be run against the test's own assertions.

Investigated and deliberately not changed: splicing before --cwd parsing
degrades one diagnostic from "Missing value for --cwd" to "Invalid --cwd:
<path>", but that is pre-existing — verified on the pre-fix build via
--json-errors, which already did it. This change joins the pattern rather than
creating it, and both forms exit 64 on malformed input either way.

Verification runs on the remote runner.

Refs #3912

* chore(#3912): backfill changeset pr numbers to 3983

* test(#3912): pin the reason-table invariant as set equality, not a count

A graph-backed review flagged the unchecked lookup in
expectedErrorCode3912. Investigated by execution: the drift guard DOES
hold — for an unmapped reason under v2 the production error() yields 1
while the table yields undefined, so the assertion fails. Not a
correctness defect, and deliberately NOT made tolerant, since a tolerant
lookup would destroy the guard.

Two real problems remained. The guard asserted the wrong invariant: it
counted the TABLE's keys at 25 rather than checking they match the
ENUM's values, so a renamed member keeps the count at 25 and slips past,
and a 26th member leaves the table at 25 and slips past too. Both were
then caught only indirectly, by an undefined mismatch producing 'must
exit undefined'. It is now a sorted set equality, so the failure names
the specific missing or extra reason.

And the comment above it described a '?? FAIL' fallback that does not
exist anywhere in the function. It now states what the code actually
does, verified by running it rather than by reading it.

Refs #3912

---------

Co-authored-by: sim <sim@local>
2026-08-28 08:09:05 -04:00
Tom Boucher
d98b55562c enhance(#3910): the raw terminator is banned by construction (#3980)
* enhance(#3910): move the last src/ terminators onto the seam

Phase 6 bans the raw terminator by construction, which it cannot do while
violations stand. A census found 12 sites the rule would flag; nine of the ten
unsanctioned ones were owned by no phase of the epic at all — a coverage hole
in the decomposition, since P0-P2 are infra, P3 the gate modules, P4 the
scanners, P5 the fragments, P7 the hooks, P8 io.cts, and P6 itself only adds
the rule. `src/**/*.cts` now holds exactly 2 raw exits, both inside
`terminateNow`, the single sanctioned site.

`io.cts`'s `error()` is the interesting one. It was first called substantive on
"dozens of callers, contract risk" — asserted, not measured, and the
measurement refuted it: 289 call sites, zero inside a try whose catch would
swallow a throw. The real obstacle was structural instead: `terminateNow`
cannot emit exit 1, because ADR-3889 §1 makes 0 and 1 unallocatable and
`nameForExitCode(1)` throws. So the only route is `ExitError` under `runMain`,
which sets exitCode and writes stderr only when the error carries a user
message — keeping the existing stderr write and throwing a message-less
ExitError is observably identical.

That census was still too narrow, and running the CLI proved it. It asked
whether the CALL sits in a try/catch; the two regressions that surfaced were
interceptors elsewhere on the stack:

- `command-routing-hub.cts`'s `dispatch()` swallowed the ExitError into a
  HandlerFailure, so the caller emitted a duplicated, wrong stderr line on
  every Hub-routed path. It now rethrows ExitError explicitly — the same shape
  `gsd-tools.cjs` already used at two dispatch sites, so this follows an
  established idiom rather than inventing one.
- the profile-pipeline router's deliberately un-awaited `.catch(e => error(...))`
  turned an ExitError rejection into an uncaught exception; it now mirrors
  runMain's handling.

`edge-probe` and `ui-consideration-probe` gained `runMain` wrappers because
probe-core's new throwing default would otherwise have escaped them.

A follow-up sweep of every dispatcher — 19 command routers, the Hub, the
gsd-tools dispatch seams — found no further swallowing catch. The admitted
bound: ~1260 non-rethrowing catches repo-wide were scanned structurally but not
individually classified. Both real regressions were found by execution, not by
reading, so the suite is the detector that matters here.

`gsd-tools.cjs:253` stays a raw exit deliberately: it is the ensureRuntimeBuild
bootstrap, which runs before cli-exit is required, so the seam does not yet
exist. It needs a second allowlist entry, which means #3910's "single allowlist
entry" criterion is unachievable as written.

Verification runs on the remote runner.

Refs #3910

* enhance(#3910): ban the raw terminator by construction

Adds local/require-registered-exit and registers it on all four globs:
src/**/*.cts, scripts/**/*.cjs, hooks/**/*.js, gsd-core/bin/**/*.cjs.

Registering on the .cts glob is load-bearing, not redundant — the emitted .cjs
mirrors are globally eslint-ignored, so a rule registered only on the emitted
globs is blind to the sources. That is the #3496 lesson, and it is how the
previous guard became invisible: n/no-process-exit was 'error' in one block yet
fired zero times on all three surfaces that mattered.

The dead n/no-process-exit: 'off' block for hooks is deleted in the same PR.
Phase 7 migrated every hook, so the exemption now protects nothing.

Two allowlist entries, not the one #3910 anticipated. terminateNow's body is
detected STRUCTURALLY — a process.exit lexically inside a function of that name
— rather than by a path and line number that rots. The second is
gsd-tools.cjs's ensureRuntimeBuild bootstrap, an inline disable with its reason
at the call site: it runs before ./lib/cli-exit.cjs is required, so the seam
does not exist yet and no migration is possible. #3910's 'single allowlist
entry' criterion is therefore unachievable as written, and is amended with the
measurement rather than quietly missed.

The rule is proven able to FAIL, per glob: four positive controls, one for each
registered glob. A guard that cannot be shown to fire is not a guard. Four
matching negative controls pin process.exitCode as never-flagged — conflating
it with process.exit is what inflated this epic's original census 2x. An
allowlist case and a near-miss (same shape, different function name) fix the
structural detection in place.

Verification runs on the remote runner.

Refs #3910

* fix(#3910): stop the detached catch from throwing, and scope the allowlist

Review findings, one of them a regression the previous fix introduced.

_handlePipelineRejection called error() from inside a DETACHED .catch().
error() now throws, so that throw became an unhandled promise rejection — and
on Node >=15 with --unhandled-rejections=throw, Node dumps a raw stack trace
with absolute paths on top of the clean Error: line. That was impossible before
this branch, because process.exit(1) terminated synchronously before any
rejection machinery could observe it. The handler now writes byte-identical
stderr itself, in both plain and --json-errors form, and sets exitCode in
place. This was the THIRD interceptor found, and like the first two it surfaced
by running the CLI rather than by reading code.

The rule's terminateNow allowlist had no path constraint, so any function
anywhere named terminateNow across all four globs inherited it. It now requires
the structural nesting check AND a cli-exit.cts basename — still no line
numbers to rot.

The four per-glob positive controls only varied a filename inside RuleTester,
which never resolves eslint.config.mjs. Since the rule is filename-agnostic,
all four exercised identical logic and none proved the rule was WIRED — this
epic's own failure mode. A registration test now asserts the rule resolves for
a real path in each glob, and it is proven able to fail: removing one glob's
registration flips the resolved value from [2] to undefined.

Three evasions the rule cannot catch (computed member, aliasing, .call/.apply)
are documented in its header and pinned by tests, labelled as known limits
rather than endorsed, so a future change that starts catching them is a
deliberate diff.

Refs #3910

* docs(#3910): document the raw-terminator ban

Reference and Explanation via a new docs/features fragment (FEATURES.md is
generated from it, not hand-edited). How-To:
docs/how-to/resolve-a-raw-terminator-finding.md, indexed from docs/README.md —
a contributor whose code trips the rule picks among three replacements by
surface (runMain/ExitError for a CLI path, terminateNow for a hook,
process.exitCode where the process should drain), and needs to know why
process.exitCode is correct and never flagged, since conflating the two is what
inflated this epic's original census 2x.

The page also names the three patterns the rule cannot catch and says plainly
that using one to dodge it is a review finding, not a fix — documenting them
without that sentence would read as a sanctioned workaround.

docs/INVENTORY.md deliberately untouched: eslint-rules/ is not a tracked family
in the manifest (verified — a regen produced a zero diff), so a hand-written row
would desync the table from the family it claims to belong to.

Refs #3910

* fix(#3910): a catch that sniffs the message swallows an ExitError

The remote run returned 41 failures, and one of them was a live production
regression rather than a test artifact.

`cmdMilestoneComplete`'s unstarted-phase guard re-threw only when
`e.message.startsWith('Cannot mark milestone complete:')`. `error()` used to
`process.exit(1)`, uncatchable, so the guard always fired. It now throws an
ExitError carrying no message, the string test fails, and the ExitError was
silently swallowed — the guard stopped blocking milestone completion entirely.
Proven against the real CLI: pre-fix, a milestone with an unstarted phase
archived at exit 0; post-fix it is blocked at exit 1 with the intended message.

That is a guard that silently stopped guarding, which is this epic's thesis
appearing inside the phase meant to enforce it. Worth stating plainly: an
earlier census DID examine this site, saw a `throw e`, and classified it as
rethrowing. It was wrong — the rethrow is conditional, and a conditional
rethrow on an inspected message is indistinguishable from an unconditional one
unless you read the predicate.

So the class was swept rather than patched where it was tripped over. An AST
census of every CatchClause across src/, gsd-core/bin/ and scripts/ found 38
conditional rethrows. Two more had the same defect and are fixed the same way:
`config.cts`'s `'No config.json'` sniff and `gsd-tools.cjs`'s
`e.name === 'WindowsError'`. The remaining 25 are provably unreachable — every
one wraps a bare fs, YAML, manifest-require or git-exec primitive that cannot
throw ExitError — and two were scanner false positives, both explained. Each
fix is an unconditional `instanceof ExitError` rethrow placed BEFORE any
inspection, matching the idiom command-routing-hub and gsd-tools already used.

Residual bound, stated rather than implied: zero known-reachable unfixed sites,
contingent only on error() never later being called inside one of those 25
primitive try blocks.

The remaining failures were harness artifacts, and the harnesses were corrected
to the new contract rather than the assertions weakened. Tests that mocked
`process.exit` to observe termination now catch ExitError and assert its code;
tests parsing stderr as a single JSON object still assert exactly that, with
their ad-hoc `node -e` scripts wrapped in runMain so it is true. milestone and
phase-resolution-parity needed no test change — they were correctly written
against the real bug and are what caught it.

Verification runs on the remote runner.

Refs #3910

* chore(#3910): backfill the changeset PR number

Also reframes the fragment to lead with the user-visible change — the
milestone guard blocking again — rather than the narrowest of the three fixes.

Refs #3910

---------

Co-authored-by: sim <sim@local>
2026-08-28 03:15:39 -04:00
Tom Boucher
9f219d05ba fix(#3964): route init's waiting-signal, codebase, and skill-manifest paths through the project-aware resolver (#3971)
* test(#3964): waiting_signal, codebase_dir, and skill-manifest must be project-scoped

* fix(#3964): route waiting_signal, codebase_dir, and skill-manifest through the project-aware resolver

* chore(#3964): changeset fragment (pr number backfilled after PR creation)

* chore(#3964): backfill changeset PR number (3971)

* test(#3964): assert on the POSIX-normalized codebase_dir across platforms

---------

Co-authored-by: sim <sim@local>
2026-08-28 02:25:52 -04:00
Tom Boucher
5ad708aa4e fix(#3972): one worktreesOptedOut ladder in planning-workspace, shared by resolver and guard (#3979)
* test(#3972): the guard fallback must share the opt-out ladder

* fix(#3972): one worktreesOptedOut ladder in planning-workspace, shared by resolver and guard

* test(#3972): without a workstream the root config is the effective config

* fix(#3972): guard the whole ladder body — malformed env values degrade, never throw

* chore(#3972): changeset fragment (pr number backfilled after PR creation)

* chore(#3972): backfill changeset PR number (3979)

---------

Co-authored-by: sim <sim@local>
2026-08-28 01:53:23 -04:00
Tom Boucher
e9868a92ba fix(#1875): route installer settings/defaults writes through atomic, lock-guarded primitives (#3966)
* fix(#1875): route writeSettings through atomicWriteFileSync

writeSettings is the sole writer of settings.json/settings.local.json for
six runtimes and wrote them with a naked fs.writeFileSync. Hosts discard the
entire settings file on any parse failure, so a crash mid-write cost the user
every hook, permission, env var, and statusline they had — not just GSD's.

Route it through the atomicWriteFileSync (temp+rename) already used elsewhere
in the installer and already bound in this file.

withWriteFailure in the migration integration harness matched only the final
destination path, so an atomic write bypassed the injection entirely and turned
a rollback assertion into a vacuous pass. It now also matches the .tmp- sibling.

Refs open-gsd/gsd-core#1874 (F5)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#1875): add changeset fragment

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#1876): honor the readSettings null contract in the #338 local-merge leg

readSettings returns null only for an unparseable file — its documented
"preserve existing, don't touch" signal. The #338 migration coerced that null
to {} and wrote the result back, so a settings.local.json with one stray comma
lost all its non-GSD content on the next install.

The guard stands the whole migration down rather than just the local write:
skipping the merge while still stripping the shared file would destroy the GSD
entries outright instead of relocating them.

Aborting here reaches a pre-existing latent crash that the clobber had been
masking. Both are fixed, with their own regression test:

- the unparseable-settings guard returned bare `undefined` while all five
  sibling early exits return the full result shape, so installAllRuntimes'
  statusline lookup (results.find(r => r.runtime)) threw;
- handleStatusline dereferences result.settings, which is null on every early
  exit, so the call site now falls through to the banner branch.

Both crashes reproduce on unmodified next with a malformed settings.local.json
and no migration involved.

Refs open-gsd/gsd-core#1874 (F6)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#1876): add changeset fragment

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#1877): lock and atomically write the machine-global ~/.gsd/defaults.json

Every non-Claude install read-modify-writes ~/.gsd/defaults.json with no lock
and two separate naked whole-file writes. The file is read by every runtime and
project on the machine, so concurrent installs lost each other's key, and a
crash in either write window truncated it — silently, because the read path
swallows parse errors and treats a corrupt file as absent.

Take the existing acquireInstallMigrationLock around the read-modify-write and
apply both mutations in one atomicWriteFileSync. An install that changes nothing
no longer rewrites the file at all.

Existing semantics are unchanged: the explicit resolve_model_ids:true opt-in
(#1569) and an existing "omit" are preserved, non-canonical values still default
to "omit" (#1156), a pre-existing runtime string is preserved (#2395), the
malformed-non-object recovery (#1657) stands, and both console lines still print
when both keys change.

The #2834 structural test sliced a fixed 1200-character window from the function
source; the added lock comment pushed an asserted token past it. The window now
tracks the function body, so a comment or guard cannot red it spuriously.

Refs open-gsd/gsd-core#1874 (F18)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#1877): add changeset fragment

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#1874): preserve target mode and create temp files exclusively in atomicWriteFileSync

* fix(#1874): write the migration-lock payload through the exclusive descriptor

* chore(#1874): bold-led changeset fragments + fragment for the hardening pass

* chore(#1874): rename the shadowed lock-release catch binding

* test(#1874): fix source-grep/try-finally violations, add missing fault-injection cases

- drop the /TypeError/.test(stderr) source-grep assertion (exitCode already
  proves the installer didn't crash)
- convert inline try/finally fs-mock restoration to t.after() across the F5/F18
  test suites, per this repo's no-try/finally-in-test-body rule
- add a rename-failure fault-injection case for atomicWriteFileSync
- add a read-only-.gsd-directory fault-injection case for the F18 lock+write path
- extract MAX_TEMP_FILE_ATTEMPTS constant, dedupe the partial-write-then-throw
  mock into a shared tests/helpers.cjs helper

Found during this session's own Standards-axis code-review pass on resurrected
PR #3385.

* chore(#1874): reset changeset fragments to pr:0 placeholder

The resurrected fragments carried the closed PR's number (3385). This is a
new PR, so reset to the pr:0 placeholder and backfill the real number once
gh pr create returns it, per CONTRIBUTING.md's PR Number Handling.

* chore(#1874): backfill changeset PR number (#3966)

---------

Co-authored-by: Richard Spiers <1355479+richardspiers@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: sim <sim@local>
2026-08-27 23:27:43 -04:00
Tom Boucher
15af0f5536 enhance(#3951): B6+B7 — widen two unreachable lint rules and make the guard ledger true (#3965)
* fix(#3951): two lint rules that could not reach the code they govern

B6 names two widenings. Measuring them first turned up a defect the criterion did
not know about, and refuted the reason it gave for one of them.

1. no-adhoc-markdown-parsing self-gates on its own filename.

   Lines 107-110 short-circuit create() to {} unless the path matches
   /(?:^|\/)src\/[^/]+\.cts$/. B6 says to widen the files: glob in
   eslint.config.mjs - but doing only that ships an INERT rule, because the gate
   still returns {} for every new path. Both halves have to change, and the gate
   is the load-bearing one.

   That same regex hides a live hole: [^/]+ is FLAT-ONLY, so it requires the file
   to sit directly in src/. The registered glob is src/**/*.cts, which includes
   subdirectories. 28 .cts files - health-diagnostic-rules/ (10),
   installer-migrations/ (11), observability/ (3), host-integration-adapters/ (2),
   vendor/ (2) - are inside the registered glob and silently skipped.

   Measured with the gate neutralized: 0 violations there today. The hole is
   hiding nothing right now, and is fixed anyway, because "no violations today" is
   not a property that keeps holding.

   The fix is not invented: require-subprocess-timeout.cjs:196 already carries the
   correct form of this guard, /(?:^|\/)src\/.*\.cts$/ with .*, one directory over.
   Checked the other 21 rules for the same bug - no-adhoc-regex-escape and
   no-private-binary-resolution short-circuit only to exempt their own seam file,
   which is the right shape, and no-crlf-fragile-split has no filename gate at
   all. This bug is unique to the one rule.

2. no-adhoc-regex-escape could not see the shape that actually occurs.

   Line 396 gated the whole UNSAFE-NEW-REGEXP arm on arg.type === 'Identifier'.
   Every check below it - the _SOURCE provenance check, the
   isSoleReturnOfOwnParameter shape - lives inside that branch, so
   new RegExp(obj['key']) and new RegExp(cfg.pattern) were never examined at all.
   Runtime data arrives as a property access far more often than as a bare
   identifier, which is exactly why this rule never fired on the #3477 ReDoS.

   Widened to MemberExpression, measured by AST walk across all five registered
   blocks rather than by grep. 27 sites, zero TSAsExpression:

     18  safe new RegExp(X.source, flags)  -> exempted, keyed strictly on the
         PROPERTY being `source`, never on the object. Keying on the object would
         wave through X.anything and buy nothing. B6 estimated ~10; that was an
         undercount.
      3  _SOURCE-suffixed constants reached through a required module namespace
         (phaseId.BRACKET_PHASE_TOKEN_SOURCE) -> the same provenance-exempt class
         the rule already recognizes for bare identifiers, extended to reach them.
         Without this the widening produces 3 false flags.
      6  real findings -> marked, each a test extracting a pattern from a shipped
         file at test time, where the runtime contract IS the product.

   Deliberately the NARROW MemberExpression form. The rule's own
   isSoleReturnOfOwnParameter doc comment records that an earlier broad
   "any non-literal identifier" heuristic produced ~25 false positives and was
   rejected; a re-run of the census after this change flags exactly the 6 above
   and nothing else.

Verified by execution, not by reading: the gate now accepts src/<subdir>/x.cts,
still accepts flat src/x.cts, and still exempts paths outside src/ - each pinned
by a test proven to fail against the old regex. build:lib, lint and lint:ci all
exit 0.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3951): give no-adhoc-markdown-parsing its reach, and fix the 80 parses it finds

The rule self-gates on filename AND is registered on one glob, so widening either
half alone is inert. Both move here: the gate now accepts tests/**/*.cjs and
scripts/**/*.cjs alongside src/**/*.cts, and eslint.config.mjs registers it on the
same two.

A test pins that the gate and the registration AGREE, in both directions. The
original defect was a gate narrower than its registration; the failure mode of
this fix is a gate wider than its registration. Both are silent, so the test
asserts the pair rather than either half.

80 violations across 43 files, all in tests/, zero in scripts/. 70 are routed
through the existing seams - scanFencedBlocks, collectSection, stripFencedCode,
tokenizeHeadings from markdown-sectionizer; splitTableRow, parseMarkdownTable,
findTableWithColumns from markdown-table. Headerless STATE.md tables use
splitTableRow per line, because parseMarkdownTable needs a real delimiter row.

10 are suppressed, 12.5%, well under the third that would have meant the rule is
mis-scoped for tests/ rather than the tests carrying debt. Each names its reason:
three regression guards (#3873 / bug-#21) are deliberately independent of the
generator's own fence handling, and routing them through the seam would have them
test the generator against itself; one is a negative-text probe that extracts
nothing; six are a shell-pipe-to-jq detector whose regex coincidentally matches the
table fingerprint and is not markdown parsing at all.

All ten sit in tests whose subject is .md content, which is normally a reason to
prefer the seam. The marker used is allow-adhoc-markdown, distinct from
no-source-grep's allow-test-rule, and lint:ci's lint-allow-test-rule-refs reports
the same 280/280 unverified count as before - checked rather than assumed, because
those two markers are easy to conflate.

The widening earned its keep immediately: it found a test that passed for the
wrong reason.

  tests/config-field-docs.test.cjs asserted notEqual(<cell>, '600') against the
  TYPE column instead of the DEFAULT column. notEqual('number', '600') is true
  forever, so the guard against workflow.subagent_timeout regressing to the old
  seconds default could never fire. docs/CONFIGURATION.md:434 is
  `| workflow.subagent_timeout | number | 300000 | ... |`, so the default is cell
  index 2; the assertion is now row-scoped through splitTableRow and reads 300000.

That is the argument for the widening in one case: the violation was invisible to
lint, the suite was green, and the assertion was vacuous. A rule that cannot reach
a file cannot tell you the file is lying.

Not fixed here, and recorded rather than assumed: #3426/#3239 are NOT reachable by
this widening. tests/package-legitimacy-gate.test.cjs yields zero violations even
with the gate bypassed - its hand-rolled scans are real, but built from line
filters and split('|') rather than the regex-literal fingerprints this rule
detects. They need new detectors. The epic assumed a wider glob would catch them.

build:lib, lint and lint:ci all exit 0; the post-fix census across tests/** and
scripts/** is 0 violations.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3951): B7 — and #3356's defects were still live in the code

B7 asks that each closed child be driven fail-first with a behavioral identity
test at the CONSUMER's output. Four of eleven children had no test citing their
issue number. Auditing them by BEHAVIOR rather than by number-grep changed the
answer for three of the four.

#3364 and #2540 — traceability only. Both were implemented by #3941 and their
consumer-output tests exist and were shown failing-first; neither cited its
originating issue, so an audit that greps for the number reports them uncovered.
Tagged the specific asserting test in each file, following the citation form those
files already use.

#3372 — covered, but only at helper level, and the triage narrowed it. Of the four
commands the issue names, only estimate-cli's collectCalibrationSamples actually
enumerates phase dirs from disk; smart-entry, audit and roadmap-upgrade derive from
ROADMAP/body text and never reach the sentinel path, so they are benign by
construction and were left alone rather than "fixed" into churn. The existing #3882
rows asserted the helper's return value. Added a consumer-output test driving
`query estimate-calibrate` and asserting sample_count and the persisted document.
RED proof: reverted collectCalibrationSamples to a raw readdirSync and ran the real
CLI - sample_count 3, sentinel leaked; restored - sample_count 2.

#3356 — NOT covered, and BOTH halves of the defect were still live in source. The
issue is closed; the bug was not fixed. Fixed here rather than writing tests that
document a bug as correct.

  Defect 1, the contradicted row. quick.md:627 claimed
  `quick-tasks-append` performs "the equivalent write" to the Step 7c row. It did
  not: the `#` cell was a positional ordinal and `Directory` read `—`, because the
  route had no way to receive a quick id or task directory. Added OPTIONAL
  `--quick-id` / `--slug` / `--directory`. A caller with neither - fast.md, the
  original #2133 caller - omits them and gets the byte-identical prior row, so
  nothing existing changes. A caller that HAS a real id and directory now gets the
  canonical row quick.md:632 renders. The false-equivalence sentence itself is
  corrected rather than left to mislead the next reader.

  Defect 2, the forced re-derive. The route called readModifyWriteStateMd with no
  options, so a body-only append to the Quick Tasks table triggered a full
  re-derive of the disk-derived progress.* frontmatter. Every other body-only
  writer passes { resync: false } - src/state.cts's own docstring prescribes it -
  and this route was the lone outlier. RED proof: reverted the option, seeded a
  project with 2 real phase dirs and a curated total_phases of 25, ran
  quick-tasks-append; total_phases collapsed to 2. Restored; it stayed 25.

That second one is the shape this epic exists to close: a silent write that
replaces curated state with a re-derivation nobody asked for, exit 0 throughout.

build:lib, lint and lint:ci all exit 0.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3951): amend B6's ledger to what was measured, and document the new flags

The ADR gains a ledger amendment in its own correction style - the sixth wrong
premise it records, found the same way as the other five, by measuring before
building.

B6 says the net guard count must fall. It rose: 62 -> 69, +7, measured from the
epic's filing commit to origin/next. The attribution is the point, though. Five of
the seven came from PRs unrelated to this epic, one was added by a phase of it, and
the epic did retire something sub-file - #3884 removed a detector with an explicit
"net: -1 detector, 0 added" ledger. Every named casualty is load-bearing, two
already carry retractions in this same document, and a sweep of all 22 rules plus
every scripts/lint-* found no provably dead guard. There is no honest way to make
the count fall; forcing it would trade coverage for a number, which is the Goodhart
outcome Decision 6 exists to prevent.

The amendment also records that B6's own prescribed fix for one widening was inert.
no-adhoc-markdown-parsing self-gates on its filename, so widening only the files:
glob - which is what the criterion says to do - ships a rule that still returns {}
for every new path. And #3426/#3239 are not reachable by that widening at all;
their scans use line filters and split('|'), not the regex fingerprints the rule
detects. The roster row tracked them against the wrong mechanism.

Three roster rows updated from aspiration to fact: the two widenings are DONE with
their measured counts, and lint-phase-enumeration-drift is marked RETAINED rather
than "expected casualty - verify before retiring", because Phase 5 verified it and
kept it.

The rule Decision 6 should carry forward is stated plainly: a guard ledger is a
claim about COVERAGE, not about COUNT. "Net count must fall" is measurable and
wrong. "Every guard is reachable, and each retirement names what makes its defect
unrepresentable" is the property that was actually wanted.

CLI-TOOLS.md documents the optional --quick-id/--slug/--directory flags and says
plainly that omitting them keeps the pre-#3356 row byte-identical, plus that the
append no longer re-derives progress frontmatter.

New features fragment (id 3951); FEATURES.md regenerated rather than hand-edited.
Changeset is Changed, pr:0 pending backfill.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3951): correct four rows that pinned the lint rule's old narrow reach

The remote suite came back RED with 5 failures, all in tests/eslint-rules.test.cjs.
They are stale tests, not a regression: four rows assert that
no-adhoc-markdown-parsing is inert outside src/*.cts, which is exactly the
contract this deliverable changes.

Confirmed by reading rather than inferred from the names - the row at :1981 used
filename: 'tests/some.test.cjs' and filename: 'scripts/helper.cjs', the two roots
the rule now covers on purpose.

Worth recording WHY local gates missed this. npm run lint and lint:ci were green,
and the touched test files passed standalone. Lint only reports violations in real
files; these rows assert the rule's REACH using synthetic RuleTester filenames, so
nothing but the full suite could see them. Local green on a rule change says
nothing about the rule's own tests.

Each row is rewritten with BOTH halves rather than flipped from valid to invalid:

  - the same fingerprint under tests/ or scripts/ is now flagged, with the right
    messageId
  - the negative space is preserved - the same fingerprint under a path outside
    all three roots (gsd-core/bin/lib/foo.cjs) is still NOT flagged

The second half is the one that matters. Without it the rule has no boundary and
nothing would catch an over-wide gate later, which is the mirror image of the bug
this deliverable just fixed.

Each row is renamed to state the current contract; the old names said
"non-src/*.cts ... is not flagged" and would have been actively misleading once
the bodies changed.

Proven to test the widening rather than restate it: every flagged half was run
against HEAD~2's pre-widening rule and does NOT fire there, then against the
current rule and does. 12/12 on that probe; the full file is 178/178.

Swept for the same staleness elsewhere and found none.
require-subprocess-timeout's own "inert outside src/*.cts" row is untouched -
that rule's gate was not widened here - and no-adhoc-regex-escape's test file
already carries correctly-targeted rows.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3951): acknowledge the quick.md growth the attribution guard reported

The full suite came back RED with one failure, and it is mine:

  1 file(s) grew without an acknowledgment:
    quick.md grew 364 bytes

gsd-core/workflows/quick.md is runtime-loaded emitted content, so correcting
its false 'performs the equivalent write' claim trips emitted-attribution by
construction. This is the acknowledgment, not a workaround - there is nothing
to regenerate.

The fragment names ONE path, which is the only one the guard reported. The four
spent acknowledgments it also listed (audit-uat, plan-phase, progress, review)
belong to other fragments whose ripple the base already absorbs; they are inert,
not failures, and are deliberately NOT copied here - naming paths I did not
change would make this record false in the other direction.

Byte figure corrected before committing: the guard reported 37220 -> 37584
(+364), but origin/next has since moved and quick.md is 37232 there now, so the
measured delta is +352. The reason text says so and names the base as a moving
figure rather than pinning a number that is already stale.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3951): move the quick.md growth ack to a trailer, delete the obsolete fragment

The acknowledgment mechanism changed under this branch. Merging next brought in
the redesign - it also deleted .github/workflows/ack-fragment-sweep.yml, which
was in the merge status and which I did not register at the time - and the guard
now says so directly:

  Add a trailer to a commit in this PR (never a new file).
    Emitted-Drift-Ack-Growth: quick.md - <why this growth is deliberate>

So tests/emitted-drift-acks/3951-quick-append-equivalence.json is obsolete on
arrival. A fragment file is no longer read by anything, and leaving it would be a
dead record that looks like an active one. It is deleted here rather than kept
"just in case".

The byte figure moved again with the merge: 37232 -> 37596, +364. The earlier
fragment said +352, measured before the merge auto-merged quick.md itself. The
trailer carries no number, which is the better design - the figure was stale
twice in two attempts.

Refs #3951

Emitted-Drift-Ack-Growth: quick.md — #3356/#3951 replaces a false claim with an accurate one. Line 627 said the `quick-tasks-append` shortcut "performs the equivalent write" to the Step 7c row rendered above it; it did not, and that was the documented half of #3356 — with no quick id or task directory the route emitted a positional ordinal in `#` and an em-dash in `Directory`, a visibly different row. The corrected sentence has to carry three facts the original elided: what the shortcut actually writes when it has neither input, that this is honest behavior for its real caller (`fast.md`, which has neither), and how a caller with both now gets the byte-identical canonical row via the new optional `--quick-id`/`--slug`/`--directory` flags. Prose is the product here — an executing agent reads this line to decide whether the shortcut is safe for its case, and a shorter correction would either drop the flags (leaving the reader unable to act on the fix) or drop the limitation (recreating the false claim in gentler words).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3951): backfill changeset pr number

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 23:10:49 -04:00
Tom Boucher
2ea5efc151 enhance(#3911): hooks declare their crash policy (#3960)
* enhance(#3911): give hooks an exit seam that needs no build

ADR-3889 Phase 7 foundation. The 19 shipped enforcement hooks hold 91 of the
epic's 128 terminators and cannot reach `terminateNow` today.

The obvious route — requiring `gsd-core/bin/lib/cli-exit.cjs`, as
gsd-agent-isolation-guard.js already does for two other modules — is rejected.
That precedent carries its own warning (#3582): those files are tsc output,
gitignored and absent on a raw plugin-marketplace or git-clone install, so the
hook must first call ensureRuntimeBuild() to self-heal. Making the module a
hook needs IN ORDER TO TERMINATE depend on a build inverts the dependency, and
its failure mode is precisely the fail-open this phase exists to remove: a
guard that cannot terminate cannot deny. `lint-hooks-runtime-build-seam`
already encodes that concern, and Design B would have had to add an
ensureRuntimeBuild() call to all 19 hooks to satisfy it.

So `hooks/lib/` becomes a third emit location for cli-exit and a fifth for the
registry, preserving the invariant `src/cli-exit.cts`'s own header states: it
imports nothing but node:fs and its sibling registry, and the generator
dual-emits that sibling alongside each copy so a relative require resolves next
to whichever copy loaded it. Shipping needed no change — build-hooks.js already
declares HOOKS_SUBDIRS_TO_COPY = ['lib'].

Proven, not asserted: the two files are copied into an otherwise-empty tmpdir
and a child process requires them and terminates — PASS exits 0, HOOK_DENY
exits 2 with the payload on both stdout and stderr. That test fails the moment
the hooks copy gains a require reaching outside hooks/lib/.

Also fixed inline: the registry's fifth target let any `--write` test overwrite
the real committed hooks/lib/exit-code-registry.js, because the test helper
derived only three of the other output paths. It now redirects all five, and a
regression test asserts every committed artifact is byte-identical after a
redirected write.

Install-tree goldens pick up the two new shipped paths across 11 runtimes —
insertions only, no removals. lint:ci was green while they were stale, so this
was found by regenerating rather than by a gate.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): declare a crash policy, and migrate the write guard

Adds `hooks/lib/hook-exit.js` — the hook-facing vocabulary over `terminateNow`,
hand-written because the cli-exit copy beside it is generated:

  allow(payload)          exit 0
  deny(payload, stderr?)  exit 2
  crash(onCrash, payload) whichever the hook DECLARED

`crash()` takes the policy as a required argument with no default, which is the
whole mechanism: fail-open by accident stops being expressible. A hook must
name ALLOW or DENY at the call site, and an unrecognized value terminates
INTERNAL rather than guessing. Fail-open stays legal; fail-open by omission
does not.

`gsd-write-guard.js` is the first hook migrated, all 12 sites, and it exposed a
gap in the seam. `terminateNow`'s doc comment justified its fd-2 write by
citing this hook's `emitBlock` — but modeled it as sending the same bytes to
both streams, when `emitBlock` actually sends full JSON to stdout and only the
bare `reason` string to stderr, because Kimi's hook bus feeds stderr verbatim
back to the model. Migrating as written would have turned a readable sentence
into a JSON blob for Kimi-backed agents.

#3911 requires both "all 19 hooks terminate through terminateNow" and "no
hook's effective default changes". Those are jointly satisfiable only by
teaching the seam to carry a distinct stderr payload, so `terminateNow` gains
an optional third argument: omitted, behavior is byte-for-byte what it was; a
string is written raw, which is exactly the Kimi case. The doc comment's
inaccurate claim about emitBlock is corrected in place.

Proven rather than asserted: the pre-migration file is reconstructed from HEAD
and driven with the same catastrophic-shrink payload as the migrated one —
exit code, stdout and stderr all byte-identical.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): all 19 hooks terminate through the seam

Migrates the remaining 18 enforcement hooks onto allow/deny/crash. An AST walk
now reports zero `process.exit(` call sites across every `hooks/*.js` — down
from the 91 the census measured.

Each hook with an outer catch declares its policy once, at module top, with the
reason that policy is right for that specific guard: a read guard that cannot
scan must not block the read; a statusline that renders every prompt must
degrade rather than crash; an injection scanner must not retroactively block a
result already returned. Those sentences are the deliverable — they are what
turns fail-open-by-accident into fail-open-on-purpose. No hook's effective
default changed.

Wiring exposed two defects, both fixed here rather than noted.

A SECOND stdout/stderr-splitting site turned up in `gsd-workflow-guard.js`'s
`emitForceAddBlock`, matching the pattern already known from the write guard —
full JSON to stdout, bare reason to stderr for the Kimi bus. It uses the
`stderrPayload` argument added in the previous commit, which is now carrying
its second real caller rather than one special case.

More seriously, `terminateNow` emitted both streams inside ONE try, so a
payload that failed to serialize aborted before the stderr write ever ran. The
two windsurf guards write nothing to stdout on a block and only a reason string
to stderr, so `deny(undefined, reason)` exited 2 with EMPTY stderr — a deny
that silently loses its reason, which is the exact "fails with success" class
this epic exists to close. The streams are now emitted independently, each with
its own guard, and `undefined` means "nothing to write for this stream" rather
than an error. Regression tests inject a throwing write on one fd and assert
the other still receives its payload; they fail against the single-try version.

Byte-identity was proven per hook, not assumed: each pre-change file is
reconstructed from HEAD and driven side by side with the migrated one across
its normal path, its deny path, malformed stdin and empty stdin — exit code,
stdout and stderr compared.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): harden the three shell hooks, and pin every hook's policy

`gsd-phase-boundary.sh`, `gsd-session-state.sh` and `gsd-validate-commit.sh`
gain `set -euo pipefail`.

The expected hazard did not materialize, and that is worth recording: every
intentionally-non-zero command in all three is already the condition of an
`if`/`elif`, which `set -e` never fires on, and none of them reads a
possibly-unset variable or pipes through a grep that may legitimately match
nothing. No `|| true` guards were needed. Each hook was still checked
command-by-command before the flags went in rather than after.

Twenty-one before/after cases across the three hooks — disabled and enabled,
planning and non-planning, missing STATE.md, malformed JSON, the Kimi payload
shape, quoted and unquoted `-m`, valid and over-long Conventional Commits —
all match on exit code, stdout and stderr.

The hardening is shown to actually fire, not merely added: with a stubbed
`node` that fails at the JSON-emit step, phase-boundary and session-state go
from silently exiting 0 with empty stdout to failing visibly with the error
surfaced. No such case could be constructed for `gsd-validate-commit.sh`,
whose every statement already sits inside an if-condition — recorded as
unproven rather than claimed.

`tests/hooks-crash-policy.test.cjs` adds the per-hook coverage the issue asks
for, table-driven over all 19 hooks rather than 76 hand-written cases: normal
allow, deny where a deny path exists, crash-honors-the-declared-policy, and an
unclosed-stdin case — the one `process.exitCode` structurally cannot serve. The
deny assertions encode each hook's ACTUAL stream split rather than a uniform
shape, since four of the six deliberately differ. A drift guard enumerates
`hooks/*.js` and fails if a terminating hook is ever added without a row.

Writing those tests surfaced two hooks that emit a block decision in their JSON
body and exit 0. Both were checked rather than assumed, and neither is a
fails-with-success: `gsd-read-injection-scanner.js` is PostToolUse, where the
tool has already run and exit 2 has no meaning, and `gsd-cursor-subagent-start.js`
follows Cursor's JSON-body protocol. They are deliberately left alone — a
mechanical sweep to `deny()` would have broken exactly these two.

Verification runs on the remote runner.

Refs #3911

* fix(#3838): the commit validator says when it could not validate

#3911 claims to subsume #3838. Measurement said otherwise, so this closes it
for real rather than by assertion.

`set -euo pipefail`, added earlier on this branch, does NOT fix #3838: bash
exempts a command used as an `if` condition from `set -e`, and all three of the
hook's swallow-and-pass sites are exactly that shape. Verified against the
hardened hook with a node shim that fails only the classifier call — a
non-conforming commit still exited 0 with empty stdout AND empty stderr,
indistinguishable from "your commit conforms". That is the defect verbatim.

All three sites named in #3838 now capture the real exit status instead of
consuming it as a condition, and each distinguishes its genuine negative from
"could not run":

- the classifier: 0 = is a git commit, 1 = genuinely not one, anything else =
  could not classify. Its `node -e` now wraps the require and the call in
  try/catch and exits 3 on a throw, so a broken require chain can never be
  mistaken for `isGitSubcommand` legitimately returning false — which is the
  arm that matters, since `token-scanner.cjs` is a gitignored build artifact
  and a fresh checkout lands there.
- the opt-in config read and the JSON command extraction get the same
  treatment.

On "could not run" the hook emits a diagnostic to stderr naming which check
failed and why, then exits 0. The issue confirms this is safe — it is a
PreToolUse hook, so stderr does not disturb the JSON protocol — and ranks it
the smallest sufficient fix. The gate still fails open, but it can no longer do
so silently, which is the whole complaint: a validator that disables itself
quietly costs more than one that is absent, because it is trusted.

Both controls are unchanged and pinned by tests: a conforming commit still
passes silently, a non-conforming one still exits 2 with its existing block
payload. The defect test asserts stderr is non-empty and names the failure; it
fails against the pre-fix hook.

Verification runs on the remote runner.

Refs #3911, #3838

* docs(#3911): document the hook crash-policy contract

Reference and Explanation via a new docs/features fragment (FEATURES.md is
generated from it), INVENTORY rows for the three new hooks/lib files, and an
ARCHITECTURE note on the hooks section.

How-To: docs/how-to/declare-a-hook-crash-policy.md, indexed from docs/README.md
— a hook author now has to choose and declare a crash policy, which is more
than one step and crosses into which harness protocol their hook speaks. It
covers allow/deny/crash, writing an ON_CRASH reason that is actually useful,
when a deny needs a distinct stderr payload, the two hooks whose harness reads
a JSON-body decision and must NOT use deny(), and what to do when a check
cannot run at all — with #3838 as the worked example.

Refs #3911

* test(#3911): prove the seam actually ships, and stop hand-rolling temp cleanup

Two review findings.

The acceptance criterion 'hooks/dist/** stays in parity via the build seam
(lint:hooks-runtime-build-seam)' was misstated and unmet: that lint checks
something else — that a hook requiring a compiled gsd-core/bin/lib module also
calls ensureRuntimeBuild(). Nothing exercised that the three new hooks/lib
files reach hooks/dist/lib at all. That gap is not theoretical: #770 is a
recorded ship-blocking bug where a new hook never shipped because a copy list
missed it. The suite now builds dist through the repo's own ensureBuiltHooks(),
byte-compares each shipped copy against its source, and spawns a child that
requires the SHIPPED dist copy and denies — which is what catches a copy that
exists but cannot resolve its sibling registry.

gsd-validate-commit.sh hand-duplicated mktemp/run/rm three times; one idempotent
trap on EXIT replaces them, guarded so cleanup cannot alter the exit status.
Behavior-neutral across five cases, with temp-file counts taken before and
after each run.

Refs #3911

* fix(#3911): stage transitive hook lib requires, not just one level

The remote run returned 7 failures across 3 real causes.

The important one is a PRODUCTION bug this phase exposed rather than caused.
`writeCursorHooksJson` scanned each hook script for `./lib/X` requires exactly
one level deep and never re-scanned the lib files it staged for their own
sibling requires. Nothing had a transitive lib dependency before, so the gap
was invisible. Adding hook-exit.js -> cli-exit.js -> exit-code-registry.js
made real Cursor installs ship a bundle that dies at require time with
MODULE_NOT_FOUND. It now walks to a fixed point, and a real installed Cursor
hook runs to completion.

The staging harness in shared-hooks-dir-resolution hand-copied its fixture, so
the injection scanner crashed at require time and its exit-1 was being read as
a policy decision. Migrated to copyScriptWithDeps, which walks the require
graph — the repo's recorded rule for this class, since adding another
copyFileSync keeps it alive for the next person.

The missing-lib-source test in cursor-hook-workspace-roots hardcoded which lib
file it expected to be named in the abort message; the same throw now fires for
a different file first. Its assertion is unchanged in substance — staging still
must abort rather than ship a broken hook — only the name is no longer pinned.

The last one was my own test asserting an uppercase reason code. Measured
against origin/next: the pre-change hook emits the same lowercase
'config_unreadable', so the test was wrong, not the migration. Corrected to the
real value rather than making the code match the test.

Verification runs on the remote runner.

Refs #3911

* chore(#3911): regenerate the cursor install-tree golden

The staging fix means a Cursor install now correctly carries the two
transitive lib files it was silently missing. Additive only — no path was
removed. The golden diff is the evidence the packaging defect was real.

Refs #3911

* chore(#3911): backfill the changeset PR number

Refs #3911

* fix(#3911): a git probe that timed out is not a negative

A macOS CI lane failed three deny cases at 2084ms, 2112ms and 2177ms — just
past the 2000ms budget these hooks give their git probes. The three that passed
took 72ms, 595ms and 651ms. Under shard contention `git rev-parse` overruns,
the hook reads the non-zero result as "not a git repo", and allows with exit 0
and empty stdout AND empty stderr. Under load, the guards silently stop
guarding. That is ADR-3889's thesis exactly, sitting inside the security hooks
this phase is about.

The repo had already recognized the class in one place — gsd-cursor-subagent-start.js
fail-closed-denies on `git_timed_out` (#3045) — but nowhere else.

`hooks/lib/git-probe.js` classifies a probe's outcome, distinguishing a real
non-zero exit from ETIMEDOUT, a signal kill, and a spawn failure, rather than
folding all four into `status !== 0`. Three guards route their eight git probes
through it.

The resolution is the same shape #3838 took, and the same one that issue
endorsed as smallest-sufficient: fail open, but loudly. **No exit code changes
on any path** — a developer on a loaded machine is still not blocked, which
keeps #3911's declaration-pass contract intact for exit codes. What changes is
that the hook now says on stderr which probe could not answer, instead of
presenting silence as a clean verdict.

Scope was checked across every hooks/*.js, not just the three that failed:
gsd-agent-isolation-guard spawns no git; gsd-statusline's two probes gate only
a cosmetic display segment, not an allow/deny decision, and are left alone.

The C2 deny assertion was a real-race test — it demanded exit 2 while a slow
git legitimately yields 0. It now requires the hook to either deny, or allow
with a diagnostic naming the probe that could not run; a silent allow still
fails, so the assertion is not vacuous. A deterministic regression stubs git on
PATH to sleep past the budget rather than waiting for load to reproduce it.

Verification runs on the remote runner.

Refs #3911

* test(#3911): a PATH shim cannot intercept the hooks' git spawn on Windows

The deterministic timeout regression stubbed git on PATH and asserted the
guard reports rather than silently allows. It passes on Linux and macOS and
failed on Windows in 83ms and 176ms — the stub was never invoked at all.

Mechanism: the hooks call spawnSync('git', args) with no shell:true, so on
Windows CreateProcess resolves git.exe only and never a PATH .cmd shim. The
git.cmd branch could not have worked and is removed rather than left implying
a Windows path that does. Adding shell:true to the hooks to serve a test would
change product behavior and widen an injection surface, so the case is skipped
on win32 only, with the mechanism written into the skip reason so a future
reader does not 'fix' it that way.

Linux and macOS keep the coverage, and macOS is where the underlying fail-open
was actually caught.

Refs #3911

---------

Co-authored-by: sim <sim@local>
2026-08-27 22:21:10 -04:00
Tom Boucher
ddeb141bd2 fix(#3749): route migrate-config, project_exists, and health repairs through the project-aware resolver (#3955)
* test(#3749): GSD_PROJECT-scoped migrate-config, project_exists, and repair paths

* fix(#3749): route migrate-config, project_exists, and health repairs through the project-aware resolver

* chore(#3749): changeset fragment (pr number backfilled after PR creation)

* chore(#3749): backfill changeset PR number (3955)

* test(#3749): assert on the POSIX-normalized project_path across platforms

---------

Co-authored-by: sim <sim@local>
2026-08-27 17:57:05 -04:00
Tom Boucher
8b41d855e0 fix(#3742): path-shaped comment channel keys and post-restore propagation (#3952)
* test(#3742): frontmatter comment survival must not depend on the body or indentation

* fix(#3742): path-shaped comment channel keys and post-restore channel propagation

* chore(#3742): changeset fragment (pr number backfilled after PR creation)

* chore(#3742): backfill changeset PR number (3952)

* test(#3742): direct mutation-shard coverage for the nested comment channel

---------

Co-authored-by: sim <sim@local>
2026-08-27 16:36:34 -04:00
Tom Boucher
03b7125293 enhance(#3909): a probe that could not run no longer asserts a verdict (#3944)
* test(#3909): failing-first suite for the fabricated probe fallbacks

Binds the four fabrication sites found by executing the surfaces (ADR-3889
failure class (c)), each with a positive control so an over-firing fix goes red:

- the blocking api-coverage.verify-pre gate certifying "no external-API
  integration" from a zero-byte phase scope
- the assumption-delta query route scanning an unresolvable phase section as
  the empty string and reporting it as an examined negative
- both capability fragments' probe fallbacks, which append a fabricated
  verdict rather than replacing, and fire on the legitimate exit-1 negative

Verification runs on the remote runner.

Refs #3909

* enhance(#3909): a probe that could not run no longer asserts a verdict

ADR-3889 Phase 5. Four sites turned a failed or unexamined probe into a
confident negative; each now reports what it could not establish.

- check api-coverage.verify-pre: a phase with no plan body and no roadmap
  section ran detection over zero bytes and PASSED the blocking seal gate,
  certifying "no external-API integration" from input it never read. It now
  holds with scope_unavailable. The discriminator is bytes examined, never
  signals found, so a phase whose plans are real and simply carry no API
  vocabulary passes exactly as before.
- query assumption-delta scan: an unresolvable phase section was scanned as
  the empty string and reported as an examined negative. It now returns
  {skipped, reason: phase_unresolved}, still at exit 0 — an ADR-2980 degraded
  result in the payload, leaving the gsd-tools exit projection to P8.
- both capability fragments: `|| echo '{"detected":false}'` appended rather
  than replaced, and fired on the legitimate exit-1 negative, so a correct
  answer and an honest skip both arrived as two concatenated objects. They now
  keep the probe's own payload and manufacture only an explicit
  probe_unavailable skip when the probe produced nothing at all.

Every registered outcome is more restrictive on a blocking gate, so this can
turn a false green red and never a red green.

Docs: FEATURES 156, CONFIGURATION (both keys), references/api-coverage.md
seal-time outcome table, and a new how-to for the reason-code vocabulary.

Verification runs on the remote runner.

Closes #3909

* test(#3909): correct the stale unknown-phase assertion

`unknown phase → detected:false, no throw (graceful)` scanned phase 999
against a two-phase roadmap and asserted `detected === false`. That pinned
the fabrication as intended behavior: the phase does not exist, so the
detector was handed the empty string and its "no core assumption changed"
answer described nothing that was ever read.

It now asserts the skipped-with-reason shape. The graceful-degradation
contract the test was actually protecting — the query succeeds and does not
throw on an unknown phase — is unchanged.

Found by code review, not by the author.

Refs #3909

* docs(#3909): author the FEATURES entry in its generator source

`docs/FEATURES.md` is generated by `scripts/gen-features.cjs` from the
per-feature fragments in `docs/features/`. The API-coverage entry was edited
in the generated file, so the next regeneration silently dropped it.

The text now lives in `docs/features/api-coverage-gate.md` and
`docs/FEATURES.md` is regenerated from it, leaving the shipped file
byte-identical and its content actually derivable.

Caught by `lint:generated-sync`.

Refs #3909

* test(#3909): bind the skip to "not found", and pin the discriminator

The first verification run went red on one case, and the case was wrong
rather than the code.

`getRoadmapPhaseWithFallback` returns `null` for an unknown phase and for a
missing ROADMAP.md, but for a section whose body is whitespace-only it returns
the heading line alone — which is not empty. So a body-less section WAS found,
and reporting `detected:false` over its heading is a real negative, not a
fabrication. The test had assumed the resolver yielded `''` there.

Correcting the test rather than the resolver keeps `skipped` bound to the
distinction the issue asks for — found versus not found — and avoids diverging
`assumption-delta scan` from `roadmap.get-phase`, which the fragment documents
as sharing one resolver.

Also adds the seeded property the test matrix had promised: for any plan body,
the scope read back is whitespace-only exactly when the body was. That pins the
gate's discriminator to bytes examined, so it cannot quietly become "no signals
found", across unicode whitespace and CRLF.

`docs/INVENTORY.md` picks up the reference doc's new seal-time outcome table —
surfaced by the co-change gate, not by a lint failure.

Refs #3909

* chore(#3909): backfill the changeset PR number

Refs #3909

---------

Co-authored-by: sim <sim@local>
2026-08-27 15:50:12 -04:00
Tom Boucher
9410f7e6e6 enhance(#3897): ADR-3473 §8.3 rungs 2-4 — runtime marker, derived Codex sandbox, short-form depends_on (#3941)
* test(#3897): failing-first coverage for §8.3 rungs 2-4

ADR-3473 §8.3 has four rungs; #3883/PR #3896 shipped the first. This pins the
other three RED before any fix.

Rung 2 — the install marker has four readers and resolveRuntime is not one.

  resolveRuntime resolves GSD_RUNTIME > config.runtime > 'claude' and reads no
  marker at all, while bin/install.js writes one (#2297) and FOUR hand-rolled
  readInstallRuntimeMarker copies exist: src/model-resolver.cts:65 (cached, with
  test seams), hooks/gsd-agent-isolation-guard.js:112, and TWICE in
  hooks/gsd-cursor-subagent-start.js at :346 and :355. Four copies of one rule.

  Fixtures and seam names mined from PR #3382 rather than re-derived; it
  implemented this rung and was closed "not on the merits".

Rung 3 — the sandbox map, and the fallback that was the real defect.

  Measured across all 35 files in agents/, deriving workspace-write iff tools:
  declares Write or Edit:

    - all 11 CODEX_AGENT_SANDBOX entries derive to their mapped value exactly,
      zero disagreements — the map carries nothing the contract does not
    - 24 roles fall through `|| 'read-only'`, of which 16 declare Write or Edit

  So the map is redundant and the silent fallback is the defect. The maintainer
  chose to derive but hold those 16 at read-only pending the question of whether
  Codex enforces sandbox_mode or merely advises; HALT.md records it.

  T20 asserts the emitted sandbox_mode PER ROLE against a captured baseline, not
  in aggregate — an aggregate passes while one role silently widens, which is
  the proxy-instead-of-identity shape this repo names. T24 and T25 fail on a
  stale hold, so the hold list cannot rot into the subset map being deleted.

Rung 4 — shortFormToId, recovered rather than invented.

  I nearly reported this as another wrong §8.3 claim: `git log -S shortFormToId`
  returns only documentation commits. That was the wrong instrument. Direct
  inspection of sdk/src/query/phase.ts at 11918dcc3^ shows five occurrences, and
  the tests match that code rather than a guess at its semantics — including
  first-write-wins on a duplicate short form.

  T43 asserts at the consumer's output: the emitted `waves` map from the real
  CLI, which pre-fix collapses to {"1":[...]} because every short-form edge is
  dropped. A unit assertion on resolveDependencyId would have passed throughout
  this defect's life.

Observed RED, this tree:
  rung 2   11/11 fail — no marker rung, no seams
  rung 3   T23,T24,T25,T26,T30 fail; T28 fails (validate agents passes a TOML
           whose sandbox_mode disagrees — it checks presence only)
  rung 4   T42,T44 fail; T43,T49 fail with waves collapsed to a single wave 1

Green and staying green: T20/T21/T22/T27 as captured baselines, #3885's
unresolvable-token warning and wave-verdict suppression, and #3785's
display-mapping passthrough. If the third tier over-reaches, those go red — that
is their job.

Disclosed weakness: T45 (a canonical id with no dash is not short-form indexed)
cannot be isolated behaviorally, because planMap always masks it. It is a
non-crash boundary pin, weaker than the other rows, and is recorded as such
rather than presented as equivalent.

Design:      .gsd/phase/feat-3897-adr3473-83-rungs/40-design.md
Test matrix: .gsd/phase/feat-3897-adr3473-83-rungs/50-test-matrix.md
Decision:    .gsd/phase/feat-3897-adr3473-83-rungs/HALT.md

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3897): §8.3 rungs 2-4 — one marker reader, a derived sandbox, the third depends_on tier

ADR-3473 §8.3 has four rungs. #3883/PR #3896 shipped the first. These are the
other three.

Rung 2 — the install marker had four readers, and resolveRuntime was not one.

  resolveRuntime resolved GSD_RUNTIME > config.runtime > 'claude' and read no
  marker, while bin/install.js writes one (#2297) and four hand-rolled
  readInstallRuntimeMarker copies existed: src/model-resolver.cts (cached, with
  seams), hooks/gsd-agent-isolation-guard.js, and twice in
  hooks/gsd-cursor-subagent-start.js.

  model-resolver's was already the house idiom, so it was promoted rather than
  replaced: src/runtime-slash.cts now owns it, and model-resolver plus both
  hooks delegate. The hooks reach it through ensureRuntimeBuild(), the seam
  lint-hooks-runtime-build-seam enforces. No import cycle existed - checked
  both directions before moving anything.

  The marker is the THIRD rung: env > project config > marker > 'claude'.

  N1 was checked rather than assumed, and my first reading of it was wrong. A
  marker holding an unknown name comes back essentially verbatim, which looked
  like a validation gap. Measured against the env rung with the same inputs -
  including "../../etc/passwd" and "claude;rm -rf /" - the two are identical,
  because they share resolveRuntimeNameFromCandidates. N1 asks for exactly that,
  and it is met. The residual (the shared normalizer normalizes shape, it does
  not validate against the known-runtime set) is pre-existing on the env rung
  and plausibly deliberate, since a new runtime should not need a code change.
  The marker also does not widen the trust boundary in any real sense: it lives
  inside the install tree beside the code, so anyone who can write it can write
  runtime-slash.cjs itself.

Rung 3 — the map was redundant; the silent fallback was the defect.

  Measured across all 35 files in agents/, deriving workspace-write iff tools:
  declares Write or Edit: all 11 CODEX_AGENT_SANDBOX entries derive to their
  mapped value exactly, zero disagreements. The map carried nothing the contract
  did not already have, so it is DELETED rather than reconciled. What was
  actually broken is `|| 'read-only'`, which silently under-granted 24 of 35
  roles.

  16 of those 24 declare Write or Edit and would widen under derivation. Per the
  maintainer's decision (HALT.md), they are held at read-only pending the
  question of whether Codex enforces sandbox_mode or merely advises. Emitted
  TOML is therefore byte-identical for all 35 roles - asserted per role, not in
  aggregate, because an aggregate passes while one role silently widens.

  The hold list self-invalidates. A hold whose role no longer derives broader
  fails, and so does a hold naming a role with no agents/<name>.md. Without
  that it would rot into exactly the hand-maintained subset map being deleted,
  and this commit's own ledger claim would become false over time. Both cases
  were proved by injecting them and watching them throw.

  Two committed tests asserted the deleted map's existence and contents. They
  were pinning the thing being removed, so the tests moved rather than the
  production code: the 11 role-value pairs survive as a test-local
  PRE_3897_CODEX_AGENT_SANDBOX baseline, and the assertions now drive the real
  derivation against real agents/*.md. The coverage is preserved; only its
  source moved out of production code.

  validate agents gains checkCodexSandboxPosture, mirroring the existing
  checkCodexModelPosture: each installed TOML's sandbox_mode must equal the
  role's expected value, failing with role, expected and found. It previously
  checked file presence and manifest completeness only, so a TOML whose
  sandbox_mode disagreed passed.

Rung 4 — shortFormToId, recovered rather than invented.

  I nearly reported this as another wrong §8.3 claim: git log -S returns only
  documentation commits. Wrong instrument. sdk/src/query/phase.ts at 11918dcc3^
  carries five occurrences, and the implementation here matches that code rather
  than a guess at its semantics - including first-write-wins on a duplicate
  short form, deterministic from the sorted plan order.

  It resolves the bare plan number: depends_on: ["01"] now reaches
  26-01-auth-hardening. That is a control-flow change, not a diagnostic one -
  plans that silently collapsed into a single wave 1 now execute in their
  declared waves, and execute-phase.md consumes those wave values.

  In-phase only, by construction: the map is built from this phase's rawPlans,
  so a same-named short form in another phase does not resolve.

  #3785's display-mapping passthrough and #3885's unresolvable-token warning and
  wave-verdict suppression are untouched and stay green. If the third tier had
  over-reached, those are what would have caught it.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): close a fail-open I introduced, and wire the posture check to its command

Two blockers from review. Both are mine, and one is a security regression my own
change created.

1. A held role could escape its hold by editing its own frontmatter.

  The Codex install loop set the sandbox identity from the agent's frontmatter
  `name:` field rather than from its filename, so the hold lookup keyed off a
  value the file itself declares:

    deriveCodexSandboxMode('gsd-doc-writer',   <real file>)          -> read-only
    deriveCodexSandboxMode('gsd-doc-writer-x', <same file, name: edited>) -> workspace-write
    deriveCodexSandboxMode('GSD-Doc-Writer',   <same file, name: recased>) -> workspace-write

  What makes this a blocker rather than a nit is the DIRECTION. The deleted
  CODEX_AGENT_SANDBOX map had the identical lookup-key quirk, but it was an
  allowlist: an unmatched key fell back to read-only, which is safe. The new
  scheme derives workspace-write from the tool contract and uses the hold as a
  subtraction, so the same mismatch fails OPEN. I converted a fail-closed quirk
  into a fail-open one and did not notice; the isolated reviewer proved it by
  execution.

  Neither safety net caught it. validateCodexSandboxHolds only checks that
  <key>.md exists, never that a file's derived identity matches its key.
  checkCodexSandboxPosture looks the canonical source up by the installed TOML's
  filename, finds nothing for a renamed agent, and treats it as a custom
  non-roster agent — silently no violation.

  The identity is now the FILENAME STEM, which is what validateCodexSandboxHolds
  already validates and what an attacker editing frontmatter cannot change
  without renaming the file — at which point the existing validator catches it.
  The lookup is case-insensitive so a recase does not slip past either. The
  frontmatter name still drives the TOML body and filename, unchanged; only the
  sandbox identity moved.

  All 35 roster files were checked: name matches filename stem everywhere, so a
  stricter "they must agree or throw" invariant would have been safe against real
  content. It is deliberately NOT added — it would abort an install on a tampered
  file where emitting a correctly-derived read-only TOML is the safer outcome.
  Recorded as a fork rather than decided silently.

2. checkCodexSandboxPosture was exported and never called.

  cmdValidateAgents (src/verify.cts) called checkAgentsInstalled and
  checkCodexModelPosture only; grep for the sandbox check in that file returned
  nothing. So criterion 3 — "validate agents fails on semantic drift, not only on
  missing files" — was unmet, and `validate agents` behaved exactly as before.
  That is ADR-3473 Decision 2's named shape: a declared policy with no executor.

  It also meant the T28 test asserted at the helper's return value while the
  COMMAND stayed broken — the ADR-3180 Decision 4(b) failure this epic exists to
  close, committed by me while enforcing it elsewhere in the same epic.

  Now wired as an additive `sandbox_posture` field beside `codex_posture`,
  following the sibling precedent exactly. Drift is report-only, not a non-zero
  exit, because that is what checkCodexModelPosture does — two sibling posture
  checks disagreeing about whether a violation is fatal would be its own defect.
  The choice is recorded in a comment rather than left implicit. A consumer-output
  test now drives the real CLI and asserts on the emitted JSON, and was shown
  failing before the wiring and passing after.

Also corrected a stale artifact: the design's Known limit L1 still claimed rung 3
was not in this deliverable, written while it was halted and false once the
maintainer unblocked it.

Verified after both fixes: the three bypass probes all return read-only, the
per-role table is 35/35 byte-identical, and both hold self-invalidation cases
still throw.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3897): the marker rung, the derived sandbox, and the bare plan-number depends_on

Reference: the runtime precedence ladder in docs/CLI-TOOLS.md gains the install
marker rung; docs/COMMANDS.md documents validate agents' new sandbox_posture
field; docs/reference/plan-md.md documents that depends_on accepts the bare plan
number.

Explanation: a docs/features fragment keyed id 3897, so it cannot collide with a
concurrent PR hand-allocating a section number, regenerated into FEATURES.md.

ADR-3473 §8.3 gains an ANSWER blockquote in the document's own correction style,
recording what was measured and built against the section's 2026-08-26 correction
- including the qualification that checkAgentsInstalled itself still checks
presence only, and the semantic assertion lives in a sibling wired into validate
agents rather than folded into it.

No how-to. Both user-visible changes are zero-step: a non-Claude install resolving
its own runtime, and plans executing in their declared waves, both happen without
the user doing anything. docs/how-to/control-the-reported-host-runtime.md covers a
DIFFERENT ladder (resolveReportedRuntime / agent_runtime) that this change does
not touch, and was deliberately left alone rather than edited by association.

No tutorial - nothing multi-step to walk through. docs/AGENTS.md unchanged: it
documents Claude-side tools frontmatter, never Codex sandbox_mode, and the
emitted tools contract did not change.

The prompt layer documents depends_on only by example, not by schema, so nothing
there needed editing - and few-shot-examples/plan-checker.md already showed
depends_on: ['01'], which now actually resolves.

Translated copies of plan-md.md are untouched; the project treats translations as
community-maintained.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): move the sandbox derivation out of the installer, off the install path, and off a third parser

The full suite came back with 26 failures across four files. Three distinct
causes, mapped individually rather than assuming the first explained the rest.

A. Requiring bin/install.js printed the GSD banner to stdout and corrupted
   `validate agents` JSON.

     Unexpected token '', "[36m   ██"... is not valid JSON

   checkCodexSandboxPosture reached deriveCodexSandboxMode by lazily requiring
   bin/install.js, whose module load prints the ASCII banner. So the command
   emitted banner bytes before its JSON and every JSON consumer broke, including
   ten tests that predate this branch. src/ reaching into bin/ was backwards
   layering that happened to also be loud.

   The derivation now lives in src/codex-agent-toml.cts - the existing Codex TOML
   domain module, no new module and no six-gate ripple - and both bin/install.js
   and src/agent-install-check.cts import it. One owner, which is §8.3's rule
   applied to the fix for §8.3.

B. The stale-hold throw fired on a legitimate partial source dir, and masked a
   security assertion.

   validateCodexSandboxHolds treated "this hold's .md is absent from the install
   SOURCE dir" as a stale hold and threw. A test fixture, or any partial install
   source, legitimately contains a couple of agents. Worse, it threw BEFORE the
   path-escape check, so a test asserting that a `../../evil` frontmatter name is
   rejected got my unrelated error instead of the traversal rejection it was
   written for. A fail-closed check of mine was hiding a real security check.

   The "no stale holds, shrink-only" invariant is a property of the repo's
   canonical agents/ roster, not of whatever directory an install happens to read.
   It is off the runtime path and enforced where it belongs, in the tests that
   already existed for it. A partial source dir now installs cleanly, and the
   evil-name case throws with its own escapes-configHome message again.

C. T8 depended on ambient process.env state.

   The marker/env parity assertion round-tripped through live process.env. It now
   compares against resolveExplicitRuntime's already-exported dependency-injection
   parameter - deterministic and hermetic, same claim. Proven still falsifiable
   rather than assumed: with the marker rung's normalization temporarily bypassed
   the two rungs diverge ("codex\n../../etc/passwd" vs "codex-../../etc/passwd")
   and the assertion fails, then passes again once reverted.

One correction folded in along the way. The first version of the move added
private _extractFrontmatterAndBody/_extractFrontmatterField helpers to
codex-agent-toml.cts - a THIRD copy of frontmatter extraction, where the graph
already shows two (bin/install.js:2348, runtime-artifact-conversion.cts:893).
Adding a third inside the epic whose thesis is one implementation per rule is not
defensible. deriveCodexSandboxMode no longer parses anything: it takes
(identity, toolsValue) and each caller supplies the tools value using the
extractor it already has. Both helpers are deleted. The identity argument is
still the filename stem, so the fail-open fix is untouched.

Verified after all three: `validate agents --raw` emits parseable JSON with no
banner and both posture fields; the four hold-bypass probes still return
read-only; the per-role table is 35/35 byte-identical at 26 read-only / 9
workspace-write; the hold list is still 16.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): drop a dev-only transitive dep, make the derivation total, retire a stale fallback test

Suite down to 7 failures from 26. Three more causes, mapped individually.

A. My extractor import dragged in a script that does not exist in an installed
   tree.

     Cannot find module '../../../scripts/fix-slash-commands.cjs'

   Chain: src/agent-install-check.cts imported runtime-artifact-conversion.cjs,
   which requires command-roster.cjs, whose line 36 requires
   ../../../scripts/fix-slash-commands.cjs. That path exists in the repo and not
   in an install, so every test exercising a synthetic install dir died at module
   load. I picked that extractor for convenience without checking what it pulls
   in - the same mistake that produced the banner bug, one layer further out.

   agent-install-check now uses a single-purpose extractToolsLine on
   codex-agent-toml.cts. That is deliberately NOT a general frontmatter parser:
   we deleted those helpers a commit ago for good reason, and this reads one
   line. Verified from outside the repo root that requiring either module prints
   nothing and does not throw.

B. A test pinned the deleted name-based fallback.

   'defaults unknown agents to read-only' called generateCodexAgentToml with a
   fixture declaring tools: Read, Write, Edit. Under derivation an unknown agent
   with a writing contract correctly derives workspace-write - design row S6, a
   new writing role gets the contract, not the pin. The behavior it asserted was
   the silent fallback this rung deleted; identity no longer decides the sandbox.

   Replaced with two rows rather than a flipped string: no tools declared ->
   read-only (absence is not a grant), and Write/Edit declared -> workspace-write.
   Strictly more coverage than the row it replaces.

C. The stale-hold check still threw per derivation call.

   Last commit took the roster-existence check off the install path, but
   deriveCodexSandboxMode itself still threw when a hold's role did not derive
   broader FOR THE CONTENT IT WAS HANDED - so it fired on any synthetic fixture
   for a held role.

   The throw is gone, and it cost nothing: if a held role's content does not
   derive broader, the hold pins read-only and derivation returns read-only
   anyway, so the hold is a no-op and there is nothing to fail about. The
   staleness invariant is a property of the real agents/ roster, and
   validateCodexSandboxHolds still enforces it there - confirmed against the real
   roster after the change, not assumed.

   deriveCodexSandboxMode is now total: every (identity, toolsValue) including
   undefined and null returns read-only or workspace-write, never throws.

Verified: validate agents emits parseable JSON; the four hold-bypass probes
return read-only; the per-role table is 35/35 at 26 read-only / 9
workspace-write.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3897): put the rung-3 decision in the shipped docs instead of pointing at an ignored path

The ADR entry and the feature fragment both ended their rung-3 explanation with
"see .gsd/phase/feat-3897-adr3473-83-rungs/45-decision-rung3-sandbox.md". That
directory is gitignored (.gitignore:55), so the rationale for holding 16 roles at
read-only was reachable only from the machine that produced it. A reader of the
ADR got a pointer to nothing.

Both now carry the reasoning inline: the criterion asks both that the sandbox
derive from the declared tool contract and that no role gain a broader sandbox,
and those cannot both hold, because a faithful derivation widens 16 roles the
deleted map never listed and that fell through its silent read-only default. The
resolution is derive-and-hold - the derivation owns the rule now, each hold is
released as its enforcement question is answered, and a hold is reversible where
a widened sandbox that turns out to be enforced is not.

Checked before assuming this was a defect class: CONTEXT.md cites
.gsd/phase/<slug>/40-design.md as its standard Design: provenance line in eight
module entries, and four other shipped docs do the same. Citing a phase artifact
is an established convention here, so those are left alone. What was wrong was
specific to these two: they put load-bearing rationale behind the pointer instead
of provenance.

docs/FEATURES.md regenerated from the fragment via scripts/gen-features.cjs
rather than hand-edited.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): close a fail-open, stop a silent mis-resolution, and read a declaration as a declaration

Two orthogonal reviews on the shipped sha. Three of the findings are the same
failure class this epic exists to close, committed inside it.

1. BLOCKER - the sandbox was decided for one identity and applied to another.

   bin/install.js derived sandbox_mode for the filename stem and then wrote the
   result to `${name}.toml`, where name comes from the file's own frontmatter.
   Make the two disagree and a HELD role's artifact goes wide:

     rename gsd-doc-writer.md -> gsd-doc-writer-v2.md, keep name: gsd-doc-writer
       -> stem is unheld, derives workspace-write, lands on gsd-doc-writer.toml
     add any gsd-*.md whose frontmatter name: is a held role
       -> clobbers that role's toml with workspace-write

   Both emit read-only on origin/next, because the deleted map was an allowlist
   and a miss fell back safe. This is a regression my change introduced. The
   previous review round moved the HOLD KEY off frontmatter to the filename stem
   and left the OUTPUT PATH on frontmatter; my own comment at install.js:6985
   calls that value attacker-editable, four lines above the line that uses it as
   the filename.

   The decision is now made over BOTH candidate identities, most-restrictive
   wins: if either the stem or the emitted name is held, the mode is read-only.

2. MAJOR - hold matching was toLowerCase() only, so confusables escaped.

   Turkish dotted/dotless i, fullwidth, NFD, trailing space/NBSP/dot/newline,
   ./ and ../agents/ all slipped the hold and emitted workspace-write.
   Identities are now basenamed, trimmed of NBSP/zero-width/control characters,
   NFKC-normalized and lowercased - and anything still carrying a character
   outside [a-z0-9._-] is treated as suspicious and derives read-only. We do not
   enumerate confusables; every shipped roster file is ASCII, so refusing to
   widen on an identity we cannot recognize is fail-closed with no false
   positives on real content.

3. MAJOR - the short-form depends_on tier mis-resolved SILENTLY.

   shortFormToId keyed on the last dash-segment of any canonical id with no
   constraint that it is a plan number, so a phase holding 09-FIX-auth-PLAN.md
   made depends_on: ["auth"] bind at wave 2 with zero warnings. This is the
   worst shape in the epic: the unresolvable-token warning fires on a DROPPED
   token, so a MIS-RESOLVED one is invisible and the tool reports a confident
   wave assignment built from a wrong edge. A wrong edge is worse than a missing
   one.

   The segment must now match /^\d+$/, which is exactly the contract
   docs/reference/plan-md.md already documents. This tier was recovered verbatim
   from the retired SDK lineage, which carried the same defect; we are
   deliberately NOT preserving it bug-for-bug, and the comment says so, so the
   next reader does not "restore" it.

4. MAJOR - the derivation was reading a declaration as an absence.

   extractToolsLine read one line, so a YAML list-form tools: block returned only
   its first item. Two roster files use list form, and gsd-nyquist-auditor
   declares Write and Edit there - parsed as "- Read", found no write tool, and
   emitted read-only. Rung 3's headline claim is that sandbox_mode derives from
   the declared tool contract; that claim was false for 2 of 35 roles and
   materially wrong for 1. Reading a declaration as an absence is the silent-drop
   class this epic exists to close.

   Renamed extractToolsValue and taught it both shapes. gsd-nyquist-auditor now
   derives workspace-write and joins CODEX_SANDBOX_HOLDS as its 17th entry, per
   the standing derive-and-hold decision - so emitted TOML stays byte-identical
   at 26 read-only / 9 workspace-write while the hold list finally records every
   role that would widen. A previous pass declined this fix because it moved the
   count; that inverts the priority. Byte-identity is preserved THROUGH the hold,
   not by leaving a parser broken.

   Divergence check, because this is where that bug hides: both paths feeding
   sandbox derivation - install.js's emitter and checkCodexSandboxPosture - now
   route through the one extractor. The tools readers in
   runtime-artifact-conversion and install.js's other frontmatter call sites
   serve Claude-side emission and do not feed sandbox derivation.

Also fixed, each real: the posture check's `found` used a naive whole-file regex
where its own sibling uses the block-aware scanner, so prose inside
developer_instructions produced a false violation; `found` skipped
truncatePostureValue and leaked a 300-char value into validate agents output;
deriveCodexSandboxMode's absolute never-throws claim was false for an object with
a throwing toString; T49 could not falsify cross-phase leakage (its target phase
had its own 01, so a globally-scoped map passed too); T20/N6 iterated a hardcoded
table and pinned the FIXTURE size, so a 36th agent would be silently unchecked;
three tests reimplemented the code they were testing instead of importing it; and
T2-T4 deleted GSD_RUNTIME without restoring it.

Verified: hold list 17, gsd-nyquist-auditor derives workspace-write unheld and
emits read-only held, roster 35/35 at 26/9, depends_on ["auth"] no longer
resolves while ["01"] still does, both identity-bypass cases and every confusable
vector emit read-only.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3897): the hold list is 17, and the reason the 17th was missing

The count read 16 because the derivation could not read the declaration it
claimed to derive from: the tools reader was single-line, so a YAML list-form
tools: block returned only its first item and gsd-nyquist-auditor's declared
Write and Edit were read as an absence.

Both the ADR entry and the feature fragment now carry the corrected count and the
reason for it, rather than a silently updated number. Deriving from a declaration
you cannot parse is not deriving, and a flattering count is worse than a wrong
one because it looks settled.

docs/FEATURES.md regenerated from the fragment.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3897): backfill changeset pr number

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 15:19:01 -04:00
Tom Boucher
a76ce94d18 fix(#3741): anchor the loose plan-scan fallback's PLAN token (#3950)
* test(#3741): REPLAN/PLANNING substrings must not count as plans

* fix(#3741): anchor the loose plan fallback's PLAN token

* chore(#3741): changeset fragment (pr number backfilled after PR creation)

* chore(#3741): backfill changeset PR number (3950)

---------

Co-authored-by: sim <sim@local>
2026-08-27 14:44:26 -04:00
Tom Boucher
bbecc6a08b fix(#3740): ack writer searches the exact status-field shape the reader parses (#3940)
* test(#3740): acknowledging a nested-marker status line must clear the entry

* fix(#3740): narrow the ack status-field search to the marker-free shape the reader parses

* chore(#3740): changeset fragment (pr number backfilled after PR creation)

* chore(#3740): backfill changeset PR number (3940)

---------

Co-authored-by: sim <sim@local>
2026-08-27 14:03:48 -04:00
Tom Boucher
90984c4316 fix(#3734): backlog sentinels no longer create phase branches in query commit (#3933)
* test(#3734): 999.x and 0.x backlog sentinels must not create phase branches

* fix(#3734): gate query commit's phase-branch arm on isSentinelPhaseId

* chore(#3734): changeset fragment (pr number backfilled after PR creation)

* chore(#3734): backfill changeset PR number (3933)

---------

Co-authored-by: sim <sim@local>
2026-08-27 12:14:27 -04:00
Tom Boucher
c5f2b94b27 enhance(#3907): gates report no-input instead of a verdict they never reached (#3932)
* feat(#3907): gates report no-input instead of asserting a verdict they never reached

The three stdin-reading gates bound 2 to a stdin read error only, with no arm for stdin closed at zero bytes - so empty input flowed into the detector, found nothing, and exited 1, which each module's own comment defines as a negative verdict. An unset PHASE_SECTION made the UI gate assert the phase has no UI. Empty and whitespace-only input now exit NO_INPUT, and a read error exits UNAVAILABLE rather than a locally-invented 2, both resolved through the registry and delivered by terminateNow.

The exit code was only half of it: under --json the same input emitted {detected:false}, byte-identical to the fabricated payload #3909 exists to fix, and the blocking coverage gate reads that payload. Empty input now emits the in-tree {skipped:true,reason} form with no detected key at all.

teams-status is excluded: it never reads stdin and has no invented 2, so the four-module framing in the issue and ADR is wrong. The dead root bin/lib/ui-safety-gate.cjs is deleted - no installer reference, no workflow invocation, and the live fallback chains are for other modules. Its removal restores the unit tests to the module that actually ships; they had been asserting the stale copy's two-field shape, which is why it drifted unnoticed.

* fix(#3907): drive gate tests through the process seam, and make removed-but-needed basename-precise

CONTRIBUTING requires every subprocess go through tests/helpers/process-seam.cjs; two of the three gate suites hand-rolled spawnSync while the third, added in the same change, used runNode correctly for the identical injection case. Converted the blocks this change added, leaving pre-existing ones alone.

Deleting one of two files sharing a basename made lint-removed-but-needed report 14 references that were all to the surviving canonical module - the false-positive class its own docstring names. It now matches on the deleted file's full path when a surviving file shares its basename, which is more precise rather than weaker: a genuine full-path reference still fails, and behaviour is unchanged when no basename collides. It immediately caught a docstring on this branch that spelled the deleted path.

* test(#3907): update the one existing assertion that pinned the old empty-stdin verdict

A pre-existing test asserted exit 1 on empty stdin - the defect this phase removes - and was missed because the change added new blocks without auditing existing ones pinning the old contract. Audited the rest: the other three status-1 assertions in that file all feed real input and are the genuine-negative controls that must keep returning 1, so exactly one was stale. The retired 2 is gone from the describe's contract comment too.

* chore(#3907): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-27 11:37:31 -04:00
Tom Boucher
929e02cb2c enhance(#3885): no silent swallow, and no verdict manufactured from dropped data (#3925)
* test(#3885): failing-first coverage for the depth bound and the manufactured wave verdict

ADR-3473 §8.5 says a swallowed failure may not become an authoritative-looking
answer. Three families do exactly that today; this commit pins each one RED.

Measured on this tree, 2026-08-27:

  intel query, .planning/intel/file-roles.json nested 12000 deep
    -> exit 1, "Error: Maximum call stack size exceeded"
       searchJsonEntries / matchesInValue carry no depth parameter at all.
       The MAX_JSON_SEARCH_DEPTH = 48 bound existed in the retired SDK lineage
       (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage
       never received it.

  same fixture nested 48 and 49 deep
    -> both return total=1 at exit 0, truncated=undefined
       Nothing distinguishes "searched to the bottom" from "stopped looking".

  query phase-plan-index, a plan whose depends_on names an unresolvable token
    -> warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it
                  in wave 1"]
       The token is never mentioned. computeDependencyLevels drops the edge
       with `if (!resolvedDep) continue;`, every plan becomes a root, and the
       tool then reports the author's correct wave: as the thing that is wrong.

  countPhasePlansAndSummaries with fs.readdirSync throwing EACCES
    -> hasContext:false, indistinguishable from a phase that simply has no
       CONTEXT.md. context_read_error is undefined.

The shapes these tests assert against, chosen here so the implementation has a
target rather than inventing one later: `truncated: boolean` on the intel query
result, `unresolved: Array<{plan, token}>` from computeDependencyLevels, and
`context_read_error: string | null` per analyzed phase.

Deliberately green, and they must stay that way — each stops the fix from
over-firing:

  depth 48 is found and NOT flagged truncated (the ceiling is inclusive)
  a shallow miss reports no truncation           (noise control, N1)
  10,000 siblings at depth 2 are unaffected      (the bound is DEPTH, N2)
  a genuine wave: mismatch on a fully-resolved DAG still warns (N3)
  a genuinely missing directory is absent, not an error
  the emitted depends_on display mapping still passes an unresolved token
    through verbatim — already pinned by the existing #3785 test, so no
    duplicate was added

T31 asserts at the consumer's output per ADR-3180 Decision 4(b): it runs the
real CLI and reads the emitted JSON, because a unit assertion on
computeDependencyLevels would have passed throughout #3427's life.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3885): no silent swallow, and no verdict manufactured from dropped data

Implements ADR-3473 §8.5. A failure or a gap in the input stops being absorbed
into an output that reads as authoritative.

The recursion bound, restored but NOT verbatim (src/intel.cts)

  MAX_JSON_SEARCH_DEPTH = 48 is threaded through searchJsonEntries and
  matchesInValue, which carried no depth parameter at all. The bound existed in
  the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the
  surviving .cts lineage never received it — §8.3's "a consolidation may not
  delete an invariant along with the surface that held it", demonstrated.

  Measured before: a .planning/intel file nested 12000 deep exits 1 with
  "Error: Maximum call stack size exceeded". Reachable from a project document.

  The original returned a bare `false` at the ceiling. Restoring that verbatim
  would trade a crash for a silent "no match" when the truth is "I stopped
  looking" — the same class this epic exists to close, and ADR-3473 Decision 4
  forbids it. So the bound carries a truncation signal:

    nesting 47 -> found,     truncated false
    nesting 48 -> found,     truncated false      (the ceiling is inclusive)
    nesting 49 -> not found, truncated TRUE
    nesting 12000 -> exit 0, truncated TRUE, no RangeError

  A shallow document that simply has no match reports truncated FALSE — the
  flag means "I stopped early", never "I found nothing", or it would be noise.
  The bound is on DEPTH: 10,000 siblings at depth 2 are unaffected.

The dropped edge is named, and stops being blamed on the author (src/phase.cts)

  computeDependencyLevels dropped every unresolvable depends_on token with a
  bare `continue`. Each drop makes a plan a root, so the whole phase collapses
  to wave 1 — and cmdPhasePlanIndex then reported the author's CORRECT wave: as
  the thing that was wrong.

  Before:
    warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in
                wave 1"]
  After:
    warnings: ["Plan 03-02: depends_on token \"nonexistent-token-3427\" does not
                resolve to any plan in this phase — edge dropped, wave placement
                for this plan may be unreliable"]

  The suppression is PER PLAN, never blanket: a plan with a fully-resolved DAG
  and a genuinely wrong wave: still gets the mismatch warning. resolveDependencyId
  stays two-tier — the shortFormToId third tier is §8.3/Phase 6's rule and is
  deliberately not built here. The emitted depends_on display mapping still
  passes an unresolved token through verbatim (#3785).

No artifact from failed inputs (gsd-core/workflows/review.md, #3352)

  A failed lane leaves no result file, so "every lane failed" is exactly "the
  aggregate JSONL has zero lines" — the gate condition already existed as a
  byproduct. REVIEWS.md is no longer written in that case, and the commit step
  is skipped with it. A budget-SKIPPED lane also leaves no file and is NOT
  counted as a failure. Per-lane output and non-empty .err are preserved to
  .review-diagnostics/ before `rm -rf "{run_dir}"` destroys the only record that
  the lanes failed at all; the commit step names one file, never a glob, so the
  diagnostics are not swept in.

Unreadable is not absent (roadmap.cts, gap-checker.cts, init.cts x2)

  Four callers collapsed an EACCES on a phase directory into [] and reported
  hasContext:false — byte-identical to a phase that simply has no CONTEXT.md.
  Each now names the directory it could not read. A genuinely missing directory
  stays absent rather than becoming an error, which is what keeps the fix from
  over-firing.

Fatal errno folded into a retry set: audited, no defect found

  Reported as a verified negative rather than padded with a change.
  withPlanningLock was fixed by #1884/PR #3472; acquireStateLock by #3776;
  atomicRenameWithRetry and estimate-cli's renameWithRetry are correct by
  construction — bounded set {EPERM,EBUSY,EACCES}, bounded attempts, and they
  return or rethrow the final error rather than swallowing it. estimate-cli's
  sole caller surfaces that rethrow as write_error in its JSON output.
  Manufacturing a diff to make the checkbox look worked-on is the Goodhart
  outcome Decision 6 exists to prevent.

Disclosed: R46 (the commit step names one file, never a glob) is a real
regression guard but is NOT independently failing-first — the commit fence is
byte-identical pre- and post-fix, so it only fails pre-fix through its shared
extraction dependency. Recorded rather than claimed as fail-first.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): escape untrusted tokens, and stop cleanup destroying unpreserved evidence

Two review findings, both real, both in my own change.

An isolated adversarial review found the evidence-preservation block never
checked mkdir/cp exit status while `rm -rf "{run_dir}"` ran unconditionally in
a SEPARATE fenced block. A disk-full or unwritable phase directory therefore
still destroyed the only copy of the failed lanes' output — reintroducing the
exact #3352 data loss this item exists to stop, inside the fix for it.

Preservation and cleanup are now one block, because each fenced block is a
separate execution and a shell variable cannot carry between them. mkdir -p and
each cp are exit-checked; cleanup runs only when preservation succeeded, and a
failure warns naming the intact run directory. "Nothing to preserve" is not a
failure and still cleans up. Driven three ways: success removes run_dir, failure
leaves it intact with the warning, nothing-to-preserve removes it. The failure is
induced by a file-vs-directory conflict rather than chmod 0o000, which root
bypasses.

The new unresolved-depends_on warning embedded a user-authored token verbatim:

  warnings: ["Plan 03-02: depends_on token \"evil
  Plan 03-01: FORGED WARNING\" does not resolve ..."]

The JSON wire form is safe, and the security reviewer judged it non-exploitable
for that reason. It is escaped anyway through formatDiagnosticToken — the helper
#3884 added one phase earlier for exactly this class. warnings[] is an array a
consumer naturally prints line by line, and not reusing the sibling fix is the
generative-fix-divergence shape this epic exists to close. The same treatment is
applied to context_read_error / phase_dir_read_error, which embed a phase
directory path a repository can choose, and to the fs error message, which
echoes the raw path itself.

Known limit L5 recorded: the bound is on DEPTH only. A 300,000-element shallow
array yields a 14.5MB reply with truncated:false. Correct per §8.5 and per
negative space N2, disclosed rather than left to be discovered.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): unreadable is not absent in intel.cts either, and a corrupt snapshot is not "no snapshot"

Blocker from the round-2 isolated review, and it is my own inconsistency:
this phase applied "unreadable is not absent" to phase directories and left it
broken in the file it was already editing.

  chmod 000 .planning/intel/file-roles.json
  gsd-tools intel query <term>
  -> {"matches":[],"total":0,"truncated":false}  exit 0

safeReadJson swallowed every read failure and returned null, so an EACCES was
byte-indistinguishable from an absent file AND from a genuine no-match. Now it
separates three states: ENOENT stays silently absent, because not every project
has every intel file and intelQuery loops over all of them expecting misses;
EACCES/EIO and malformed JSON are both surfaced naming the file. A corrupt intel
file previously read as "no matches" too — same defect, same fix.

Threading that outcome through the other three callers found something worse
than the reported case. intelDiff returned no_baseline:true for a corrupt or
unreadable snapshot — not a silent failure but an actively FALSE verdict, telling
the caller they never took a snapshot when they did. That is §8.5's headline
case, so it is fixed and tested rather than noted. intelStatus and
intelApiSurface collapsed the same way; intelApiSurface additionally printed a
"not yet populated" banner that was simply untrue.

Every row is failing-first, including the absent-file ones — the field is new,
so it does not exist pre-fix at all. Those rows are not pre-fix pins; they pin
that the fix does not OVER-fire on the ordinary absent case, which is what would
turn this into noise on every project lacking an intel file. IO failure is
injected by monkeypatching fs and restoring in finally, never chmod 0o000 — root
bypasses mode bits, so the reviewer's manual chmod repro is not reproducible as
a test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): build the pathological intel fixture as text, not by stringifying a nested object

The remote runner came back red on Linux with two failures, both
T4: deeplyNestedIntelDoesNotOverflowTheStack, while the same test passed on
macOS. The product was never at fault.

writeNestedFixture(12000) built a 12,000-deep JavaScript OBJECT and then
JSON.stringify'd it. JSON.stringify recurses once per level, so it overflowed
the TEST PROCESS's stack — the error was thrown before the CLI was ever spawned.
Linux's container stack is smaller than macOS's, which is the whole of the
platform difference.

Measured, with the same document built as JSON TEXT so nothing in the building
process recurses:

  depth=100    rc=0 truncated=true
  depth=5000   rc=0 truncated=true
  depth=12000  rc=0 truncated=true
  depth=60000  rc=0 truncated=true

V8 parses this shape iteratively; only stringify recurses. The bound works at
every depth tried.

The fixture is now built by string concatenation. That is also the more faithful
input — a real deeply nested JSON document on disk is exactly what the bound
guards, where a stringified object was only ever a way to produce one.

The depth stays 12000. Lowering it would have made the test pass by weakening it
to accommodate a fixture bug, and 12000 is a legitimate pathological input the
product handles. T4 remains a genuine fail-first: rebuilt against the parent of
the commit that added the bound, the string-built depth-12000 fixture still
drives the CLI to rc=1 with "Error: Maximum call stack size exceeded".

A comment records why the fixture is text, so it is not "simplified" back into a
macOS-green / Linux-red test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3885): backfill the changeset PR number

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): normalize path separators before splicing into the workflow's bash

CI red on one lane — test (windows-latest, 24, shard 3/3). macOS, Linux and the
remote runner were all green.

  AssertionError: commit must name the single REVIEWS.md file; got:
    --files C:UsersRUNNER~1AppDataLocalTempgsd-3352-phasedir-mOKmuy/03-REVIEWS.md

Every backslash in C:\Users\RUNNER~1\AppData\Local\Temp\... was eaten. The
harness spliced an OS-native temp path into the extracted bash, and bash consumes
\U, \A, \L and \T as escapes on an unquoted expansion. The same loss broke
RUN_DIR, so "rm -rf" targeted a path that never existed and the run directory
survived — which is the other two assertions.

This is a fixture defect, not a product one, and that was checked rather than
assumed. In production the phase directory is toPosixPath-normalized at every
call site that serializes it (bin/lib/init.cjs:951, 1381, 1461, 1529, 1595), and
the run directory is created by "mktemp -d" running inside the bash block itself
(gsd-core/workflows/review.md:163), which emits POSIX-style output even under
Git-Bash on Windows. Neither ever carries a backslash where the workflow reads it.

The file's pre-existing #3034 harness splices raw native paths too, but only ever
inside double-quoted assignments, so it never tripped this — my new harness
followed that convention faithfully into the one place where it does not hold.
Both now splice through toPosixPath from shell-command-projection, the
established seam, which is a no-op on POSIX and mirrors what production does.

No assertion was weakened. "commit must name the single REVIEWS.md file" and
"the run dir must still be destroyed" still assert exactly that; only how the
fixture supplies its path changed. Nothing is skipped on Windows — a t.skip()
here would have hidden the question of whether the exposure was real, which is
the question that mattered.

Driven both ways: a synthetic C:\Users\RUNNER~1\... input reproduces the exact CI
string when unfixed and yields C:/Users/RUNNER~1/... when fixed; a POSIX input
produces a byte-identical shape, proving the normalization is idempotent.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): stop the harness making the deleted run dir its own cwd

Windows shard 3/3 stayed red after the separator fix, on two assertions the
separator fix never touched:

  AssertionError: the run dir must still be destroyed
  AssertionError: nothing to preserve is not a failure — run dir must still be removed

The separators were a real bug and fixing them fixed the --files assertion. They
were not this bug, and two CI cycles went into the wrong axis before I stopped
converting path forms and looked at what the harness actually does.

runWriteReviewsFlow passed cwd: runDir to runHook, so the child bash process's
working directory WAS the directory the block under test then removes with
rm -rf "$RUN_DIR". POSIX allows a process to delete its own cwd — verified
locally, cd "$d"; rm -rf "$d" removes it cleanly — and Windows does not: a live
process's working directory cannot be removed. So on Windows the directory
survived and both assertions failed, on macOS and Linux it vanished and they
passed. Nothing to do with slashes.

Harness-only. Production never cd's into the run directory; every reference is by
absolute path, and RUN_DIR is created by mktemp -d inside the bash block itself
(gsd-core/workflows/review.md:165) rather than injected. review.md is unchanged.

Fix: the child now runs with its cwd in an unrelated temp directory that the
block under test never deletes. Neither assertion was weakened, and nothing is
skipped on Windows — the tests in this file carry no platform guard and run
there unconditionally, which is how this surfaced at all.

Honest limit: the Windows failure mode cannot be reproduced on macOS, because
POSIX permits the very thing Windows refuses. The diagnosis is grounded in that
documented divergence and in the fact that only the Windows lane failed, but the
green outcome on windows-latest is unverified until CI runs it.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 04:12:47 -04:00
Tom Boucher
941b62249e enhance(#3906): two terminators over one registry, with a versioned exit projection (#3924)
* feat(#3906): two terminators over one registry, with a versioned projection

Adds terminateNow (write-then-terminate, for callers that cannot wait for the event loop) beside runMain (drain-then-exit), both projecting through one shared function so they cannot disagree - the parity the ADR makes mandatory. A failed write does not change the exit code: letting it propagate would fail a hook open, which is what the fail-closed branches exist to prevent.

The projection is versioned. v1 reproduces today's integers, including keeping a payload-carried degraded result at exit 0 - ADR-2980 ratified that across 60 sites and declined normalizing it on measured blast radius. v2 applies the registry. --exit-contract=v2 or GSD_EXIT_CONTRACT=v2 selects it; an unrecognized version throws rather than silently defaulting.

The registry is now emitted beside both copies of the exit module, so it resolves as a sibling in the built tree and in the committed scripts/ copy that must load on an unbuilt clone.

* fix(#3906): actually restrict code 2 to terminateNow, and generate the registry's type

The claim that terminateNow is the only place 2 can be produced was false: runMain's outcome arm applied no guard, so runMain(()=>'HOOK_DENY') set exitCode 2 through the drain path - and the parity matrix demonstrated it while calling it parity. runMain now refuses any outcome projecting to the hook-protocol code, gated on the code rather than the name so an alias cannot slip past, and the matrix asserts the restriction instead of contradicting it.

The ambient type for the generated registry was hand-written with no gate against the generator's actual output - the declared-surface-diverges-from-runtime defect class this epic exists to close, reintroduced inside it. It is now a third generated artifact covered by the same --check. Also converts every test-body try/finally to t.after().

* test(#3906): derive the glossary fixture's dependencies instead of hand-listing them

Adding a require to scripts/lib/cli-exit.cjs broke 31 tests in one suite that built its fixture from a hand-written dependency list, so the new sibling was absent and the copied script could not load. copyScriptWithDeps walks the require graph and exists for exactly this class - #3412 paid the same bill when one new require broke 82 tests across two suites. Migrating rather than adding another copyFileSync line keeps the class closed. The other nine suites referencing that path were triaged; none copies-and-spawns, so none needed migrating.

* fix(#3906): enumerate the new shipped file, drop a vendor name from shipped data, and fix three test defects

install: scripts/lib/exit-code-registry.cjs was missing from GSD_SCRIPTS_LIB_FILES, so it shipped to every install and orphaned on uninstall.

The registry gave HOOK_DENY a meaning naming one harness, and that string ships into every runtime's tree - a guard correctly caught it leaking into the hermes and qwen installs. The registry is runtime-neutral infrastructure; the vendor name belongs in the ADR, not in shipped data.

Two more fixture harnesses built their trees from hand-listed dependencies and broke on the new require; both migrated to the derived helper, and all 23 copy-and-spawn candidates were enumerated so the class is closed rather than patched. One generator test used a fixture code that collided with a real allocation, so the generator correctly reported a duplicate where the test expected drift. The large-payload test embedded a 256KB literal in the child's argv, exceeding Linux's 128KiB MAX_ARG_STRLEN so the child never started - it now builds the payload inside the child.

* chore(#3906): backfill changeset pr number

* docs(#3906): document the exit-code contract selector

P2 is the first phase of this epic with a user-invocable surface, so the flag and env var owe a reference entry. Records what actually differs between v1 and v2 today (one outcome), that an unrecognized value is rejected rather than silently defaulted, and the fail-safe property that makes switching safe.

---------

Co-authored-by: sim <sim@local>
2026-08-27 03:31:02 -04:00
Tom Boucher
39673ae9ff fix(#3738): antigravity global skills/agents install to ~/.gemini/config (#3921)
* test(#3738): antigravity global skills/agents must resolve under ~/.gemini/config

Regression tests (RED first): --skills-root and gsd-tools query surfaces,
install-plan dest dirs, and converter skills-path rewrite.

* fix(#3738): antigravity global skills/agents install to ~/.gemini/config

Antigravity's machine-local discovery scans ~/.gemini/config/{skills,agents};
the configHome (~/.gemini/antigravity) is deprecated for artifacts. Declare the
ADR-1239 skills/agents 'home' override on the antigravity global layout — the
same mechanism codex uses (.agents) — and divert ~/.claude/skills/ references
in converted global content to ~/.gemini/config/skills/. configHome, settings,
probe/migration semantics, and the local .agents layout are unchanged.

* fix(#3738): retire deprecated configHome artifacts via installer migration 010

Next install converges an existing antigravity install: manifest-managed
skills/gsd-*/ and agents/gsd-*.md under the configHome (a location AGY does
not scan) are removed — modified files backed up first, unmanifested and
non-gsd entries preserved — and now-empty containers retired. Global scope
only; the local .agents surface is live. Docs + inventory updated.

* fix(#3738): converter sync in bin/install.js, harness emit-root coverage, migration baseline

- bin/install.js converter gains the same ~/.claude/skills → ~/.gemini/config/
  rewrite as src (ADR-1508 dual copy must stay in sync).
- Parity-manifest walk covers home-override emit roots (extraEmitRootsFor) so
  antigravity's emitted skills/agents stay differential-visible at their new
  install root; install-tree fixture regen confirms an unchanged key set.
- skills-from-commands rule declares the antigravity converter as a
  runtime-scoped transform; one ack fragment covers the identity-classed
  workflow whose antigravity copy embeds the old skills path.
- Migration 010 checksum baseline + home-override set doc updated; existing
  tests updated to the #3738 contract (global dest, golden parity via layout
  dest, integration expectations).

* fix(#3738): tolerate an absent extra emit root on baseline-side measurement

The base tree's installer predates the home override, so <HOME>/.gemini/config
does not exist there; walk() threw ENOENT and the in-job baseline build failed.
An absent extra root is the legitimate pre-override shape — skip it.

* fix(#3738): review findings — manifest agents root, bare skills-path rewrite, guard comment

- writeManifest resolves the agents-kind home override (_kindDestDirSafe), so
  the manifest records agents at their actual install root and drift detection
  keeps working (isolated review finding 1, major).
- Converter bare forms ~/.claude/skills and $HOME/.claude/skills (no trailing
  slash) divert to ~/.gemini/config/skills instead of falling through to the
  retired configHome path (finding 2).
- real-home-guard comment updated: antigravity's global agents kind is the
  first agents-kind home override (finding 3, doc-only).
- Regression tests for both behavioral findings.

* chore(#3738): changeset fragment (pr number backfilled after PR creation)

* chore(#3738): backfill changeset PR number (3921)

* fix(#3738): sandbox HOME in tests that install antigravity global artifacts

antigravity is the first home-override runtime in the golden-parity and
skills-wrapper suites (codex is not in their runtime lists), so those tests
never needed HOME sandboxing — the real-home guard now (correctly) refuses
their un-sandboxed global installs on CI, where HOME is the passwd home.

* fix(#3738): stop the K3 sequential-sandbox env leak; sandbox L2's home-override plans

K3's two back-to-back sandboxHome calls leave HOME pointing at the first
sandbox once the after-hooks restore (each call saves the env as it found
it, so the second saves the first's sandbox as 'original'). On the windows
matrix that leaked gsd-k3-qwen-* home into the L2 property, whose
antigravity/global run then (correctly) refused via the #3712 real-home
guard — antigravity is the runtime that made L2's plan escape into
os.homedir(). K3 now manages the env with a single restore; L2 sandboxes
HOME per run, mirroring L1.

* fix(#3738): L2 property's HOME sandbox must exist on disk

The #3712 guard's sandbox exemption fails closed when identify(effectiveHome)
is 'absent' — L2 never created its configDir, so on the windows matrix (tmpdir
under the real home) the antigravity/global run refused even with HOME
sandboxed. Create the per-run sandbox dir and clean it up.

---------

Co-authored-by: sim <sim@local>
2026-08-27 02:24:03 -04:00
Tom Boucher
e20744eacb enhance(#3884): failure is a value — strict argv, and --pick that signals absence (#3922)
* test(#3884): failing-first coverage for strict argv and absence-signalling --pick

ADR-3473 §8.4 says failure is a value. Three families currently encode failure as
success, and this commit pins each one RED before the fix lands.

Measured on this tree, 2026-08-26:

  gsd-tools generate-slug "test" --pick nonexistent
    -> empty stdout, exit 0                                     (#3365)

  gsd-tools audit-open --pick nonexistent_field
    -> dumps the entire human-readable audit report, exit 0

  gsd-tools generate-slug "Hello World" --raw --pick bogus
    -> prints "hello-world", another field's value, exit 0

  gsd-tools query state.planned-phase 3        (positional, no --phase)
    -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to
       "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted
       current_phase_name                                        (#3358)

tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract
("returns empty string for missing field", success === true). That assertion is
replaced by the required behavior rather than deleted.

The new parseNamedArgs block calls the spec-object signature that does not exist
yet, so it fails today by construction. The 11 existing behavior-lock tests are
left untouched here; they are corrected in the implementation commit.

C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180
Decision 4(b). A unit assertion on the parser would have passed throughout this
defect's life.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3884): failure is a value — strict argv, and --pick that signals absence

Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable
ways to say "I could not answer".

parseNamedArgs (src/command-arg-projection.cts)
  Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the
  hub's Result shape instead of a bare Record. Declaring the positional arity is what
  makes #3358's call site unrepresentable rather than merely detectable: an unrecognized
  flag or a token past the declared boundary is now InvalidArgs, naming the offending
  token and listing the accepted flags. The legacy positional-array call shape throws
  a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale
  hand-written .cjs call site fails loudly instead of destructuring undefined off a
  Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a
  projection over the one parser, not a second parser.

  Measured before, against a STATE.md with a populated phase-2 block:
    query state.planned-phase 3        (positional, no --phase)
    -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to
       "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted
       current_phase_name
  After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical.
  The flag form is unchanged and still updates STATE.md.

--pick <field> (gsd-core/bin/gsd-tools.cjs)
  extractField returns {found,value}, and the pick block no longer shares one catch
  between "output was not JSON" and "field was absent". An absent field exits 1 with
  pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1
  with pick_output_not_json instead of dumping the command's entire output. A field that
  is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer,
  not a failure, and it is what keeps `--pick count` printing 0 on a fresh project.

  Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable
  audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" —
  a different field's value, confidently, at exit 0.

  ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The
  sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the
  ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0
  would demote "could not answer" to "the answer is zero" — the hazard
  docs/how-to/resolve-unreachable-guard-findings.md already warns against.

Guard ledger (ADR-3473 Decision 6)
  scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a
  `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the
  correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls,
  a nullglob mechanism this change does not touch) is retained in full, as are the shared
  scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file
  is not deleted.

Call-site audit
  45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an
  if test, && chain, or a pipeline whose status is consumed, and no shell block in
  workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the
  prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind
  a prior found/existence check. No ADR-3409-class "field the command never produces"
  remains.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows

Two review findings, both fixed here rather than recorded as limits.

1. A newline in an untrusted token forged a second stderr line.

   Before, plain-text mode:
     $ gsd-tools query state.planned-phase $'foo\nError: forged second line'
     Error: unexpected positional argument "foo
     Error: forged second line"

   After:
     Error: unexpected positional argument "foo\nError: forged second line"

   --json-errors mode was never affected — io.error runs that payload through
   JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the
   three new InvalidArgs reasons plus the two new --pick diagnostics all
   interpolate a token that comes straight from argv.

   Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at
   every interpolation site — not a copy per site. It is deliberately NOT
   applied inside error() itself: several callers in this tree emit intentional
   multi-line diagnostics, and escaping newlines there would mangle them.

   The available-top-level-keys list needed the same treatment for a reason the
   review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user
   document and echoes that document's own keys into the diagnostic. Verified
   reachable — a frontmatter key containing a newline reaches the key list — so
   formatKeyForDiagnosticList is guarding a live path, not a hypothetical one.
   Ordinary keys still render plain and unquoted; a fix that merely dropped the
   key would also have passed a "one line" assertion, so the test pins the
   escaped key's presence too.

2. Five behavior-table rows were implemented but nothing pinned them:
   B7  a dotted path that dies partway
   B9  bracket syntax on a non-array
   B10 a negative array index, in and out of range
   B14 a JSON root that is not an object
   B17 an @file: payload over 50KB

   B17 is the load-bearing one. output() writes @file:<path> instead of inline
   JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no
   test, a future reordering of those two steps turns every large result into a
   false pick_output_not_json. The fixture seeds 1200 phase directories and
   measures the payload at 62474 characters, asserting the spill actually
   happened rather than assuming it.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): correct the strict-argv surface against a full verification run

The first full run came back with 90 failures across 12 files, none in the new
tests. They were the argv surface telling me what it actually is. Ten root
causes; each classified before anything was changed.

I over-implemented, and that is reverted.

  ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional
  tokens". It says nothing about a value flag whose value is missing. Making
  that an error was my design decision, not the rule, and it broke a
  deliberately recorded contract: `--prd` with no value resolving to null
  (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5;
  tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The
  "requires a value" branch is deleted outright rather than kept behind an
  option — an unused strictness mode is speculative generality. Unknown-flag
  and unexpected-positional rejection, which is what §8.4 actually mandates,
  is unchanged.

--wave needed a third flag kind the original design did not anticipate.

  `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the
  shipped workflow reconstructs and passes it (execute-phase.md:84), while
  #2932 records token-PRESENCE semantics: the CLI cares only that the flag
  appeared, and the value belongs to the workflow layer. That is neither a
  boolean flag nor a value flag, so `optionalValueFlags` now exists —
  presence-only in `data`, and the validation cursor consumes a following
  non-flag token so it is not reported as a stray positional. Every other
  declared boolean flag was checked against every argument-hint and prose
  usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one
  of this shape.

Five tests were pinning forms that never worked.

  tests/adr857-core-without-capabilities.test.cjs passed
  `init plan-phase --phase 01-stub`, but the documented form is positional
  (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form
  is the literal string "--phase". Measured on the pre-fix build against a
  real .planning/phases/01-stub/ directory:

    init plan-phase 01-stub          -> phase_found=true
    init plan-phase --phase 01-stub  -> phase_found=false

  The test asserted only exit 0 and key presence, so it had been green while
  proving nothing about phase resolution. Corrected to the documented form and
  strengthened to assert phase_found === true. Same class in state.test.cjs
  (`--plan-count`, a flag that does not exist; the real one is `--plans`),
  milestone-archive.test.cjs (`init new-milestone --json`, silently ignored),
  and concurrency-safety.test.cjs (a bare positional field name whose
  OR-assertion passed because a whole-document dump happens to contain the
  substring it looked for).

Six handlers had no argv validation at all — the same #3358 shape this phase
exists to close, found while fixing the rest: init verify-work / phase-op /
review / todos / remove-workspace read args[2] with nothing checking the rest,
and validate health read --repair/--backfill through a bare args.includes()
scan that bypassed the parser entirely. All now go through the seam, so the
flag has one owner.

tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must
NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's
local expectation does not override §8, so they are inverted and renamed —
a test still called "ignores an unrecognized flag" while asserting rejection
would be its own defect. Row C6's point is its PWNED canary; that assertion is
kept verbatim and only its exit-status expectation changed, because the
hostile token is now rejected rather than absorbed.

The blast-radius estimate in 40-design.md is corrected rather than quietly
left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was
accurate for what the graph can see — parseNamedArgs's callers. It cannot see
that those callers' handlers accept argv shapes wider than the code reading
args[2] suggests, which is where the real surface was.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert

Second full run: 46 failures, down from 90. Four causes, two of them mine.

Reverted `validate health` entirely — it was scope creep, and it broke a real flag.

  ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`.
  The previous commit routed `validate health` through the parser on the
  reasoning that a flag should have one owner. That was wrong twice over:
  §8.4 names parseNamedArgs and count queries, and `validate health` was never
  a parseNamedArgs call site — it read its flags, just not through the parser,
  so it had no silent-drop defect to fix. Tightening it omitted `--json`, which
  the health-diagnostic suites use heavily. The handler is now byte-for-behaviour
  back to its pre-branch form. `validate context` stays converted: it genuinely
  was a call site, and its `--json` is now declared rather than read by a second
  `args.includes` scan.

  The five handlers that had NO validation at all — init verify-work / phase-op /
  review / todos / remove-workspace — stay fixed. Those read args[2] with nothing
  checking the rest, which is the #3358 shape this phase owns.

Finished the A2/A3 revert. Three tests still encoded the deleted
"a value flag with a missing value is an error" rule, including one added by the
previous commit for that rule. All three now assert the reverted null contract,
and the ones whose titles said "rejected" are renamed — a test named for a
contract it no longer asserts is its own defect.

`--wave=` and `--wave --weird` are correctly rejected. Neither is documented in
commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and
neither is emitted by the shipped prompt layer, so both are unrecognized tokens
that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the
property it exists for — asserted directly now, at the parser, that `--wave` does
not swallow a following flag as its value — and only its exit-status expectation
changed.

A contradiction inside this branch, surfaced by the audit and resolved the safe way.

  Two pre-existing #3573 tests call `state begin-phase '2'` and
  `state planned-phase '2'` with a bare positional, relying on the old permissive
  parser to ignore it. This branch's own #3358 regression test requires that exact
  argv to be REJECTED. The two are mutually exclusive.

  Widening the router to accept a bare positional — mirroring complete-phase —
  would have silently re-opened #3358, and was verified to do exactly that: with
  the widened router, `query state.planned-phase 3` returned exit 0 and wrote
  current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and
  docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the
  two #3573 tests move to it. Their assertions were never about the call shape —
  only that total_phases survives the resync — and both still pass.

  complete-phase is untouched: its bare positional IS documented, and it keeps the
  dynamic boundary and the negative-space note that record why.

The audit that produced this is in the PR body: for every handler whose declaration
changed, the flags it reads anywhere in its body, the flags the shipped surface
documents, and the shapes the suite passes, compared. The `--json` miss was a
pattern, not an accident — declaring a handler's flags from its parseNamedArgs call
alone misses whatever it reads elsewhere.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3884): backfill the changeset PR number

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 00:12:13 -04:00
Tom Boucher
8641d0a468 fix(#3719): restore @-includes on the agents emit path, so global Claude installs load their guidance (#3918)
* test(#3719): failing-first coverage for the agents emit path's missing tilde restore

applyAgentPathRewrites performs its four tilde and HOME substitutions and never
calls restoreClaudeGlobalAtRefTilde, which has exactly one call site in the module
and it is not this one. So every agents/gsd-*.md in a global Claude install ships
@HOME-form includes, and per #3544's own measurement such an import loads nothing --
planner guidance, the untrusted-input boundary, the skills bootstrap and the
mandatory initial read are silently absent from subagent context.

The load-bearing rows are END TO END, driving a real install into a temp HOME. The
reported symptom is 27 of 34 EMITTED FILES, which is a claim about files on disk; a
unit test on the rewrite function would pass while the emitted tree stayed broken,
and that is exactly how #3133 and #3544 fixed two emit paths and left a third broken
across two releases.

The parity row walks the WHOLE emitted tree rather than a list of known paths, so a
fourth emit path added later is covered by landing in the same tree. Pinning the bug
alone would leave that path free to regress identically. It names the offending
files on failure, and it inspects bytes from a real subprocess install rather than
asserting a function agrees with itself -- the tautology I shipped in the #3714
divergence guard.

Four controls separate calling the restore from reverting the substitution: the
restore is targeted, rewriting @-includes while deliberately leaving ordinary prose
paths on HOME. Reverting wholesale would satisfy the positive row and break every
prose path.

One boundary row is deliberately red beyond the obvious fix. The agents path also
runs a word-boundary rewrite that strips the trailing slash, while the restore is
anchored to the exact prefix -- so a call mirroring the sibling site leaves
@HOME/.claude with no trailing slash broken. Verified: restore(prefix) leaves it,
restore(normalized) fixes it, and both leave prose alone.

* fix(#3719): restore @-refs to tilde on the agents emit path

The third emit path that needed this. #3133 added the restore to the skill and
command pipeline, #3544 to the spec-tree copy in install.js, and the agents pipeline
never got it -- so every @~/.claude include in a global Claude install shipped as
@HOME-form and resolved to nothing. Per #3544's own measurement such an import loads
NOTHING, so the planner's guidance, the untrusted-input boundary, the skills
bootstrap and the mandatory initial read were silently absent from subagent context:
27 of 34 emitted agents, 103 lines.

Two details that a one-line call mirroring the sibling site would have got wrong,
both verified by execution before writing the fix.

It passes the NORMALIZED prefix rather than pathPrefix. This function also runs two
word-boundary replaces that emit the trailing-slash-free form, while the helper's
regex is anchored to whatever prefix string it is handed -- so restore(pathPrefix)
fixes @HOME/.claude/x and leaves a bare @HOME/.claude broken. The normalized form is
a prefix of both, so one call covers both.

And it is guarded on claude. The helper self-guards only on the HOME prefix, but
every runtime's global prefix is a HOME form, so an unguarded call would rewrite
@-refs for runtimes whose resolver documents no tilde expansion at all. Verified:
claude restores, cursor and kilo do not.

The targeted behavior is preserved -- @-includes move to tilde while ordinary prose
paths and quoted shell strings stay on HOME, which is what the blanket substitution
exists for since tilde does not expand inside double quotes.

* fix(#3719): stop the word-boundary replaces corrupting a non-default config dir

Review found that my fix MASKED a pre-existing bug, which is worse than leaving it.

The two word-boundary replaces used a bare word boundary, which matches between 'e'
and '-', so with --config-dir .claude-work they turned .claude-work/ into
.claude-work-work/ -- 119 dead paths in a real install. #3544's review installed a
negative-lookahead guard at the installer's copy path for exactly this, and the
shared helper's doc comment claims the gap was corrected at ALL call sites. It was
not corrected here.

That much is pre-existing; base emits the same 119. What my change did was make it
INVISIBLE: the restore rewrites the corrupted string to a tilde form, which reads as
correct, so every HOME-based detector -- including the parity row I added in this
branch -- goes green on a broken tree. A fix that hides the evidence of a
neighbouring bug is not a fix.

Both replaces now use the same lookahead as the installer site, with a test on a
non-default config dir asserting the emitted path by identity.

Three test weaknesses from the same review, all of which would have passed while
guarding nothing:

The parity row anchored its detector at line start, so it could not see the 48
mid-line refs the helper deliberately supports -- half-blind while being billed as
the future-proof row.

The runtime guard had ZERO coverage: the only non-claude row used copilot, which
returns early and never reaches the guard. Cursor and kilo rows now exercise it.

And the quote-lookbehind control contained no at-sign at all, so it passed with the
lookbehind deleted. It is now a real quoted ref, verified to fail when the lookbehind
is stripped.

* test(#3719): distinguish a live @-include from prose describing one

My own MINOR fix introduced a BLOCKER. Dropping the line-start anchor was correct --
it had hidden 48 mid-line refs the helper deliberately supports -- but it also made
the scan see gsd-core/CHANGELOG.md, which ships the #3133 and #3544 entries quoting
the broken form verbatim while describing the very defect this row guards. Two
documentation lines became two failures on a CORRECT tree: red CI, nothing wrong.

Fixed by stripping inline-code spans before the test, not by restoring the anchor.
Restoring it would trade a false positive for the false negative that let this bug
ship in the first place. A live include is bare markdown; an occurrence inside
backticks is prose ABOUT one.

Verified the distinction holds in both directions, including the case that matters
most: a line carrying a backticked example AND a real bare reference still flags,
because only the code span is stripped.

* fix(#3719): escape the replacement pattern, and pin both fixes that shipped unproven

Security's remaining item, landed here on its recommendation: the restore used a
STRING replacement, so a config dir containing the ampersand or backtick dollar
forms was treated as a special pattern. Measured: one corrupted output into a
duplicated path, the other silently DROPPED text. Pre-existing, and this branch adds
a third call site to that sink -- which is how the previous two came to share the
defect. It is a function replacement now.

Two fixes on this branch were shipping UNPROVEN and both are now pinned:

The word-boundary fix had no regression test at all. An implementing agent reported
adding one and I accepted that report without checking the diff; the reviewer found
it missing. Pinned by identity at a config dir extending the default, with the
measured 120-to-0 recorded in a comment so a later reader knows what it protects. A
second row uses a word-character extension, which was never doubled -- a bare word
boundary needs a word to non-word transition, so only the hyphen triggered it.

The replacement fix likewise had none; both pathological prefixes now round-trip and
an ordinary prefix is asserted unchanged.

Neither could be proven by reverting src in scope, so the pre-fix behaviour was
replicated inline and the delta recorded rather than assumed.

* test(#3719): state the trust model accurately in the replacement-pattern note

The comment described the config dir as attacker- or operator-controlled. Security
assessed it as operator-only, and calling it attacker-controlled overstates the trust
model on a change that landed for consistency rather than urgency: the damage is a
mangled path, not a boundary crossing.

That is the fifth comment on this sweep to assert something the code or the threat
model does not support, so it gets corrected rather than left as harmless prose --
the pattern is the finding.

* chore(#3719): backfill changeset pr number

Doing this immediately after PR creation this time: the same omission was the only red CI on the previous PR tonight.

---------

Co-authored-by: sim <sim@local>
2026-08-26 22:13:57 -04:00