Commit Graph

1087 Commits

Author SHA1 Message Date
Tom Boucher
17e163f15c docs(#4333): document the ADR Amends/Amended-by convention (#4334)
* docs(#4333): document the ADR Amends/Amended-by convention

Two patterns for amending an accepted ADR are established practice —
an in-place `## Amendment (YYYY-MM-DD)` section, and a separate ADR
that declares `Amends` with a reciprocal `Amended by` back-link — but
only the first was ever written down. #4030 shows the cost: a
contributor concluded no ADR owned a contract that ADR-857 already
covers, because nothing said the second pattern (used by ADR-1244 and
ADR-2782 to extend ADR-857 itself) existed.

Document both patterns in docs/contributor-standards.md, note the
Amends/Amended-by reciprocity rule in docs/adr/README.md alongside the
existing Supersedes/Subsumes rule (and that it isn't yet gated by
scripts/gen-adr-index.cjs the way those are), and point CONTRIBUTING.md's
new-ADR process at the amendment path for revisiting an existing one.

Closes #4333

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4333): fix imprecise Amends/Amended-by precedent citations

Orthogonal review caught two inaccuracies: PR #1643 doesn't match the
in-place dated-section pattern (it rewrites the original Decision text
rather than appending an untouched dated section), and ADR-1244's
relationship to ADR-857 is prose ("extended by"), not the structured
Amends/Amended-by header field. ADR-2782 is the verified precedent for
the structured field pair — its one Amends field names four targets
(857, 894, 1016, 1244), all four carrying the reciprocal back-link.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 16:45:57 -04:00
Dennis Alexis Valin Dittrich
1017898cb9 fix(#3771): make remediation examples non-binding and surface revision conflicts (#3916)
* fix(#3771): separate the binding property from the advisory remediation

Checker findings fused "what property failed" with "how to fix it" into a
single `fix_hint` and never marked which half binds. The checker rendered
every hint under a "must fix" heading, the orchestrators injected the issues
verbatim and ordered targeted updates, and the shared revision references
mapped each hint to a prescriptive strategy — so a contract-following planner
applied a hint literally even when a smaller mechanism satisfied the same
property, or when the hint contradicted a locked decision. There was no
channel to report that conflict, and every attempt burned a revision
iteration.

Checker side: every issue now carries a binding `required_property` (the
invariant that failed) plus its evidence and severity, and `fix_hint` is
labelled non-binding wherever it appears — including the human-facing blocker
rendering, so "must fix" unambiguously names the property and never the
example.

Planner side: revision re-checks locked decisions, capability guidance and
existing plan constraints before editing; satisfying a blocker through a
smaller valid alternative counts as addressing it; and a hint that conflicts
with any of those returns `REVISION_CONFLICT` carrying the conflict and the
alternatives considered. Orchestrators route that to user choice or the
configured plan-review convergence loop without consuming retry budget.

Also applied to the UI-spec revision loop and the gap-plan hint, and the
generic pattern's stray `suggested_fix` field name is reconciled to the
plan-checker's `fix_hint`.

Nothing legitimately binding is weakened: blockers still block, severity
still gates, iteration caps and stall escalation still fire, and required
task fields and decision coverage still hold.

Refs #3771

* test(#3771): pin the binding/advisory split across the revision chain

Locks the separation at every link that carries it: the checker's issue
schema and blocker rendering, the planner's constraint re-check and
REVISION_CONFLICT return, the generic pattern's reconciled field names, and
each orchestrator's conflict routing without retry-budget consumption. Also
pins what must not have been weakened — blockers, severity gating, iteration
caps and stall escalation.

Red against the pre-fix prose: 32 of 34 assertions fail (the 2 that pass are
the preservation checks, correctly).

Refs #3771

* chore(#3771): add changeset fragment for the remediation-binding fix

* chore(#3771): acknowledge the remediation-binding growth

Five runtime-loaded files grow: the two checkers carry the binding/advisory
split where the model reads it (a `required_property` on every dimension
example, since a schema the examples contradict teaches the examples), and
the three orchestrators carry the REVISION_CONFLICT route, which has to live
with the `iteration_count`/`revision_count` state it declines to spend.

Deletes tests/emitted-drift-acks/3172-stated-failing-direction.json: it is
fully spent on next and still owned plan-phase.md, so it walls off a key it
can no longer clear (#3078). Its removal is the documented remedy for the
duplicate-key collision, not drive-by cleanup.

* fix(#3771): close the review gaps in the conflict contract

Adversarial review (Codex, Antigravity) found four real defects in the first
pass, each confirmed against the source before acting:

- The UI checker's structured return still ordered `Fix: {exact fix required}`
  and "list each BLOCK dimension with exact fix required". The dimension
  examples had been marked non-binding but the rendering the researcher
  actually reads had not — the same omission this issue is about.
- `ui-phase` and the canonical `revision-loop` flow incremented their counter
  BEFORE dispatching the reviser, so "do NOT increment on REVISION_CONFLICT"
  was unreachable prose: the iteration was already spent. The increment now
  sits on the return path in both.
- The conflict gate offered "accept as-is", which is an early exit from a
  still-failing blocker — a weakening the brief explicitly forbids. The three
  options are now adopt an alternative / override the constraint / amend the
  constraint; every one resolves the conflict. Accepting an unaddressed blocker
  remains available only at the unchanged iteration-cap escalation.
- The convergence route was declarative: nothing in
  plan-review-convergence.md could receive a conflict. plan-phase now records
  it in REVIEWS.md — the channel that loop already consumes — convergence
  refuses to declare convergence over an open entry, and routing back into a
  run convergence itself started is explicitly excluded as a cycle. `quick` has
  no REVIEWS.md and no phase, so its convergence branch was dead prose and is
  deleted in favour of asking the user.

Also reconciles the last two drifted field names (`finding`, `affected_field`)
to the plan-checker schema, and repairs a silent no-op: the few-shot
`required_property` insertion never applied because those lines are
blockquoted, and the test's own block filter was anchored on indentation only,
so a vacuous loop passed over zero blocks. Both are fixed and the filter now
asserts it found blocks.

Refs #3771

* chore(#3771): extend the growth acknowledgment for the review round

plan-review-convergence.md joins the list: the conflict route needed a
receiving end, and it lands on the seam that loop already reads (REVIEWS.md)
rather than a new mechanism. The plan-phase, ui-phase and gsd-ui-checker
entries gain the second-pass reasoning — an executable convergence branch, the
increment moved onto the return path, and the structured return that still
ordered an exact fix.

* fix(#3771): make the conflict route bounded, ordered, and owned

Round-2 adversarial review found five more defects, each confirmed in the
source before acting:

- The convergence gate sat AFTER `gsd_run state planned-phase` and the success
  banner, so a run could write and announce convergence over an unresolved
  conflict. OPEN_CONFLICTS is now read from REVIEWS.md and is part of the
  converged CONDITION, evaluated before any write.
- plan-phase's cycle-exclusion ("unless this run was invoked by convergence")
  was not a question the orchestrator can answer at runtime. plan-phase now
  never invokes convergence at all — it records the conflict when a phase
  REVIEWS.md exists and resolves it with the user in-place, which removes the
  cycle instead of describing it.
- Closure had no owner. plan-phase writes the row, so plan-phase strikes it
  resolved; convergence only reads. An open row is a live blocker, never a
  stale artifact.
- Declining to increment the counter removed the only bound on the conflict
  path: an agent returning the same conflict forever would loop unattended. A
  conflict naming the same `required_property` twice in a row is now a stall
  and escalates through the existing gate.
- verify-work's gap-plan revision loop hands `<revision_context>` to
  gsd-planner and so inherits the whole contract, but stated none of it and
  could not handle the conflict return. It is now covered like the others, and
  is in the test's orchestrator table.

Refs #3771

* chore(#3771): acknowledge the round-2 growth

verify-work.md joins the list — the flow the second review found missed — and
the plan-phase, ui-phase and plan-review-convergence entries gain the
round-2 reasoning: the gate moved ahead of the state write, the convergence
hand-off replaced with a runtime-checkable record-and-resolve, and the
recurrence bound that replaces the counter the conflict path stopped spending.

* fix(#3771): make the convergence gate countable and stop the conflict fall-through

Third adversarial round (Antigravity) found three defects:

- The OPEN_CONFLICTS pipeline had no `grep -v '~~'` despite its own comment
  claiming one, and `grep -c '^| '` also counts a markdown table's header and
  separator rows — every resolved conflict would have read as open and
  convergence would have deadlocked instead of converging. plan-phase now
  records each conflict as a `- [ ]` checklist line and flips it to `- [x]`, so
  the gate is an exact fixed-string match with no table parsing.
- "then continue below" fell through to the checker spawn, so a SECOND
  REVISION_CONFLICT would have been handed to the checker as though it were a
  revised plan. plan-phase, quick and verify-work now re-evaluate the return
  from the top of the conflict handler; ui-phase already looped back.
- revision-loop.md still described plan-phase routing a conflict to the
  convergence loop instead of asking — the behaviour round 2 removed. Recording
  is now stated as being in addition to asking, never instead of it.

Refs #3771

* chore(#3771): bring the changeset in line with what shipped

Two review rounds widened the change after the fragment was written:
verify-work's gap-plan revision and the convergence loop are covered, two
more drifted field names are reconciled, and the conflict path carries an
explicit recurrence bound.

* fix(#3771): declare and emit the REVISION_CONFLICT marker

check:contract-drift on CI caught what local lint never reached: four
workflows dispatch on `## REVISION_CONFLICT`, but no agent declared or emitted
it — an orphan consumer, matching a marker nothing produces. The shared
reference (planner-revision.md Step 7b) described the return; the agent
definitions did not carry it.

gsd-planner and gsd-ui-researcher now emit the marker in-fence alongside their
other return markers, and both registry rows in agent-contracts.md declare it.
gsd-planner's Consumed by gains the two workflows that dispatch on it and were
missing from the row.

The gate is right: a return contract belongs where the agent is defined, not
only in a reference the agent happens to load.

Refs #3771

* chore(#3771): acknowledge the return-marker growth

gsd-planner.md and gsd-ui-researcher.md each gain the REVISION_CONFLICT
marker that check:contract-drift requires them to emit.

* fix(#3771): hoist the shared conflict protocol out of the workflows

Two CI failures, both correct gates:

- tests/few-shot-calibration.test.cjs pins the plan-checker calibration file
  at exactly 4 examples (2 positive, 2 negative). The example added in the
  first pass broke that balance — and described PLANNER behaviour in the
  CHECKER's calibration set, which is the wrong surface for it. Removed; the
  smaller-alternative rule is already normative in gsd-plan-checker.md and
  planner-revision.md, and pinned by the regression suite.
- tests/phase6-capstone-conformance.test.cjs (ADR-857 phase 6, #1168) requires
  plan-phase.md to stay BELOW its pre-phase-6 baseline of 94519 bytes. The
  inline conflict block pushed it to 94988.

The fix for the second is the one that should have been made first: the
record/resolve/close protocol and the recurrence bound were identical in four
workflows, and revision-loop.md — which plan-phase already @-imports — is what
a shared contract is for. The protocol now lives there once; plan-phase states
only its bindings (which counter, which artifact, which next step) and points
at it. plan-phase.md: 94988 -> 92739, under the ratchet with headroom, and the
four-way duplication is gone.

quick, ui-phase and verify-work do not import the reference, so they keep their
inline statements. The suite asserts each rule against what the runtime
actually loads for that orchestrator, not against the file in isolation.

Refs #3771

* docs(#3771): state the shared-protocol relationship accurately

Three of the four revision-bearing workflows do not @-import revision-loop.md,
so 'follows it verbatim' overstated the coupling. Only plan-phase defers; the
others restate the rules inline and this section is the authority they must
agree with.

* refactor(#3771): name the authority instead of restating it four times

Self-review finding: the same ~700-byte reviser paragraph was inlined in four
prompts while gsd-planner already loads planner-revision.md whenever
<revision_context> is present (agents/gsd-planner.md:555, :588) — a fifth copy
of the same contract. Four places to edit in lockstep is precisely the drift
class this PR exists to fix; the generic pattern calling the field
suggested_fix while the checker emitted fix_hint is what that looks like after
a year.

Each prompt now carries only the load-bearing clauses and names the authority
it summarises. Deliberately NOT reduced to a bare pointer: these are LLM
prompts, and a contract stated only in a file the reader is supposed to fetch
is the failure mode of this very bug. The saving is modest (~50 bytes each) —
the point is the named source of truth, not the bytes.

Suite: dropped seven assertions that pinned heading text and bold-lead
phrasing a reword would break without changing what the runtime is told; their
neighbours already pin the same contract by content. 519 -> 499 lines, 61
tests. Red gate against origin/next: 56 of 61 fail.

Refs #3771

* fix(#3771): sanitize agent-authored conflict text and bound total conflicts

Cross-AI review (agy/Gemini) raised two security findings. I confirmed the
first by running the gate's own awk against a crafted file and reproduced it
exactly; the second is a real hole in my bound.

1. FAIL-OPEN in the convergence gate. The conflict text is agent-authored and
   is appended to REVIEWS.md, which the gate scans with an awk that stops at
   the next '## ' heading. One line of agent text beginning '## ' ends that
   scan early, so conflicts below it are never counted and convergence declares
   success over a live blocker. Measured: 3 open conflicts, awk returned 2.

   Fixed at the write boundary, which is the trust boundary: every field has
   newlines and tabs collapsed to spaces and a leading '#', '-', '|' or fence
   stripped, so one conflict is exactly one line. Both producing agents now
   declare their fields single-line plain text, and the reader states the
   invariant it depends on so a later edit cannot silently break it. Verified:
   3 open + 1 resolved now counts 3; missing file and absent section count 0.

2. The recurrence bound was 'same required_property twice in a row', which an
   agent alternating property names never trips, leaving the un-incremented
   conflict path unbounded. Now bounded twice: the repeat rule catches the
   common case, and the THIRD conflict return of a loop escalates whatever
   property it names. A conflict still never consumes a revision iteration;
   this cap is separate from and additional to the revision cap.

Rejected from the same review: deleting 'a planner that reaches
required_property by a smaller or different mechanism has addressed the issue
in full' from the CHECKER prompt as misplaced. It is load-bearing exactly
there. A checker that does not know a different mechanism counts will re-flag
the issue on re-check, which is the revision loop that never terminates. The
argument offered for deleting it, that the checker evaluates the new state
independently, describes the failure mode.

Refs #3771

* fix(#3771): fail closed on an unverifiable convergence gate

Second cross-AI pass (agy, this time with the full files rather than the diff)
found two more, both real:

1. The gate read REVIEWS_FILE with `2>/dev/null || echo 0`, so an unreadable or
   empty path counted as ZERO open conflicts and converged. That path is
   resolved a few lines earlier by a pre-existing unquoted
   `ls ${phase_dir}/${padded_phase}-REVIEWS.md` (line 346, not touched by this
   PR), which yields an empty string rather than an error when the path
   contains a space. Unverifiable is not clean: the gate now tests -z and -r
   first and BLOCKS. Verified both branches.

   The unquoted ls itself is left alone deliberately — it predates this change
   and belongs to the reviews lookup, not the conflict gate. Fixing it at my
   own boundary removes its effect on this gate without widening scope.

2. REVIEWS.md is writable by the review agent, which could flip a `- [ ]` to
   `- [x]` or delete the section and forge the state of a blocking gate. The
   section now declares a single writer: /gsd:plan-phase appends and closes,
   every other agent leaves it byte-for-byte alone, readers read.

Also trimmed a clause that explained the increment ordering by reference to
what the file said before this PR. Commit history is not instruction, and
these files are prompts.

Rejected: the claim that quick's conflict gate deadlocks autonomous pipelines
by asking the user. Its existing max-iteration escalation in the same file
already asks the user the same way; this adds no new interaction class.
Noted but out of scope: the per-dimension YAML example blocks and the shim
boilerplate duplicated across agent prompts both predate this change.

Refs #3771

* fix(#3771): count conflicts by line shape, not by section

CodeRabbit review on the rehearsal PR. Five findings, all valid, all applied.

The best one is a deletion. The convergence gate scanned between
'## Plan-Revision Conflicts' and the next '## ' heading, and that scan stops at
the FIRST heading it meets — so one stray '## ' line hid every conflict beneath
it and returned 0, converging over a live blocker. Reproduced: section-scan 0,
shape-scan 1. Sanitizing at the write boundary does not cover a hand-edited,
legacy, or foreign-written REVIEWS.md, so the reader needed its own guarantee.

It now matches the conflict line SHAPE anywhere in the file:

  grep -c '^- \[ \] .*required_property:'

No section bookkeeping, nothing a heading can truncate, and it composes with the
writer's sanitization (which strips a leading '-' from agent text, so agent prose
cannot forge the shape). Verified: injected heading -> 1, all resolved -> 0.

The other four:

- Both checkers told the author never to emit a contradictory fix_hint, then
  offered an escape hatch that put the forbidden route in the hint anyway. They
  now name NO route in that case and state only that the property conflicts with
  the constraint. A hint carrying a forbidden route is applied by anyone who
  trusts hints.
- The REVISION_CONFLICT marker description in gsd-planner.md was narrower than
  planner-revision.md: it covered a contradictory hint but not an unreachable
  required_property. A planner reading only the agent file would have burned
  retry budget on the case the reference routes to a conflict.
- The few-shot calibration examples used uppercase BLOCKER/INFO while the schema
  defines blocker/warning/info. Pre-existing, but it is the same schema-vs-example
  disagreement this PR exists to end, and the file was already being edited.
- verify-work's re-entry instruction existed but sat after the Bounded clause, so
  the paragraph read "re-spawn ... stop re-spawning ... after re-spawning". The
  re-entry now immediately follows the re-spawn, and states that only a
  non-conflict return may reach the checker or increment iteration_count.

Refs #3771

* fix(#3771): resolve the contradictory scope_sanity severity examples

Sixth CodeRabbit finding, posted outside the diff range and missed on my first
read — I had claimed all findings were addressed after reading only the five
inline comments. This one was in the review body.

agents/gsd-plan-checker.md carried TWO scope_sanity examples with identical
metrics (5 tasks, 12 files) and OPPOSITE severities: warning in Dimension 5,
blocker in <examples>. Line 872 states "2-3 tasks/plan good, 4 warning, 5+
blocker" and the severity table lists warning as "Scope 4 tasks (borderline)",
so the warning example contradicted both.

ADR-2629 Decision 5's "over budget is a WARNING, never a blocker" does not
excuse it: that rule governs the smart-zone TOKEN estimate (the estimate-check
verb, lines 299-306), which is a different axis from task count. Verified in
source before touching it.

The contradiction is pre-existing but this PR made it binding and visible:
severity is now declared part of the binding payload, and both examples were
given the same required_property, so they now disagree on the severity of an
identical finding about an identical property.

Deviating from the proposed correction, which was warning -> blocker: that
would duplicate the <examples> entry outright (same tasks, files, severity).
The Dimension 5 example is instead made a genuine 4-task borderline warning, so
the file keeps one worked example per severity and the thresholds, the severity
table and both examples finally agree.

Refs #3771

* fix(#3771): stop laundering a grep error into zero open conflicts

Seventh CodeRabbit finding — from a SECOND review round my own CR-4 push
triggered, which I had not looked for. This one is a regression I introduced
while fixing the previous fail-open.

CR-4 replaced the truncatable section scan with:

  OPEN_CONFLICTS=$(grep -c '^- \[ \] .*required_property:' "$REVIEWS_FILE" || true)

`|| true` masks every grep failure. grep exits 1 for "no matches" (a legitimate
zero) but 2 for a read error, and `|| true` turns both into an empty capture
that `${OPEN_CONFLICTS:-0}` renders as 0. If REVIEWS.md is removed or becomes
unreadable between the -r check and the scan, the gate reports no conflicts and
convergence proceeds. Proven: unreadable file -> captured empty -> 0.

The status is now inspected, and only exit 1 counts as zero; anything else
blocks.

My first attempt at this fix was itself wrong and my own harness caught it: I
wrote `if ! grep ...; then grep_status=$?`, but `!` inverts the status, so `$?`
in that branch is 0 and every failure reads as success — the clean-file case
printed "BLOCKED (grep exit 0)". The status must be read in the ELSE branch of a
non-negated `if`, which is what CodeRabbit proposed. Both traps are now pinned
by tests.

Verified end to end: all resolved -> 0, no conflicts at all -> 0, injected
heading -> 1, unreadable file -> BLOCKED with grep exit 2.

Refs #3771

* test(#3771): execute the conflict gate instead of reading it

CodeRabbit round three: 0 actionable, 1 nitpick — "these assertions inspect
Markdown source only; they do not prove that grep status 1 produces zero
conflicts or that a scan error exits before convergence." Rated Trivial. It is
the most valuable finding of the three rounds.

This gate has been wrong three times: a section scan a heading could truncate, a
`|| true` that laundered grep's error status into zero, and an `if !` whose `$?`
reported the negation rather than the command. Every one of those passed the
text assertions that existed at the time. I proved each fix by hand in a shell,
and none of that proof lived in the suite.

The gate is one self-contained fenced block, so the test now extracts it from
the workflow — located by content, not line number — writes it to a script and
RUNS it against fixtures: two open plus one resolved counts 2; no matches counts
0 and does not fail; a conflict below an injected `## ` heading still counts; an
unreadable path and an empty path both BLOCK with a non-zero status and no zero
count on stdout.

Non-vacuity proven by mutation rather than asserted. Reverting the gate to each
of its three historical broken forms reds the suite:

  section-scan awk  -> 7 failures (5 in the gate cases)
  || true           -> 4 failures (3 in the gate cases)
  if ! (negated $?) -> 4 failures (3 in the gate cases)
  restored          -> 69 pass, 0 fail

The prose assertions stay: they are the right instrument for a prompt. This
covers the one part of the change that is real shell an orchestrator executes.

Refs #3771

* test(#3771): route the gate harness through the shared test helpers

ESLint's project rules caught three violations in the new harness: an unbounded
execFileSync (DEFECT.UNBOUNDED-SUBPROCESS — an unbounded spawn is an indefinite
hang, and on macOS CI that is how a stuck run stops reporting instead of failing)
and two raw fs.rmSync calls, which skip the Windows-EBUSY retry budget that
helpers.cleanup carries.

Now uses createTempDir/cleanup from tests/helpers.cjs and passes an explicit
30s timeout. Suppressing the rules was available and would have been the wrong
call: both exist because of real CI failure modes on platforms I am not testing
on.

* chore(#3771): backfill the changeset PR number

The pr: field is drift-checked against the PR event payload, so it cannot be
written before the PR exists. Set to 3916.

* fix(#3771): close revision conflict persistence gaps

Use the authoritative review artifact, keep conflict and normal retry paths disjoint, and enforce one writer-reader grammar so malformed state fails closed.

Emitted-Drift-Ack-Growth: diagnose-issues.md — #3771 marks the gap-plan remediation hint non-binding while keeping root_cause authoritative
Emitted-Drift-Ack-Growth: gsd-plan-checker.md — #3771 separates binding required_property evidence from advisory fix_hint examples across the checker contract
Emitted-Drift-Ack-Growth: gsd-planner.md — #3771 declares the REVISION_CONFLICT return used when remediation contradicts governing constraints
Emitted-Drift-Ack-Growth: gsd-ui-checker.md — #3771 applies the same binding-property and advisory-hint split to UI review findings
Emitted-Drift-Ack-Growth: gsd-ui-researcher.md — #3771 defines the UI revision producer's structured REVISION_CONFLICT return
Emitted-Drift-Ack-Growth: plan-phase.md — #3771 routes and persists bounded revision conflicts before spending the normal retry budget
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #3771 adds the fail-closed owned-block parser and prevents convergence over open conflicts
Emitted-Drift-Ack-Growth: review.md — #3771 emits and preserves the canonical writer-owned conflict block across review regeneration
Emitted-Drift-Ack-Growth: ui-phase.md — #3771 routes UI revision conflicts to resolution before consuming revision_count
Emitted-Drift-Ack-Growth: verify-work.md — #3771 gives gap-plan revision the same bounded conflict route before iteration_count

* test(#3916): guard rebases against schema drift

Load the current-base progressive-disclosure examples so every integrated issue remains bound by required_property after branch reconciliation.

* test(#3916): skip the extracted-gate suite's bash spawns on win32

Third review round's sole survivor: runConflictGate()/withReviews() spawn
bash against a Node-native temp path built by createTempDir(), which is
backslash-separated on the Windows CI lane and not a path Git Bash is
guaranteed to accept (DEFECT.WINDOWS-TEST-PORTABILITY, matching the
observed CI failure at revision-remediation-binding.test.cjs:844,
ENOENT on a path Windows read as a directory separator). No eslint rule
catches it since the call has neither a chmod nor a `bash -c` form.

Guards the four call sites with the repo's existing skipOnWin32
convention (describe/test `{ skip: IS_WINDOWS }`) rather than
normalizing the harness path to forward slashes, which would defeat the
one test whose purpose is proving the production gate does NOT rewrite
a literal backslash in a POSIX filename.

* fix(#3916): backfill changeset pr field to the fork validation PR number

* fix(#3771): forbid silently accepting an open plan-revision conflict at max-cycles escalation

The max-cycles escalation prompt only surfaced HIGH_COUNT and ACTIONABLE_COUNT; an open
plan-revision conflict (OPEN_CONFLICTS > 0) was never disclosed there, and "Proceed anyway"
could exit successfully over it — exactly the failure mode this PR exists to close (a success
banner over an unresolved conflict nobody resolved). Blockers still block: withhold "Proceed
anyway" and route to Manual review whenever a conflict is open.

* fix(#3771): do not hard-block REVISION_CONFLICT persistence when no REVIEWS.md exists yet

A phase's first-ever revision cycle can return REVISION_CONFLICT before any REVIEWS.md has been
written — REVIEWS_PATH is then legitimately empty, not a corrupt or deleted file. The persistence
gate's own accompanying prose already says the record channel applies 'when REVIEWS_FILE is
non-empty', but the bash condition never checked that, so it hard-blocked every conflict on a
brand-new phase regardless of whether persistence was even expected to run. Require a non-empty
REVIEWS_FILE before treating a missing file as an error.

* chore(#3771): raise the plan-phase.md ADR-857 host-loop ceiling to 96700

The frozen pre-phase-6 ceiling (94519) collided on rebase: this PR's own
REVISION_CONFLICT persistence/routing gate is core planner control flow, not an
un-extracted optional feature, and landed alongside an unrelated, already-merged
same-file growth (the #4.6 context-drift pre-check) already on next. Same
rationale #1298 already established for execute-phase.md's ceiling.

* chore(#3916): backfill changeset pr field to the upstream PR number

* fix(#3771): make the writer-side REVISION_CONFLICT sanitize step real shell

The Conflict Return record channel sanitized agent-authored fields via a
prose instruction ("Sanitize each agent-authored field before appending")
for the orchestrator LLM to apply by hand, while the reader-side gate in
plan-review-convergence.md parses the same slot with real, executed awk.
Flagged Minor across two review rounds (round 4, round 6) since no code
performed the sanitize anywhere.

plan-phase.md's Conflict Return step now runs a real bash gate: sanitize
each field (collapse newline/tab to space, strip a leading #/-/|/fence),
build the one-line record, skip the append if an identical line already
exists (idempotent), insert before the writer-owned end delimiter, and
fail closed if that delimiter is missing rather than silently dropping
the conflict.

tests/revision-remediation-binding.test.cjs extracts and RUNS the new
fence (matching how the reader gate is already tested), composing it
with the existing reader gate: hostile-field fast-check fuzzing, a
repeated-conflict idempotency check, and a missing-delimiter fail-closed
check that the file is left byte-for-byte unchanged on failure.

Emitted-Drift-Ack-Growth: plan-phase.md — #3916 turns the writer-side
REVISION_CONFLICT sanitize+insert step into real, executed shell instead
of a prose instruction, matching the reader gate's existing rigor

* fix(#3916): backfill changeset pr field to the fork validation PR number

Fork CI's changeset-lint reads the real PR number from its own event
payload; the fragment still carried the upstream number from the last
sync, so the DEFECT.CHANGESET-PR-FIELD-DRIFT check failed on this fork
PR. Re-backfill to the upstream number before the final push.

* fix(#3771): close the awk -v forgery and same-session close gaps agy found

Adversarial review (gemini-3.8-flash-high via the internal agy review
lane) on the full PR found two BLOCKERs against the just-added
writer-side conflict gate:

1. `awk -v line="${LINE}"` decodes escape sequences in its argument, so
   a literal two-character `\n` in agent-authored text became a real
   newline inside awk, splitting the appended record across two
   physical lines. `tr` only strips actual control bytes, so it never
   saw this — it defeated the exact forgery the gate exists to
   prevent, both the reader's zero-count and the writer's own
   idempotency check. Fixed by passing LINE/END through awk's
   ENVIRON, which is not escape-decoded.

2. A conflict resolved and re-spawned within the same plan-phase
   session was never flipped from `- [ ]` to `- [x]` — the record
   channel bullet said "plan-phase closes it," but no step did. Only
   a *separate* `--reviews` re-entry (line ~622, still prose-only)
   closes conflicts; the in-session resolve path left them open
   forever, permanently blocking convergence. Fixed by carrying the
   just-written line in `PENDING_CONFLICT` and closing it in the
   `Otherwise` branch before the checker re-spawns.

Also fixed a MAJOR: docs/COMMANDS.md described the `--max-cycles`
escalation gate as uniformly offering "proceed or review manually,"
but the code (this PR's own change) withholds "Proceed anyway"
specifically when a plan-revision conflict is open — only manual
review is offered in that case. Docs now say so.

Not applied: the reviewer's `\r` truncated to plain tr from a MINOR
that also asked for temp-file permission preservation across `mktemp`.
Applying `chmod --reference` is not portable to macOS/BSD `chmod`, so
this is left as a documented low-severity tradeoff — the temp file now
sits alongside REVIEWS.md (same filesystem, atomic `mv`), which was
the same finding's more substantive half. Also not applied: a
suggested `gsd_run review record-conflict` CLI subcommand to
deduplicate the two `awk` blocks — a new command plus wiring is out of
scope for a review-remediation fix.

tests/revision-remediation-binding.test.cjs adds regression coverage
for both BLOCKERs: a literal-backslash-n hostile field composed with
the reader gate, and a close-gate extraction that verifies the flip to
`[x]`, the reader's count dropping to 0, and a fail-closed path when
the pending line is missing.

Emitted-Drift-Ack-Growth: plan-phase.md — #3916 fixes an awk -v escape-
decoding forgery and adds the missing same-session conflict-close step
an adversarial review found in the writer-side gate

* chore(#3916): backfill changeset pr field to the upstream PR number

Fork-validation CI needed pr: 1 to pass its own changeset-lint; restore
pr: 3916 before this push reaches open-gsd/gsd-core.

* fix(#3771): trim plan-phase.md prose back under the XL byte cap

Merging origin/next's unrelated growth pushed plan-phase.md 473 bytes
past the workflow-size-budget XL cap and the ADR-857 phase-6 baseline,
both tripped by CI after review approval. Removed an unpinned inert
bash comment and tightened connective prose in three REVISION_CONFLICT
bullets; no executable shell or test-pinned substring changed.

* chore(rehearsal): pin changeset pr field to fork rehearsal PR #25

Scratch-only commit for the rehearsal branch's own CI. Will not be
carried onto the branch backing upstream #3916 — that keeps pr: 3916.

* fix(#3771): address CodeRabbit findings on the REVISION_CONFLICT protocol

Fork rehearsal PR #25's first CodeRabbit pass surfaced 7 findings against
the already-approved #3916 diff; each verified against current code
before fixing (none hallucinated):

- plan-phase.md: writer-side awk gates now strip a trailing \r before
  comparing lines, matching the reader gate (plan-review-convergence.md)
  -- a CRLF REVIEWS.md previously made both writer gates fail closed.
- plan-phase.md: the close-fence's REVIEWS_FILE/PENDING_CONFLICT/
  CONFLICT_RESOLUTION were read without ever being (re)defined in that
  fence -- shell state does not survive across separate fenced blocks
  (same convention already documented in review.md). Added the explicit
  recompute/set instruction.
- revision-loop.md: previous_conflict_property was never reset after a
  normal (non-conflict) revision, so a later, unrelated conflict on the
  same property could be misread as a repeat and escalate prematurely.
- gsd-plan-checker.md / few-shot-examples/plan-checker.md: two example
  required_property strings were unconditionally binding in a way their
  own dimension's rules aren't (no-analog RESEARCH.md fallback; tasks
  that create no functions), now scoped to match.
- quick/steps/plan-checker-loop.md: added the same disjoint
  "Otherwise (not REVISION_CONFLICT)" branch plan-phase.md already had,
  closing an ambiguity between the conflict and non-conflict return paths.
- revision-remediation-binding.test.cjs: the REVIEWS_PATH init-order
  assertion used indexOf() without checking for -1, so it would pass
  vacuously if either anchor were renamed away.

Also restores an "Export the row's CONFLICT_*" instruction I had cut in
the prior byte-budget trim -- checked non-pinned by tests, but it was the
only text telling the agent to set those vars before the awk block reads
them via ENVIRON.

Net growth from these fixes required reclaiming bytes elsewhere in
plan-phase.md (verified against every pinned substring in
revision-remediation-binding.test.cjs) to stay under the XL tier's
hard 98304-byte cap; final size 98245 bytes.

* fix(#3771): resync the #4079 shrink-only mirror to the current PRE_PHASE6 line

tests/plan-phase-background-wait-wakeup.test.cjs (landed on next via an
unrelated #4079 PR, merged in by this branch's next-sync) mirrored
plan-phase.md's phase6 shrink-only ceiling as a hardcoded local constant
(94519) rather than reading tests/phase6-capstone-conformance.test.cjs's
PRE_PHASE6 value. That value has since been legitimately raised twice
during this PR's own review (94519 -> 96700 -> 98300) to accommodate the
REVISION_CONFLICT persistence/routing gate. The two branches' independent
histories left the mirror stale post-merge -- not a textual git conflict,
but the same class of thing. Resynced to 98300.

* fix(#3771): address round-2 CodeRabbit findings on the conflict gates

CodeRabbit's re-review of the previous remediation commit found two real
issues in what it had already flagged:

- Both writer-side awk CRLF fixes used \`sub(/\r$/, "")\` directly on \`\$0\`,
  which mutates it in place -- \`{ print }\` then emitted the CR-stripped
  copy for every passed-through line, silently rewriting an unrelated
  CRLF REVIEWS.md to LF on any insert or close. Now compares against a
  separate \`cur\` copy and prints the original, untouched \`\$0\`.
- The close-fence's "recompute REVIEWS_FILE/PENDING_CONFLICT" prose
  implied in-fence derivation, but the fence has no such code and the
  test harness (\`runCloseGate\`) deliberately supplies all three as
  pre-set env vars -- matching how the open fence's "Export the row's
  CONFLICT_*" instruction already works. Reworded to "export ... in the
  same invocation", matching that established, test-verified pattern
  instead of promising logic that isn't there.

Added a regression test proving the CRLF fix no longer touches
passthrough lines (red against the mutate-in-place version, green now).

* fix(#3771): use a CRLF-safe check in the new passthrough regression test

local/no-crlf-fragile-split forbids splitting readFileSync content on a
literal \n (Windows git-autocrlf checkouts yield \r\n). My CRLF
passthrough-preservation test from the previous commit did exactly that
to inspect the first line. Replaced with a direct startsWith() check
against the known CRLF-terminated header, which needs no split.

* test(#3771): assert the record itself is inserted in the CRLF passthrough test

CodeRabbit nitpick (round 3): the passthrough-preservation test checked
gate status and the pre-existing line's CRLF ending, but never asserted
the new REVISION_CONFLICT record was actually written.

* fix(#3771): address agy/gemini-3.8-flash-high adversarial review findings

Full-PR adversarial review (internal /gsd-review antigravity lane,
gemini-3.8-flash-high) surfaced 9 findings; each verified against current
code before fixing (none hallucinated):

HIGH:
- quick-batch/steps/plan-checker-loop.md never received the
  required_property/fix_hint binding language or REVISION_CONFLICT
  handling this PR added everywhere else -- a genuinely unmigrated
  producing context. Migrated to match quick/steps/plan-checker-loop.md,
  and added it to the ORCHESTRATORS consistency battery in
  revision-remediation-binding.test.cjs so future drift is caught
  automatically.
- The close-fence's PENDING_CONFLICT was an agent-supplied env var that
  had to exactly reconstruct a five-field sanitized line across a
  multi-minute subagent dispatch -- fragile, and a scalar var also meant
  a second simultaneous conflict silently dropped the first on overwrite.
  Redesigned to match the open conflict by CONFLICT_DIMENSION/
  CONFLICT_PLAN identity instead: the agent re-supplies two short,
  already-tracked identifiers rather than reconstructing the full
  sanitized text, and each conflict resolves independently regardless of
  how many are open. Updated the test harness's runCloseGate contract to
  match, and added a two-open-conflicts regression test.

MEDIUM:
- plan-phase.md's `--reviews` replanning path told the reader to "flip
  the matching line to [x]" in prose only, with no executable path to
  it -- pointed it at the same close gate used in step 12.
- plan-review-convergence.md's reader-gate awk tolerated a blank line
  before the opening delimiter but not before the heading that follows
  it; a formatter or LLM writer inserting one would hard-abort
  convergence on an otherwise well-formed REVIEWS.md. Added the same
  tolerance already granted above it, with a regression test.

LOW:
- Clarified that the escalation destination for a stalled conflict is
  the same iteration/revision-count cap gate already defined in each of
  quick, quick-batch, ui-phase, and verify-work, rather than an
  undefined "stall" concept.
- Clarified "twice in a row" means no successful revision intervened,
  matching revision-loop.md's now-explicit previous_conflict_property
  reset.
- Fixed gsd-ui-researcher.md's stale rationale text, copied verbatim
  from planner-revision.md: ui-phase presents the conflict table
  directly to the user, it does not persist to a shared file scanned by
  heading.

Net growth again required reclaiming bytes in plan-phase.md (verified
against every pinned substring in revision-remediation-binding.test.cjs)
to stay under the XL tier's hard 98304-byte cap; removed a now-dead
PENDING_CONFLICT assignment in the process. Final size 98258 bytes.

* fix(#3771): scope row 48's quick/steps guard away from plan-checker-loop.md

tests/gsd-quick-batch-quick-regression.test.cjs's row 48 (#3676) flagged
this branch's quick-batch/steps/plan-checker-loop.md migration (the agy
HIGH finding) as a violation, because it also edits
quick/steps/plan-checker-loop.md for the same underlying #3771 protocol
fix.

Verified against git history before scoping: 2f64e6230 (#3676's own
landing commit) CREATED quick-batch/steps/plan-checker-loop.md as a new,
independent 119-line file, never a call-site into quick/'s copy. Row
48's "shared primitives, never edits the ordinary quick command" premise
was never about this specific file -- it was always meant to carry its
own per-flow copy of whatever revision-loop contract applies, same as
ui-phase.md/verify-work.md throughout this PR. This is the same
false-positive class the row's own comments already document scoping
away twice (#3730, #2529 round 40); excluded plan-checker-loop.md from
its touched-quick-steps check with the same evidence trail.

* chore(#3771): point changeset pr field at upstream PR 3916

---------

Co-authored-by: davdittrich <davdittrich@gmail.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 15:16:38 -04:00
Michel Moreira
86b745b48b fix(#4270): forward Codex spawn model routing (#4281)
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 14:20:26 -04:00
Tom Boucher
7ff196c505 fix(#4096): honor --dry-run in todo complete and write completion keys inside the frontmatter fence (#4325)
* fix(#4096): honor --dry-run in todo complete and upsert completion keys inside the frontmatter fence

* review(#4096): tighten todo complete flag rejection to any dash-prefixed token

* chore(#4096): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 13:49:58 -04:00
Tom Boucher
2e1ede6d99 fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline (#4318)
* test(#4093): regression matrix for advance-plan zero-labeled-fields decline

* fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline

* refactor(#4093): collapse IIFE to a plain block (review finding)

* docs(#4093): document the advance-plan recovery decline + changeset

* chore(#4093): backfill PR number in changeset

* fix(#4093): budget lint-compiled-artifact-sync's tsc compile as a compile, not a probe

---------

Co-authored-by: sim <sim@local>
2026-09-05 10:46:37 -04:00
Behruz Nassre Esfahani
5d804dd287 fix(#3709): clear the context-monitor warn sentinel on PreCompact (#3808)
* fix(#3709): clear the context-monitor warn sentinel on PreCompact

The monitor's per-session warn sentinel survived a compaction, so once the
first CRITICAL of a session had fired, `lastLevel` stayed pinned at 'critical'
for the rest of the run. The hook was already wired to PreCompact (#772), but
read the event only at the very END, and solely to pick an output envelope.

Two documented behaviours died as a result:

  - "First warning always fires immediately" — the first warning of the
    post-compaction cycle was debounced instead.
  - "Severity escalation (WARNING -> CRITICAL) bypasses debounce" — computed as
    `lastLevel === 'warning'`, which can never be true again, so every later
    CRITICAL waited out the full five-tool-use debounce, exactly when an
    immediate warning matters most.

`criticalRecorded` was equally sticky: a session that compacted and later truly
ran out kept a /gsd:resume-work breadcrumb (#1974) describing the earlier
near-miss rather than the exhaustion that ended the run.

Reproduced first, with the issue's own literal repro, including the detail that
the compaction consumed a debounce slot (callsSinceWarn 0 -> 1).

The reset runs BEFORE the metrics read, deliberately: a post-compaction reading
is healthy again, so the ENOENT / stale / above-threshold branches would all
exit first and never reach it. Returning early also stops the compaction from
eating a slot of the cycle it was meant to restart. The event name is now read
once through a shared `readEventName()` helper, so this reset and the #2289
output allowlist cannot drift on what counts as "no event name".

Seven rows against a real sequence (the defect is state carried ACROSS calls, so
they need their own driver — the existing helpers delete the sentinel after each
invocation). Reverting the reset turns SIX of them red; the seventh is the
non-vacuity row asserting a NON-compaction event must not clear the sentinel,
which correctly passes either way.

AC4 initially passed with and without the fix — asserting `criticalRecorded ===
true` is vacuous when the seeded stale sentinel already carries it. It now seeds
a `staleProbe` marker that can only survive if the sentinel survives, so its
absence is what proves the state was rebuilt.

hooks/dist/ is gitignored and regenerated by build:hooks, so no committed dist
copy needs syncing.

Verified: `npm run lint:ci` exit 0; acceptance criteria 1-6 driven end-to-end
against the real hook.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3709): reset ahead of the config gate, and pin the placement itself

Codex review of the #3709 fix, before opening the PR. Three findings, all in
this change's own new code.

1. `context_warnings: false` prevented the reset. The config early-exit sits
   ABOVE where the reset was placed, so a session that disabled warnings,
   compacted, then re-enabled them mid-session resurrected the stale sentinel
   and the original bug with it. Config is re-read per invocation, so that
   sequence is supported rather than hypothetical. The reset now runs ahead of
   the config gate: clearing the sentinel is CLEANUP, not a warning — state that
   must not outlive a compaction should not outlive it merely because warnings
   are switched off right now. It cannot emit anything from there, so the
   disabled contract is untouched.

2. Nothing pinned the "before the metrics read" placement. Every row wrote a
   fresh metrics file, so the reset could have been moved below the metrics
   read, the stale check, or the healthy-threshold exit with all seven rows
   still green — while a REAL PreCompact, which carries no fresh metrics and
   follows a recovery to healthy usage, silently kept its sentinel. Three rows
   now pin it: no metrics file at all, usage recovered to healthy, and warnings
   disabled. Each catches a distinct wrong placement — moving the reset below
   the config check reds the third; below the metrics read reds all three.

3. The absent-sentinel row proved nothing. `assert.doesNotThrow` was vacuous
   because the driver caught every child exit, so a hook that exited 1 on the
   ENOENT unlink would still have passed. The driver now returns the exit code
   and the row asserts it is 0.

Also corrected the `readEventName` comment: it said the event is "read once",
which is not literally true — there are two call sites. The point is one
DEFINITION of what counts as an event name, so the reset and the #2289
allowlist cannot drift; the comment now says that.

Verified: 60 rows in tests/perf-317-context-monitor-fs.test.cjs, 0 fail, with
both placement mutations driven to red and reverted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3709): backfill changeset pr number

The fragment shipped with the documented `pr: 0` placeholder, which the
changeset lint treats as always-silent, because the PR number does not exist
until the PR is opened. Backfilled to 3808 now that it does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3709): a compaction clears the stale reading too, not just the state

Review round 1. Major 1 was right and it mattered: clearing only the sentinel
traded a warning that never fires for one that fires when it must not.

The statusline bridge still holds the PRE-compaction reading, and STALE_SECONDS
is 60, so for up to a minute it still reads fresh and still says the context is
exhausted. With the sentinel gone, firstWarn is true, so the next PostToolUse
emitted a spurious CONTEXT CRITICAL immediately after the compaction that FREED
the context — and flipped criticalRecorded, spawning a false context-exhaustion
breadcrumb. That is the same breadcrumb inaccuracy #3709 exists to fix, re-entered
from the other side. Reproduced before fixing, exactly as the review described.

A compaction now invalidates the warning state AND the reading that produced it.
Removing the bridge loses nothing: the statusline owns that file and rewrites it
on every render, and its absence is already the "no reading yet" state a fresh
session starts in, which exits silently.

Two things my own verification caught while fixing it:

  - The first attempt did NOTHING. metricsPath was declared below the PreCompact
    block, so referencing it hit the temporal dead zone, threw, and the outer
    catch swallowed it into a silent exit 0. The probe still printed "silent",
    which looked like success but was the old debounce. metricsPath is now
    hoisted beside warnPath.

  - The new Major 1 row was VACUOUS. The driver's `metrics: false` DELETES the
    bridge, but the defect is a bridge that is still there and still reads fresh,
    so the row passed on the ENOENT early-exit rather than on the fix. Only the
    sentinel-only mutation exposed it. The driver grew a `metrics: 'keep'` mode
    that leaves the stale file in place; both Major 1 rows now red under that
    mutation.

Also from the review:
  - Minor 1 — the compaction-abort path is now stated in the source rather than
    left silent, including why a conditional reset (SessionStart source "compact")
    is out of scope for this fix.
  - Minor 2 — docs/context-monitor.md completed: PreCompact wiring and the early
    return under How It Works, a table of all three things the reset clears, the
    breadcrumb guard, the warnings-disabled interaction, and the never-block
    property under Safety.
  - Minor 3 — changeset trimmed from ~1,400 chars of implementation narration to
    the user-visible change.
  - Nit 1 — a failed unlink (Windows EPERM/EBUSY) no longer leaves the bug
    silently intact: the file is neutralised in place instead, with a shape safe
    for each (an empty sentinel, a timestamp-0 bridge).
  - Nit 2 — reviewer-process narration removed from shipped test source. The
    remaining "Codex" mentions are pre-existing and name the RUNTIME.
  - Nit 3 — the debounce-slot row now asserts the observable consequence (the
    first post-compaction warning fires) rather than repeating AC1's assertion.
  - Nit 4 — the file docblock now lists the folded-in blocks and asks the next
    contributor to extend it.

Verified: `npm run lint:ci` exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3709): the unlink-failure fallback truncates to empty, matching deletion

The fallback wrote well-formed neutral values, and neither was equivalent
to the deletion it stood in for: '{}' parses, so firstWarn was false and
the first post-compaction warning was debounced — AC2 undone on exactly
the path the fallback exists for — and '{"timestamp":0}' was never stale
(the guard is `metrics.timestamp && ...`), so the flow reached emit with
remaining === undefined and injected a literal 'Usage at undefined%'.
Truncating to '' makes JSON.parse throw on both reads: the sentinel read
keeps firstWarn true, the bridge read falls to the outer catch and exits
0 silently (review of #3808, Blocker 1).

The branch is now executed for real: an EPERM is injected into the
child's fs.unlinkSync via --require preload — method monkeypatching,
never chmod 0o000, which root bypasses under Docker/CI (Blocker 2). Both
rows proved failing-first against the neutral-value fallback. The
boundary trios at WARNING=35 / CRITICAL=25 are completed on the emit
path with 34, 26, and 24 (Major 3); 36/35/25 were already pinned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3709): the truncation fallback refuses to follow a planted symlink

The per-session files live in a shared sticky tmpdir, where an unlink
failing EPERM is exactly what another user's planted file produces — and
a planted SYMLINK would make the fallback's plain truncating write empty
out its TARGET, weaponising the hook against any file its own user can
write. Open with O_WRONLY|O_TRUNC|O_NOFOLLOW instead: a symlink fails
ELOOP into the same give-up arm. On Windows the constant is absent and
'|| 0' keeps the fallback alive there, where the held-handle case it
exists for occurs and temp dirs are per-user. Found by Codex review;
the new row proved failing-first against the writeFileSync fallback
(victim file truncated to zero bytes).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3709): refuse non-regular files everywhere, not only where O_NOFOLLOW exists

Codex round 2: '|| 0' removed the no-follow protection exactly where it
cannot be expressed as an open flag — Windows, whose tmpdir is NOT
guaranteed per-user (TEMP/TMP overrides, system-temp fallback). An
lstat isFile() guard now rejects symlinks and every other non-regular
shape on all platforms before the truncating open; O_NOFOLLOW stays, as
the lstat->open substitution-race backstop where the platform has it.
The symlink row additionally asserts the planted link SURVIVES the call,
so a preload match that stops engaging can no longer pass the row
vacuously off a successful unlink.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(#3709): tolerate the Windows give-up, still outlaw neutral values

Both windows-latest CI lanes fail the two EPERM rows deterministically:
the runners hold freshly written files with a share mode that allows
DELETE (every real-unlink row passes) but refuses a truncating
write-open, so the fallback's give-up arm engages — which is the
fallback working as designed, not the defect the rows exist to catch.
The rows are now platform-aware: POSIX still requires exact truncation
and the behavioural follow-ons; Windows accepts truncated-or-untouched
but still rejects the Blocker-1 regression class (a parseable neutral
value is never legal anywhere), with the follow-ons gated on the
truncation actually landing. Also corrects the hook comment: libuv
defines O_NOFOLLOW as 0 on Windows — a no-op, not an absent constant.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3709): a compaction watermark closes the window bridge deletion only narrowed

Round-3 Major 1: the statusline is an uncoordinated process that
re-writes the bridge on every render, so a render landing between the
PreCompact clear and the compaction's completion re-created the
PRE-compaction reading under a CURRENT timestamp — past STALE_SECONDS,
into a spurious post-compaction CRITICAL and a false exhaustion
breadcrumb: the exact failure the deletion was added to prevent.
PreCompact now also writes claude-ctx-<id>-compacted.json ({at}) and the
metrics read drops any reading not STRICTLY newer than it — which also
covers unstamped/zero timestamps once a compaction happened. Written
unlink-then-O_EXCL so a planted file or symlink is never followed;
failure degrades to the old narrowing. Docs and changeset now describe
the watermark instead of overclaiming for the deletion.

Round-3 Major 2: DEBOUNCE_CALLS and STALE_SECONDS get their trios — the
gate increments BEFORE comparing, so seeds 3/4/5 pin 4-debounced,
5-emits, 6-emits; ages 59/60/61 pin the strict >. The child's clock is
pinned via a --require preload (a wall-clock boundary row would flip on
one second of startup delay). timestamp-0's falsy bypass is pinned
directly as characterized behaviour. Mutation-proven: dropping
O_NOFOLLOW, <= for <, and >= for > each red exactly one row.

Minors: the symlink row's comment now names the lstat guard it actually
pins, and a preload-blinded-lstat row drives the O_NOFOLLOW substitution
-race backstop for real (3); absence assertions use warnRaw so a
corrupt leftover cannot pass as deleted (4); the Windows give-up is an
explicit t.skip, never a silent if (5); readEventName is total via
String(), keeping #2289's side-effects-always-run contract for
malformed event names, with a row (6); the PreCompact rationale lives
once in docs/context-monitor.md with the code keeping only line-level
constraints (9); the changeset is release-note-sized (10).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3709): the grace window covers the compaction's duration, not just its start

Codex on the first watermark cut: the watermark stamps the compaction's
START, so a statusline render one second later — still mid-compaction,
still the old reading — passed 'strictly newer' and re-fired the false
CRITICAL. Readings inside COMPACT_GRACE_SECONDS (60) past the watermark
are now dropped: the window covers the compaction's own duration, a
healthy reading dropped there behaves identically to an accepted one
(it exits above-threshold anyway), and a genuine exhaustion warning is
delayed at most one window after a compact. A watermark stamped ahead
of the reader's clock is ignored — a clock step backwards or a stray
file must degrade to plain staleness, never mute the monitor
indefinitely. Both proven failing-first.

readEventName is strict about TYPE, not coerced: String() rendered
['PreCompact'] as 'PreCompact' and would run the reset off a malformed
payload. typeof: every non-string is 'no event' — silent, side effects
intact — with rows for the number, hostile-object, and array-wrapped
cases. The lstat-claim preload arm now writes an engagement marker the
substitution-race row asserts on, so a match string that silently stops
matching can no longer let the row pass off the real lstat guard.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: retrigger CI — the previous wave was cancelled by an Actions outage, zero job failures

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3709): drive the compaction rows on the clock, not on a future stamp

Round-4 review raised three majors, all in the test scaffolding around the
fix rather than in the fix itself.

Major 2 (taken first — it is the cheapest and it unblocks Minor 6): call()
passed process.env to the child unmodified, so two rows depended on ambient
GEMINI_API_KEY. The preserved Gemini fallback is `eventName === "" &&
!!process.env.GEMINI_API_KEY`, and readEventName returns "" for every
malformed name, so with the key set the malformed-event row's `stdout === ''`
assertion failed outright — reproduced by running it under GEMINI_API_KEY=x.
call() now takes an explicit env, the way the sibling runMonitor helper in
this file always has, and both rows pin the variable unset. (The array row
survived an ambient key only because its reading was debounced — incidental,
not independence, so it is pinned too.)

Major 1: the AC2/AC3 rows drove the hook with a bridge stamped 62 seconds in
the FUTURE — a shape hooks/gsd-statusline.js cannot produce, since it always
stamps Math.floor(Date.now()/1000) on the same clock. They proved "the
sentinel was cleared" while their assertion messages claimed the documented
immediate-warning behaviour, which is gated behind the grace window and went
unexercised. Both rows now run the real sequence on the clock-pinning preload
this PR already added for the STALE trio: PreCompact at a fixed instant, then
a normally-stamped render one second past the window. Verified non-vacuous —
stubbing the sentinel unlink reds both.

Major 3: COMPACT_GRACE_SECONDS, the one constant this PR introduces, was the
only threshold without a limit-1/limit/limit+1 trio, in a PR that adds full
trios for four pre-existing ones. The seeded offsets were +0, +1 and +61; the
boundary itself (+60) and limit-1 (+59) were untested. Added, driven by
advancing the reader's clock rather than post-dating the reading, so the
reading is never ahead of the reader and only the grace gate can drop it.
Verified against three mutations — `>` to `>=`, the constant to 59, and the
constant to 61 — each of which reds exactly one row of the trio.

No production code changed. Verified: 85/85 in this file, lint:ci exit 0,
and the two Minor-6 rows now pass under GEMINI_API_KEY=x as well as unset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG

* fix(#3709): harden the watermark read and pin the thresholds it introduces

Codex review of the full PR found four majors. All reproduced here against
the real hook before fixing.

MAJOR — the watermark was write-hardened but read-untrusted. PreCompact
already refuses to follow or overwrite a planted object (unlink-then-O_EXCL),
but the read was a bare readFileSync, so anything the write side gave up on
was followed by every later invocation. In a shared sticky os.tmpdir() that
is a mute primitive — a planted recent watermark suppresses monitoring — and
a symlink to a FIFO stalls a synchronous read. Measured against the
pre-hardening file: a symlink to a planted watermark WAS honored and muted
the monitor. The read now uses the same lstat + O_NOFOLLOW pair the sentinel
path uses, plus a size bound; symlink, directory and oversized cases are all
refused, with a plain-file control proving watermarks still work.

MAJOR — the `now + 5` skew tolerance was an unnamed, untested threshold. It
is now WATERMARK_SKEW_SECONDS with a +4/+5/+6 trio, verified against two
mutations (`<=` to `<`, and the constant to 6), each of which reds one row.
This is the same class as round 4's Major 3, one layer up.

MAJOR — the malformed-event row shared one session across both subcases, so
the hostile-object iteration's `assert.ok(s.warn())` passed off the sentinel
the `42` iteration left behind. A regression throwing before the bookkeeping
would have kept it green — vacuous for exactly the subcase it exists for.
Fresh session per subcase, with an explicit no-sentinel precondition.

MAJOR — the stale-reading row's non-vacuity is an artifact of call()'s future
stamp: with a production stamp the watermark suppresses the same reading, so
the row cannot isolate bridge deletion. The two guards genuinely overlap
inside the window, so no end-to-end row can separate them; the comment now
says so and points at the direct pin (s.metrics() === null) instead of
claiming an isolation it does not have.

Docs corrected where measurement contradicted them: the window NARROWS the
race rather than covering the compaction's duration, and the delay is not
bounded by the window alone — first recovery is watermark+61s with no skew
but watermark+66s at the accepted +5s skew. Aborted compactions are muted
the same way. The truncation fallback is documented as best-effort, which is
what the code and the Windows rows already do.

Verified: 89/89 in this file, lint:ci exit 0, symlink/directory/oversize all
refused where the pre-hardening file honored them, both new trios
mutation-checked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG

* fix(#3709): move the PR's two new exits onto the declared-policy vocabulary

#3911 / ADR-3889 migrated this hook off raw process.exit() while this PR was
in review, replacing every exit with hooks/lib/hook-exit.js's allow(), which
forces each call site to name its crash policy. The PreCompact reset and the
watermark gate are added by THIS PR, so they did not exist to be migrated and
came through the merge as the only two raw exits left in the file — caught by
the new local/require-registered-exit rule. Both are ALLOW: a compaction is
never blocked by this hook, which is the policy the rest of the file declares.

Caught only in CI, not locally: `npm run lint` runs eslint with --cache, and
the cached entry for this file predated the new rule, so a warm local cache
reported clean. Re-verified with the cache cleared.

allow() terminates rather than throwing, which matters for the watermark call
site because it sits inside a try/catch — a throwing helper would unwind into
that catch and silently drop the grace-window mute. Verified behaviourally,
not by reading: the grace trio, the skew trio and the non-regular-file rows
all still pass.

Verified: lint:ci exit 0 with a cold eslint cache, full suite exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG

* fix(#3709): harden the routine sentinel writes and the read beside them

Round 7 ruled that the three routine debounce-accounting writes to the warn
sentinel must match the three writes this PR already hardened: leaving the
fourth unhardened beside them is the asymmetry that invites the defect back.
They now go through one writeSentinel() helper using the compaction
watermark's own unlink-then-O_EXCL shape, rather than a second policy — the
unlink removes any existing object, and O_EXCL then refuses to create through
one, so a write can only land on a fresh regular file this process made.

The routine READ beside them was the last bare readFileSync on warnPath, and
the same rationale applies to it verbatim; the watermark's read was hardened in
round 4 for exactly this reason. Same lstat + O_NOFOLLOW + size bound. Its
scope is stated in the test rather than overclaimed: lstat establishes that the
sentinel is a plain regular file, not that it is trustworthy, so a cross-owner
regular file at the predictable path is still read and is left as a disclosed
pre-existing residual.

Also fixes an accept-direction regression this PR introduced and six rounds of
review missed. readEventName collapsed an ABSENT event name and a MALFORMED one
onto the same '', and the preserved Gemini fallback keys off eventName === "",
so with GEMINI_API_KEY set a malformed payload began emitting an AfterTool
envelope. At the merge-base, data.hook_event_name.trim() threw on a truthy
non-string after the side effects and nothing was ever emitted. Measured
base-vs-head with a fresh sentinel per run: 42, ['PreCompact'] and {} all went
silent -> EMITS, while an absent name and 'PostToolUse' were unchanged.
readEventName now returns '' only for an absent name and null for a
present-but-non-string one; both call sites compare for equality only, so every
well-formed payload behaves identically.

Five new rows, each proven fail-first with the mutations attributed separately:
reverting the writes reds the write-through and non-regular rows, reverting the
read reds the mute and non-regular rows, and reverting the absent/malformed
split reds the Gemini row. The changeset's "behaves like a fresh session" is
narrowed to name the 60-second suppression window and the best-effort reset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018FUAVz49BghqxoJgwt7EW9

* test(#3709): pin both 4096-byte read bounds at their boundaries

Round 8 asked for limit-1/limit/limit+1 coverage on the size bound the
round-7 sentinel read-hardening introduced (gsd-context-monitor.js:335).
The existing refusal row pads to 8192 -- a full 4096 bytes clear of the
fence -- so `>` vs `>=`, or an off-by-one in the constant itself, was
invisible to it.

Covers the sibling bound too. The identical check guards the round-4
WATERMARK read at :278 and its refusal row pads to 8192 in exactly the
same way; the review's own rationale (this file already holds
WATERMARK_SKEW_SECONDS to a boundary trio, so an uncovered bound is the
odd one out) applies to it unchanged. That half is a class sweep of a
pre-existing bound and is test-only -- say the word and it comes out
without touching the rest.

Both trios assert on observable hook output rather than an internal
error. Sentinel: an honored {callsSinceWarn:1,lastLevel:'warning'} keeps
the debounce arm taken at remaining=30, so nothing is emitted, while a
refused one falls back to first-warn defaults and emits. Watermark:
honored mutes (stdout empty), refused leaves the warning. The 4097 row is
the non-vacuity control for the two accept rows. Payloads are sized by
measurement, with Buffer.byteLength asserted to equal the target, not by
arithmetic on an assumed prefix width.

Proven fail-first in both directions, with the hook restored after:
`> 4096` -> `>= 4096` reds both trios (94/96); `> 4096` -> `> 4097` reds
both trios (94/96); restored, 96/96. Under both mutations only the two
new rows fail -- the pre-existing 8192-padded rows stay green, which is
the review's fencepost claim demonstrated rather than assumed.

* fix(#3709): correct the changeset's mute-window claim and a superseded comment

Both from the pre-push Codex pass on the full PR.

The changeset said readings are "suppressed for up to 60 seconds after a
compaction starts". That is false at the accepted skew boundary, and this
repo's own docs/context-monitor.md already carried the accurate figure:
first recovery is watermark+61s with no skew and watermark+66s for a
watermark at the +5s skew limit. Measured independently at +64 silent,
+65 silent, +66 warning. The changeset now states the window plus the
accepted skew, matching the doc rather than contradicting it.

A comment in the malformed-event row still described readEventName as
returning "" for every malformed name. Round 7 superseded that: a
present-but-non-string name returns null and only an ABSENT one returns
"", so a malformed payload can no longer reach the Gemini fallback at
all. Marked as historical and corrected. The GEMINI_API_KEY pin stays --
the row is about readEventName's typing, not the fallback, and an ambient
key would still change what it measures.

Codex's three Major findings are not taken, on attribution rather than
logic; the reasoning is in the PR reply. In short: the watermark does not
exist at the merge-base at all (0 occurrences), so "base emits, HEAD
mutes" compares a new feature against its absence rather than showing a
regression; and the base sentinel read is a bare readFileSync, which
blocks on a planted FIFO exactly as the hardened read would, so the
TOCTOU stall is not introduced here. The underlying limits -- watermark
provenance, and lstat->open races on a non-symlink substitution -- are
real, pre-existing, and already offered to the maintainer as follow-ups.

* fix(#3709): read both sentinels through one hardened helper; state the two limits precisely

Round 9's Major, with a correction to its premise, and both Minors.

The review names "watermark read/write helpers this PR adds" that a call
site at :238-250 duplicates inline. There are no such helpers: this PR
adds readEventName and writeSentinel, the latter a write-side primitive a
read cannot call, and :238-248 is base code the diff never touched. What
IS duplicated is the hardened READ. The watermark read (round 4) and the
warnPath read (round 7) are the same ten lines twice -- lstat, isFile and
a 4096-byte bound, O_RDONLY|O_NOFOLLOW, readSync, close -- differing only
in the path variable and the error string, and that is two copies to keep
in step by hand. Now one function, readSentinel(target), beside
writeSentinel. Refusal throws; both callers already wrapped the read in a
try/catch that degrades to "no file", so behaviour is unchanged by
construction.

Proven rather than assumed: with the helper replaced by a bare
readFileSync in a complete scratch tree, exactly the five hardened-read
rows in tests/perf-317-context-monitor-fs.test.cjs go red -- round 7's
symlinked and non-regular sentinel and its size bound, round 4's
non-regular watermark, round 8's watermark size bound -- so the helper
carries both call sites' guarantees and the rows pin it. 96/96 with the
helper in place.

Minor, drop vs delay: the grace-window comment said "dropped" on one line
and "delayed" three lines later, and docs/context-monitor.md said
"delayed". A genuine exhaustion reading inside the window is skipped, not
queued: its warning and its #1974 breadcrumb both fire on the next reading
after the window, so both are delayed when a later reading comes and lost
when none does -- a session ending inside the window records neither.
Comment and docs now say exactly that, and that the loss is accepted over
trusting a reading that may be the pre-compaction value under a fresh
timestamp.

Minor, ordering: the PreCompact unlink and the debounce
writeSentinel(warnPath) are two writers with nothing serialising them; a
debounce invocation that read pre-compaction state and lands its write
after the unlink would resurrect the sentinel the reset removes. The hook
relies on the host dispatching a session's hooks one at a time, which
Claude Code does and the other runtimes are assumed to. Stated at the
reset as an assumption, with the lock-file alternative named and not
taken.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

* fix(#3709): write the compaction watermark through writeSentinel

Review of #3808, round 10. The PreCompact watermark write was the block
writeSentinel was lifted from in round 7, and it kept its own inline copy
of unlink-then-O_EXCL a few lines below the helper. Round 9 flagged that
write-side duplication; the round-9 reply misread it as the read side and
unified only the reads. The write now calls the helper too, so the hook
holds one copy of the hardened write, not two.

Behaviour is unchanged: same unlink-then-O_EXCL sequence, same flags,
same best-effort outer catch. The one difference is that writeSentinel
closes the descriptor in a finally, where the inline copy leaked it if
writeSync threw before closeSync.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R7bQLXAKubb4EFLtCPiLiL

* fix(#3709): read the statusline bridge through the same hardening as the sentinels

Review of #3808, round 11. `metricsPath` is built one line from `warnPath` and
`watermarkPath` — same tmpdir, same predictable `claude-ctx-{sessionId}` shape,
same threat model this PR documents at length for its siblings — and it is the
only one of the three read on EVERY invocation. It was also the only one still
reached by a bare `readFileSync`, so the symlink-follow and the symlink-to-FIFO
stall that rounds 4 and 7 closed on the other two stayed reachable here, on the
file's highest-traffic path. It now goes through `readSentinel` like the rest.

The 4096-byte bound is ample for it: the statusline writes four fixed fields
(`gsd-statusline.js`), about 140 bytes with a UUID session id, so no legitimate
bridge approaches it. A refusal lands in the same rethrow an unreadable or
malformed bridge already did.

The comment introducing `readSentinel` claimed the warn sentinel was "the one
bare readFileSync". Read as scoped to `warnPath` that was true, but it reads as
a claim about the file and it is not one — the bridge kept its own until this
round. Corrected rather than left to mislead the next reader.

Round 11 Minor: `readSentinel` discarded `fs.readSync`'s return value and
assumed the buffer was full, so a file truncated between the `lstat` and the
read left a zero-filled tail. It now refuses a short read. Stated plainly
because it was measured: this guard has NO observable behavioural delta —
deleting it leaves the new row green, because the NUL tail makes `JSON.parse`
throw one line later and both paths degrade to "no sentinel". It is a
consistency fix in a function whose purpose is refusing to trust what it read,
and the test comment says exactly that rather than implying coverage it lacks.

Five rows added: the bridge refusing a planted symlink (with an attacker-chosen
reading that WOULD warn if followed, so silence is proof), a non-regular bridge,
an oversized bridge, the shrink path end to end, and the direction that matters
most — a healthy bridge still warns, so the hardening is not a mute. Proven by
mutation: reverting the bridge to `readFileSync` reddens two rows. The shrink
injection carries an engagement marker for the same reason the lstat-claim one
does, learned the same way: the hook rewrites the sentinel later in the
invocation, so a size check afterwards passes whether the truncation landed or
not.

An independent full-PR pass on this round added two more, both taken:

`writeSentinel` discarded `fs.writeSync`'s return value, and a short write is
permitted by the syscall — so a truncated sentinel could reach disk and every
later read would reject it, silently losing the debounce accounting or the
watermark this write exists to record. It now loops until the payload is
written, as Node's own `writeFileSync` does, with an explicit no-progress guard.
Pinned by a row that injects a one-byte first write; reverting the loop reddens
it.

The directory row's comment claimed it pinned the `lstat` isFile() check. It
does not — measured: deleting that condition leaves the row green, because
reading a directory fails on its own a line later. The comment now says the row
pins the outcome, and names the symlink row as the one that pins isFile().

DISCLOSED, NOT FIXED HERE — a session id long enough to push the derived
filenames past NAME_MAX. The bridge is `claude-ctx-{id}.json`; the sentinel and
watermark add longer suffixes, so on a 255-byte limit the watermark stops fitting
at a 230-character id and the sentinel at 233. Measured base-vs-HEAD at 233+:
base is SILENT, HEAD emits the warning, because the bare `writeFileSync` base
used threw ENAMETOOLONG out of the warning path while `writeSentinel` degrades
best-effort and lets the warning through. That is an accept-direction delta and
it is in the delivering direction — base swallowed a warning the user should
have seen, which is this issue's own failure class. The underlying limit is a
property of the per-session filename scheme, shared by two files that predate
this PR, and bounding session ids belongs to whatever writes them
(`gsd-statusline.js`), not to the sentinel logic. Happy to fold a length guard
in here if you would rather have it in this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 06:29:06 -04:00
Dennis Alexis Valin Dittrich
5869febb16 enhance(#4155): invalidate verification results when covered inputs change (#4290)
* enhance(#4155): invalidate verification results when covered inputs change

readVerificationStatus() now recomputes a deterministic sha256 fingerprint
over a VERIFICATION.md's declared covered_files (phase PLAN/SUMMARY,
requirements, implementation files in the verified change set) and returns
stale on any mismatch, fail-closed when a covered file is missing,
unreadable, or escapes the project root. Legacy reports with no fingerprint
metadata keep the prior SUMMARY-mtime staleness check unchanged.

The verifier computes covered_digest via the new verification.fingerprint
CLI command rather than by hand, since a digest is deterministic math, not
an LLM-estimated value.

* chore(#4155): backfill fork PR number in changeset

* fix(#4155): trim gsd-verifier.md fingerprint instructions to fit LARGE tier byte cap

* fix(#4155): address CodeRabbit findings on fingerprint fail-closed behavior

Partial fingerprint metadata (one of covered_files/covered_digest present,
the other missing or malformed) now fails closed to stale instead of
silently downgrading to the legacy mtime-only check. computeCoveredDigest
also canonicalizes with realpathSync before re-confining, so an in-root
symlink whose target escapes the project root can no longer produce a
matching digest. gsd-verifier.md restores the completeness requirement and
checklist item trimmed by the earlier size-budget fix, within the LARGE
tier byte cap.

* chore(#4155): acknowledge gsd-verifier.md growth for the #4155 fingerprint instructions

Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the covered-input fingerprint instructions and frontmatter fields the #4155 verification staleness mechanism requires; trimmed to stay within the LARGE tier byte cap

* fix(#4155): address gemini adversarial review findings

computeCoveredDigest now threads the caller-supplied opts.fs seam through
its confinement and read paths instead of always using raw node:fs — a
caller like planning-inspect.cts's containmentEnforcingVerificationFs (GAP
2, #2790 follow-up) was silently bypassed for covered-input reads. The
project-root anchor itself still canonicalizes through real fs (it is a
trusted value the caller derived, not attacker-influenced covered-input
data); only per-file candidate reads go through the injected seam.

Covered-file paths are now canonicalized (./ prefixes, redundant slashes,
internal .. segments) before becoming dedup/sort/hash keys or confinement
subjects — closes both a spurious-stale false positive (two spellings of
the same file hashing differently) and a confinement gap (an internal ..
segment that doesn't start the string).

gsd-verifier.md now states covered-file paths are project-root-relative,
not phaseDir-relative, closing an ambiguity that would have made a real
verifier agent's first fingerprint invocation fail closed.

defaultFsImpl's methods now late-bind through fs.<method> rather than
capturing function references at module load — the earlier direct-capture
form was invisible to existing tests' t.mock.method(fs, 'statSync', ...)
seams, a real regression caught by the full suite (not the reviewer).

* fix(#4155): catch a plan/summary added to the phase dir after verification but never declared

The content digest only recomputes hashes for paths the verifier actually
declared in covered_files — it had no way to notice a plan or summary
added to the phase directory after verification if that new file was
never declared, silently regressing behind the legacy mtime check it
replaces (which scans the live directory, not a declared list).

findUncoveredCurrentArtifact re-scans the live phase directory for every
current *-PLAN.md/*-SUMMARY.md and requires each to be represented in
covered_files, closing that gap; a directory scan failure fails closed to
stale rather than silently skipping the check.

CONTEXT.md's Verification Module entry corrected to describe the
fingerprint path's stricter fail-closed FS-error contract (routes to
stale) instead of the module's original degrade-to-safe one (missing /
not-stale), which only the legacy path still keeps.

* refactor(#4155): extract canonicalizeCoveredFiles, add real nested-project e2e test

computeCoveredDigest and cmdVerificationFingerprint each normalized/deduped/
sorted covered_files independently — one shared helper now backs both
(gemini review's ponytail-lens finding).

Adds one CLI-to-readVerificationStatus test against a genuine
.planning/phases/NN-x/ project with an implementation file outside
.planning/ entirely, closing the review finding that prior #4155 unit
fixtures put phaseDir directly under an ownerless tmpdir (findProjectRoot
falls back to phaseDir itself there) and never exercised real multi-level
path resolution.

* fix(#4155): route computeCoveredDigest through real fs, fail closed on unreadable plans/

Two independent review rounds (opus critical-reviewer + opus ponytail +
agy, run twice) found two instances of the same fail-open class:

- computeCoveredDigest's per-file reads routed through the caller's
  injected fsImpl. planning-inspect.cts passes a `.planning/`-confined
  containment fs into readVerificationStatus's opts.fs, so any covered
  implementation file outside `.planning/` (mandatory per the issue)
  made the confinement wrapper throw, which was caught and turned into
  a stale digest -- reporting every fingerprinted phase permanently
  stale via `planning.inspect`, regardless of actual drift. Per-file
  reads now always use real node:fs, matching the pre-existing
  treatment of root canonicalization; the realRel-vs-realRoot check is
  the real confinement boundary for this data and needs no seam.

- allCurrentArtifactsCovered's try/catch never fired (scanPhasePlans
  reports readdir failures via a `scope` field, it never throws), so
  an unreadable nested plans/ dir was silently treated as "zero
  artifacts, all covered" instead of failing closed. Now branches on
  scope !== SCOPE.COMPLETE.

Also, per ponytail's second-round findings: reverted an unwarranted
FINGERPRINT_VERSION bump and digest length-prefix from the first fix
(no v1 digest has ever existed -- the feature is unreleased -- and the
prefix closed a collision that grants no capability beyond what a
writer of covered_files already has more cheaply); removed a
verifier-facing escape-hatch instruction whose own example was a case
that should trigger staleness, not bypass it; corrected CONTEXT.md
references to the renamed allCurrentArtifactsCovered and a stale
"unconditional" rescan claim; simplified the isStale derivation,
removed dead FsLike members, and tightened test coverage.

Regression tests for both fail-open bugs are included and were each
confirmed to fail against the pre-fix code before the fix landed.

full test suite: 2558/2560 pass, 2 skipped, 0 fail

* fix(#4155): trim gsd-verifier.md under the LARGE size cap

Fork CI caught what my local runs missed: the superseded/nested-plans
instruction added earlier pushed gsd-verifier.md to 49299 bytes,
147 over the LARGE tier's 49152-byte hard cap
(tests/agent-size-budget.test.cjs). Tightened the #4155 instruction's
wording and dropped a redundant inline comment tag; no content lost.

* chore(#4155): point changeset at the upstream PR number

pr: 19 was the fork PR opened for internal review-lane CI; now that
open-gsd/gsd-core#4290 exists, the changeset field must match it per
CONTRIBUTING.md's release-notes convention.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 05:42:52 -04:00
Behruz Nassre Esfahani
5ad9a36f35 fix(#4255): resolve reviewer-lane effort from the lane, not from gsd-plan-checker (#4275)
`review-lane plan` resolved every cross-AI reviewer lane's reasoning effort by
spawning `query resolve-execution gsd-plan-checker --host <slug>`. The agent id
was a hardcoded literal, so `--host` chose only the argv RENDERING while the
LEVEL always came from the installed plan-checker's frontmatter — `low` under
every shipped model profile. Every prompt-fed lane therefore ran at a fast
structural verifier's effort, and because the rendered argument is a CLI config
override it silently beat the effort the operator had configured for that CLI.
At `low` a large source-grounded prompt makes a model end its turn with no final
message, so the lane came back empty and its stub read as a crash.

Effort is a property of the review, so the lane declares it. Two new fields on
ReviewerLane — `effortConfigKey` (`review.effort.<slug>`) and `defaultEffort` —
carried through each capability manifest and the generated registry, set on the
three lanes with an argv effort channel and null on the other nine. A new pure
`resolveLaneEffort()` resolves config key -> lane default -> nothing, where
"nothing" emits no effort argument at all and the reviewer CLI's own
configuration decides; `inherit` selects that path explicitly and an
unrecognized level falls back to the lane default rather than being forwarded to
a CLI that would reject it. The host's negotiated effortSurface still gates the
rendering, so ADR-1239/#2481's trust boundary holds on this path too. Resolving
in-process also removes up to twelve subprocess spawns per review.

The empty-output stub now names the effort the lane ran at and distinguishes a
clean exit from a timeout kill, a non-zero exit, and a process that never ran —
`status` is null for both a timeout and a signal, so those were indistinguishable
before. The hint is hedged: a clean empty exit is most often a model stopping
short, but it is also consistent with a CLI writing its output elsewhere.

Also: the capability validator now knows both fields, rejects a malformed key or
an out-of-vocabulary default, and rejects a default declared without a config
key (a level the operator could never override). An existing end-to-end row in
tests/effort-surface-axis.test.cjs asserted the old coupling; it now configures
the lane's own key and pins the decoupling in the same real spawn, with the
agent execution tier set to a level that must not appear.

Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort; leaving the new key undocumented there is the same invisibility that made the plan-checker coupling survive this long.

Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort, so leaving the new key undocumented there is the same invisibility that let the plan-checker coupling survive.

Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 05:25:44 -04:00
Dennis Alexis Valin Dittrich
925a363879 enhance(#4032): apply configured agent tool grants (#4238)
* test(4032): add failing installed-agent grants contract

Cover global and project agent_tools precedence at the real Claude installer seam before adding implementation.

* feat(4032): apply configured agent tool grants during staging

Resolve selector-level global and project config once per staging call, then append validated grants before runtime conversion.

* test(4032): cover host grant and quoted MCP contracts

Exercise installed host artifacts and prove ZCode must treat quoted MCP scalars like plain MCP grants.

* feat(4032): apply configured agent tool grants across runtimes

Move augmentation and scalar identity into the converter seam so every staged artifact preserves host policy.

* fix(4032): register agent tool grants in configuration

Accept documented agent_tools config without unknown-key warnings.\n\nKeep installer fixtures on the shared temporary-directory helper.

* fix(4032): translate configured MCP grants for Kilo

Reuse the converter-owned scalar decoder so quoted canonical grants reach Kilo's native permission keys without altering other host policies.

* fix(4032): decode YAML-escaped tool grants

* fix(4032): emit valid inline agent tool grants

* fix(4032): reject invalid trailing-colon grants

* test(#4032): cover cross-review remediation gaps

* fix(#4032): close cross-runtime grant gaps

* test(#4032): expose Kimi global project context

* fix(#4032): preserve Kimi project config context

* chore(#4032): add release note

* test(#4032): expose fork review regressions

* fix(#4032): address fork review findings

* test(#4032): make byte-stability assertion portable

Compare repeat installs at one root so platform-specific path rendering cannot
masquerade as an agent_tools behavior change.

* chore(#4032): bind changeset to upstream PR 4238

* fix(#4032): address trek-e review findings (2,3,4,5,6,7,8)

Fixes fail-closed decode-failure handling in ZCode's mcp__ stripper,
a comment-only `tools:` header mis-parse that silently dropped
configured grants, and a naive comma-split that could tear a quoted
scalar containing a literal comma. Documents Kilo's inherent
`{server}_{tool}` MCP-permission-key collision (external, fixed
format — not ours to widen) and locks the existing first-seen-wins
resolution in with a regression test.

Opts kimi/kimi-code out of the ADR-1235 pre-converter path-rewrite
step: routing Kimi through that pipeline (needed so project-scoped
agent_tools selectors reach it) was short-circuiting Kimi's own
neutralizeKimiAgentPrompt, which expects the original ~/.claude/gsd-core
text rather than a pre-rewritten Kimi path.

Extends the fast-check token pool and per-runtime install coverage
with the missing comment/comma/broad-runtime cases the prior review
flagged as untested.

* docs(#4032): add CONTEXT.md glossary entries for agent_tools resolver + pre-converter step

Documents readGsdEffectiveAgentTools (Install Model Override Resolver
Module) and the appendAgentTools pre-converter pipeline step (Runtime
Artifact Conversion Module), per contributor-standards.md's
new-seam glossary requirement (finding 1).

* fix(#4032): address agy adversarial review findings

An agy (gemini-3.8-flash-high) adversarial pass over the prior review-fix
commit found the fixes for findings 3, 4, 6 and 8 had unfixed sibling gaps,
plus a genuine new regression and two CONTEXT.md inaccuracies:

- ZCode's comment-only `tools: # note` header matched the inline-value
  branch instead of falling through to the block-list scan, so a following
  mcp__* item leaked through unstripped — the exact defect finding 4 fixed
  in appendAgentTools, unfixed in this sibling function.
- Reverted capabilities/kimi-code/capability.json's noPathRewrite: true.
  kimi-code uses the standard 'agents' kind with converter: null (not
  kimi-agents — confirmed by reading the descriptor, not its prose
  description), so it never went through the pipeline change finding 5
  fixed, and disabling its path rewrite broke every ~/.claude/ embed in
  its shipped agents instead.
- decodeToolScalar never stripped a trailing ` # comment` from a bare
  (unquoted) scalar, so a comment after a block-list item, or after an
  appended grant on an inline line, became part of the "tool name" —
  fixed at the source (one call site fixes every consumer).
- appendAgentTools's comment-index scan wasn't quote-aware, so a `#`
  inside a quoted scalar (`"mcp__server #1"`) was mistaken for a comment
  start and corrupted the quote.
- parseFrontmatterTools (Kimi/Qwen's tool-list reader, downstream of
  appendAgentTools's own output) had the same naive comma-split and
  comment-only-header gaps as findings 4 and 6, unpatched.
- The all-runtime smoke test's presence assertion was built on a guessed
  omit-list; empirically only 7 of 17 runtimes keep an arbitrary mcp__
  grant recognizable, replaced with a verified allowlist.
- CONTEXT.md claimed a `project:<agent>` selector prefix that does not
  exist (project override is a same-key merge across two config files)
  and mislabeled stageAgentsForRuntimeWithConverter's module.

* fix(#4032): address full-PR review (Opus critical/ponytail + agy)

A whole-PR pass (critical-code-reviewer + ponytail-review on Opus, plus a
second agy full-source adversarial pass) surfaced defects the earlier
finding-scoped passes couldn't reach:

- appendAgentTools corrupted a `tools:` line whose ENTIRE value is a
  leading quoted scalar (`tools: "Read"` -> `tools: "Read", Write`,
  invalid YAML) — there is no safe line-surgical rewrite here, so it now
  refuses to touch that shape instead of emitting broken frontmatter.
- decodeToolScalar's malformed-trailing-quote check ran BEFORE comment
  stripping, so a bare tool name with a quote inside its own trailing
  comment (`Bash # note: "internal"`) was wrongly rejected. Reordered.
- findUnquotedCommentIndex (added in the prior remediation commit) was
  built on a wrong model of YAML: a `#` after whitespace starts a real
  comment in a plain scalar regardless of nearby quote characters —
  verified against the actual parser. The one case that DOES need
  protection (a leading quoted scalar) is now refused outright above, so
  the quote-tracking scan was dead weight solving a problem that no
  longer reaches it. Removed; reverted to the plain `[ \t]#` scan.
- Kilo has a SEPARATE agent-frontmatter parser (convertClaudeToKiloFrontmatter,
  distinct from the buildKiloAgentPermissionBlock fixed earlier) with the
  same comment-only-header and naive-comma-split gaps as findings 4 and 6
  — unfixed in both its src/ and bin/install.js copies. Fixed in both,
  exporting splitToolScalars for bin/install.js to reuse rather than
  reimplementing it.
- Pipeline docstring in stageAgentsForRuntimeWithConverter still listed 5
  steps, omitting appendAgentTools (now step 3 of 6).
- docs/CONFIGURATION.md didn't state that a --global install still
  discovers agent_tools from the cwd's .planning/config.json (confirmed
  intentional and already covered by a dedicated test, not a bug).
- Removed install-engine.cts's deps.cwd injection seam: zero callers or
  tests ever populated it.

Two claims from this round were verified and rejected, not fixed:
prototype pollution via a `__proto__` selector key (empirically confirmed
`Object.prototype` is never touched — only reassigns the resolver's own
local object's prototype, with no observable effect), and a `*` grant
value crashing YAML parsing as an alias reference (empirically confirmed
it parses as plain scalar text, no crash). A pre-existing, unrelated
defect (extractFrontmatterField returns null for block-list `tools:` on
Copilot/Antigravity/Cursor/Codex/Qwen, affecting two shipped agents
today) was filed as a follow-up rather than fixed here — it predates
#4032 and isn't caused or worsened by this PR.

* fix(#4032): update stale slug-derivation-drift-guard fixture line

normalizeKimiSkillName's real closing brace moved from line 616 to 635 as a
side effect of this PR's edits to runtime-artifact-conversion.cts; the
MAJOR-1 fixture's hardcoded realEndLine had gone stale.

* fix(#4032): address CodeRabbit findings on projectDir threading and flow-sequence tools

bin/install.js's installAgentsKindStandalone call site omitted the projectDir
argument the function already supports, so a global install through this
legacy branch silently fell back to the runtime config dir instead of
process.cwd() when resolving project-scoped agent_tools grants — inconsistent
with the sibling installOpencodeFamilyArtifacts call site, which already
threads it correctly.

appendAgentTools' leading-quoted-scalar bailout did not cover a YAML flow
sequence (`tools: [Bash, Read]`): splitToolScalars tore it apart on the
in-sequence commas and appended past its closing bracket, producing invalid
frontmatter. Extended the bailout regex to also refuse a value starting with
`[`, matching the same "whole node, nothing may follow" reasoning already
applied to quoted scalars.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 04:52:45 -04:00
Dennis Alexis Valin Dittrich
e8800287d5 enhance(#4153): fail closed unresolved update targets (#4237)
* test(#4153): cover unresolved update target

* fix(#4153): fail closed unresolved update target

* test(#4153): require a concrete recovery installer

* fix(#4153): use concrete unresolved recovery command

* chore(#4153): bind changeset to fork PR

* test(#4153): cover portable update diagnostics

* fix(#4153): keep update diagnostics portable

* fix(#4153): harden update version diagnostics

* test(#4153): reject jq in update version checks

* test(#4153): expose step-local parser gap

* fix(#4153): keep JSON parsing step-local

* docs(#4153): align update target guidance

* test(#4153): expose workflow runtime fallback

* test(#4153): expose resolver runtime fallback

* fix(#4153): leave unknown workflow runtime empty

* fix(#4153): stop inferring Claude for unknown targets

* test(#4153): preserve Claude workflow targeting

* test(#4153): preserve known runtime directory identity

* fix(#4153): recognize Claude workflow paths

* fix(#4153): reuse known runtime directory identities

* chore(#4153): acknowledge emitted workflow growth

The fail-closed diagnostic and known-runtime preservation deliberately add 48 emitted bytes.

Emitted-Drift-Ack-Growth: update.md — explicit unresolved-target diagnostics and known-runtime preservation

* test(#4153): expose missing Windsurf workflow contract

* docs(#4153): document Windsurf update targets

* chore(#4153): bind changeset to upstream PR

* fix(#4153): gate unresolved-target exit before the VERSION-missing fallback

The VERSION-missing bullet in get_installed_version sat before the
UPDATE_TARGET_UNRESOLVED exit and shared its trigger condition (version
0.0.0). An LLM agent reading the workflow top-to-bottom could satisfy
"proceed to install" without ever reaching the fail-closed exit this
PR adds, reopening the ill-defined mutating path #4153 closes. Reorder
so the unresolved-target gate runs first and scope the VERSION-missing
bullet to require an already-resolved target.

Also drop two vacuous mutationSpies entries: they checked '--sync'/
'--reapply' (commands/gsd/update.md content) against `step`, a slice of
workflows/update.md — always -1 regardless of correctness. Those routes
bypass get_installed_version entirely and are already covered by
install.test.cjs, reapply-patches.test.cjs, and
skill-frontmatter-contract.test.cjs.

* chore(#4153): point changeset pr field at fork PR #10 for fork CI

* test(#4153): guard RUNTIME_DIRS/update.md table parity, confirm narrowing intent

Nit 1: update.md's PREFERRED_RUNTIME prose and RUNTIME_DIRS
(src/update-context.cts) are two independently maintained copies of the
same runtime->dir mapping with no parity check; add one so a future
edit to either surface without the other fails loudly instead of
silently drifting.

Nit 2: call out in the changeset that a custom --config-dir matching no
known runtime, marker file, or env var now resolves unresolved instead
of silently defaulting to claude -- this narrowing is intentional, it's
the fail-closed behavior #4153 asks for.

* fix(#4153): drop dead $UC fallback in check_latest_version's uc_field, cover unresolved-runtime fast path

agy (gemini-3.8-flash-high) adversarial review of the full PR:

1. check_latest_version's uc_field() copy-pasted get_installed_version's
   `${2:-$UC}` fallback, but every call site here passes $2 explicitly and
   $UC does not exist in this step's scope -- dead, misleading reference.
   Use $2 directly.
2. No unit test covered resolveUpdateContext's preferredConfigDir fast path
   returning runtime: '' for a custom --config-dir matching no RUNTIME_DIRS
   suffix, marker file, or env var (the exact fail-closed case #4153 adds).
   Added.

A third finding (update.md:90 using /gsd:update vs docs using /gsd-update)
was investigated and rejected: /gsd:update is the actual registered
Claude Code command name (commands/gsd/update.md name: gsd:update) and is
locked by this PR's own test (tests/update-workflow.test.cjs); /gsd-update
is a separate, pre-existing, intentional prose convention used in
audience-facing docs (README/INVENTORY/FEATURES). Not a defect.

* chore(#4153): backfill changeset pr field to upstream PR #4237

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 04:17:21 -04:00
Cody Anderson
77e2472ca0 enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration

Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on
Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets
(the .env.example/.sample/.template/.dist templates stay readable).
Read checks file_path; Grep checks an explicit path and judges the glob
per brace alternative; Bash runs a two-pass token scan (quotes, comments,
redirects with fd digits, separators, $( )/backtick/<( ) recursion,
heredoc bodies never scanned as commands, nested bash -c/eval rescans,
git <ref>:<path> shapes) with a closed non-reading exemption set for
existence checks. Fail-open crash policy; 1 MiB commands are denied as
command-too-large; more than 64 glob alternatives as glob-too-complex.

Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt
for approval whenever any Read() deny rule exists, even in auto mode. A
hook denial is not a permission rule and never arms that check. The
installer-written deny rules are retired in the follow-up commit.

Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks
HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking
guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell),
shell-command-projection managed sets, installer-migration-report,
OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch),
docs tables in five locales, ADR-766 always-on list, regen:derived
fixtures, and a new table-driven unit suite.

* test(#4221): pin the secret-read guard in existing hook gates

Register gsd-secret-read-guard.js in every existing hook gate: the
hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest
REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity
EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal-
hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS,
kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and
typed-payload floors, the OpenCode adapter (grep mapping, include ->
glob, three dispatch tests) and a Kimi TOML matcher assertion.

* fix(#4221): retire installer Read() deny rules (legacy filter)

Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS
and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets)
strings. mergeClaudePermissions now only filters them out of an existing
permissions.deny: an absent deny key stays absent, a malformed one is
still repaired to [], and an array emptied by the filter is deleted so
no `"deny": []` residue is left. Uninstall filters the same legacy list
and, symmetric with the Antigravity branch, drops an emptied allow or
deny key and an emptied permissions object.

Unlike the #2278 allow-side migration there is no surviving current
deny list, so the constant is renamed rather than mirrored. Removal is
byte-exact: a hand-written identical rule is indistinguishable from the
installer's and is removed too (the manifest never recorded permission
strings). USER-GUIDE and CONTEXT.md updated.

* test(#4221): flip install-regressions deny-rule assertions to the retired shape

The fresh-merge, non-destructive merge, idempotency, end-to-end install,
reinstall and uninstall assertions now expect no Read(.env*) deny rules
and no permissions.deny key on a fresh install; the deny:null repair case
is kept. A new describe block covers the legacy filter: retired strings
removed with a user entry kept, partial sets, near-miss strings
untouched, idempotency, GSD-only deny array deleted, a pre-existing
empty deny preserved, and uninstall symmetry for allow/deny/permissions.

* chore(#4221): add changeset fragment for PR #4236

* fix(#4221): case-fold names; scan shell stdin and xargs pipes

Review round 1 (trek-e):

- Blocker: secret-name matching is now case-insensitive in the Read,
  Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a
  case-insensitive filesystem are recognized as the same secret file.
- Major: a shell interpreter's script is now scanned wherever it comes
  from. The tokenizer keeps heredoc bodies as per-segment tokens and
  records separator operators; pass 2 groups by segment id and resolves
  bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined
  `-lc`) scans the script operand, a file operand is checked as a file
  (a `<( )` operand's echo/printf output is reconstructed), otherwise
  stdin is the script and heredocs, here-strings and a piped echo/printf
  source are scanned. `eval` joins all its operands; `source`/`.` handle
  process substitution. Data heredocs (`cat <<EOF`, the commit-message
  shape) stay unscanned.
- Major: `… | xargs <cmd>` checks the upstream segment's operands as
  file names when the sub-command reads (`echo .env | xargs cat`,
  `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the
  inference; a shell sub-command's `-c` script is scanned.

Header, USER-GUIDE bullet and changeset updated; documented gaps now
include piped scripts from non-echo sources and `exec`/`timeout`
wrappers. 60 new suite cases pin the block and allow shapes.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 04:00:08 -04:00
Dennis Alexis Valin Dittrich
a262ad6b61 fix(#4148): dispatch wave-pre step hooks (#4185)
* fix(#4148): dispatch wave-pre step hooks

External capabilities can render step hooks before a wave, but the execute workflow consumed only contributions and silently skipped every step. Reuse the shared dispatch contract before executor spawning and pin the capability-validator boundary with a red-first regression.

Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning

* test(#4148): pin wave-pre dispatch ordering

* test(#4148): pin wave-pre dispatch contract

* chore(#4148): bind upstream changeset PR

* chore(#4148): restore fork changeset identity

* fix(#4148): align wave-pre dispatch contract

Mirror the sibling wave-post all-shapes clarification while pruning redundant prose so the rebased workflow remains below its frozen byte ceiling.

Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning

* chore(#4148): restore upstream changeset identity

* fix(#4148): align wave-pre capability guidance

* docs(#4148): identify wave-pre manifest input

Name the third-party manifest trust origin at the wave-pre dispatch boundary so the reviewer-requested validation guidance matches wave-post.

Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning

* docs(#4148): preserve execute-phase byte budget

Remove a redundant advisory label while retaining the non-blocking contract, keeping the reviewer-required trust-boundary wording at the enforced 93,400-byte ceiling.

* fix(#4148): mark wave-pre manifest-input validation as security-relevant

Reviewer nit on PR #4185: wave-pre's step-dispatch sentence had the
(third-party manifest input) parenthetical but dropped the ⚠ marker
that wave-post's parallel sentence (execute-phase.md:1044) carries,
losing the visual flag that this validation is security-motivated.

Trims the redundant "of one" from "not one shape of one" to reclaim
the 4 bytes the marker adds — the ADR-857 byte-margin gate
(tests/claude-orchestration.test.cjs) leaves zero slack at the
93,400-byte ceiling.

* fix(#4148): trim wave-pre step-dispatch prose to clear ADR-857 byte ceiling

Merging next's unrelated growth (#3990's TDD_APPLICABLE conditional) pushed
execute-phase.md 116 bytes past the 93,400-byte ceiling, failing CI on all
three platforms. The security-relevant ⚠ marker and ref.command validation
call-out (added per prior reviewer nit) are preserved verbatim per the
pinned regression test in capability-registry.test.cjs; only the
non-pinned connective prose is trimmed.

* fix(#4148): recalibrate execute-phase.md self-imposed margin, restore security marker

next grew execute-phase.md by ~230 bytes across two unrelated merges during
this fix (#3990's TDD_APPLICABLE conditional, then a further step-extraction
commit), consuming this test's own self-imposed 93,400 safety buffer under
ADR-857's actual, unmodified 93,600 ceiling (docs/adr/857-capability-system.md:22).
The wave-pre step-dispatch sentence cannot shrink further without dropping one
of the pinned substrings this same test file asserts on (kind=="step",
loop-hook-dispatch, never blocks or redirects executor spawning, Validate
`ref.command`).

Raises the self-imposed margin to 93,550 (still 50 bytes under the real,
untouched ADR ceiling) and restores the ⚠ marker the prior reviewer round
required for the ref.command validation call-out, which byte pressure had
dropped.

* fix(#4148): restore full ref.command validation wording, drop self-imposed margin

Adversarial review (agy/gemini-3.8-flash-high) flagged two issues in the prior
CI-recovery commit:

1. Trimming "in-context before any shell use" from the step-dispatch warning
   weakened the inline operational instruction (the reader is told WHAT to
   validate but not the specific in-context-not-shell mechanism the referenced
   loop-hook-dispatch.md:45-51 threat model requires). Restored it - the merge
   with next since the last commit freed enough real margin (77 bytes under
   the untouched 93,600 ADR-857 ceiling) to afford it without any margin
   change.

2. The prior commit self-imposed margin bump (93400 to 93550) was, on
   reflection, the wrong lever: it is a number this PR invented, not an ADR
   value, and re-bumping it every time next grows execute-phase.md is a
   losing pattern (already needed twice in one session). Removed the
   redundant assertion; the same line existing bytes-under-93600 check
   against the real, frozen ADR-857 ceiling (docs/adr/857-capability-system.md:22)
   is the actual invariant and is untouched. workflow-size-budget.test.cjs
   tier hard cap (98304 bytes, extract-not-bump by design) remains the
   correct backstop for runaway growth.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 02:56:37 -04:00
JusticeWay
0fca71eaae enhance(#2529): cover every workflow with response-language directives + CI lint (#2558)
* enhance(#2529): cover every workflow with response-language directives + CI lint

Every workflow now carries response-language coverage in one of three forms,
and a CI lint keeps it that way.

- 43 workflows load the new shared reference,
  `gsd-core/references/response-language-directive.md`, by eager `@`-import.
- Lazy-loaded modes/steps/templates, which cannot rely on an eager import,
  carry an exact inline directive; 35 such paths are pinned by exact path.
- Fragments dispatched by a covered parent inherit coverage, proven per file
  rather than granted per directory.

The 45 workflows whose directive covered only "questions, prompts, and
explanations" now name inter-tool narration, which is the defect #2529
reports: the running commentary between tool calls stayed English while the
answers around it were translated.

`scripts/lint-response-language-coverage.cjs` enforces it and fails closed on
three independent discovery failures (unreadable catalog, empty catalog,
unfollowed symlink). It resolves which reference a workflow imports and applies
the same four-predicate test to that file, so a weakened shared reference
uncovers its importers instead of passing silently, reported once as a systemic
failure rather than 43 times. The walk follows symlinked subtrees with a
realpath cycle bound. `lint:ci` invokes it by name.

REQ-LANG-03 and REQ-LANG-04 state the contract in docs/FEATURES.md;
REQ-LANG-04 names the two forms that satisfy it ("narration", "between tool
calls") rather than enumerating class members an author cannot use verbatim,
and a test pins that text to what the matcher accepts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): register the coverage test in the docs-guard lane

`107eb8c1` (#3787) landed the docs-guard lane on `next` while this PR was
open: a test that reads a `docs/` path must be named in
`scripts/docs-guard-registry.cjs` or carry a `docs-guard-exempt` marker,
so the guards that read a doc run on the PR that changes it.

`tests/response-language-coverage.test.cjs` reads `docs/FEATURES.md` -- it
extracts every form REQ-LANG-04 offers an author and runs each through the
matcher that enforces it. Registration, not exemption, is the correct side
of that gate: a reword of the requirement with no code change is precisely
the diff this test exists to catch, and it is the diff the lane would
otherwise skip.

Registered narrowly (`['docs/FEATURES.md']`) rather than with the `'*'`
sentinel, so an unrelated docs change does not pull this test into the lane.

Verified: lint-docs-guard-registration 0 violations, tests/ci-docs-guard-registry.test.cjs
51/51, lint:ci exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): consolidate this PR's emitted-growth acks into its own fragment

This PR ripples emitted bytes across 85 workflow paths. Until now each ripple
was acknowledged by appending to whichever live fragment owned that path,
because two ack sources may never name the same path.

`a84f7563` (#3078) swept all 45 fully-spent fragments off `next`. Forty-two of
the paths this PR grows were owned by swept fragments, so those keys are now
unowned and this PR's own fragment declares them directly -- one path, one
source, and no dependence on a fragment that no longer exists. Each adopted
entry keeps its measurement and records where it came from.

Two paths are handled differently, because the sweep did not free them:

- `review.md` is now owned by `3034-parallel-reviewer-lanes.json`, which
  landed on `next` after the sweep. Its entry is live, so the old route still
  applies: this PR's note is appended to that entry rather than declared a
  second time.
- `plan-review-convergence.md` keeps the arrangement made in round 24.

Result: 3 fragments in the directory, 85 keys in this PR's own,
0 cross-source duplicates. `lint-emitted-drift-ack` exit 0,
`tests/emitted-attribution.test.cjs` green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): move REQ-LANG-03/04 into the feature fragment that now generates them

`36375513` (#3845) made docs/FEATURES.md a generated projection of
docs/features/*.md, marked "do not edit by hand". This PR wrote REQ-LANG-03
and REQ-LANG-04 straight into the generated file, so the rebase left the
requirement present in the projection and absent from its source -- the next
regeneration would have deleted both, and `tests/features-index-gate.test.cjs`
was already red on the mismatch.

Both requirements now live in docs/features/response-language-config.md
alongside REQ-LANG-01 and -02. Regenerating produces a docs/FEATURES.md that is
byte-identical to the committed one, so the text this PR shipped is unchanged --
only its source of truth moved to where #3840 put it.

The docs-guard registration is widened to name the fragment as well as the
projection. The requirement's source is the fragment now, and an edit there
that skips regeneration would otherwise reach this guard through neither path.

Verified: features-index-gate 68/68, lint-docs-guard-registration 0 violations,
ci-docs-guard-registry + response-language-coverage 142/142.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): hand the plan-phase ack back to its new live owner

`c933184b` (#3825) landed `3172-stated-failing-direction.json` on `next` after
fragment had adopted that path when the sweep left it unowned, so the merged
tree named it from two sources -- a hard failure in
`scripts/lint-emitted-drift-ack.cjs`.

The path has a live owner again, so the append route applies: this PR's note
joins that entry, carrying its own measurement, and the key is dropped from
this PR's fragment (84 keys left, the others untouched). The provenance
sentence written for the swept-fragment case is removed rather than reused --
this path was never orphaned, so that account of it would be false.

Same shape as `review.md` and `plan-review-convergence.md`: ownership is a
property of the merged tree, and a fragment landing upstream after a push can
reclaim a key no local check would have flagged.

Verified: lint-emitted-drift-ack exit 0, lint:ci exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): state byte figures that are true against the tree

The reference claimed `execute-phase.md` has "2 bytes of headroom under the
ceiling named below". That was true when the sentence was written -- the file
sat at 93398 against the 93400 comfort assert -- and upstream has since shrunk
it to 91493 against a 93600 hard ceiling, so the figure now understates the
headroom by three orders of magnitude. The rationale the sentence supports does
not depend on the number, so the number is gone rather than refreshed: a
restated figure would go stale again on the next upstream edit, and nothing
parses it.

Audited every other numeric claim this PR ships the same way, mechanically
against the merge base: all 82 FILE-delta claims in the ack fragment match the
real per-file delta exactly, and the 1,629-byte reference and 63-byte import
line check out. One class was imprecise: the 41 notes for workflows whose
inline directive was rewritten in place quoted the conversion counterfactual as
"+1,692 bytes more loaded context", which is the reference form's whole weight,
not the increase over the inline directive those files already carry. Each now
names both quantities and the net (+1,605 / +1,609 / +1,584).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): one rule for pinned vs inherited coverage, and the docs to pick it

Review measured that 14 of the 35 pinned fragments would pass by inheritance
anyway, and that the PR asserted both readings at once: inheritance is real
coverage (so those 14 pins are noise) or it is not (so 30 inheriting fragments
are green-but-uncovered). Only one can be true.

Inheritance is real: the predicate proves it per file -- the parent must
dispatch this exact path from a read/execute context AND be covered itself --
so the parent's directive is in the loaded context by the time the fragment is
read. The 14 pins are therefore removed along with the directive lines they
pinned, and those files inherit like the 30 structurally identical ones. The
rule is now stated where the set is declared, and enforced from the other side
by a test: no member of the pinned set may be one that would have inherited.
That is what decides the form for the next fragment.

- pinned set 35 -> 21; 14 workflow files revert to their base content
- `findViolations` no longer returns early on a pinned path: a file that becomes
  eagerly loaded and takes the shared reference is strictly better off, and the
  gate must not red that. The reference form is admitted because its own wording
  is validated in turn; an arbitrary reworded inline line still fails.
- the reference-directive cache is keyed by size and mtime, not by path alone,
  so a rewritten reference re-asked in one process no longer returns the stale
  verdict
- `carriesInlineDirective` names its negation blindness: four independent hits
  read vocabulary, not polarity
- the real-tree scan asserts each source produced files instead of `> 152`, a
  constant that read as the workflow count and would have passed a scan that
  lost one of its two directories
- the pinned-set size assertion goes the same way: the size follows from the
  rule, so the rule is what the suite asserts

Docs, for the gate that now governs every future workflow:
- `docs/contributing/response-language-coverage.md` -- why the narration class
  is the discriminator, the four coverage forms, the decision order that picks
  one, the pinned line, and what each failure message means
- a row in CONTRIBUTING.md's CI checks table, matching the docs-guard row
- `docs/CONFIGURATION.md` points at it from the `response_language` entry

Also: the changeset said 45 reworded workflows; it is 44 (42 @-reference + 21
pinned + 44 rewritten = 107 touched). That text ships to CHANGELOG.md.

`3707-parse-gap-reporting.json` landed on `next` reclaiming `audit-uat.md` and
`progress.md`; both handed back by the append route, leaving 82 keys here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): correct the reference-taker count, 43 -> 42

The ack notes said the import line is byte-identical "in each of the 43
workflows that take the reference" and that the alternative would be "43 inline
copies". The shared reference has 42 importers; the 43rd file in review's table
is `execute-phase.md`, which imports the OTHER reference. Corrected in all 41
notes that carry the sentence, across this PR's fragment and the two it appends
to.

Found by re-running the numeric audit from the previous round after the rebase,
which also re-verified all 84 FILE-delta claims against the new base -- all
exact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2529): migrate the emitted-drift ack from a fragment to commit trailers

ADR-3942 (#3954) landed while this PR was open: the acknowledgment is now a commit
trailer and tests/emitted-drift-acks/ no longer exists. The fragment is deleted and
each key it declared becomes one trailer, reasons unchanged.

The four keys this PR had handed to 3034-*, 3172-* and 3707-* under the one-source
rule come home here. That rule was the whole reason for the hand-backs, and the
trailer model has no shared namespace to collide in -- five of this PR's rounds were
spent on exactly those collisions.

Emitted-Drift-Ack-Growth: add-backlog.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: add-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: add-tests.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: add-todo.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: ai-integration-phase.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by ddf85287 (fix(#3357), #3513), and no fragment on `next` declares this path now. The ack therefore returns to this PR's own fragment, which is the only live source for it — the change to the path is this PR's.
Emitted-Drift-Ack-Growth: analyze-dependencies.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: audit-fix.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 19: this PR declared the path in its own fragment, and `3602-workflow-subagent-model-resolution.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3602-workflow-subagent-model-resolution.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: audit-milestone.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed from `2962-zsh-nomatch-for-glob-portability.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: audit-uat.md — A live/archived split was added and then reverted on this branch (see `$comment`): the split's extra rule in `initialize`, the narrowed Unparsed-table filter, and the separate 'Unparsed UAT Files in Archived Milestones' informational section are all removed, so the file settles at origin/next 5582 -> 7124 bytes (+1542, final). #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: autonomous.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 16: `3210-autonomous-precondition-gate.json` landed on `next` in 8fc88f66 (fix(#3210), #3528) and declares this path today. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3210-autonomous-precondition-gate.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: check-todos.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: cleanup.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 18: this PR declared the path in its own fragment, and `2142-quick-task-archival.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `2142-quick-task-archival.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: code-review-fix.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 13: `3190-code-review-fix-auto-rewrite-review.json` landed on `next` in 1d5d7795 (fix(#3190), #3434) and declares this path too. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3190-code-review-fix-auto-rewrite-review.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: code-review.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 19: the fragment that carried this sentence (`3503-diff-base-scope-anchor.json`) was retired on `next` by 2fca0e17 (enhance(#2554), #3695), and `2554-code-review-depth-overrides.json` declares this path today. One path takes exactly one ack source, so the sentence follows the path to its live owner. Re-homed from `2554-code-review-depth-overrides.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: complete-milestone.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 15: the fragment carrying it (`3458-audit-open-acknowledge-wiring.json`) was retired on `next` and the path is declared by `3409-unreachable-guard-arms.json` today. One path takes exactly one ack source, so the sentence follows the path to its live owner. Re-homed from `3409-unreachable-guard-arms.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: debug.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 14: the fragment that carried this sentence (`3149-init-debug-entry-point.json`) was retired on `next` by 26f8015c (fix(#3448), #3476), and `3448-debug-autoresume-next-action.json` declares the path today. One path takes exactly one ack source, so the sentence follows the path to its live owner rather than being dropped or re-armed under a retired number. Re-homed from `3448-debug-autoresume-next-action.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: diagnose-issues.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: an inline copy in every workflow would be that many places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 19: this PR declared the path in its own fragment, and `3602-workflow-subagent-model-resolution.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3602-workflow-subagent-model-resolution.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: discuss-phase.md — #2529 round 36: this workflow's inline directive was rewritten in round 10 to name inter-tool narration, but in the compressed form, and that rewrite came to −1 byte against `next` — so it declared no growth and this key was absent from this PR's ack set until now. Round 36 replaces the compressed clause with the same enumeration the other rewordings carry — narration between tool calls, status updates, progress notes, findings, questions, prompts, and explanations — because `discuss-phase.md` started from the identical upstream sentence as `verify-work.md` and `new-milestone.md` and those two took the full list, so the shorthand was an inconsistency rather than a decision. +87 bytes against `next`, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. The directive stays INLINE rather than becoming an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have cost 1,692 bytes of loaded context against the 87 this sentence costs. `commands/gsd/discuss-phase.md` dispatches this workflow lazily (`Read and execute ...`) rather than `@`-importing it, so the 87 bytes land in the installed file and are read once the workflow is dispatched, not on every command invocation.
Emitted-Drift-Ack-Growth: discuss-phase-assumptions.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 15: `3409-unreachable-guard-arms.json` landed on `next` in #3558 and declares this path too. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3409-unreachable-guard-arms.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: discuss-phase-power.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: do.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: docs-update.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +83 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +83 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 83 bytes this inline directive costs, a net +1,609. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,609 bytes more loaded context per invocation. Re-homed in round 19: this PR declared the path in its own fragment, and `3602-workflow-subagent-model-resolution.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3602-workflow-subagent-model-resolution.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: edit-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 13: `3262-editphase-milestone-scope-guard.json` landed on `next` in fd4715f8 (fix(#3262), #3446) and declares this path too. One path takes exactly one ack source, so the sentence moves here and the key leaves this PR's fragment. Re-homed from `3262-editphase-milestone-scope-guard.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: eval-review.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by ddf85287 (fix(#3357), #3513), and no fragment on `next` declares this path now. The ack therefore returns to this PR's own fragment, which is the only live source for it — the change to the path is this PR's.
Emitted-Drift-Ack-Growth: execute-plan.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 14: the fragment that carried this sentence (`2652-quick-diagnose-dispatch-isolation.json`) was retired on `next` by 362d0434 (fix(#3370), #3478), and `3370-execute-phase-gate-conflation.json` declares the path today. One path takes exactly one ack source, so the sentence follows the path to its live owner rather than being dropped or re-armed under a retired number. Re-homed from `3370-execute-phase-gate-conflation.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: explore.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed from `2229-explore-claim-disposition.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: extract-learnings.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: fast.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Re-homed in round 18: this PR declared the path in its own fragment, and `3585-planning-commit-guard.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3585-planning-commit-guard.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: forensics.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: graduation.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: health.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 13: the fragment that carried this sentence (`2573-state-head-freshness.json`) was retired on `next` by 7ddcc198 (fix(#3309)), and `3309-health-docs-generated.json` declares the path today. One path takes exactly one ack source, so the sentence follows the path to its live owner rather than being dropped or re-armed under a retired number. Re-homed from `3309-health-docs-generated.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: help.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: import.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 18: this PR declared the path in its own fragment, and `3576-references-canonical-cites.json` landed on `next` declaring it too. One path takes exactly one ack source, so the sentence moves here and the key leaves ours. Re-homed from `3576-references-canonical-cites.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: inbox.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation.
Emitted-Drift-Ack-Growth: ingest-docs.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed from `2658-trae-instruction-file-path.json` in round 26: `a84f7563` (#3078) swept that fragment as all-spent, so this path is unowned and this PR's own fragment declares it directly.
Emitted-Drift-Ack-Growth: insert-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: list-phase-assumptions.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: list-seeds.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed.
Emitted-Drift-Ack-Growth: list-workspaces.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 …

* fix(#2529): read the catalog-relative dispatch spelling, and the plural of "output"

Two false positives in the coverage lint, both surfaced by this round's work
rather than by a red gate finding them for us.

#3552 landed `execute-phase/steps/protected-branch.md` on next while this PR was
open, dispatched from execute-phase.md's `"none"` arm with the path written
RELATIVE to the catalog. `namesFragmentAsEntryPoint` only ever looked for the
`gsd-core/workflows/`-rooted spelling, so it read a live dispatch as no dispatch
and the new fragment as uncovered. It now accepts both spellings and matches the
relative one on a path boundary, so `vendor/<path>` cannot vouch for `<path>`.

Recognizing that spelling makes one pin redundant: execute-phase.md dispatches
executor-isolation-dispatch.md the same way, so the fragment inherits and its
own copy of the sentence comes back out. That is the rule round 29 encoded,
enforced by the test that measures it rather than by hand.

`output` was the one term in USER_OUTPUT_RE without an `s?`, so "translate all
outputs, including narration between tool calls" read as uncovered. The new
property tests caught it on their first run.

Those properties pin the rule the hand-written cases are instances of: four
signals on ONE line accept, dropping any one rejects, spreading them across
lines rejects. The vocabulary is written out in the test rather than read back
from the script's regexes, per CONTRIBUTING.md "Fixture provenance (#2371)" -- a
generator seeded from the matcher can only re-derive what the matcher already
believes, and that independence is what caught the plural. Both new properties
are mutation-verified: dropping the narration predicate reds the necessity
property, and collapsing the document to a single line reds the cross-line one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#2529): spell the narration enumeration out in the last two shorthand directives

discuss-phase.md and plan-phase.md were the only two of this PR's 44
rewordings that abbreviated the inserted clause to "narration between tool
calls included" instead of naming the output classes the way the rest of them
do. Both forms satisfy the lint's four predicates, so nothing was broken --
but the point of #2529 is that an author reading one workflow should not have
to infer what the neighbouring one means by "included".

Both abbreviations were size decisions rather than wording ones, and both
reasons have since expired because next shrank the files. discuss-phase.md
sat 25 bytes under the 32,000-byte #717 dispatcher budget and now has 1,825;
plan-phase.md sat 87 bytes under the 94,519-byte ADR-857 capstone ratchet
against a +108 clause and now has 3,180. Neither budget is raised here and no
unrelated prose is trimmed; workflow-size-budget and
phase6-capstone-conformance both pass.

discuss-phase.md started from the identical upstream sentence as verify-work.md
and new-milestone.md ("All user-facing questions, prompts, and explanations in
this workflow"), and those two received the full enumeration; it now matches
them exactly. plan-phase.md keeps its own scope word ("orchestrator output") and
its subagent pass-through instruction, both upstream's, and only trades the
shorthand for the enumeration.

The shorthand now appears nowhere in the catalog. The two remaining variants
(plan-review-convergence.md, spec-phase.md) keep upstream's own verb and scope
and end on "report prose", which is what those workflows actually emit --
rewriting those would change a directive's strength, not its wording.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2529): match the workflow extension case-insensitively in coverage discovery

`findMarkdownFilesRecursive` filtered on `entry.name.endsWith('.md')`, so
`SETTINGS.MD` — the same file to Windows and macOS, a different one to Linux —
was skipped on the only platform whose verdict gates the merge. The direction of
that failure is the problem: a workflow the walk declines to see is a workflow
this lint certifies by omission, which is the same vacuous pass `main()` already
refuses when discovery returns nothing at all.

The filter is now an allowlist keyed on the lowercased `path.extname`.
`.mdx` stays out on purpose: admitting an extension states what a workflow IS,
and that claim has a second half — `inheritsParentCoverage` resolves a
fragment's parent as `<workflow>.md`. An `.mdx` entry belongs here next to the
parent resolution it would have to move with, not ahead of it.

Two tests: an uppercase-extension file is discovered AND lands as a violation
rather than an exemption, and every admitted extension is spelled so the
lowercasing match can reach it (an uppercase or dotless entry would be dead
configuration that reads like coverage).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#2529): cover the quick-batch workflow that #3676 landed uncovered

`next` gained `quick-batch.md` and nine step fragments in 2f64e6230 (#3676,
PR #4212) with no response-language directive, so the merge result reds this
PR's own lint with 10 violations. The lint is doing exactly what it exists to
do; the coverage is what has to move.

`quick-batch.md` takes the shared @-reference on line 1, the same as the other
42 top-level workflows, and eight of the nine fragments then inherit through
its `read and execute` stubs. The ninth does not:
`quick-batch/steps/plan-checker-loop.md` is dispatched by a SIBLING fragment
(`planner-wave.md:134`) and named in the parent only inside a parenthetical
with no dispatch verb, which is the shape round 29's rule already covers for
`execute-phase/steps/regression-gate-run.md` and
`plan-phase/steps/prd-express-path.md`. It carries the pinned inline directive
and joins `EXACT_INLINE_DIRECTIVE_WORKFLOWS`; the comment above that set now
names four such fragments instead of three. Coverage: 163 workflows.

`FULL_BUDGET` in tests/skill-frontmatter-contract.test.cjs moves 844 -> 846.
The same commit grew `help/modes/full.md` from 834 to 844 lines, landing it
exactly on the ceiling with zero slack, and the two lines this PR adds there
are its pinned directive and the blank separating it. That is a coverage
contract every workflow carries, not the content creep the budget guards.
The #597 ratchet rule holds: actualMax 846, slack 0, well inside LARGE_GRACE.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: quick-batch.md — #2529: the workflow arrived on `next` in 2f64e6230 (#3676, PR #4212) with no response-language directive, so this PR's lint reds on the merge result; covering it is the PR's whole contract, not an optional extra. It gains the shared directive as a single eager `@`-reference line, the identical form the other 42 top-level workflows take. FILE delta: +62 bytes, as the gate measures it. LOADED-CONTEXT delta: +1,691 bytes — the import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates read the FILE and not the transitive inline, so they see 62 of those 1,691 bytes; the remaining 1,629 are declared here because no gate reads them. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Nine `quick-batch/steps/*` fragments are covered without a byte of their own — eight inherit through the parent's dispatch stubs, and the ninth takes the pinned inline sentence, which the emitted surface does not measure.

* fix(#2529): scope row 48 by what a diff says, not by which paths it names

`tests/gsd-quick-batch-quick-regression.test.cjs` treats any branch touching a
`quick-batch` path as #3676 phase work, then forbids it from editing ordinary
`quick.md`. This PR covers EVERY workflow with the shared response-language
directive — quick-batch.md and its fragments included — so the scope check
turned true, and the row read this PR's one-line directive on `quick.md` as a
phase violation.

That is the false positive the row's own #3730 note already scoped away from,
arriving by the other door: not an unrelated branch that misses the surface,
but a catalog-wide sweep that touches all of it. A path now counts as phase
work only when its diff says something other than the coverage contract, and
the two accepted directive forms are read from
`scripts/lint-response-language-coverage.cjs` rather than restated, so a
reworded contract cannot leave the carve-out matching prose the lint no longer
recognizes. A file the branch ADDED still counts — every line is new, which is
what a real #3676-phase branch looks like.

The invariant is unweakened in the direction that matters: a phase branch that
edits `commands/gsd/quick.md`, `gsd-core/workflows/quick.md` or anything under
`quick/steps/` for any reason other than the directive still fails the row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-04 21:12:12 -04:00
Tom Boucher
d29b50d696 fix(#4051): route specific intents first and confirm before dispatch in --do (#4289)
* test(#4051): pin freeform routing specificity contract in do.md

* fix(#4051): order freeform routing specific-first, confirm before dispatch, argument-aware forwarding

* fix(#4051): regenerate FEATURES.md, satisfy docs-guard on new routing test

Emitted-Drift-Ack-Growth: do.md — deliberate growth: specific-first routing table (code-review, plan review, ui-review, secure-phase, audit, docs-update, phase CRUD rows), a REQ-DO-03 confirm step, and argument-hint-aware dispatch.

* chore(#4051): fold regression into non-bug-prefixed test filename per lint-regression-test-names

* fix(#4051): review fixes — em-dash description style, split audit-fix route

* chore(#4051): sync skill mirrors of execute-phase/phase descriptions

* chore(#4051): add changeset (pr backfill to follow)

* chore(#4051): backfill PR 4289 in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-04 17:19:16 -04:00
Tom Boucher
8249ebcf6e fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate

RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do
not exist yet; every row fails on require. Per #3770 only an intentional
target-test failure may authorize GREEN; zero-test discovery, fixture crashes,
unrelated failures, and unexpected green are INVALID_RED.

* fix(3770): require intentional RED evidence before GREEN

Only an intentional failure of the TARGET test (distinctly named, TAP-reported
assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test
discovery, fixture/load crashes (file-named failures), nonzero exits without a
failing test, unrelated failures, unexpected greens, and malformed/missing
records are INVALID_RED and block GREEN.

- src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses
  the prohibition-enforcement TAP primitives; fail-closed, never throws)
- check tdd-red-evidence <record.json>: validates the persisted record
  (command, exit code, failing test, expected, actual)
- gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now
  requires the evidence record + gate verdict, not a nonzero exit or a RED: tag

* chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs

* fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib

- gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B <
  49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid)
- tests: the row-6 fixture used String.replace (first-occurrence), so the
  `not ok` line still named the target test and the classifier was right to
  accept it; replaceAll makes the failure genuinely unrelated
- eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs
  (lint the src/*.cts source, per ADR-457 migration rule)

Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line

* chore(3770): add changeset

* chore(3770): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-04 16:44:38 -04:00
Tom Boucher
2f4f7538e9 fix(#4264): wire both TDD dispatch backends to phase.tdd-applicable (#4284) 2026-09-04 16:08:42 -04:00
Xiangfang Chen
b1b7cabfb5 docs(#4123): add gsd-qoder EoS registry entry (#4278)
* docs(registries): add gsd-qoder EoS entry

Adds one `type: "eos"` entry for a Qoder host integration and regenerates
docs/registries/eos-registry.md.

Qoder is Alibaba's AI coding product family (Qoder CLI and Qoder Desktop).
The integration depends on @opengsd/gsd-core, negotiates the ADR-1239
host-integration handshake, and projects GSD's agents, skills, and hook
scripts into the Qoder config directory (~/.qoder, or ~/.qoder-cn for the
China edition), merging GSD's lifecycle hooks into settings.json.

Every axis is sourced from Qoder's own docs per the never-infer rule.
`dispatch.isolation` is `none`: Qoder documents `isolation: worktree` as a
frontmatter-declared, per-agent-definition property, and GSD's two
isolation negotiation models both assume a per-dispatch injection point
Qoder does not expose.

Re-homes the Qoder runtime work from #860 / PR #2005, which was closed in
favor of the EoS path.

Closes #4123

* docs(#4123): backfill changeset pr field
2026-09-04 15:09:10 -04:00
Tom Boucher
75ee7b0214 enhance(#4273): add phase.tdd-applicable single-owner predicate (#4277)
* enhance(#4273): add phase.tdd-applicable single-owner predicate

One query verb computes TDD-applicability for a plan (CLI flag, plan
type: tdd frontmatter, a task's tdd="true" attribute, or the
workflow.tdd_mode config default), mirroring phase.mvp-mode's
precedence-cascade shape. Foundation for epic #4272 Phase 2, which
wires both dispatch backends to consume it instead of restating the
predicate independently.

Also fixes workflow.tdd_mode, workflow.research, and
workflow.nyquist_validation, which never reached
cmdInitExecutePhase/cmdInitPlanPhase/cmdInitDebug/cmdInitNewMilestone
because loadConfig() never populates config.workflow — a dead
accessor found while wiring this verb's own config read, fixed inline
per the no-defer rule rather than left alongside it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4273): document phase.tdd-applicable's FEATURES.md entry

Add a docs/features/ fragment for the new phase.tdd-applicable query
verb and regenerate docs/FEATURES.md. docs/COMMANDS.md is left
untouched: it documents /gsd-* slash commands only, and the sibling
verb phase.tdd-applicable mirrors (phase.mvp-mode) has no formal CLI
reference entry anywhere in docs/ either -- only inline prose mentions
in docs/reference/workflow-fragments.md -- so there is no COMMANDS.md
precedent to extend.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): use PHASE_NOT_FOUND reason code, remove try/finally from tests

Two orthogonal code reviews flagged a mistyped error reason and a CONTRIBUTING.md-banned try/finally pattern in the phase.tdd-applicable change; both are corrected here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): stop whitelisting capability-owned config keys centrally

workflow.tdd_mode, workflow.research, and workflow.nyquist_validation are
each already owned by their own first-party capability's federated config
schema (the tdd/research/nyquist capabilities declare them under their own
capability.json `config`), resolved via isCapabilityConfigKey. Adding them
to gsd-core/bin/shared/config-schema.manifest.json's central validKeys, as
the prior commit in this branch did (mirroring workflow.mvp_mode, which
genuinely is central-only), declares the same key in two places at once.
That collision breaks capability-loader.cts's loadRegistry composition:
gsd-test caught this as 84-85 unrelated failures across
capability-cli/capability-command-dispatch/capability-lifecycle test files,
every one showing "unknown capability: <id>" for a freshly-installed
third-party capability that should have resolved fine.

Verified directly (not asserted): reverting only this file, keeping the
config-loader.cts tdd_mode/research/nyquist_validation flattening and the
init.cts call-site fixes from the prior commit, and re-running the exact
capability install + capability set repro from
tests/capability-cli.test.cjs's "issue-2322" test locally reproduces the
failure with the whitelist entries present and clears it without them.
loadConfig() still surfaces all three flattened values correctly with no
central whitelist entry (confirmed directly against the compiled module) —
the whitelist additions were never required for the #4273 fix to work; they
were an incorrect over-application of the mvp_mode precedent to keys that
aren't central.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): use getNested for tdd_mode (no legacy top-level fallback), allowlist new test file

Both fixes address defects found by a gsd-test bench run: tdd_mode routed through get() invented an undocumented top-level alias that silently outranked the canonical workflow.tdd_mode key, and the new phase-tdd-applicable test file was missing from the file-count allowlist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4273): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 14:14:56 -04:00
Adnan
f4bf449296 fix(#3850): surface gaps_found VERIFICATION files in audit-uat (#3879)
* fix(#3850): surface gaps_found VERIFICATION files in audit-uat

cmdAuditUat admits `human_needed` OR `gaps_found`, but
parseVerificationItems had a body only for the first and returned an
empty array for the second — standing on a comment deferring to
`plan-phase --gaps`, a different command audit-uat never reaches. Since
cmdAuditUat pushes a file into `results` only when `items.length > 0`, a
`gaps_found` report did not under-report: it vanished, taking its
phase's `by_phase` row with it, so a clean-looking total gave the reader
no cue anything was skipped.

Eligibility now has one owner (the caller) and parseVerificationItems
reports what the file says.

The closed-entry filter could not be built on extractFrontmatter: its
array-item parser keeps only each `- ` entry's FIRST line and has no
notion of nested key/value objects, so an entry's `status:`/
`resolution:` siblings never reach its output and a closed entry is
indistinguishable from an open one downstream. Rather than grow a
competing object-list parser — or change extractFrontmatter, whose blast
radius is every frontmatter consumer in the repo — this reads the raw
segment BEFORE the flattening, via the existing anchored
sliceTopLevelFrontmatterSegments, and hands it to the `## Gaps`
machinery that already parses exactly this `- `-opened, indentation-
continued shape.

The human_needed path is byte-for-byte unchanged: same reader, same
display names, same numbering, no resolved-entry filtering — pinned by
a test and verified by identical CLI output on base and head.
parseGapsItems keeps its narrower `status: resolved` rule so no
*-UAT.md behaviour moves.

Closes #3850

* chore(#3850): backfill changeset pr number for #3879

* fix(#3850): one parse per entry, one fence parser, one resolved-entry rule

Adversarial review on #3879: B1, B2, M3, m5, m8 and n9.

B1 — `sliceFrontmatterArrayEntries` hand-rolled a second frontmatter fence
regex, which re-asserted the byte-0 rule #2977 removed: a BOM'd file (PowerShell
5.1 `>`/`Out-File` writes one by default) sliced nothing, so a `gaps_found`
report vanished from the audit exactly as it did before this fix — this issue's
own symptom, on a platform the repo already has a named defect class for.
`extractFrontmatter`'s BOM+fence logic is now factored out as
`frontmatterRegion` and shared. One fence parser, not two.

B2 — the resolved-entry skip paired two DIFFERENT parsers by array index:
`parseYamlRegion` is indent-blind, `splitGapsEntries` is indent-anchored. A
block sequence written at its key's indent — ordinary, legal YAML — makes them
disagree about entry count, and from the first disagreement every index names a
different entry, so an OPEN entry inherits a CLOSED one's resolution and is
silently dropped. That is the defect this PR exists to fix, reintroduced inside
the fix. Display name and sibling fields now come from ONE parse of the raw
slice; `frontmatterEntryDisplayName` applies `parseQuotedScalar` exactly as
`parseYamlRegion` does, so the string is byte-identical to what
`extractFrontmatter` produced. The flattened array remains the #2286 GATE, but
is no longer the source of items. `sliceFrontmatterArrayEntries` also takes the
LAST duplicate key, matching `parseYamlRegion`'s last-wins assignment.

M3 — `frontmatterEntryToUatItem` is the single entry->UatItem mapper both
readers use, rather than two copies differing only in `result`.

m8 — closed entries are skipped on BOTH statuses. The earlier asymmetry cited an
acceptance criterion #3850 does not contain: the issue has no AC section, and
its suggested fix (2) states the skip unconditionally, naming a file with 14 of
16 entries resolved. That file is `human_needed`, so the asymmetry left the
reporter's own scenario over-reporting by 14.

m5 — `sliceTopLevelFrontmatterSegments`' contract doc names both consumers and
says the column-0 boundary rule is now a cross-module contract.

n9 — the vestigial bare block is gone and its body de-indented.

Tests: the B1 BOM case, B2's nested-sequence and bare-bullet repros, a CRLF
fixture (M4 — it survived by accident, now pinned) and the unified skip rule.
Fail-first verified by running the new tests against the pre-fix build: the BOM,
nested-sequence and unified-skip cases are red there.

* fix(#3850): read the entries as objects, not as re-parsed display text

Rebased onto `next`, which changed the ground this fix stood on. ADR-3473 §8.1
(#3881) replaced the hand-rolled frontmatter scanner with the vendored js-yaml:
`parseQuotedScalar` and `parseYamlRegion` no longer exist, and an object entry
now flattens to `test: A, resolution: R` rather than to its first line.

The original mechanism existed ONLY to work around that lossy first-line
flattening — it sliced the raw frontmatter segment and re-parsed each entry by
hand so a `resolution:` sibling was visible at all. With a real parser upstream
that workaround is obsolete, so it is deleted rather than repaired:
`sliceFrontmatterArrayEntries`, `frontmatterEntryDisplayName`, the
`splitGapsEntries`/`extractGapEntryFields` reuse and the second fence regex are
all gone.

`frontmatter.cts` instead exposes `frontmatterObjectListEntries(content, key)` —
the same parse `extractFrontmatter` runs (same BOM strip, same byte-0 fence,
same anchor/alias and sentinel guards, same ambiguous-colon repair), stopping
one step before the display flattening. `flattenObjectListItem` is exposed
alongside it so a caller deriving a display name produces the byte-identical
string `extractFrontmatter` would have.

That collapses the review's blockers into properties of the parse rather than
things this fix has to get right:

- B1 (BOM) — shares `extractFrontmatter`'s strip; verified through the CLI.
- B2 (index pairing) — there is no second reader. Display name and sibling
  fields come from one object.
- M3 (duplicate mapper) — one `frontmatterEntryToUatItem` for both readers.
- M4 (CRLF) — js-yaml's, not ours; verified through the CLI.

Also confirmed on the rebased base, per review: #3850 still reproduces on `next`
after #3707 landed (`total_files: 0`, `total_items: 0` on a `gaps_found`
fixture), so this PR is still doing work #3707 did not do. Nothing was dropped
as redundant.

One behaviour note: `entryField` returns a present value verbatim and treats
only whitespace-only as absent. Trimming would rewrite an author's `truth:` on
its way to becoming the display name.

* fix(#3850): keep every frontmatter list entry at its own row

Review round 3's Blocker. `frontmatterObjectListEntries` filtered its result
to objects, and filtering COMPACTS: `parseHumanVerificationItems` then
numbered the survivors by their position in the compacted array. On a list
mixing object and non-object entries the non-object rows disappeared outright
and the rest were renumbered — #3850's own vanishing-row defect, reached
through entry SHAPE instead of file STATUS. Base never had it: it walked the
display array, so every row surfaced at its own position.

Renamed to `frontmatterListEntries` and it no longer filters (the name now
matches what it returns). Deciding what a non-object entry MEANS is a
caller's judgement; dropping it is nobody's.

Both readers now walk the DISPLAY array — one element per row, the array
#2286 already gates on — and consult the parsed array only for "does this
entry carry a closure field?". `parsedEntriesFor` owns that pairing and
checks the two lengths agree before trusting an index; all-null is the
correct degradation, since over-reporting a closed row is recoverable and
closing the wrong one is not. Names stay byte-identical to base for every
entry shape, including a nested sequence (`[nested]`, not `["nested"]`).

Same class closed in the gaps reader: a non-object `gaps:` entry surfaced
nothing at all and now surfaces as `unknown`, which is this module's
documented fail-safe direction (`parseGapsItems`) on a false-negative bug.

Also restores the shared fence parser round 2 accepted. The ADR-3473 rebase
dropped `frontmatterRegion` and left the BOM strip and byte-0 fence rule
inlined twice; `extractFrontmatter` now routes through it, so "one fence
parser" is enforced rather than asserted in a comment.

Minors: `frontmatterEntryToUatItem`'s dead `forcedResult` option deleted and
its "shared by both readers" comment corrected — it has one call site, and
the two readers differ deliberately, each mirroring its own established
sibling (`parseGapsItems` vs #2286). Documented at the divergence.

Tests: `B2` asserted a name substring, so it passed while the row was
mis-numbered and would have passed through outright loss; it now asserts
positions and count. B2b pins the reviewer's 6-entry mixed fixture verbatim,
B2c the survivors' file positions across skipped rows, B2d the gaps reader.
All four fail-first against the reviewed head; 332/332 green with the fix.

* fix(#3850): make status authoritative, and let the two gaps readers agree

Round 4 review, all five findings.

Major. `isFrontmatterEntryResolved` treated a non-empty `resolution:` as
closure regardless of `status:`, so `status: failed` + `resolution:
"attempted retry, still failing"` vanished from the report — the
silently-vanishing-item defect #3850 exists to close, reached by field
combination instead of file status.

Closure is now per key, because the two keys have different conventions
and one rule cannot serve both:

  `gaps:`               `status: resolved` only, byte-identical to the
                        rule `parseGapsItems` applies to a `## Gaps`
                        markdown section, so one authored entry cannot
                        read closed in one reader and open in the other.
  `human_verification:` a bare `resolution:` still closes, since that is
                        how verifier-written entries record it — but a
                        readable `status:` that contradicts it wins.

A single unified rule was the first draft and is wrong: it closes a
frontmatter `gaps:` entry carrying `resolution:` and no `status:`, which
`parseGapsItems` surfaces, and `parseVerificationGapsItems`' own docstring
claims it mirrors that reader's fail-safe status handling.

The contradiction guard is not a judgment call about YAML. It is the rule
this codebase already applies to the same field pair: `validateResolution`
(probe-core.cts) rejects a populated `resolution:` on a non-resolved status
outright — "a populated payload is an authoring mistake ... Reject it so
the mistake surfaces." A reporter cannot throw, so it surfaces the item.

Minor 1. Direct unit tests for `frontmatterListEntries` and
`flattenObjectListItem` in `tests/frontmatter.unit.test.cjs`, the file that
historically co-changes with `frontmatter.cts`. They were reachable only
through `uat.cts`' readers before.

Minor 2. `parsedEntriesFor`'s degrade-to-all-null branch is asserted
directly. Verified unreachable through content rather than assumed: both
readers enter through `frontmatterRegion`, `extractFrontmatter`'s only
extra argument gates a warning, and `normalizeParsedValue`'s `value.map`
is 1:1. It is a drift alarm for a future edit to either parser, so the
helper is exported for tests rather than left as the one unpinned branch.

Minor 3. The vestigial `const skipResolved = true` and its dead
conditional are gone.

Minor 4. `frontmatterEntryToUatItem` no longer reads `test:`. A `gaps:`
entry has no `test:` in its vocabulary — the template's entries carry
truth/status/reason/artifacts/missing — so it was speculative support for
a field the shape does not have, and it collided with the 1..N row numbers
`parseHumanVerificationItems` assigns by array position. Not reading it
makes the collision impossible; an offset would have rewritten an authored
value, against `entryField`'s verbatim contract.

Docs, changeset and the dispatcher docstring all stated the unconditional
rule and are corrected — three prior rounds here were comment/code drift.

Fail-first proven: restoring the universal rule reddens all three new unit
tests and both rewritten properties.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-04 17:47:01 +00:00
Tom Boucher
585a8b7f1b fix(#3747): correct antigravity matrix evidence and pin the CLI-only skills install path (#4274)
* test(#3747): fail-first regression — matrix must not cite configHome skills path for antigravity

* fix(#3747): correct disproven antigravity stateIO evidence; pin CLI-only probe branch install path

* fix(#3747): scope doc evidence claim to skills discovery per adversarial review

* chore(#3747): add changeset

* chore(#3747): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-04 11:48:38 -04:00
Tom Boucher
1fe85cd43e chore(#4244): ESLint rules for the #4220 Windows dirname-walk / TMPDIR-triad bug class (#4246)
* fix(#4244): repoint TEMP/TMP alongside TMPDIR and fix the sweepProtectSet fixed-point walk

Repo-wide sweep (ahead of adding lint rules for these exact bug classes)
found both incident patterns still live and unfixed on `next`:

- scripts/run-tests.cjs's sweepProtectSet walk stopped on
  `cur !== runTempRoot && cur.length > 1` — a POSIX-only sentinel.
  win32 dirname('D:\') is a fixed point (length 3, never satisfies
  `> 1`... wait, it does satisfy length>1), so a selected file living
  outside runTempRoot (the common case) spins the walk forever on
  Windows. Extracted a pure, exported computeSweepProtectSet helper
  that terminates on dirname(cur) === cur instead, with in-process
  RuleTester-style coverage for both win32 and posix paths.

- tests/run-tests-temp-root.test.cjs's own #4020 regression test set
  only TMPDIR on its runNode(...) child env. Node's os.tmpdir() never
  reads TMPDIR on Windows (only TEMP, then TMP), so the redirect
  silently no-oped there — masked because Windows CI died in the
  dirname-walk hang above before ever reaching this test.

- tests/config-schema.property.test.cjs's fallow config-set test had
  the same TMPDIR-only pattern, direct process.env assignment this
  time, restored in its own finally block.

Origin: #4220 and its shared root cause #4020.

* feat(#4244): require-full-tmpdir-triad and no-unbounded-dirname-walk ESLint rules

Two custom local ESLint rules catch the #4220 / #4020 Windows CI hang bug
class at author time, joining the ADR-1703 DEFECT.WINDOWS-TEST-PORTABILITY
catalog. Neither eslint-plugin-unicorn nor eslint-plugin-n has a rule for
either shape.

- local/require-full-tmpdir-triad: flags a TMPDIR environment override
  (direct process.env.TMPDIR assignment, or a TMPDIR property in a
  spawn-like call's env: object literal) not accompanied by TEMP and TMP
  in the same scope. Node's os.tmpdir() never reads TMPDIR on Windows.
  Registered on tests/**/*.cjs, matching the require-userprofile-with-home
  precedent.

- local/no-unbounded-dirname-walk: flags a while/do-while loop reassigning
  from dirname() with no fixed-point termination guard
  (dirname(cur) !== cur, or path.parse(cur).root). path.dirname() is a
  no-op at the platform root, but the value differs by platform
  (win32 'D:\' is length 3, posix '/' is length 1), so a POSIX-shaped
  length/equality bound never fires on Windows. Registered on BOTH
  tests/**/*.cjs and scripts/**/*.cjs — the real #4020 bug lived in
  scripts/run-tests.cjs, not tests/.

Both rules join the zero-escape-hatch discipline already established for
this catalog (no bespoke comment marker; PROTECTED_RULES in
tests/portability-rule-disable-ban.test.cjs independently bans
eslint-disable of either). ADR-1703 and its two companion contributing
docs get an amendment documenting the mechanism, code examples, and the
repo-wide sweep (three live instances found and fixed in the prior
commit; no others found). CI test-scope selection updated so an edit to
either rule or to scripts/run-tests.cjs re-runs the right suites.

* fix(#4244): no-unbounded-dirname-walk must analyze a single-condition loop test too

checkWhile bailed out early unless node.test was a LogicalExpression,
so a single-condition loop -- while (cur !== root) { cur = dirname(cur); } --
was silently skipped and never reported. That is the EXACT minimal
shape of the original #4020/#4220 bug, and it is literally the shape
used by this rule's own shipped RuleTester fixtures (the "equality-only
bound" invalid cases), which were failing (0 errors reported, 1
expected) until this fix -- confirmed by running RuleTester directly
against both fixtures, not just via a passing test-runner exit code.

The conjunct-collection helper already handled a non-LogicalExpression
test correctly (it pushes a single node as the sole conjunct); only the
early-return gate needed to stop requiring a compound && / || test.

Verified: RuleTester run directly against both previously-broken
fixtures plus two new sanity cases (a guarded single-condition loop
stays valid; an unrelated single-condition loop stays silent), and a
fresh `npx eslint .` across the whole repo remains clean (no other
single-condition dirname-walk shape exists in the tree).

* fix(#4244): require-full-tmpdir-triad must recognize a destructured child_process call

isSpawnLikeCallee only recognized a MemberExpression callee
(child_process.spawnSync(...)) or a bare identifier in
ENV_LOCAL_HELPER_NAMES (runNode). A destructured import called bare --
const { spawnSync } = require('child_process'); spawnSync(...) -- has an
Identifier callee named "spawnSync", which matched neither branch, so
the whole env-literal check was skipped. gsd-test caught this: both
"invalid: child_process.spawnSync with TMPDIR-only env" cases in
tests/require-full-tmpdir-triad.rule.test.cjs were failing (0 errors
reported, 1 expected).

Widened the bare-identifier branch to also match any of the known
ENV_CHILD_PROCESS_METHODS names, matched by name only -- the same
lightweight convention this repo's other eslint-rules/*.cjs use (e.g.
no-hardcoded-tmp.cjs's isFsMethodCall), not full import data-flow
tracing.

Verified: RuleTester run directly against all 11 cases in
tests/require-full-tmpdir-triad.rule.test.cjs (not just the two that
were failing), all pass; a fresh npx eslint . and npm run lint:ci
across the whole repo remain clean.

* fix(#4244): correct a stale escape-hatch reference in a test comment

The comment on the "length comparison against another expression's
length" case referenced a "// allow-dirname-walk marker" that doesn't
exist -- the rule has zero comment-based escape hatches by design
(ADR-1703), and an earlier draft's marker mechanism was removed before
this branch's first commit. Spec-axis review caught the stale
reference. No behavior change; comment-only.

* chore(#4244): backfill changeset PR number (pr:0 -> pr:4246)

---------

Co-authored-by: sim <sim@local>
2026-09-03 14:14:09 -04:00
Tom Boucher
515191f07d feat(#3677): quick-batch hardening and acceptance (#4240)
* chore(#3677): checkpoint design artifacts (gitignored, dev-only)

* test(#3677): add failing regression test for the crash-window duplicate-dispatch gap (RED)

Independently re-traces resume-mode.md/planner-wave.md/worktree-dispatch.md/
merge-wave.md and src/quick-batch.cts's resumeBatch (lines 894-899) and
confirms the prior research pass's Open Question 1: a coordinator crash
between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge)
leaves BATCH.json at "pending" with no STATE.md row yet (only written in
Step 9), so --resume's eligibility re-derivation would dispatch a second
executor into a new worktree for the same item, orphaning the first.

This test asserts worktree-dispatch.md's Step 6 excludes an item whose
SUMMARY.md already exists from the spawn set, mirroring planner-wave.md's
existing PLAN.md-existence check one layer earlier. Fails against the
current worktree-dispatch.md, which has no such guard.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1
for the full trace and fix-location rationale.

* fix(#3677): guard worktree-dispatch.md against re-dispatching an already-executed item (GREEN)

worktree-dispatch.md's Step 6 re-derives eligibility every dispatch round
via the same quick-batch resume call resume-mode.md uses, but had no check
for "did this item already finish executing" the way planner-wave.md
already checks "did this item already get planned" (PLAN.md existence)
before re-planning. A coordinator crash between Step 6 (executor commits,
SUMMARY.md written) and Step 7 (merge) left the item eligible for a second
dispatch on --resume, orphaning the first worktree's real, already-
committed work and silently losing it once the second executor's SUMMARY.md
write clobbered the first at the same item_dir path.

Adds a SUMMARY.md-existence exclusion before spawn-plan is computed,
symmetric to planner-wave.md's PLAN.md check. The excluded item is not
lost: merge-wave.md's own mergeable-wave criterion (status=pending,
SUMMARY.md on disk, not yet merged) already picks it up independently of
this eligible/spawn list.

Workflow-prose-only fix — touches no already-merged/reviewed .cts module.
See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1
for the fix-location rationale (why not resumeBatch itself).

* test(#3677): add real-git coverage for worktree-ownership tampering, scope drift, and submodules

Closes the three coverage gaps identified in 40-design.md §2/§3 (#3677,
epic #3344 Phase 5's own AC bullets: "arbitrary-worktree ownership
attempts", "scope drift", "submodules"):

- Arbitrary-worktree ownership tampering: a manifest entry naming a
  non-agent branch is silently dropped at normalization before any git
  subprocess runs; a manifest entry naming a plausible agent-branch that
  was never actually created by this repo's own worktree.create (a
  genuinely foreign repo/branch) is blocked via base_mismatch. Both leave
  the foreign location and repoRoot's HEAD provably untouched.

- Advisory scope drift: a committed path outside declared files_modified
  still merges successfully (advisory, never blocking) while surfacing a
  scope_out_of_declared warning naming the drifted path; an exact
  declared-scope match produces zero warnings (boundary case).

- Real .gitmodules submodule integration: a repo containing a real local
  git submodule merges cleanly through executeWorktreeWaveCleanupPlan for
  an unrelated plan; a real gitlink pointer bump (declared) merges cleanly
  with the superproject tree reflecting the new pinned commit; an
  undeclared bump is advisory-only and surfaces a scope warning naming
  vendor/sub, same as any other undeclared modification.

No src/*.cts changes — all three gaps were coverage-only; the underlying
primitives already behaved correctly (independently verified against real
git subprocess output before writing each assertion).

* docs(#3677): document how to diagnose a preserved quick-batch worktree

Extends the one-sentence "worktree is preserved (never deleted)" mention
into a concrete diagnosis procedure: where the preserved directory is, how
to read the executor's real commits/diff against the plan's declared
files_modified, how to read the item's own SUMMARY.md independent of merge
outcome, how to manually merge-and-clean-up or discard, and how to re-run
--resume afterward. Also documents that a SUMMARY.md-written-but-still-
pending item (the crash-window case fixed in this same PR) needs no manual
intervention — --resume routes it straight to the merge step.

* chore(#3677): checkpoint final acceptance-evidence mapping (gitignored, dev-only)

* fix(#3677): make crash-window duplicate-dispatch guard behaviorally provable and durably recoverable

Orthogonal review (Spec finding): the crash-window regression test added
earlier this phase only asserted readStep('worktree-dispatch.md') + regex
matches against the markdown prose — proving the DOCUMENTATION says the
right thing, never that the runtime condition (pending status + on-disk
SUMMARY.md + absent STATE row) is actually handled correctly. #3677's own
"Alternatives considered" explicitly rejects "document recovery without
fault injection" for exactly this reason.

Extracts the filtering decision into a pure, independently testable
function, filterAlreadyExecuted(eligibleIds, executedIds) in
src/quick-batch-dispatch.cts, wired to a new `quick-batch filter-executed`
CLI verb (src/quick-batch-command-router.cts) — the same pure-decision-
then-CLI-wired pattern computeSpawnPlan/computeMergeOrder already
establish. worktree-dispatch.md now calls this verb explicitly instead of
only describing the decision in prose. A genuine fixture-based test in
tests/quick-batch.test.cjs constructs a REAL BATCH.json (createBatch),
writes a REAL SUMMARY.md on disk at the item's real item_dir, calls the
REAL resumeBatch, and proves both that resumeBatch alone still reports the
item eligible AND that filterAlreadyExecuted (fed a real filesystem check)
correctly excludes it. The prior prose-assertion tests are kept — they now
prove the workflow markdown is correctly WIRED to the verb — but are no
longer the only proof.

Self-discovered defect while building that fixture (fixed inline, not
deferred): tracing merge-wave.md against /gsd:quick's own prior art
(QUICK_WORKTREE_MANIFEST=$(mktemp ...), quick.md:415) showed
$QUICK_BATCH_WORKTREE_MANIFEST is a fresh PER-PROCESS temp file. A resumed
coordinator correctly does not re-dispatch an already-executed item (this
fix), but nothing durably recorded that item's worktree_path/branch/base
either — Step 7 in the resumed process would have had no data to build its
cleanup-wave entry from. Adds dispatched_worktree/dispatched_branch/
dispatched_base to QuickBatchItem (src/quick-batch.cts) — deliberately NOT
a reuse of the pre-existing `worktree` field, whose loadBatch validation
requires the path to exist on disk (verified empirically: reusing it made
the batch permanently unloadable the moment a legitimately-merged worktree
was removed). worktree-dispatch.md persists the triple once a worktree is
created; merge-wave.md falls back to it when the ephemeral manifest lacks
an entry, clears it after a successful merge, and fails closed rather than
guessing if no record exists anywhere.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.1
and §9.3 for the full trace, empirical verification notes, and rejected
alternatives (reusing `worktree` directly).

* test(#3677): prove the arbitrary-worktree-ownership boundary against two real sibling worktrees

Orthogonal review (Security finding): the two existing ownership-tampering
tests didn't test ownership — one was trivially rejected by
WORKTREE_AGENT_BRANCH_RE's shape check before any git call (proves branch-
NAME filtering, not ownership), the other pointed at a wholly separate,
never-linked foreign repo, so merge-base failed immediately because the
branch didn't exist as a ref at all. Neither exercised the real scenario:
a manifest entry whose worktree_path/branch are swapped to point at a
DIFFERENT, GENUINELY-REGISTERED sibling worktree of the SAME repoRoot,
with a branch name passing the shape check and a base in allowed_bases.

Investigated executeWorktreeWaveCleanupPlan (src/worktree-safety.cts)
directly: this is NOT a reachable gap. Git enforces branch-per-worktree
uniqueness, so a swapped-in entry.branch can only match worktree_path's
ACTUAL checked-out branch if it names that sibling's own real, uniquely-
generated branch name — which manifest tampering confined to one batch's
own record has no way to know (branch names are
agent-<quick_id>[-<timestamp>]-shaped, and quick_id allocation is
collision-checked GLOBALLY across every existing quick task and batch, not
merely within one batch).

Adds a stronger test that empirically proves this: two REAL, concurrently-
alive sibling worktrees of the same repo (both via real `git worktree add`,
both WORKTREE_AGENT_BRANCH_RE-passing, both sharing one merge-base), with
worktree_path/branch swapped between them in both directions. Both attempts
are blocked via branch_mismatch; both real worktrees, their branches, and
one sibling's real uncommitted-to-main commit survive completely untouched.
Supplements (does not replace) the original two tests, which still prove
distinct, real boundaries.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.2
for the full trace, including the one explicitly-documented (not fixed)
trust boundary this investigation surfaced: the primitive defends against
fabricated data, not a caller bug that misattributes a real-but-wrong
item's own triple to a different item.

* chore(#3677): checkpoint design-doc addendum for review pass 2 findings (gitignored, dev-only)

* docs(#3677): add changeset for PR 4240

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-03 09:47:22 -04:00
Tom Boucher
2f64e6230a feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core

Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch
binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the
new pure decision-logic module (arg validation, effective concurrency,
deterministic merge order, spawn backpressure, verification/merge
routing, cleanup-entry construction — design doc rows 3-15,24,26-28,
30-36,39; property rows 51-53). quick-batch-update-items.test.cjs
covers the new updateBatchItems export on src/quick-batch.cts (rows
15,22-23, including the negative cycle-rejection case).
quick-batch-command-router.test.cjs covers the new
gsd-tools quick-batch CLI family (rows 46-47). These reference modules/
exports that do not exist yet.

* feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router

Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision
layer — CLI verbs and pure orchestration logic only; no workflow
markdown, no Agent()/git-worktree I/O.

- src/quick-batch-dispatch.cts (new): pure decision functions consumed
  by the (separate, follow-up) /gsd:quick-batch workflow markdown —
  parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder,
  computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome,
  buildCleanupManifestEntry (the last parses caller-supplied plan text
  via the existing parsePlanDocument; no filesystem access).

- src/quick-batch.cts: adds updateBatchItems, resolving the design
  doc's Open Question 1 as ONE additive export on this module instead
  of the second, independent BATCH.json writer the design doc
  originally proposed. Reuses the same withPlanningLock transaction
  shape, computeWaves, and platformWriteSync call resumeBatch/
  completeQuickItem already use; fails closed without persisting on
  an unknown item, an unknown/self dependency, or an introduced cycle.

- src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI
  family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs)
  as a first-party always-on command (like /gsd:quick), not the opt-in
  capability-registry path graphify uses. Verbs: create/update/resume/
  complete (wrap quick-batch.cts) and effective-concurrency/
  merge-eligible/spawn-plan/verification-routing/merge-routing/
  cleanup-entry/parse-args (wrap quick-batch-dispatch.cts).

Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property
rows 51-53. Rows covering workflow markdown / Agent() dispatch /
`git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the
follow-up markdown-authoring pass, per the phase brief's explicit
scope boundary.

* docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces

New-.cts-module ripple for the two Phase 4 modules (epic #3344,
ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts,
ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source,
not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json
(via `node scripts/gen-inventory-manifest.cjs --write`, after
`npm run build:lib`), and CONTEXT.md glossary entries for
"Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router
Module", plus an update to the existing "Quick-Batch Core Primitives
Module" entry documenting the new updateBatchItems export.

* test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count)

scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs
file under the quick-batch production module by longest-prefix match,
and that module is already at its 2-file cap (quick-batch.test.cjs +
quick-batch.property.test.cjs). The standalone
tests/quick-batch-update-items.test.cjs added in the prior commit
pushed it to 3 and failed `npm run lint:ci`. Fold its content into
quick-batch.test.cjs (append-only — no existing test in that file is
modified) and update the CONTEXT.md glossary reference to match.

Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after
`npm ci` (this worktree previously had no local node_modules, which
also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-*
unable to resolve typescript — resolved by npm ci, no code change
needed there). `npm run lint:ci` and
`npx tsc -p tsconfig.build.json --noEmit` are both green after this
fix.

* test(#3676): add failing tests for the quick-batch command/workflow markdown

Failing-first tests for Phase 4's markdown-authoring pass (epic #3344,
ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs
covers commands/gsd/quick-batch.md's frontmatter/objective/process,
gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR
1610 NEW_FILE_CAP) and step-fragment count, the isolation model
(rows 20-22), the executor single-writer invariant (row 18), merge
validation reusing the existing bounded primitive (row 25), the
optional research/plan-checker/verification leaves (rows 16,17,19,
30,31), planning-failure blocking execution (row 29), the submodule
guard (rows 36,44), and the new agents/gsd-planner.md quick-batch
mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers
row 48 (ordinary /gsd:quick stays byte-identical). Named
`gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's
longest-prefix bucketing doesn't fold these markdown-only tests into
the already-capped quick-batch/quick-batch-dispatch/
quick-batch-command-router production-module buckets from the CORE
pass. These reference files that do not exist yet.

* feat(#3676): author the quick-batch command, workflow, and planner mode

Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch
binding") — the orchestration layer that calls into Pass 1's CLI
verbs (src/quick-batch-command-router.cts).

- commands/gsd/quick-batch.md (new): frontmatter/objective/process,
  delegates argument validation to `quick-batch parse-args`
  (parseQuickBatchArgs) rather than re-deriving the grammar.

- gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR
  1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded
  step fragments under gsd-core/workflows/quick-batch/steps/:
  resume-mode, batch-init, research-phase (flag:--research),
  planner-wave (+ nested plan-checker-loop when --validate),
  worktree-dispatch, merge-wave, verification-wave (flag:--validate),
  completion. Covers design doc rows 3-45: capacity/isolation
  resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG-
  layer planning with full-task-catalog prompts and always-required
  depends_on/files_modified frontmatter, serialized worktree create/
  merge/cleanup via the existing worktree.cleanup-wave primitive,
  deterministic wave-order merging, verification routing
  (human_needed/gaps_found), the executor single-writer invariant,
  submodule fail-loud guard, and #1941 fork-base auto-degrade.

- agents/gsd-planner.md: additive new `load_mode_context` bullet for
  `**Mode:** quick-batch`, pointing at the new
  gsd-core/references/planner-quick-batch.md reference (documents the
  always-required depends_on/files_modified contract, reusing the
  existing frontmatter grammar — no new keys). Existing modes
  byte-identical, only a new bullet added.

- src/init.cts (+init-command-router.cts, +command-aliases.cts):
  cmdInitQuickBatch / `init.quick-batch` — model profiles,
  commit_docs, roadmap/planning existence checks, and the
  section_manifest field gating research-phase/verification-wave
  (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY
  atoms — no new atom needed).

Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the
prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43,
45 covered by construction (verb wiring, single-writer prompt
constraints, crash-window resume via unmodified Phase 3 primitives).

* docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern

npm run regen:derived output for the new command/workflow/reference
(epic #3344, ADR-1239 "Quick-batch binding"):
- skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md)
- docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md,
  planner-quick-batch.md, and the quick-batch-dispatch.cjs/
  quick-batch-command-router.cjs CLI-module rows' now-live
  `/gsd-quick-batch` cross-reference (was "(separate, follow-up)")
  + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`)
- gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`)
  — research-phase/verification-wave gsd:section entries for the new
  quick-batch workflow
- tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) —
  the new command/workflow/skill/reference files now ship to every
  runtime

scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for
gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS
word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional-
flag pattern gsd-core/workflows/quick.md already carries baselined
(e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2);
quoting would break the intended "omit this arg when the flag is
false" splitting.

* fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch

Security review pass findings, both confirmed real:

1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment
   interpolated the raw, attacker-influenced task ${description} (and
   the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw
   description into every planner's prompt in the layer) straight into
   Agent() prompt bodies with no boundary. Fixed by wrapping every such
   interpolation in a <security_context> + DATA_START/DATA_END
   boundary, matching the CONCRETE convention already implemented in
   this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md,
   gsd-core/workflows/debug.md) — commands/gsd/quick.md's own
   <security_notes> only asserts this convention in prose, so the
   debug-agent files are the real precedent followed here. Added a new
   <security_notes> block to commands/gsd/quick-batch.md (it had none)
   documenting both this fix and the one below.

2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection.
   gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both
   ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED,
   causing shell word-splitting and pathname expansion on raw task-list
   text before the parser ever saw it. Fixed at the source: added a
   `--text <string>` form to the `parse-args` verb
   (src/quick-batch-command-router.cts) that accepts the ENTIRE
   $ARGUMENTS as ONE quoted argv element and does the whitespace split
   itself, in Node — which is never glob-aware, unlike the shell.
   Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form
   is kept for direct/test callers that already have a real argv array.

The SC2086 baseline entry added for the original unquoted line is now
stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports
it) and has been removed; the two SC2046 entries for the UNRELATED,
still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style
conditional-flag splitting remain — that line only ever expands to one
of a few known-safe literal strings (never raw user text), matching
quick.md's own already-baselined convention exactly.

Tests: quick-batch-command-router.test.cjs covers the new --text form
(token splitting, glob-shaped text passing through literally
unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs
asserts the DATA_START/DATA_END boundary on every leaf prompt
(research-phase/planner-wave/plan-checker-loop/verification-wave,
including the shared task catalog) and the quoted --text call sites.

* fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35

Spec review pass findings — the test matrix claimed "yes" coverage
these assertions did not actually support:

- Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection
  only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs,
  committed alongside the security fix that touches the same file) that
  .planning/quick-batches/ is never created for any rejected value —
  createBatch is genuinely never reached.
- Row 18 (--resume <unknown-batch-id>): previously only exercised a
  hand-corrupted BATCH.json, never a genuinely nonexistent batch
  directory. Added the real nonexistent-id case (also in
  quick-batch-command-router.test.cjs).
- Row 24 (post-planning updateBatchItems racing a concurrent
  completeQuickItem for a different item, both through
  withPlanningLock): zero test existed. Added a property test
  (tests/quick-batch.property.test.cjs, appended — Phase 3's own file,
  no existing test touched) exercising both call orders and asserting
  no lost update in the final on-disk manifest — the same technique
  Phase 3's own row-15 lock-contention property test uses (sequential
  calls through the real lock; a working mutex makes any interleaving
  equivalent to some serial order, so this is the same claim a literal
  concurrent-thread test would make without OS-level threading).
- Row 34 (worktree preserved on merge_failed) and row 35 (undeclared-
  deletion detection): both were previously asserted only at the pure
  routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs
  using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs
  already establishes for executeWorktreeWaveCleanupPlan (real repo,
  real worktree, a REAL merge conflict / a REAL file deletion diffed
  against declared_deletions) — asserting the actual worktree directory
  survives on disk, not just that a pure function returns a
  preserveWorktree:true field. Named gsd-quick-batch-* so lint-test-
  file-count's bucketing doesn't fold it into any capped module bucket.

Row 48 (/gsd:quick regression) intentionally left as-is per the
reviewer's own framing: the byte-identity claim is already
mechanically proven by the changed-path diff (git diff --name-only
empty on those two paths IS byte-identity), and a genuine execution-
level regression test would require actually running the workflow —
out of scope for this repo's unit-test model (no other quick.md
regression test in this repo does that either).

* docs(#3676): add the changeset and user-facing docs the command needed

Standards review pass findings — both HARD:

- Missing changeset. None of the 6 prior #3676 commits touched
  .changeset/*. /gsd-quick-batch is a new user-facing command;
  CLAUDE.md/CONTRIBUTING.md require one. Added
  .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder —
  backfilled after the PR opens, matching CLAUDE.md's own documented
  convention and Phase 3's own precedent, #4190's
  .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen
  form `/gsd-quick-batch` throughout, never the source-artifact colon
  form (`scripts/lint-docs-command-form.cjs` confirms 0 violations;
  that check scans docs/**, not .changeset/, so it was never actually
  in scope for the fragment itself, but the wording still follows the
  doc convention for consistency, matching how Phase 3's own fragment
  named the not-yet-shipped command).
- Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis
  how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's
  existing convention for /gsd-quick /gsd-fast) covering --jobs,
  --validate, --research, --resume, --file, the capacity/isolation
  interaction, and resume/failure recovery. Cross-linked from
  docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's
  own "Related" section. Added a /gsd-quick-batch section to
  docs/COMMANDS.md (same table format as the existing /gsd-quick
  entry) and docs/features/quick-batch.md (REQ-QB-01..12, same
  frontmatter shape as docs/features/quick-mode.md) — regenerated
  docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md
  via the standard generators.

* fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught

gsd-test's real run against 155e8975b3 found 43 failures, all rooted in
this phase's own new command/workflow never being registered across
~10 independent generated/hand-maintained registries this repo keeps
in parity by convention. Root-caused each, no test weakened or
special-cased.

- help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live-
  registry.test.cjs): added a /gsd:quick-batch entry to
  gsd-core/workflows/help/modes/full.md (the real help.md content;
  gsd-core/workflows/help.md is a thin dispatcher) documenting every
  flag (--file/--jobs/--validate/--research/--resume), matching the
  existing /gsd:quick entry's format.

- gen-section-manifest.test.cjs: quick-batch.md's
  `gsd_run query init.quick-batch` invocation used inline
  `$([ ... ] && echo --flag)` substitutions, which never satisfy the
  test's exact-whitespace-token / assigned-variable detection (the
  trailing `))` glued onto `--research` in the compound substitution
  broke the "exact token" match). Rewrote to the same
  VALIDATE_PARAM/RESEARCH_PARAM two-line pattern
  gsd-core/workflows/quick.md's own Step 2 already uses.

- runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md
  fragments that call gsd_run each needed their OWN embedded copy of
  the canonical shim preamble (every workflow .md that calls gsd_run
  carries its own copy — reading one file does not persist shell state
  into another). Ran `node scripts/sync-runtime-launcher.cjs`, which
  inserted it before each file's first gsd_run call.
  plan-checker-loop.md correctly has none — it never calls gsd_run
  directly.

- Namespace routing (skill-manifest.test.cjs, install-nested-
  layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added
  `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and
  routing table (same namespace `quick` already routes through), and
  to src/clusters.cts's `utility` cluster (same cluster `quick`
  already belongs to). Verified by hand-running installRuntimeArtifacts
  + applySurface for augment/cline against a real temp install: exactly
  6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested
  under gsd-ns-workflow/skills/, never re-flattened.

- mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a
  brand-new command is a real count change, not a bug this test should
  hide).

- model-omit-when-inherit-guard.test.cjs: added the canonical
  `<!-- #2517 model-omit-on-inherit -->` marker block to
  gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/
  researcher/checker/executor/verifier — lives in a steps/ fragment,
  read combined with the host by this test's own readWorkflowCombined,
  same as quick.md's own research-phase.md carries it for its gated
  section). Also fixed a genuine pre-existing inconsistency in the
  test's own "#2711: the guarded set is derived from dispatch sites"
  check: its `nonDispatching` computation read the BARE host file while
  `derived` (the set it's checked against) reads the combined
  host+steps content — inconsistent with that same test file's own
  #2994 doc comment explaining why the combined read is necessary.
  quick-batch.md is the first workflow whose EVERY model="{...}"
  dispatch site lives in a mandatory (never gated) steps/ fragment —
  extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a
  brand-new file — which is what exposed the mismatch. Fixed by using
  the same readWorkflowCombined read in both places.

- skill-frontmatter-contract.test.cjs: shortened
  commands/gsd/quick-batch.md's frontmatter `description` from 107 to
  91 chars (<=100 budget), and added `quick-batch.md` to the hand-
  maintained KNOWN_SKILLS consolidation allowlist with a #3676
  justification comment (a genuinely new first-party command, not a
  consolidation of an existing skill).

- workflow-fragments-emission.install.test.cjs: added `quick-batch.md`
  to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is
  deliberately NOT a no-op for it — its research-phase/verification-
  wave sections are gated).

- Regenerated all downstream artifacts (npm run build:lib && npm run
  regen:derived && npm run gen:plugin-skills -- --write && npm run
  gen:features -- --write): skills/gsd-quick-batch/SKILL.md,
  skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for
  augment/cline/hermes/qwen/trae/zcode.

- emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition
  (one new `load_mode_context` bullet pointing at the new
  gsd-core/references/planner-quick-batch.md reference) grew the file
  124 bytes without an acknowledgment trailer. Acknowledged below —
  the growth is the deliberate, additive, single-bullet change from
  the earlier feat(#3676) commit, not drift.

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green
(includes lint-workflow-shellcheck, lint-test-file-count,
lint-docs-command-form). The deep install/spawn/registry tests gsd-test
actually runs (docs-parity-live-registry, gen-section-manifest,
runtime-launcher-parity, install-nested-layout,
runtime-artifact-layout-surface, skill-manifest, skill-frontmatter-
contract, mcp-server-catalog, model-omit-when-inherit-guard,
workflow-fragments-emission) are not part of lint:ci — each fix above
was independently verified by hand-invoking the exact production
function the failing test calls (installRuntimeArtifacts, applySurface,
composeWorkflow, the CLUSTERS union, the section-manifest forwarding
regex) against the real repo tree and confirming the expected shape.

Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift.

* fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget

skill-frontmatter-contract.test.cjs's "feature #3039: tiered help —
size budgets" enforces a SEPARATE line-count ceiling for
gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines,
tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's
assertTightCeiling) — independent of the skill-frontmatter description-
length budget and consolidation allowlist I touched in the prior round;
those are unrelated checks in the same test FILE, not the same check.

Root cause: the /gsd:quick-batch entry I added to full.md in the
docs-parity fix round was 17 lines, pushing the file from 834 to 851
lines — 7 over the 844 ceiling. Condensed the entry (merged the
per-flag bullet list into one dense "Flags:" line, dropped from 3
Usage examples to 1) to 844 lines exactly — at the ceiling with zero
slack, which assertTightCeiling accepts (it only fails on
actualMax > ceiling, or on slack > grace when the ceiling is too
LOOSE — zero slack triggers neither).

Verified after trimming: full.md still contains a live /gsd:quick-batch
reference (bidirectional parity) and all 5 argument-hint flags
(--jobs/--validate/--research/--resume/--file) still appear as literal
tokens (docs-parity-live-registry.test.cjs's own flag-coverage check,
re-run by hand against the trimmed content).

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green.

* docs(#3676): backfill changeset pr number to 4212

Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's
pr:0 placeholder backfilled with the real PR number now that
gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR
Number Handling convention and Phase 3's own #4190 precedent
(708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt
from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE.

* fix(#3676): resolve prompt-injection-scan false positive on test fixture

tests/quick-batch.test.cjs:232's row 11b regression proves the task-list
parser carries a prompt-injection-shaped task description through
createBatch as inert data, never interpreted. The fixture has to be a
real "ignore all previous instructions..." phrase or the test asserts
nothing, but the full-file --diff scan flagged it once unrelated edits
in the same file pulled it into the changed-file set.

Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the
sanctioned, precedented exemption already used for other legitimate
security-regression fixtures (tests/windsurf-conversion.test.cjs,
tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs)
per DEFECT.PROMPT-INJECTION-SCAN-COLLISION.

---------

Co-authored-by: sim <sim@local>
2026-09-02 22:38:31 -04:00
Tom Boucher
91ed46882a feat(#3675): quick-batch core primitives and resumable manifest (#4190)
* test(#3675): add failing tests for quick-batch core primitives

Adds the full behavioral (tests/quick-batch.test.cjs) and property-based
(tests/quick-batch.property.test.cjs) coverage for #3675's quick-batch core
primitives per the phase's 35-row test matrix — task-list parsing (inline +
--file, with path-confinement/symlink-escape/non-regular-file rejection),
collision-safe quick-id preallocation under withPlanningLock, BATCH.json
schema/validation/resume, dependency-DAG + partitionByFileOverlap wave
construction, and exactly-once STATE.md completion (including the
STATE-row-written-but-manifest-not-yet-updated crash window).

The import target (gsd-core/bin/lib/quick-batch.cjs, compiled from a
not-yet-written src/quick-batch.cts) does not exist yet — every test in both
files fails at the top-level require() before any assertion runs. Five
fast-check properties cover collision-freedom under lock contention, resume
idempotency, exactly-once STATE completion, wave totality, and DAG-respecting
wave order, per the design doc's property-based-coverage requirement.

* feat(#3675): implement quick-batch core primitives

Adds src/quick-batch.cts (ADR-457 build-at-publish, compiled to
gsd-core/bin/lib/quick-batch.cjs) implementing #3675's quick-batch core
primitives per the phase design lock — pure/state primitives and
CLI-testable core operations only, no agent dispatch, no worktree creation,
no user-facing command (Phase 4/#3676's job):

- parseTaskList / parseTaskListFromFile: inline bulleted/numbered task-list
  parsing (>=2 items required) and a --file variant strictly confined to the
  planning workspace root via requireSafePath, rejecting non-regular-file
  targets.
- allocateQuickIds / createBatch: collision-safe YYMMDD-xxx quick-id
  preallocation under withPlanningLock, checked against both on-disk
  .planning/quick/ entries and sibling .planning/quick-batches/*/BATCH.json
  manifests (never on-disk-only, which would miss another in-flight batch
  that hasn't dispatched any real quick directory yet) — replicates
  cmdInitQuick's own grammar rather than delegating to it (that function's
  2-second granularity is not batch-safe).
- computeWaves: deterministic wave construction combining dependency-DAG
  layering with partitionByFileOverlap (#3674), called per DAG layer over
  path-separator-normalized planned_files — normalization happens at this
  module's boundary, never inside the Phase 2 helper.
- loadBatch: fail-closed BATCH.json schema validation (corrupt/truncated
  JSON, wrong types, missing fields, out-of-batch dependency references,
  dependency cycles, a worktree path absent from disk).
- resumeBatch: skips complete items, never auto-retries failed items,
  propagates/reverses blocked status along the DAG to a fixed point, and
  detects a STATE.md row that already exists for a non-complete item (the
  "STATE written, BATCH.json not yet updated" crash window) — completing it
  without re-appending. Idempotent across repeated calls.
- completeQuickItem / hasQuickTaskRow: exactly-once STATE.md completion —
  appendQuickTaskRow (unmodified) is called at most once per quick id, gated
  by hasQuickTaskRow's own idempotency check re-parsing the real "Quick Tasks
  Completed" table, since appendQuickTaskRow itself carries no idempotency.

BATCH.json lives at .planning/quick-batches/<batch-id>/BATCH.json, a sibling
of .planning/quick/ — never inside it, so scanQuickTasks never misreads a
batch manifest as a broken quick task.

* docs(#3675): register the new quick-batch module

New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained
registrations beyond the code itself: .gitignore (compiled artifact),
eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs),
docs/INVENTORY.md's CLI Modules roster row plus the regenerated
docs/INVENTORY-MANIFEST.json cli_modules entry, and a CONTEXT.md glossary
entry matching the convention set by the sibling File Overlap Partitioner
Module (#3674) entry it sits beside.

NOTE: docs/INVENTORY-MANIFEST.json was updated BY HAND (alphabetically
sorted single-entry insertion into families.cli_modules, matching the
existing file's structure) rather than via
`node scripts/gen-inventory-manifest.cjs --write` — this session's
MEMTRACE-FIRST guard hard-blocks direct execution of that indexed script
path from Bash, with no available Memtrace tool to route through instead.
The orchestrator should re-run `node scripts/gen-inventory-manifest.cjs
--check` to confirm this hand-edit is byte-identical to the generator's own
output before merging.

* fix(#3675): resolve lint findings in quick-batch primitives and tests

Unsafe `any[]` assignment from `new Array(n)` in the DAG cycle-check color
array, two unnecessary `as string[]` casts TS 5.5's inferred type
predicates already narrowed, raw `fs.rmSync` in test cleanup (needs the
Windows-EBUSY retry budget `helpers.cleanup` carries), an unused `loadBatch`
import, an unbounded `mkfifo` subprocess spawn missing a timeout, and a
CONTEXT.md glossary illustration that looked like a real file reference.

* feat(#3675): close acceptance-criteria gaps found in review

Standards- and spec-axis review (plus a self-caught race) surfaced real
gaps against issue #3675's own acceptance criteria and this repo's test
conventions:

- BATCH.json was missing options, base_revision, per-item wave, and
  per-item commit — the issue's AC explicitly lists all four as things
  the manifest must track. Added them: createBatch persists caller-supplied
  batchOptions/baseRevision verbatim and assigns each item its computed
  wave index; completeQuickItem now persists the commit onto the item,
  not just the STATE.md row. All four are backward-tolerant on load (an
  older/hand-built manifest without them still validates).
- resumeBatch had no "incompatible base divergence" check at all, despite
  the AC and the ADR's own "Base divergence" section requiring one. Added
  an opt-in currentBaseRevision comparison that fails closed with a
  recoverable diagnostic on mismatch, and touches nothing on refusal.
- resumeBatch read-modify-wrote BATCH.json OUTSIDE withPlanningLock — the
  only durable write path in this module that wasn't lock-protected,
  a real lost-update race against a concurrent completeQuickItem or
  another resume. Now runs inside the same lock createBatch/
  completeQuickItem use.
- loadBatch and collectExistingBatchQuickIds used raw JSON.parse with no
  size cap (security review, Low/informational); switched to the
  existing safeJsonParse (1MB cap) for defense-in-depth.
- Parser (parseTaskList) had only example-based tests; CLAUDE.md requires
  a fast-check property test for parsers. Added one plus a companion
  reject-property for <2 items.
- The id-exhaustion fail-closed ceiling (MAX_TIME_BLOCK) was untested at
  any boundary. Exported the pure allocateIdsGivenUsed/MAX_TIME_BLOCK for
  direct limit-1/limit/limit+1 testing without needing 46k fixture dirs.
- Issue AC explicitly asks for prompt-injection-payload test coverage,
  distinct from the existing shell-metacharacter test; added one.
- Test row 9 (FIFO skip) silently returned instead of calling t.skip(),
  so an unsupported platform would report a pass rather than a documented
  skip; fixed to bind the test-context param and skip properly.
- Extracted toWaveInput to remove a 2-site production duplication of the
  QuickBatchItem -> computeWaves reshape (Standards-axis smell).
- Added the required .changeset/ fragment (CONTRIBUTING.md: editing src/
  is user-facing even though the compiled .cjs is gitignored).

* fix(#3675): restore "not valid JSON" wording in loadBatch's parse-failure reason

gsd-test caught this: switching loadBatch to safeJsonParse changed the parse-
failure message shape ("... parse error — ...") without preserving the
"not valid JSON" substring row 27's own test asserts on. Re-wrap
safeJsonParse's error into the original diagnostic phrasing regardless of
which of its three failure modes fired.

* docs(#3675): backfill changeset pr number to 4190

* fix(#3675): detect a silently-no-op mkfifo on Windows, not just a throwing one

CI caught this on windows-latest: row 9's platform-skip only caught mkfifo
throwing (command not found). On this runner mkfifo resolves to something
that exits 0 without creating a file (NTFS has no FIFO concept), so
execution fell through to parseTaskListFromFile against a path that
doesn't exist, producing an ENOENT stat error instead of the expected
"not a regular file" rejection. Check the artifact actually exists before
trusting a zero exit code, and skip with a documented reason either way.

---------

Co-authored-by: sim <sim@local>
2026-09-02 15:38:27 -04:00
Tom Boucher
acb903c2e8 enhance(#3661): make the code-review hook point configurable (#4159)
* feat(#3661): make the code-review hook point configurable

Add `workflow.code_review_point` (`execute:post` default, or
`execute:wave:post`) so a multi-wave phase can run code review once per
wave instead of once at the end, scoped to what changed since the phase's
prior review.

The code-review capability now declares its step at both loop points via a
new generic `pointFrom` step field: `pointFrom` names an enum config key,
and the step is only active at its own `point` when that key resolves to
a matching value. `_resolvePointGate` (capability-activation.cts) is the
single shared implementation consumed identically by loop-resolver.cts and
capability-state.cts, and capability-validator.cjs enforces that `pointFrom`
references an enum key whose values cover the declaring step's own point.

code-review.md's manual-invocation gate now reads `workflow.code_review`
directly instead of probing registry presence at the hardcoded execute:post
point (so manual `/gsd-code-review` keeps working regardless of which
automatic point is configured), and its file-scope tiers narrow to what
changed since the phase's last review commit when one exists.

execute-phase.md's wave-post step dispatch gets a small, precedented
carve-out so the code-review skill still receives its required phase
argument when dispatched generically (caught by the isolated spec review).

Closes #3661

Emitted-Drift-Ack-Growth: code-review.md — #3661 adds a point-aware config gate check and LAST_REVIEW_COMMIT-based incremental scoping to the file-scope tiers.
Emitted-Drift-Ack-Growth: execute-phase.md — #3661 adds one carve-out sentence so the wave-post generic step dispatch passes PHASE_NUMBER to the code-review skill.

* docs: backfill changeset PR number for #3661 (#4159)

* fix: scope tests/io.test.cjs's fs.writeSync fault-injection mocks by fd

Five fault-injection mocks in the "bug #1008" describe blocks intercepted
every fs.writeSync call regardless of file descriptor, and several threw or
truncated unconditionally on the first call. This surfaced as an
intermittent macOS CI failure: node:test's own IPC channel back to the
parent process (which also goes through fs.writeSync internally) could get
a bogus injected error or truncated write if node's internal machinery
called it while one of these mocks was active, corrupting the message
frame the parent tried to deserialize ("Unable to deserialize cloned
data.", location tests/io.test.cjs:1:1, uncaughtException — a whole-file
IPC crash, not a test assertion failure).

Root cause confirmed by a working counter-example already in the same
file: the "#3912 A6" mocks gate on `fd !== 2` before any fault injection
and were never implicated. Applied the same fd-scoped pattern to the five
unscoped mocks (four output()-targeting tests gate on fd 1, one
error()-targeting test gates on fd 2), and added a regression test proving
an unrelated fd passes through untouched while the fault-injection mock is
active.

Found while verifying #3661; unrelated to that change's own diff.

---------

Co-authored-by: sim <sim@local>
2026-09-02 11:01:54 -04:00
Tom Boucher
77dcdda534 enhance(#4014): an unreadable directory must not report as an empty one (#4163)
* test(#4014): add failing-first coverage for unreadable-vs-empty directory scope (epic #3473 B4)

* fix(#4014): an unreadable directory must not report as an empty one (epic #3473 B4)

* test(#4014): update hardcoded generateSlugInternal closing-brace line after import shift

src/core-utils.cts's new #4014 import block shifted every subsequent line by
6, moving generateSlugInternal's real closing brace from line 193 to 199.
tests/slug-derivation-drift-guard.test.cjs's MAJOR-1 fixture hardcodes that
line number to plant a synthetic violation immediately after the function's
real body; the guard script itself locates the boundary dynamically via
brace-matching and needed no change.

* docs(#4014): document the unreadable-directory scope signal and add changeset

* docs(#4014): backfill changeset PR number to #4163

* test(#4014): kill pre-existing core-utils.cjs mutation-score gap, unrelated to this issue's diff

---------

Co-authored-by: sim <sim@local>
2026-09-02 08:12:57 -04:00
Tom Boucher
b0572c0108 feat(#3674): extract shared file-overlap wave partitioner (#4166)
* test(#3674): characterize existing wave-dispatch output and add tests for the extracted partitioner

Pins resolveWaveDispatch's and emitWorkflowScript's current, unextracted
output (chain-overlap, disjoint-empty-set, and a multi-wave/multi-stage
golden script) as a regression safety net ahead of extracting
partitionStages into a standalone module. Also adds the new module's
unit and property tests (test matrix rows 1-11) against its expected
public API, which does not exist yet and is added in the next commit.

* feat(#3674): extract file-overlap partitioner into a shared, generic module

Moves partitionStages' greedy first-fit file-overlap algorithm into a new,
dependency-free src/file-overlap-partitioner.cts module (partitionByFileOverlap),
generalized over a plain {id, files}[] shape rather than claude-orchestration.cts's
Plan/Wave interfaces. partitionStages becomes a thin adapter mapping its own
Plan[] shape onto the generic input and back — behavior-preserving, no dependency
ordering, no path normalization, no filesystem access moved or added. Enables a
future consumer (quick-batch, #3675 / ADR-1239) to reuse the same primitive
without pulling in orchestration internals.

* docs(#3674): register the file-overlap-partitioner module bookkeeping

New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained
registrations beyond the code itself: .gitignore (compiled artifact),
eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs),
docs/INVENTORY.md's CLI Modules roster row (regenerated via
gen-inventory-manifest.cjs --write), and a CONTEXT.md glossary entry
matching the convention set by similarly-scoped leaf modules
(text-lines.cts, plan-dependency-graph.cts, spec-section.cts).

* fix(#3674): alphabetize INVENTORY.md row, manifest regen no-op, fast-check import already correct

- docs/INVENTORY.md: move file-overlap-partitioner.cjs row to alphabetical position
- docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest.cjs --write,
  produced no diff (manifest is keyed by content, not row order)
- tests/claude-orchestration.test.cjs's direct require('fast-check') is correct as-is:
  tests/helpers/fast-check-setup.cjs's own docstring scopes the shared-seed wrapper to
  "every *.property.test.cjs file"; claude-orchestration.test.cjs is not a .property.test.cjs
  file, and every .property.test.cjs file sampled uses the wrapper consistently. No outlier.

* fix(#3674): constrain the no-overlap property test to unique ids, fixing an ambiguous duplicate-id reconstruction

The `no two plans in the same stage share a modified file` property
reconstructs which physical item produced each output id via
`remaining.findIndex(r => r.id === id)`. Under duplicate ids (an
explicitly-supported input shape for `partitionByFileOverlap`) that
reconstruction can pick the wrong physical occurrence, producing a
false-positive overlap failure (observed counterexample: p0(f1),
p208(f1), p208([]) — correctly staged as [[p0,p208#2],[p208#1]], but
misread by id-order as [[p0,p208#1],...], which do overlap).
Properties (a) determinism and (b) totality already exercise
duplicate ids correctly and are left unchanged; only this property's
generated items are now constrained to unique ids via
`fc.uniqueArray`, where the reconstruction is unambiguous.

---------

Co-authored-by: sim <sim@local>
2026-09-02 07:33:42 -04:00
Tom Boucher
647365faf1 fix(#4011): key the TDD runtime gate on TDD_MODE alone (#4180)
* test(#4011): TDD gate keys on TDD_MODE alone, not the MVP intersection

Contract updates: no shipped line may conjoin MVP_MODE with TDD_MODE as
a gate condition, the end-of-phase escalation must not require MVP, the
executor agent's gate section triggers on TDD_MODE alone, and the gate
semantics reference loads without MVP_MODE.

* fix(#4011): key the TDD runtime gate on TDD_MODE alone

The RED-commit gate shipped as #76's MVP slice kept the paired
invocation's conjunct, so workflow.tdd_mode=true was silently inert on
every non-MVP phase, contradicting references/tdd.md's own contract.
Drops the MVP conjunct from the per-task gate and the end-of-phase
review escalation; rescopes execute-mvp-tdd.md's load condition,
gsd-executor's gate section, and mvp-concepts' intersection claim.
MVP remains free to imply TDD; the file is not renamed (stated
assumption in the PR body).

* test(#4011): scope no-conjunct detector to shell conditions; clean stale MVP+TDD phrasing

Review follow-ups: the detector now only inspects if/[ condition lines
so explanatory prose mentioning both flags cannot trip it; remaining
'under/outside MVP+TDD' phrases in execute-phase.md, the gate
reference, and docs/INVENTORY.md now describe TDD-mode semantics.

Emitted-Drift-Ack-Growth: execute-phase.md — TDD-gate decoupling comment + escalation rescoping (#4011)
Emitted-Drift-Ack-Growth: gsd-executor.md — gate section trigger rescoped to TDD_MODE alone (#4011)

* chore(#4011): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 04:24:52 -04:00
Tom Boucher
7960374d15 fix(#3962): rename the TDD-Audit trailer token to gate-status (#4174)
* test(#3962): shipped TDD-Audit trailer token must round-trip through git

Behavioral coverage: extract the trailers:key token from ship.md and
prove a real git commit carrying that trailer reads back via
%(trailers:key=token,valueonly). gate_status contains an underscore,
which git's trailer machinery cannot tokenize, so the audit read was
structurally empty.

* fix(#3962): rename the TDD-Audit trailer token to gate-status

Underscore is not a valid git trailer token character, so
%(trailers:key=gate_status,...) could never match. Renames the token
at the read, the documented aggregate write, and the section's
prose/table header. Self-suppression semantics (#2431) unchanged.

* test(#3962): compare outcome against the seam's literal, fix doc token

Review follow-ups: the round-trip test compared outcome to 'EXITED'
but the process seam emits 'exited'; and the trailer rename is
propagated to docs/ship-pr-body-sections.md's read/write examples.

* chore(#3962): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 00:01:55 -04:00
Tom Boucher
91d5fdff6f chore(#3546): migrate hook advisory assertions onto typed output surfaces (#4167)
* chore(#3546): migrate hook advisory assertions onto typed output surfaces

Add additive typed fields to 5 hook scripts' PreToolUse/PostToolUse
advisory output alongside the existing additionalContext prose:

- gsd-read-guard.js: code ('READ_BEFORE_EDIT'), fileName
- gsd-context-monitor.js: severity ('warning'|'critical')
- gsd-prompt-guard.js: findings ([{ruleId, match}], module-local RULE_IDS
  + renderFinding mapper mirroring gsd-read-injection-scanner.js's #3523
  pattern)
- gsd-read-injection-scanner.js: severity ('LOW'|'HIGH'), source (its
  findings array already existed from #3523)
- gsd-workflow-guard.js: code ('WORKFLOW_ADVISORY') on the advisory leg,
  distinct from the existing force-add block leg's code

additionalContext stays byte-identical in every hook (verified per-hook
against the pristine HEAD version across a spread of payload shapes).

Migrates all 20 assertion sites named in the issue off
additionalContext.includes(...)/assert.match(...) substring-matching
onto the new typed fields, per CONTRIBUTING.md's prohibition on raw
text matching on test outputs.

Closes #3546

* test: fix undersized commit-class timeout in gsd-statusline.test.cjs's commitN helper

Surfaced by gsd-test on the #3546 checkpoint: `commitN()`'s loop called
gitOrThrow(['add','-A']/['commit',...]) without a timeoutMs override, so
each call used DEFAULT_GIT_TIMEOUT_MS (15s) -- a bound git-fixture.cjs's
own doc comment says is sized for plumbing reads (rev-parse/branch/log),
not write-heavy add/commit spawns. That file already documents the exact
same defect class from a prior incident (PR #3323) and exports
GIT_FIXTURE_TIMEOUT_MS (60s) for fixture-construction call sites -
commitN just wasn't using it. Observed failure: `git commit -m filler 9`
timed out under normal bench load, unrelated to any of this PR's own
diff (hooks/*.js + 5 other test files).

Not a flake: root-caused to the timeout bound being sized for the wrong
call class, per this repo's no-flakes rule.

* chore(#3546): backfill changeset PR number (#4167)

---------

Co-authored-by: sim <sim@local>
2026-09-01 22:32:54 -04:00
Tom Boucher
ff0361071d feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane (#4160)
* feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane

cursor-agent exposes --model (204 selectable models) but the cursor lane
declared modelArg: null / modelConfigKey: null, so review.models.cursor
was rejected as an unknown config key and the #1517 reviewer-instances
escape hatch silently discarded a configured model at modelExpansion.

Wire the lane the same way codex already is: inject {{model}} into args
right after -p, set modelArg to --model, and declare modelConfigKey as
review.models.cursor plus its config schema entry. An unconfigured lane
still invokes byte-identically to today.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3653): add changeset for review.models.cursor

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3653): update co-change surfaces that assumed cursor has no model key

gsd-test surfaced three surfaces still hardcoding "cursor declares no
modelConfigKey", broken by wiring review.models.cursor:

- tests/reviewer-config-federation.test.cjs: the #3691-narrows-#2797
  invariant test listed cursor among lanes that must own no model key.
- tests/settings-integrations.test.cjs: the #3651 keyless-lane test
  listed cursor as keyless, including a live config-set assertion that
  now correctly succeeds instead of failing (swapped to qwen).
- gsd-core/workflows/settings-integrations.md: the settable-keys
  enumeration and two prose call-outs still named cursor as keyless.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3653): acknowledge deliberate growth of settings-integrations.md

settings-integrations.md grew 4 bytes because it now enumerates
review.models.cursor as a settable key alongside the other reviewer
lanes, matching the modelConfigKey wired for cursor in this PR.

Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3653): fix malformed Emitted-Drift-Ack-Growth trailer

The previous commit's trailer was separated from Co-Authored-By by a
blank line, splitting it into an earlier, non-trailer paragraph — git's
trailer parser only recognizes the last contiguous block. Restating it
here immediately adjacent to Co-Authored-By so both parse as trailers.

Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3653): backfill changeset pr number

pr:0 -> pr:4160 now that https://github.com/open-gsd/gsd-core/pull/4160 exists.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-01 21:40:20 -04:00
Tom Boucher
bf4485ada2 enhance(#3717): make the edge probe's shape cues language-aware via an optional text_en field (#4156)
* test(#3717): add failing-first coverage for text_en language-aware classification

Adds unit tests for the not-yet-implemented text_en field on Requirement
(fallback selection, empty/whitespace/non-string rejection, shapes-override
precedence), a SHAPE_CUES/VALID_SHAPES parity guard (RULESET.GENERATIVE-FIX),
and workflow-prose contract tests asserting spec-phase.md Step 5.5 documents
populating text_en for response_language projects. All new tests are RED
until src/edge-probe.cts and the workflow docs are updated.

* feat(#3717): make edge-probe shape classification read an optional text_en field

Requirement gains an optional text_en; classifyShape's own signature stays
untouched (a locked, directly-tested export), and the text_en ?? text
selection is pushed to proposeEdges' single call site instead. text_en is
validated fail-closed: an empty or whitespace-only value throws rather than
silently winning the ?? fallback and degrading classification to zero shapes.

This makes the #2773 doc-only translation convention an explicit,
validatable field instead of an invisible instruction, per the approved
Form-1 scope on #3717.

* docs(#3717): document the text_en field across spec-phase, reference and how-to docs

Updates Step 5.5's response_language instructions, the edge-probe reference
Inputs contract, the FEATURES.md fragment, and the non-English how-to guide
to describe the new text_en field: text keeps the requirement's own wording
in all cases, text_en (when populated) is the engine-only English rendering
the classifier prefers.

* docs(#3717): record the text_en locked-surface change in CONTEXT.md and ADR-550

Updates the Edge Probe Module glossary entry to describe the text_en field
and its fail-closed validation, and appends an ADR-550 amendment recording
why this is additive and does not re-open the #652 LLM-classifier rejection
(text_en is a plain field read by the existing deterministic regex
classifier, not a new model-dependent surface).

* docs(#3717): add changeset fragment and regenerate FEATURES.md

pr:0 placeholder — backfilled with the real PR number after the PR opens.

* docs(#3717): attribute the text_en machine check to engine-level validation, not prose tests

Code-review (Spec axis) finding: the workflow-prose contract tests and the
ADR-550 amendment overclaimed themselves as "the machine check the #2773
doc-only stopgap lacked." That check is actually engine-level
(validateRequirement/classifyShape, covered in tests/edge-probe.test.cjs) —
the prose tests are the same style of assertion #2773 already used. Reworded
both to attribute the claim correctly.

* fix(#3717): rewrap spec-phase.md so the id-unchanged sentence stays on one line

The #3717 rewrite of Step 5.5's response_language paragraph moved a line
break so "requirement `id`s" ended one physical line and "are never
translated" started the next. The pre-existing #2773 regression test
(tests/edge-probe-spec-phase-contract.test.cjs) asserts id + "never
translated" on the SAME line (no \n in between, matching git's own
line-oriented prose), so the reflow silently broke it. Rewrapped so the
sentence lands on one line again, verified against every #2773/#3717
regex assertion in that test file.

Emitted-Drift-Ack-Growth: spec-phase.md — #3717 adds text_en documentation to Step 5.5 (response_language paragraph + REQS_JSON heredoc comment); this growth is this PR's own diff, not incidental drift.

* chore(#3717): backfill changeset PR number

pr:0 -> pr:4156 now that the PR exists.

---------

Co-authored-by: sim <sim@local>
2026-09-01 21:39:53 -04:00
Tom Boucher
4dfc46bbe7 enhance(#3348): add a context-drift pre-check gate to plan-phase (#4147)
* test(#3348): add failing-first coverage for the context-drift gate

* feat(#3348): add context-drift pre-check gate for plan-phase

Compares each phase's *-RESEARCH.md/*-PATTERNS.md/*-VALIDATION.md/*-SPEC.md
effective last-changed time (git commit time, falling back to mtime for
uncommitted edits) against *-CONTEXT.md's, so plan-phase no longer silently
reuses an upstream artifact that predates a decision added to CONTEXT.md
after that artifact was derived from it. Deterministic, no model call.

New `gsd_run verify context-drift <phase>` command, sibling to the existing
verify.codebase-drift/verify.schema-drift gates in the drift capability.
Warn-only by default (workflow.context_drift_precheck), with an opt-in
workflow.context_drift_action: block escape hatch. Wired at plan:pre in
plan-phase.md, before both the RESEARCH.md and PATTERNS.md reuse decisions.

* fix(#3348): address code-review findings — raw-text-match, stale comment, import placement, duplicated phase resolution

* fix(#3859): pin the real commit's diff.ignoreSubmodules to match the empty-diff probe

The #3859 empty-diff guard decides whether a submodule bump would land using
`--ignore-submodules=dirty`, overriding the caller's `diff.ignoreSubmodules`
config. The real `git commit -- <paths>` that follows was never given the
same override, so under a bare `diff.ignoreSubmodules=all` repo config the
two calculations disagree: driven on git 2.39.5 (Debian bookworm, the
linux-node24 test-matrix image), the guard correctly stands aside but the
scoped commit itself then silently fails (exit 1, no error text) for a
gitlink bump it had just confirmed would be recorded, surfacing as
commit_failed instead of committed:true.

Pin `-c diff.ignoreSubmodules=dirty` onto the scoped commit call too, so the
probe and the commit it protects can never diverge. Harmless when no
submodule path is involved (driven: identical outcome on an ordinary scoped
file, with and without the flag).

* fix(#3348): guard resolvePhaseDirByToken's exact-match fallback against path traversal

* fix(#3348): retarget phase-enumeration-drift exemption to the consolidated resolvePhaseDirByToken helper

cmdVerifySchemaDrift's inline readdirSync was already function-scoped-exempt
in lint-phase-enumeration-drift.cjs as a single-phase LOOKUP (not a
current-milestone enumeration). This PR's refactor pass lifted that block
into a shared helper, resolvePhaseDirByToken, also used by the new
cmdVerifyContextDrift — the guard tracks exemptions by enclosing function
name, so the readdirSync now lives in an unexempted function and started
firing. Move the exemption to resolvePhaseDirByToken (same written reason,
now covering both callers) instead of migrating to listAllPhaseDirs, which
would introduce two real behavior deltas here: it catches readdirSync
failures internally (old code let them throw) and sorts results by phase
number before matchPhaseDirs picks matches[0] (old code used raw,
OS-dependent readdirSync order).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3348): satisfy lint:ci — slash form, capability registry regen

- docs/features/context-drift-gate.md used the deprecated /gsd: colon
  form; docs are never passed through the install-time slash-form
  converters, so lint-docs-command-form requires the hyphen form.
  Regenerated docs/FEATURES.md from the corrected fragment.
- Regenerated gsd-core/bin/lib/capability-registry.cjs after editing
  capabilities/drift/capability.json (lint:generated-sync).

* fix(#3859): pin the real commit's diff.ignoreSubmodules via env, not argv -c

The prior fix pinned `-c diff.ignoreSubmodules=dirty` onto the scoped commit's
argv via `commitArgs.unshift(...)`. `-c key=val` must precede the `commit`
subcommand, so this shifted `commitArgs[0]` from `'commit'` to `'-c'` for
every scoped commit call, breaking 17 position-based assertions in the
commit-files pathspec regression suite that read `a[0] === 'commit'` to find
the commit invocation among recorded git calls.

`execGit` already accepts an `env` option merged onto `process.env` before
spawning. Git honors `GIT_CONFIG_COUNT`/`GIT_CONFIG_KEY_0`/`GIT_CONFIG_VALUE_0`
as a per-invocation config override functionally identical to `-c key=val`,
expressed via env instead of argv. Passing that env alongside the existing
commitArgs (still `['commit', ..., '--', ...stagedPaths]`, argv unchanged)
fixes the real commit's effective diff.ignoreSubmodules to match the
empty-diff guard's probe without moving anything in argv position 0. Scoped
to exactly the canScope branch, matching the probe's own preconditions and
leaving no behavior change for commits the probe never evaluated.

No test file changes needed — the 17 previously-failing assertions test
argv[0] against the array passed into execGit, which never changes.

* fix(#3348): register verify-context-drift in the check subcommand router

The drift capability's new plan:pre gate declares check.query
"verify.context-drift", which normalizes to `check verify-context-drift`,
but no such subcommand was routed — phase6-capstone-conformance's
uniform-block-field test failed with "Unknown check subcommand" for
every declared gate query.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3348): extend #1592's exact-key-list snapshot for the new context-drift config keys

tests/capability-registry.test.cjs asserted an exact, hardcoded snapshot
of the drift capability's config keys. #3348 legitimately adds two new
keys (workflow.context_drift_precheck, workflow.context_drift_action)
for its own plan:pre context-drift gate — extend the expected set
(and clarify the assertion message) without weakening the test's
exactness.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3348): reconcile E2's exemption-migration pin with the resolvePhaseDirByToken extraction

#3348 (an earlier commit on this branch, e4b80ad81) extracted
cmdVerifySchemaDrift's inline phasesDir readdirSync/matchPhaseDirs block
into the shared resolvePhaseDirByToken helper (also used by the new
cmdVerifyContextDrift), and retargeted lint-phase-enumeration-drift.cjs's
function-scoped exemption from cmdVerifySchemaDrift to
resolvePhaseDirByToken accordingly — cmdVerifySchemaDrift no longer
contains a line the guard's detectors match, so it needs no exemption.

tests/phase-locator.test.cjs's E2 test still pinned the exemption to the
old name (cmdVerifySchemaDrift), unaware of the migration. Update E2 to
match the same "migrated call site's exemption must move, not
duplicate" pattern the test already applies to cmdRoadmapAnalyze and
cmdInitMilestoneOp just below it: drop cmdVerifySchemaDrift from the
still-exempt list and add symmetric assertions that it no longer
carries the exemption while resolvePhaseDirByToken now does.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3348): fix two self-contradicting/nondeterministic tests in context-drift.test.cjs

'always exits 0 (query command contract)' included the no-phase-arg
case, which contradicts the file's own earlier
'errors with usage message on missing phase arg' test (that case
legitimately exits 1 via the Usage error) — drop it from the
always-exits-0 cases.

'degrades to mtime comparison outside a git repo' and '...in a repo
with no commits' relied on real wall-clock ordering between two
back-to-back writeFileSync calls to prove CONTEXT.md is newer than
RESEARCH.md; on a fast filesystem both can land in the same mtime
tick, producing a tie that computeContextDrift's strict `<` correctly
treats as not-stale, so stale_artifacts comes back empty. Make both
tests deterministic via explicit fs.utimesSync instead of relying on
timing (CONTRIBUTING.md: never assert elapsed wall-clock time).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3348): add context_drift_precheck:false to the plan:pre all-off fixture

The "all plan:pre when-keys false" fixture explicitly disables every
known workflow.* plan:pre toggle, but didn't yet know about the new
workflow.context_drift_precheck key (defaults to true), so the new
drift context-drift gate stayed active and broke the
empty-activeHooks assertion.

Emitted-Drift-Ack-Growth: plan-phase.md — adds the #3348 context-drift plan:pre pre-check section (new ## 4.6); this PR's own diff, not incidental drift.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3348): backfill changeset PR number (pr:0 -> 4147)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-01 21:39:37 -04:00
Tom Boucher
c0fd2e3f4c feat(#3673): add dispatch.maxConcurrency axis and dispatch-capacity query (#4162)
* test(#3673): add failing tests for dispatch.maxConcurrency axis and dispatch-capacity CLI route

Extends tests/host-integration.test.cjs with negotiateHostCapabilities
maxConcurrency negotiation coverage (test matrix rows 1-15, including a
fast-check property test) and a new #3673 dispatch-capacity CLI route
describe block spawning the real gsd-tools.cjs (rows 16-25). Extends
tests/host-integration-validator-parity.test.cjs with an all-19-descriptor
maxConcurrency presence/validity sweep (row 26) and adds a hostile-input
validator test to host-integration.test.cjs (row 27).

The dispatch.maxConcurrency field does not exist yet, so these tests fail.

* feat(#3673): add dispatch.maxConcurrency axis, negotiation, validator parity, and the dispatch-capacity query

Adds a numeric dispatch.maxConcurrency sub-field to the Host-Integration
Interface (ADR-1239 Phase 1), following the existing dispatch.isolation
sub-field pattern: DispatchCapability interface, SAFE_DEFAULTS/PROFILE_BASELINES
floors, and a negotiateHostCapabilities branch that passes through a positive
safe integer and fails closed to 1 otherwise (no engine-side reduction, per
the design doc's explicit rejection of a min(host,engine) rule).

capability-validator.cjs gains parity validation for the new field (optional,
positive safe integer or the "undocumented" sentinel — mirroring isolation's
"added after existing descriptors" treatment).

gsd-tools.cjs gains a new `query dispatch-capacity` route, a pure-read sibling
of `dispatch-isolation` with no side effects: live env
(GSD_DISPATCH_MAX_CONCURRENCY) > descriptor > fallback-to-1 precedence.

All 19 capabilities/*/capability.json descriptors now declare
dispatch.maxConcurrency: claude carries the one cited value (20, per
code.claude.com/docs/en/sub-agents); the other 18 carry "undocumented"
(not yet researched for this axis).

* docs(#3673): document dispatch.maxConcurrency and add its citation row to the capability matrix

Updates docs/reference/host-integration-interface.md's dispatch struct entry
(also backfilling the previously-undocumented isolation/backgroundDispatch
sub-fields found stale in the same table) and adds fail-closed/live-transport
precedence prose for the new maxConcurrency field.

Adds a dispatch.maxConcurrency row (with citation) to all 19 host sections in
docs/reference/host-integration-capability-matrix.md — required for
tests/host-integration-descriptors.test.cjs's kimi-code matrix-parity check,
which asserts every declared dispatch sub-axis is documented there.

Updates docs/how-to/add-or-update-a-host-integration.md's dispatch checklist
and example descriptor block to mention maxConcurrency (and, likewise
backfilling a stale gap, isolation/backgroundDispatch).

* fix(#3673): extract shared maxConcurrency validator, drop dead reserved-name check

Exports isPositiveSafeInteger from src/host-integration.cts as the single
source of truth for the dispatch.maxConcurrency positive-safe-integer
contract; negotiateHostCapabilities and gsd-tools.cjs's routeDispatchCapacity
now both call it instead of independently reimplementing the same predicate.

Also removes the __proto__/constructor/prototype reserved-name branch from
capability-validator.cjs's maxConcurrency check — copy-pasted from the
string-enum fields above it, but unreachable for a numeric field (the
generic positive-safe-integer branch already rejects any string) and absent
from maxDepth, the field the code's own comment claims to mirror.

---------

Co-authored-by: sim <sim@local>
2026-09-01 21:39:12 -04:00
Dennis Alexis Valin Dittrich
b848b23861 feat(#3778): dispatch plan:pre planner contributions before quick planning (#3934)
* feat(#3778): dispatch plan:pre planner contributions in quick.md

- Add plan:pre capability gate to quick.md Step 5, mirroring plan-phase.md's
  existing render + generic contribution dispatch pattern
- Inject planner-targeted contribution fragments into the planner prompt,
  after AGENT_SKILLS_PLANNER, matching D-08 ordering
- Add tests/quick-plan-pre-capabilities.test.cjs proving the dispatch is
  generic (D-01) via real scanWiredKinds/coveredKindsInRegion functions
- Record quick.md's byte-growth rationale in this commit trailer for every
  gsd-core-verbatim runtime

Emitted-Drift-Ack-Growth: quick.md — #3778: Step 5 (Spawn planner, quick mode) gains a `plan:pre` capability gate, mirroring `plan-phase.md:420-424` and `:797`. This is shipped shell and prose read by an agent at runtime, not compiled, so the reasoning has to travel with the feature rather than being deferred to a reference doc: (1) the dispatch paragraph phrases role routing possessively ("the role each entry's `into` names") rather than as an `into ==` equality, because `coveredKindsInRegion` (scripts/gen-loop-host-contract.cjs) voids a segment's `kind == "contribution"` coverage credit when a role or capability equality shares that same segment — an equality phrasing here would silently fail the generic-dispatch proof required by D-01; (2) `activeHooks` is read directly in-context from `PLAN_PRE_HOOKS_JSON`/`HOOKS_JSON` and the unfiltered `rendered` digest is explicitly forbidden from being pasted, because `rendered` carries every kind and role — including non-planner-targeted contributions such as a `into: "checker"` twin — and pasting it would leak checker-scoped guidance into the planner's prompt (T-01-02 in the threat model); (3) the injection block sits inside `<planning_context>` AFTER `${AGENT_SKILLS_PLANNER}` and after the Project skills line, matching plan-phase's `:741` -> `:797` ordering (D-08), so agent-skills content is never shadowed by capability-contributed prose. No prose was moved into an eagerly `@`-imported reference to shrink the measured file — @gsd-core/references/loop-hook-dispatch.md already existed before this change and is deferred to for the generic contract only, exactly as plan-phase.md already does.

* test(#3778): expand quick.md plan:pre dispatch coverage to all nine locked conditions

Extend tests/quick-plan-pre-capabilities.test.cjs with D-02 (silent
omit-when-empty), D-03 (single shared planner spawn), D-06 (array-order
dispatch phrasing), D-07 (planner-only into filter), and D-08 (render call
< agent-skills placeholder < injection block < spawn ordering) assertions,
all extracted via a brace-bounded slice anchored on the literal injection
instruction rather than a naive first-brace scan (${AGENT_SKILLS_PLANNER}
and the surrounding prompt's ${VALIDATE_MODE ? ...} ternaries also contain
brace pairs).

Add a capability-registry.test.cjs describe block proving the registry-wide
D-07 exclusion is meaningful: at least one plan:pre contribution exists,
every plan:pre contribution has a non-empty into/fragment.inline, and the
registry as a whole carries at least one non-planner-into contribution.

Add a loop-host-contract.test.cjs regression pin for D-09: quick.md stays
absent from STEP_WORKFLOWS, parseLoopHostBlock still throws on quick.md's
real content, and buildContract() still yields exactly 5 entries.

Verified red-without-Task-1 by temporarily reverting quick.md to its
pre-f30de9cc content and re-running these three suites (D-08 failed as
expected), then restored via git checkout and re-confirmed green.

* docs(#3778): note quick planning also renders plan:pre in the tutorial

The tutorial's Step 6 named only /gsd-plan-phase as the trigger for the
plan:pre hook set. Since quick.md now dispatches the same hook set
(f30de9cc), the sentence understated the capability's real reach.

* feat(#3778): add changeset fragment

* chore(#3778): reference the upstream issue in the changeset fragment

The fragment was the only one of 81 in .changeset/ without a trailing
(#NNNN) reference or a bold lead-in. serializeChangelog auto-appends
only the pr: field, so the rendered CHANGELOG entry carried no link
back to issue #3778.

* test(#3778): scope the D-07 registry assertion to what it actually proves

The registry-wide non-planner check was named "D-07 exclusion is
meaningful", which overclaims: it proves only that `into` takes
non-planner values somewhere in the registry, not that anything is
excluded at plan:pre. Every plan:pre contribution is currently
into: "planner", so the filter is a forward-looking safeguard there.

Narrowing the assertion to plan:pre (as review suggested) would fail
today. Asserting plan:pre is all-planner would be brittle — it would
break the day a legitimate non-planner plan:pre contribution lands,
which is exactly when the safeguard starts doing work. So the
assertion is unchanged and only the name and comment are corrected.

* chore(#3778): point the changeset fragment at the upstream PR

The fragment carried pr: 3, the fork staging PR. changeset lint derives
the real PR number from GITHUB_EVENT_PATH, so on the upstream PR that
would read as pr-field drift. Point it at open-gsd/gsd-core#3934.

* test(#3778): require contributions in Quick revision prompts

* test(loop-host): require Quick auxiliary registration

* fix(#3778): preserve contributions in Quick plan revisions

* fix(#3778): validate Quick as a planner contribution host

* fix(#3778): tighten Quick contribution contract

* test(#3778): drop unnecessary source-contract exemption

* fix(#3778): require Quick planner target coverage

* docs(#3778): describe targeted auxiliary coverage

---------

Co-authored-by: davdittrich <davdittrich@gmail.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-01 21:16:57 -04:00
Tom Boucher
7ca5479f97 docs(#3672): amend ADR-1239 for quick-batch concurrency and single-writer contract (#4146)
Locks the architecture gate the maintainer required before any /gsd:quick-batch
implementation PR (epic #3344, condition #2): canonical terminology, the
foreground-coordinator single-writer invariant over BATCH.json/STATE.md/roadmap
metadata, the dispatch.maxConcurrency sub-field (fail-closed to 1), the
GSD_DISPATCH_MAX_CONCURRENCY live-capacity transport and its precedence over the
descriptor value, the gsd-tools query dispatch-capacity contract, the
dispatch.isolation interaction, and backpressure/recovery/deterministic-merge
rules. Docs-only; closes its own Phase 0 sub-issue, not the epic.

Co-authored-by: sim <sim@local>
2026-09-01 15:53:17 -04:00
Adnan
bdfc62889b fix(#3784): read the hybrid Current Plan: N of M shape, keep zero-padding, and name the accepted shapes on failure (#3791)
* fix(state): read the hybrid "Current Plan: N of M" shape

advancePlanCore derived the value FORMAT from the field NAME, so it
handled the legacy pair (`Current Plan` + `Total Plans in Phase`) and the
compound `Plan: N of M`, but not the hybrid of the two: the legacy field
name carrying a compound value with no Total Plans sibling. `legacyTotal`
is null so the legacy branch fell through, and the compound branch reads
the `Plan` field through a `^Plan:`-anchored pattern that never matches
`Current Plan:`. Both produced NaN against a file whose plan numbers are
plainly readable.

The shape is not exotic. An agent wrote it unprompted into a project's
STATE.md, believing it was the parseable form, and every subsequent run
in that project inherited the failure and worked around it by hand.

Track the field name and the value shape separately (`planSourceField`,
`planRawValue`) so write-back targets whichever field the value came
from. The legacy pair still takes precedence when both fields exist, so
a stray "of N" inside Current Plan cannot override an explicit Total
Plans — covered by a new test.

Also replace the caller's catch-all error. It reported "Cannot parse
Current Plan or Total Plans" for ANY transition failure, and named no
accepted shape, so a reader learned neither what failed nor what to
write. It now distinguishes "no result" from "unreadable plan position"
and lists all three shapes. The existing test asserted the literal
"cannot parse"; it now asserts the message names the shapes, which is
the property that makes it actionable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* fix(state): keep zero-padding when advancing a compound plan value

The compound write-back rewrote only the leading half of "N of M", so a
padded value drifted lopsided: "04 of 06" advanced to "5 of 06". Cosmetic
on its own, but a plan line that looks wrong is one the next writer tidies
by hand, and hand-tidying this particular line is what produced the hybrid
shape the previous commit had to teach the parser to read.

Pad the incremented number to the width it was written with. padStart never
truncates, so a value that outgrows its padding widens correctly: 09 of 12
advances to 10 of 12. Unpadded values are untouched — 2 of 6 still advances
to 3 of 6.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* fix(state): pass a literal field name to the compound write-back

The previous commit passed `planSourceField` — a variable — as the field-name
argument to `stateReplaceField`, which trips the state-write-path drift guard's
`unstripped_content_write` axis (ADR-3408 §8.3(b)). The guard is right to care:
a Title-Case literal cannot collide with a lowercase or snake_case frontmatter
key, so it is safe whatever the content argument is, while a variable could
hold anything and therefore requires its content to be demonstrably
frontmatter-stripped first.

The content argument here IS stripped — `body` is `stripFrontmatter(content)` —
but the guard does a narrow backward scan rather than dataflow tracking, by
design, and the nearest preceding assignment to `body` is another
`stateReplaceField` result. Rather than baseline a bypass or ask a future
reader to re-derive that the invariant holds, dispatch on the discriminator and
pass the literal.

Guard goes from 1 finding to 0; its own 32 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* chore(3784): add changeset fragment for #3785

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* test(#3784): cover the maintainer's AC1 write-back and AC6 reader-anchoring

Triage published six acceptance criteria; two were only half-covered.

AC1 asks that the hybrid write back to the SAME field with padding preserved.
The existing hybrid test used an unpadded value and asserted only `result.data`,
so it proved the parse but never the write. Now asserts the written content is
`05 of 06` on the original field, and that no separate `Plan:` field appears as
a side effect.

AC6 asks that the shared field reader not be loosened. Reading the hybrid is the
transition's job; `stateExtractField('Plan')` is line-anchored and has 13+
callers, so teaching it to match a name merely ENDING in "Plan" would be the
wrong fix and would silently change what those callers read. This holds by
construction here — the reader is untouched — but nothing locked it in. The new
test fails if anyone later reaches for that shortcut.

Also drops the changeset fragment written against the auto-closed PR number.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* chore(#3784): add changeset fragment for #3791

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* fix(#3784): write the advanced plan back to the field it was read from

Review findings 2-6 on #3791 were one defect seen from several angles: the
read path learned the hybrid `Current Plan: N of M` shape, the write path
did not follow it.

- `bumpLeadingNumber` now owns the increment for all three parse branches.
  Only the leading digits belong to this transition; the padding width and
  everything after it (` of M`, and the `\r` of a CRLF file) are the
  author's text and are preserved. The legacy branch wrote `String(newPlan)`,
  which turned `2 of 99` into `3` and `04` into `5`.
- `mutateCurrentPositionForAdvance` takes the plan field NAME. Its plan arm
  only ever looked for `Plan:`, so on a hybrid file the `## Current Position`
  section was never reached; combined with the body-level write being
  single-shot and bold-preferring, a file carrying the field at both sites
  advanced the header and left the section a plan behind. The parameter
  defaults to `Plan`, so the two callers that pass no plan are unchanged.
- Tests: both-sites-advance (fails without the section arm), legacy
  write-back content assertions (the previous test read only `data` and so
  could not see the lossy write), hybrid boundary at limit-1 and limit+1, a
  CRLF fixture, and an fc property pinning the padding-width contract.

Two characterization tests pinned `**Current Plan:** 02` advancing to `3`.
That dropped padding is the defect #3784 reports, so the expectation is
corrected to `03` rather than the fix being narrowed around it.

* fix(#3784): drop the unreachable advance-plan error branch, sync the doc

Findings 1 and 8 on #3791.

The `!resultData` arm could not fire: the transform callback assigns
`resultData` unconditionally, only runs once STATE.md is known to exist (the
missing-file case returns "STATE.md not found" upstream), and every
`advancePlanCore` return path sets `data`. It was a speculative second
failure mode with a message no caller could receive, and the comment beside
it claimed to distinguish two things that were never two. `!resultData`
stays in the condition as a type guard, which is all it ever was.

`docs/json-errors.md:142` quoted the old error literal verbatim and was the
sole occurrence in the tree; it now quotes the emitted one.

* chore(#3784): describe the write-back fix in the changeset

* fix(#3784): anchor the plan grammar and widen the schema row to match

Review round 3 on #3791: B1, B2, M1, M2, M3, M4 and the planSourceField nit.

B1 — `STATE_FIELD_SCHEMA.current_plan.acceptedShapes` widens to
`['N', 'N of M']`, which is what `src/state-md-schema.cts`'s own comment
instructed this PR to do on merge. `'N/M'` stays undeclared so row 23 keeps a
non-vacuous undeclared candidate to probe. Verified `gen-state-md-docs
--check` exits 0 and `--write` rewrites 0 of 6: the generated artifacts do not
surface this row, so there is nothing stale to regenerate.

B2/M4 — the discriminator was `/of\s+(\d+)/`, unanchored, so a total could be
read out of prose. `Current Plan: 4 — blocked on review of 2 PRs` parsed as
`4 of 2`, took the `currentPlan >= totalPlans` branch and WROTE
`Status: Phase complete — ready for verification` into the user's file. Both
shapes are now anchored at the start and every number comes from a capture
group via `planNumberFrom`, which rejects anything past `Number.MAX_SAFE_INTEGER`
rather than letting `data` and the persisted string disagree. Nothing on this
path calls `parseInt` on a raw field value any more.

The grammar keeps a trailing remainder after the total, because
`Plan: 2 of 5 in current phase` is a real tested shape. The refusal comes from
requiring `of <total>` to follow the leading number immediately, not from
forbidding a suffix.

M1 — `bumpLeadingNumber` is total. It returned its input unchanged when there
were no leading digits, so `+2` reported `advanced: true` while writing the
file untouched.

M2 — both section arms use replacer functions. File-derived text was being
spliced into a `String.replace` replacement string, where `$&` / `` $` `` /
`$'` expand: a value of `04 of 06 $&` spliced part of the document into itself.
`stateReplaceField` already used a function; these now agree with it.

M3 — the section arm targets the name the SECTION carries, and the body write
now writes both spellings, each with its own rendering. Keying off the header's
name left the other name stale in both directions: a legacy header beside a
`Current Plan:` section line, and a `**Plan:**` header beside one.

* fix(#3784): derive the shape error from the schema, widen the test coverage

Review round 3 on #3791: B3, m1, m2, m5 and the two test nits.

B3 — the accepted-shape set had two owners: the parser branches and an English
list hand-written beside them in `state.cts`. Nothing coupled them, so adding a
branch left the message stale and removing one left it advertising a shape that
errors, with no test able to see either. The message is now built from
`STATE_FIELD_SCHEMA.current_plan.acceptedShapes`, and the CLI test walks the
schema instead of restating the list. `Plan: N of M` is still spelled out
explicitly because no schema row owns the body-only `Plan` field —
`buildStateFrontmatter` never reads it into frontmatter, so it has no key to
hang a row on.

m1 — the property drove only the pre-existing `**Plan:**` branch, i.e. not the
branch under review. It now drives both compound spellings and ranges past 99
so the width transition is covered by the property rather than one example. A
second property covers the legacy pair's own preservation contract. Both were
mutation-checked: dropping the padStart turns 9 tests red.

m2 — degenerate boundary fixtures around the threshold (`0 of 0` is
phase-complete, not an error; `0 of 3` advances) plus the shapes the anchored
grammar must refuse, including Arabic-Indic digits.

m5 — `docs/json-errors.md` described rather than quoted the message, since it
is now schema-derived and a verbatim quote would be a third owner.

Nits — the CRLF assertion could not see a `\n` at index 0; the
`!/^Plan:/m` presence proxy is now an identity assertion on the whole
`## Current Position` body.

* fix(#3784): give the section plan write its own flag, and stop narrowing what parses

Review round 4 on #3791: Blockers 1-4, Majors 1-2, Minors 1-2.

B1 — the section fallback was guarded by `!mutated`, and `mutated` is
FUNCTION-wide, already set by the phase/status/lastActivity arms that
`advancePlanCore` always populates. A section spelling the field bold or as a
pipe-table row therefore skipped its fallback because an UNRELATED field had
been refreshed, and stayed a plan behind the header — the split-brain document
this arm exists to prevent. The arm now tracks its own `planWritten`.

Worth recording: the reviewer's fixture does not reproduce. The body-level
status write lands on the section's own `Status:` when the document has no
header `Status:`, so `mutated` is still false by the time the plan arm runs and
the fallback fires. The discriminating shape needs a header `Status:` to absorb
that write AND a bold section plan line. The mechanism was right; the example
was not, and the regression test uses the shape that actually fails.

B2 — `fallbackName` chose one name by ternary. In the legacy shape both values
are populated, so it always chose `Current Plan` and a `**Plan:**` section line
— which base did write — got nothing. Each name is now attempted independently
with its own fallback.

B3 — `PLAN_SHAPE_N` was anchored harder than `PLAN_SHAPE_N_OF_M`, so values base
parsed via `parseInt` began to hard-error: `Total Plans in Phase: 5 phases`,
`Current Plan: 3 (blocked)`. #3784's brief puts normalizing plan numbers beyond
this transition's read/write out of scope, so that narrowing was not licensed.
Both grammars now carry the same trailing tolerance. The prose defect stays
closed by the START anchor, not by forbidding suffixes.

Major 1 — the whole-body `Plan` write is scoped to documents that declare a
`Plan` field, instead of firing unconditionally where `stateReplaceField`'s
first match could be prose outside `## Current Position`.

Major 2 — the error message names both `Plan` spellings the parser accepts; it
previously omitted the sibling-paired form, which is the same
message-disagrees-with-parser drift the derivation exists to close.

B4 — the changeset claimed a guarantee B1 broke; it now describes what ships.

Minors — safe-integer boundary coverage at limit-1/limit/limit+1, and the CRLF
comment states the real mechanism (`stateExtractField`'s `(.+)` stops before the
CR; the trailing group is belt-and-braces, not the primary defence).

All three blocker regression tests verified red against the pre-fix source.

* test(#3784): pin the hybrid shape against #3807's ambiguity refusal

#4028 landed `advance-plan`'s multi-`Phase:` refusal on `next` after this
branch's last run, on the same function. The guard sits above the parse, so
a refused document is never parsed and the shape #3784 adds cannot reach the
mutation — but that is a property of source ordering, so assert it as
behaviour instead.

Fail-first proven, not assumed: with `phaseCandidates.length > 1` disabled,
the ambiguous hybrid document advances its FIRST entry's
`Current Plan: 04 of 06` to `05 of 06` and writes it — #3807's exact defect,
reached through #3784's shape. Both tests go red; both go green with the
guard restored.

The control pins the other direction: an unambiguous hybrid section still
advances, and its zero-padding still survives.

* fix(#3784): advance every spelling from its own text, refuse when they disagree

Round 6 review. B1 and M1 are one defect, so they are one fix.

`advancePlanCore` picked one field to parse from, computed `newPlan`, then
wrote BOTH spellings from that field's numbers. Two symptoms:

  B1  With `Plan` as the parse source, `Current Plan` was re-stamped with
      the number just derived from `Plan`. `Current Plan: 7` beside
      `Plan: 2 of 5` silently became `Current Plan: 3` — a value nothing
      derived for that field, no error, no diagnostic.
  M1  With the legacy pair winning, the `Plan:` line was re-rendered from
      a bare `${newPlan} of ${totalPlans}` built out of the sibling field.
      `Plan: 2 of 9` became `3 of 5`; `Plan: 03 of 05` became `4 of 5`.
      The changeset's claim that padding and everything after it survive
      was true only for whichever field happened to be the parse source.

Now: every spelling is advanced from its own raw text via
`bumpLeadingNumber`, so each keeps its own padding, its own total and its
own trailing annotation. Differing TOTALS are preserved, not reconciled —
`Plan: 2 of 9` beside a `Total Plans in Phase: 5` advances to `3 of 9`.

Differing CURRENT numbers are refused, with `reason:
"ambiguous_plan_position"` and both candidates named. Same posture as
#3807's multi-`Phase:` guard one field over: name the conflict, let the
caller resolve it, never pick. The guard sits immediately after the parse,
BEFORE the phase-complete branch — guarding only the normal advance would
let `Current Plan: 7` beside `Plan: 5 of 5` write a terminal "Phase
complete" into a document whose two spellings never agreed.

A field present but unreadable (`Plan: TBD`) is left exactly as authored.
Refusing the whole document because an unrelated line cannot be read would
be a narrowing #3784 does not license; writing a derived number over it is
the fabrication B1 was filed for.

The `planSourceField`/`planRawValue`/`useCompoundFormat` tracking is gone.
It existed only so the write path could ask which field the value came
from, and the write path no longer asks.

M2. The bare `Plan: N` + `Total Plans in Phase: M` shape is dropped. A
revision of this PR added it; base refused it. It cannot be given the
schema-row + forcing-test coupling the other shapes have, because `Plan`
is body-only and `buildStateFrontmatter` never reads it into frontmatter,
so there is no `current_*` key to hang a row on. Parser, the spelling in
`advancePlanShapeError`, and the lockstep test move together — the
invariant is the lockstep, not the length of the list.

N1. The whitespace narrowing (`5phases` no longer parses where `parseInt`
read 5) is documented in the changeset beside the other deliberate
narrowings, rather than loosened. Loosening restores the half-parse this
change exists to remove.

Tests: eight new cases plus a property that crosses the two spellings with
agreeing and disagreeing numbers — the review noted the existing
properties never did. Fail-first proven: restoring the old write path
reddens seven of the eight, both new property arms, and two pre-existing
padding tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

* fix(#3784): report Current Plan as updated only when it was written

The write became conditional in the previous commit — a `Current Plan:`
that is present but unreadable is left as authored — but the `updated`
push stayed unconditional, so `transitionCore` reported a field it had not
touched. `reconcileReportedFields` would have caught it against the
persisted bytes at the `state.cts` caller, but `transitionCore`'s own
`updated` is consumed directly (milestone-lock, the transition tests) and
has to be true on its own.

Covers the mirror of the unreadable-spelling case: `Current Plan: TBD`
beside a readable `Plan: 2 of 5`, where `Plan` is the parse source and the
legacy field is the one that cannot advance. Fail-first proven.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

* test(#3784): account for the new refusal in the output({error}) census

`tests/io.test.cjs`' A3 census asserts the exact population of
`output({error})` call sites in `src/`, per module. The
`ambiguous_plan_position` refusal added a 27th to `state.cts`, so the
census went red at 26/65.

Updated the way #3807 updated it when it added the ambiguous-POSITION
error one line above: bump the count and name the addition inline, so the
next person reads why the number is what it is. The alarm did its job —
it is the only gate that noticed a new user-visible error path had been
introduced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-01 15:03:20 -04:00
0xdhx
7c116b1c17 fix(#3697): warn when the phase-complete Requirements-line tokenizer under-selects REQ-IDs (#3744)
* fix(#3697): warn when the Requirements line under-selects REQ-IDs

`cmdPhaseComplete` tokenizes ROADMAP's `**Requirements**:` line by splitting
on `[,\s]+` and keeping tokens matching the anchored REQ-ID shape. That is
correct for the canonical comma list the template ships, and silently wrong
for every other form:

  `RANGE-01 … RANGE-05`  ->  the two ENDPOINTS only; the interior IDs are
                             never considered, yet `requirements_updated`
                             reports true with zero warnings
  `RANGE-01…05`          ->  ZERO IDs; the whole line is inert

The silence is structural: the only cross-check, `ghostReqIds`, is itself
`citedReqIds.filter(...)`, so an ID the tokenizer dropped is invisible to it
by construction — and to `traceabilityWriteMisses` and `requirements_updated`
with it.

Warn on both paths. This does not add range support: the selected set is
unchanged, so no existing ledger write changes. The trigger is ID-SHAPED
EVIDENCE only — an ID-shaped substring the tokenizer did not select, or a
range operator joining two IDs — with parenthetical citations and HTML
comments stripped before the scan, so the #2334/#2339 over-warning on
`None`, on the shipped `<!-- brackets optional -->` template comment, and on
annotated lines cannot return.

Regression tests extend the #2316/#2334 fixture family in tests/phase.test.cjs
(10 cases: 4 defect, 2 canonical controls, 4 negative-space controls).

Fixes #3697

* fix(#3697): rework under-selection detection onto tokens, not a free-text scan

Round 2, driven by the P4.6 cross-AI review (codex, gpt-5.6-sol) of b3ce71cb.
That review refuted 5 of 9 claims; three were false-positive classes in exactly
the category #2334/#2339 had to REMOVE:

  `RANGE-01, RANGE-02 - 3 points`        the bare-hyphen alternative read
                                         `RANGE-02 - 3` as a range
  `REQ-01, REQ-02 — locked per ADR-7.`   the trailing period kept `ADR-7.` out
                                         of the anchored filter, so the
                                         unanchored substring scan reported it
                                         as unparsed
  `REQ-01, REQ-02 (see (ADR-7), then ADR-8)`
                                         nested parens left `ADR-8)` behind

Replaces the free-text substring scan + loose range regex with three narrow,
token-based rules (R1 range-shaped token, R2 pure range operator flanked by two
selected IDs, R3 zero-selection with ID-shaped text). Also fixes the review's
CLAIM 9: the warning said IDs were "marked complete" when a ghost range marks
nothing — it now says "selected".

Side effect: the two false NEGATIVES the same review found are now covered —
`RANGE-01 through RANGE-05` and a parenthesised `(plus RANGE-02..RANGE-05)`.

NOT YET DONE (see the handoff prompt): regression tests for the four false
positives, the two new true positives, and the #3697-4 tightening the review's
CLAIM 8 asked for (it currently filters on the warning's phrasing rather than
asserting silence). Verified so far: tsc clean, the 10 existing #3697 tests
green, and a 20-case standalone harness covering every case above.

* test(#3697): pin the v2 token-detector boundary end-to-end

Six new cases + two hardenings for the review findings against v1:

- #3697-1 gains the worded spaced range (`RANGE-01 through RANGE-05`) —
  the operator set's `to|thru|through` arm was previously untested.
- #3697-5 (new): a tight range hidden inside balanced parentheses
  (`RANGE-01 (plus RANGE-02..RANGE-05)`) warns, names the range token,
  and ticks exactly RANGE-01 — the paren shave must not hide it.
- #3697-4 gains the four false-positive classes a free-text detector
  produced: numeric estimate (`- 3 points`), date annotation, em-dash
  citation with trailing period (`— locked per ADR-7.`), and nested
  parenthetical citations.
- #3697-3 and #3697-4 now assert the ENTIRE warnings channel is empty,
  not that one phrase is absent — a re-worded over-warning cannot pass.

Negative control: against the merge-base with its lib rebuilt, all 6
defect tests fail and all 10 controls pass.

* fix(#3697): close round-2 review findings — annotation false positives

Round 2 of the adversarial review (against 822a72a04) refuted five
claims; this closes the false-positive class and the cheap misses:

- R1's bare-hyphen arm now demands a full ID on BOTH sides
  (`REQ-01-REQ-05`): `LETTERS-\d+-\d+` is also a date-like annotation
  (`FY-2026-08`) and a sub-numbered ID, and warning on those is the
  expensive class. Tight hyphen shorthand with a live selection is the
  disclosed false negative; at zero selection R3 still catches it.
- R2 requires the endpoint pair to imply an INTERIOR (same prefix,
  gap > 1): `REQ-02 - REQ-03` selects both endpoints and can drop
  nothing, so an annotation hyphen between adjacent IDs stays silent.
- R3 skips placeholder-led lines: `None (per ADR-7)` is a declared-empty
  line citing its rationale, not unparsed residue.
- Token shave: quotes/backticks now shaved from alphanumeric tokens
  (`` `RANGE-02..RANGE-05` `` warns); punctuation-only tokens get a
  bracket-only shave so `(..)` surfaces its operator.
- 256-char token cap bounds the quadratic unanchored substring test.
- Warning text mentions range expansion only when a range rule fired.

Tests: 6 new cases (22 total). Negative control against the merge-base:
8 defect tests fail, 14 controls pass.

* fix(#3697): close round-3 review findings — half-spaced ranges, cross-prefix annotations, markdown wrappers

Round 3 of the adversarial review (against 2eb92dd0e) refuted four
claims; this closes them:

- Half-spaced ranges (`REQ-01 -REQ-05`, `REQ-01- REQ-05`) split at the
  tokenizer before R1's `\s*` can see them and under-selected silently.
  A glued-fragment rule warns when an operator is glued to a full ID
  with an ID-shaped neighbour on the open side and the endpoint pair
  implies an interior.
- Cross-prefix pairs around a separator no longer read as ranges:
  `REQ-02 - (ADR-7)` and `REQ-02 (...) (ADR-7)` are annotations, and
  real ranges are same-prefix by nature. `impliesInterior` now returns
  false on prefix mismatch and computes the gap with BigInt (parseInt
  lost precision past 2^53).
- The token shave now removes markdown emphasis markers and curly
  quotes, so `**None** (per ADR-7)` reaches the placeholder gate and
  `**RANGE-02..RANGE-05**` reaches R1.
- The unanchored-substring cap rises to 2048 (a markdown-link range
  with a long URL cleared 256); the anchored range regexes scan
  linearly and drop their cap.

Tests: 6 new cases (28 total; 377/377 file-wide). Negative control
against the merge-base: 11 defect tests fail, 17 controls pass.

* fix(#3697): round-4 review finding — word operators excluded from glued-fragment rule

`TOREQ-05` is a valid prefix-agnostic REQ-ID, and the glued-fragment
rule read it as `to` + `REQ-05`, warning on the canonical two-ID list
`REQ-01, TOREQ-05`. Glued fragments are now SYMBOL-operator-only
(`..`+, ellipsis, dashes): a word operator glued to an ID is an ID,
not a range spelling.

Tests: word-operator-prefixed ID control (misparse channel silent; the
fixture's ghost-ID warning legitimately fires, so the whole-channel
assertion stays with the registered controls) and an underscore-wrapped
tight-range defect case. 30 targeted cases; 379/379 file-wide; negative
control: 12 defect tests fail on the merge-base, 18 controls pass.

* fix(#3697): round-5 review findings — trailing word-op glue, dot shave, honest wording

- The glued-fragment TRAILING arm takes the word operators back: an ID
  must end in digits, so `REQ-01through` can never be an ID — the
  round-4 TOREQ collision was leading-arm-only, and symbol-only on both
  arms lost the `REQ-01through REQ-05` typo class.
- A trailing run of 2+ dots survives the punctuation shave: `REQ-01..`
  is a glued range operator, not sentence punctuation, and the shave
  was silently eating the `REQ-01.. REQ-05` form.
- The warning now says the line "could not be parsed as" a
  comma-separated REQ-ID list: `**REQ-01**, **REQ-05**` IS such a list
  — the selector just cannot parse decorated tokens — and a warning
  that misstates the input teaches readers to distrust it.

Tests: two new trailing-glue defect cases (32 targeted; 381/381
file-wide). Negative control: 14 defect tests fail on the merge-base,
18 controls pass.

* test(#3697): use t.after for cleanup per CONTRIBUTING test ruleset

CONTRIBUTING bans try/finally inside test bodies (it masks failures);
the approved shape is `t.after(() => cleanup(tmpDir))`. All seven
converted tests are this PR's own additions; the file's pre-existing
instances are untouched.

* chore(#3697): add changeset fragment for the Requirements-line under-selection warning

changeset-lint fails on this PR (fail_missing_fragment): src/phase.cts is a
user-facing surface and the branch carried no .changeset/*.md. Adds the Fixed
fragment via `npm run changeset -- --type Fixed --pr 3744`, symptom-led per
the house format, with the (#3697) backlink.

* refactor(#3697): extract the Requirements-line detector to a testable surface

Round-3 review Blocker 1 requires a fast-check property test over this
detector (`RULESET.TESTS.property-based-testing`: modules implementing
parsing contracts must include at least one), and Blocker 2 requires
limit-1/limit/limit+1 fixtures on its 2048-char token cap
(`RULESET.TESTS.boundary-coverage.fixtures`). Neither is expressible while
the logic is a closure inside `cmdPhaseComplete`: every existing #3697 test
reaches it by spawning the CLI, and a property test cannot pay a subprocess
per generated case.

So the selector and the three detection rules move to module scope as
`analyzeRequirementsLine` (pure, exported) plus
`formatRequirementsLineWarning`, and `cmdPhaseComplete` calls them. This
commit changes NO behaviour: `tests/phase.test.cjs` is untouched here, and
the pre-round suite passes against it unmodified (403/403).

Two things the move makes explicit rather than incidental. The selector and
the detector tokenize the SAME line DIFFERENTLY — the selector strips only
`[` and `]`, the detector also shaves quotes, emphasis and trailing sentence
punctuation — and that gap is deliberate: it is why `ADR-7)` is not selected
while `ADR-7` is still nameable in a warning. They now sit adjacent with the
reason written down, so they cannot drift apart silently.

And the stale citations in the moved comment are corrected. It pointed at
src/phase.cts:833,920,1078 for the `**Requirements**: TBD` seeds, which had
drifted to 1132/1237/1413, and at `templates/roadmap.md:32`, which is
`gsd-core/templates/roadmap.md:32`. Both are now anchored by content.

* fix(#3697): stop the warning claiming a misparse that did not happen

Round-3 review Major 3 and Minor 4. Both are the same defect: the warning
asserted more than the evidence supported.

MAJOR 3 — a correct comma list such as `RANGE-01, RANGE-02 — RANGE-05
deferred` warned "could not be parsed ... Range forms are not expanded;
rewrite the line". Reproduced: it selects RANGE-01, RANGE-02 AND RANGE-05,
i.e. every ID written on the line. Nothing was dropped, and the pinned
control only stayed silent because its pair was ADJACENT (gap == 1), so the
control was passing by accident of the fixture rather than by the rule.

The obvious fix — go silent — is not available. `RANGE-02 — RANGE-05` as a
range and as an annotation separator are textually identical, and no
token-level rule separates them; staying quiet re-opens the exact silent
under-selection #3697 is about. Deciding the ambiguity by assertion in
either direction is wrong. So it is DISCLOSED: the warning now has two
channels, chosen by whether any ID-shaped token was actually left unselected
(`droppedIdShaped`).

  * something was dropped (tight range, glued fragment, inert residue)
    -> "could not be parsed as a comma-separated REQ-ID list", as before.
  * nothing was dropped (only the spaced-operator rule fired)
    -> "contains what reads as a range between two cited REQ-IDs", stating
    both readings and saying explicitly that an annotation separator means
    the line is already correct.

This retires the "could not be parsed" wording for the four #3697-1 spaced
cases too, and that is a deliberate expectation change rather than a fix
counted twice: those lines never failed to parse either. They still warn,
still name the selected IDs, and still assert the endpoint-only marking is
unchanged; #3697-1 now also asserts the misparse channel stays SILENT.

MINOR 4 — `Deferred (see ADR-7)` reported `Unparsed text: ADR-7`, naming a
citation as requirement content it had failed to read. The trigger is
correct and stays: #3697's acceptance criterion asks for a warning "when it
selects zero IDs from a line that is non-empty and is not the `TBD`
placeholder", and inferring placeholder-ness from arbitrary prose is the
free-text heuristic this detector exists to avoid. What was wrong is the
wording, so the non-range arm now says "ID-shaped text that was not
selected" and names the escape the author actually has (`TBD` / `None`).

Tests: #3697-9 (three spaced forms — must warn, must NOT claim a misparse,
must offer both readings) and #3697-10 (`Deferred (see ADR-7)`, `N/A
(tracked in ADR-12)` — must warn, must not say "Unparsed text", must not
diagnose a range, must name the placeholder escape).

Reversion control: reverting the ambiguous channel fails #3697-9 (3 named
tests); reverting the R3 wording fails #3697-10 (2 named tests).

* fix(#3697): cap every token predicate, complete the dash set, cover the boundary

Round-3 review Blocker 2 and Nit 6, plus one self-found finding. All three
are about the detector's own predicates, so they land together.

BLOCKER 2 — the 2048-char budget had no boundary coverage.
`RULESET.TESTS.boundary-coverage.fixtures` requires limit-1 / limit /
limit+1 for any budget parameter. #3697-B1 and #3697-B2 now exercise 2047 /
2048 / 2049 against BOTH predicate families the cap guards, and each asserts
its fixture's exact length before asserting behaviour, so a mis-built
fixture fails loudly rather than passing at the wrong size. Clause (d) of
that rule — an input pushed within reserve-distance of the limit — has no
referent here: this is a hard cap with no reserve constant beside it, and
the test comment says so rather than leaving the omission to be re-derived.

NIT 6 — the cap guarded only the unanchored ID-substring regex. The
anchored range regexes were left uncapped, justified by a comment asserting
they scan linearly. The finding is right that this is informational (they
are anchored; the input is a local ROADMAP.md), but an asserted property is
cheaper to enforce than to defend, so all three predicates now share one
`short()` guard. #3697-B2 is what pins it: at 2049 the anchored scan must
now decline to classify.

SELF-FOUND (RV4 guard-shape census) — the range-operator set is a list this
code fixes at author time over a domain that grows without it, so the round
owes a census of what the enumeration reaches.

  reached:     `..`+, U+2026, U+2013, U+2014, ASCII `-`, to/thru/through
  NOT reached: U+2010 hyphen, U+2011 non-breaking hyphen, U+2012 figure
               dash, U+2015 horizontal bar, U+2212 minus sign
  consequence: a range spelled with any of those is SILENTLY under-selected
               — #3697's own defect, in the code that exists to fix it

Those five close. They are the same operator at a different codepoint and
carry none of the ASCII hyphen's collision risk, because they are not the
REQ-ID separator: `FY-2026-08` is date-shaped only with ASCII hyphens, so a
U+2010 never reaches the ID shape. They therefore join the NOHYPHEN arm
beside `—` and `–`; the strict full-ID-both-sides shape the bare hyphen is
held to is untouched, and #3697-12 pins that.

Still NOT reached, declined with reason rather than left unstated: `→`, `~`,
`..=`, `..<`, `until`, and `up to` (two tokens, so never one operator
token). Each is a symbol or word with an independent non-range use between
two REQ-IDs — the over-warning class #2334 cost three rounds.

Reversion control: reverting the uniform cap fails #3697-B2 (limit+1);
reverting the dash set fails #3697-11 (5 named tests).

* test(#3697): add the fast-check property coverage the parser rule requires

Round-3 review Blocker 1. `RULESET.TESTS.property-based-testing` (CONTEXT.md)
requires modules implementing parsing contracts to carry at least one
fast-check property test asserting a domain invariant, and the round-2 diff
had zero occurrences of `fc.` across its +311 test lines. Five properties,
1,900 generated cases:

  P1  soundness of silence (boundary containment) — for ANY canonical comma
      list of well-formed REQ-IDs, the selected set EQUALS the written set
      and nothing warns. This is the #2334 over-warning invariant and the
      #3697 under-warning invariant asserted as one statement, over
      generated IDs rather than hand-picked ones. It generalises #3697-4b:
      a prefix beginning with a word operator (`TORANGE-05`) is an ID, and
      P1 covers that class rather than the single example.
  P2  completeness — a same-prefix pair with an interior between them,
      separated by any of the nine spaced operators, ALWAYS warns.
  P3  the #2334 invariant — an ADJACENT pair around a separator can drop
      nothing, so it stays silent however it is annotated.
  P4  totality + idempotency — total over arbitrary strings, deterministic,
      and the formatter agrees with the analysis on whether there is
      anything to say (a warn with no text, or text with no warn, is a
      channel that can go silent or noisy on its own).
  P5  containment — every selected ID is ID-shaped and appears verbatim in
      the input.

Honest scoping, since a property test is easy to overclaim: P1, P3, P4 and
P5 hold against the round-2 code as well as this one — they are regression
guards, not bug-finders, and their value is that the invariants are now
stated and generatively checked rather than implied by examples. P2 is the
one that would have failed before the dash enumeration was completed.

fast-check v4 removed `fc.stringOf`, so the ID-prefix tail is built from
`fc.array(...).map(join)` with the alphabet pinned to the selector's own
`[A-Z0-9]` class.

These live in tests/phase.test.cjs rather than a new
`phase.property.test.cjs`: `lint-test-file-count` caps a production module
at 2 test files and phase.cts is already at its allowlisted entry, so a new
file would trade one gate for another.

* docs(#3697): document the ROADMAP Requirements-line grammar

Round-3 review Minor 5 — the change adds net-new user-visible warning output
for a grammar constraint documented nowhere under docs/. `type: Fixed` is
docs-exempt so this does not block, but a warning about a rule the reader
cannot look up is not actionable, and that is worth fixing whether or not a
gate demands it.

Added as a subsection of `phase complete` in docs/CLI-TOOLS.md, beside the
existing SUMMARY artifact-check advisory it is a sibling of: the supported
comma-list form, why ranges are deliberately not expanded, that `TBD` and
`None` are the entire placeholder vocabulary, and what each of the two
warning voices means — including that the range/annotation one may be
reporting a line that is already correct.

Existing file rather than a new one, deliberately: docs/ carries generated
indexes and zh-CN / ja-JP trees, and a new top-level page invites a parity
or index gate this change has no reason to touch.

* fix(#3697): rule-scope the warning-channel discriminator

Self-found at the round's pre-push review, against the Major 3 fix two
commits back. That fix chose the channel from a LINE-GLOBAL question — "was
any ID-shaped token left unselected?" — while the rules that produce the
warning are not line-global. The two disagree as soon as the line carries an
ID-shaped token no rule fired on:

  `RANGE-01, RANGE-02 — RANGE-05 deferred per (ADR-7)`

`(ADR-7)` survives the selector's bracket strip, so the global test called it
a drop and sent the line to the assertive channel — putting the false "could
not be parsed ... rewrite the line" claim back on a correct line. That is
review finding Major 3 returning through a side door, and it directly
contradicts #3697-4, which pins a parenthetical citation as NOT unparsed
residue.

The discriminator is now rule-scoped: R2 is the only ambiguous rule, so the
ambiguous channel requires that R2 fired, that no other rule did, and that
every endpoint R2 fired on was actually selected. The last conjunct is not
redundant — the detector shaves brackets and the selector does not, so R2 can
fire on a `(RANGE-02)` that was never selected, and that IS a drop:

  `RANGE-01 (RANGE-02) — RANGE-05`   -> assertive, correctly

`droppedIdShaped` is replaced by `spacedRangePairs` (R2's hits, so the
channel can ask about the endpoints the rule fired on) and the
`rangeReadingOnly` verdict.

Reversion control: against the line-global rule, #3697-9b fails. #3697-9c
passes under both rules — there the dropped token IS the R2 endpoint, so the
two agree; it is a regression guard, not a bug-finder, and is recorded as
such rather than counted as a second control.

* fix(#3697): hold every dash to the strict range shape, not just ASCII

Self-found at the round's pre-push review, and it CORRECTS a claim made two
commits back. That commit widened the range-operator set by five Unicode
dashes and asserted they "carry none of the ASCII hyphen's collision risk,
because they are not the REQ-ID separator". That reasoning was wrong. The
collision is a property of the SHAPE — `PREFIX-\d+ <dash> \d+` is also a date
(`FY-2026-08`) and a sub-numbered ID (`API-2-01`) — and the shape does not
care which dash sits in the operator slot, because the ID's own separator is
still ASCII either side of it. Measured:

  RANGE-01 (target FY-2026-08)   silent   <- pinned by #3697-4
  RANGE-01 (target FY-2026‐08)   WARNED   <- same line, U+2010

So the widening reintroduced the #2334 over-warning class on a date
annotation. It also exposed that the inconsistency PREDATES this PR: U+2013
and U+2014 were already in the loose arm at ce71dd399, so the en- and em-dash
forms of that same date annotation warned before round 3 ever ran.

One rule for every dash: a tight range spelled with any of the eight must
carry a FULL ID on both sides, exactly as the bare hyphen already had to.
`..`, `…` and the word operators stay loose — no date or sub-number reading
exists between two numbers, so the strict shape would cost them coverage for
nothing.

The cost is a false negative, and it is one the design already accepts:
`RANGE-01, RANGE-02-05` is silent today, deliberately, and now
`RANGE-01, RANGE-02–05` is too. That removes an inconsistency rather than
opening a gap, and a bare `RANGE-02–05` still warns — it selects nothing, so
R3 catches it.

Tests: #3697-13 (date annotation AND sub-numbered ID silent for all eight
dashes), #3697-13b (full-ID tight range still warns for all eight),
#3697-13c (loose operators keep their numeric endpoint), #3697-13d (the
accepted false negative is symmetric, and the bare zero-selection line still
warns).

* fix(#3697): close the round's own pre-push review findings

An adversarial cross-AI review of this round refuted 4 of its 10 claims. All
four were real. Every fix below is to code THIS round introduced.

1. THE SOFT VOICE CLAIMED TOO MUCH (refuted CLAIM 1).
   `REQ-01, (REQ-02), REQ-03 — REQ-05` took the range-reading voice and told
   the author "the line is already correct and nothing needs to change" — while
   `(REQ-02)` had been dropped by the selector, which does not strip
   parentheses.

   The channel choice is still right, and deliberately so: `(ADR-7)` and
   `(REQ-02)` are the SAME shape, so routing on "was anything unselected?" puts
   the false "could not be parsed" claim back on a line carrying a citation —
   the misroute fixed two commits ago. No rule can adjudicate this; the author
   can. So the voice stops asserting the line is correct (it now speaks about
   the SEPARATOR, which is all it has evidence about), and BOTH voices gained a
   factual clause naming ID-shaped text the selector skipped, with the reason
   (brackets are not stripped) and no verdict attached.

2. THE CAP SILENCED A LINE THAT USED TO WARN (refuted CLAIM 2).
   A 2049-char range token warned before this round and went silent after it:
   the "uniform cap" commit bounded the predicate and, with it, the warning.
   That is #3697's own defect, introduced by the fix for a nit.

   The cap bounds the WORK, not the warning. An over-cap token carrying `-` is
   now recorded as unclassified (a linear `includes`, never the unanchored
   regex the cap exists to keep off it) and gets its own voice: "could not be
   checked ... the REQ-ID selection on this line is unverified". Unclassified
   is reported, never treated as clean.

3. THE CAP WAS NOT UNIFORM (review MISSED finding).
   R2 capped the operator token but not its neighbours, so
   `<2049-char ID> .. <2049-char ID>` still ran REQ_ID_SHAPE_RE and BigInt over
   both endpoints unbounded. The glued rule had the same hole. Every
   participant is capped now.

4. PROPERTY P5 WAS VACUOUS (refuted CLAIM 5).
   It drew from a bare `fc.string()`, which over 500 samples produced max
   length 10 and ZERO inputs containing a REQ-ID — the loop body never executed
   an assertion. A containment property that never contains anything is a green
   test measuring nothing. The generator now interleaves real IDs with noise
   and the property ASSERTS it saw them (>50/500), so it can never silently go
   vacuous again. The free-form coverage it was actually providing survives,
   honestly labelled, as #3697-P6.

   The same finding refuted this round's claim that P2 distinguishes pre-round
   behaviour: every operator P2 uses was already in the pre-round operator set.
   P2 is a regression guard, and its comment now says so.

Also: docs/CLI-TOOLS.md repeated the broken channel claim verbatim (review
MISSED finding) and is corrected with the code.

Tests: #3697-9d (soft voice names the skipped ID, never claims the line is
correct), #3697-9e (over-cap token reported as unclassified, still warns),
#3697-9f (R2 and the glued rule cap their neighbours). #3697-B1/B2 now key the
boundary on the PREDICATE's verdict with `warn` asserted true at every length —
asserting `warn === false` at limit+1 was itself finding 2.

* docs(#3697): describe the third voice and the dash rule

Follow-on to the review-findings commit: that commit corrected the docs' claim
about the soft voice but left two things the code now does undescribed.

- There are THREE voices, not two. The over-cap voice ("could not be checked
  ... unverified") arrived with the fix for the review's CLAIM 2 and had no
  entry.
- Dash spellings require a full ID on both sides, and `..` / `…` / the word
  operators do not. That asymmetry is deliberate and load-bearing —
  `PREFIX-<digits><dash><digits>` is date- and sub-number-shaped — so a reader
  hitting `REQ-01-05` and getting silence has no way to find out why. The
  accepted cost (`REQ-01, REQ-02-05` unreported, bare `REQ-02-05` still
  reported) is stated rather than left to be discovered.

Documentation only; no behaviour change.

* fix(#3697): close the continuation review's findings

A continuation of the same adversarial reviewer, run against the reworked
round, refuted 6 of 7 claims. Four were real defects in this round's own work
and are fixed here; the other two are answered rather than changed, below.

1. THE SKIPPED-TEXT CLAUSE WAS ON ONE VOICE, NOT BOTH (refuted CLAIM A).
   The previous commit's message said both voices gained it. Only the soft
   return appended it. The assertive voice now carries it too — and, because
   that voice already names range tokens and inert residue under its own
   clauses, the note is filtered to what those did not already name. A warning
   that says the same token twice is one readers learn to skim.

2. THE CLAUSE'S WORDING WAS FALSE (also CLAIM A).
   It read "brackets and parentheses are not stripped". Square brackets ARE
   stripped by the selector — `[REQ-01, REQ-02]` is the documented form — so
   only parentheses qualify. Corrected in the message and in docs/CLI-TOOLS.md,
   which had inherited the same error.

3. THE OVER-CAP RULE STILL SILENCED A LINE (refuted CLAIM B).
   `oversizedTokens` filtered on `includes('-')`, which misses an over-cap
   OPERATOR: `REQ-01 <2049 dots> REQ-05` warned before this round, R2 declined
   to classify it once capped, and nothing reported it. That is the exact
   regression the field was added to close, one input over. Any token past the
   cap now counts — what it contains is irrelevant when we could not read it.

4. AND THEN OVER-REPORTED ONE (review MISSED finding).
   With (3) in place, a 2049-character CANONICAL REQ-ID was selected by the
   uncapped, fully-anchored selector AND flagged "REQ-ID selection on this line
   is unverified" — a contradiction inside one warning. A token the selector
   took was examined end to end, so it is excluded.

Two findings are answered, not changed:

  CLAIM C — the selector's own `REQ_ID_SHAPE_RE.test` is uncapped. True, and
  deliberate: this round does not touch what gets MARKED, and the pattern is
  anchored at both ends with no nested quantifier, so it is linear. The claim
  that "all predicate paths are capped" was too broad; the DETECTOR's are.

  CLAIM E — `REQ-01, REQ-02<dash>05` is silent for every dash. That is the
  documented, deliberate cost of holding dashes to the strict shape, and it is
  symmetric with ASCII, which behaved that way before this PR. The reviewer is
  right that "without losing a range spelling that should be detected" was too
  strong; a bare `REQ-02<dash>05` still warns.

Tests: #3697-9g (clause on the assertive voice, no repetition, bracket claim
true), #3697-9h (over-cap operator does not silence the line), #3697-9i (a
selected over-cap ID is never called unverified). #3697-9f is rebuilt — its
first version used the SAME id twice, so R2 could not have fired even uncapped
and it proved nothing; it now uses endpoints with a gap and fails when the
neighbour cap is removed.

Reversion control: all four fixes fail a named test when reverted in isolation
(#3697-9h, #3697-9i, #3697-9g, #3697-9f).

* fix(#3697): scope the over-cap exemption to what could actually pair

A second continuation of the same reviewer, against the reworked round,
confirmed the two claims that matter most and refuted three. This closes the
one real defect; the other two are answered below.

CLAIM J / CLAIM K (one defect, found from both directions). The previous
commit exempted EVERY selector-accepted token from `oversizedTokens`, on the
reasoning that the selector is uncapped and anchored so it examined the whole
token. True of that token's SELECTION — and not the same as "no rule was
suppressed by it". Two over-cap valid IDs either side of `..` are both
selected, so both were exempted, and R2 is capped: a line that warned before
this round went silent.

That is the third appearance of one class in this round — the cap suppresses a
check, and the suppression is not reported. Each fix for it over-corrected in
the opposite direction, which is why the rule is now stated in terms of what
was actually suppressed rather than in terms of the token: an over-cap token is
exempt only when it was selected AND nothing beside it could have paired with
it into a range (no range operator, no glued fragment, no second over-cap
token). Everything else is unexaminable and says so.

Two findings are answered, not changed:

  CLAIM M — the reviewer demonstrated, with driven evidence, a contextual rule
  that catches `REQ-01, REQ-02-05` while leaving `FY-2026-08` and `API-2-01`
  silent: recognise `PREFIX-a<dash>b` only when another SELECTED id on the line
  shares that prefix. That refutes this round's claim that the strict-dash
  trade was FORCED, and the claim is withdrawn — it is a design choice. The
  choice stands for this PR: the conservative rule is what ASCII already did
  before #3697, adopting a new contextual heuristic unreviewed at the end of a
  round is how the last three defects in this round were made, and #3697 asks
  for a warning rather than better range inference. Named here so the
  alternative is on the record rather than lost.

  Docs MISSED — CLI-TOOLS said every token over 2,048 characters "is not
  classified at all" and warns. Selection is not bounded; only range detection
  is. Corrected.

Confirmed by the same pass, and worth recording because they are the PR's
load-bearing promises: a 20,000-input comparison of the pre-extraction selector
against HEAD found `mismatches=0` (nothing about which REQ-IDs are MARKED has
changed), and the uncapped selector regex was measured linear from 100k to 800k
characters.

Tests: #3697-9f now asserts the range case is reported rather than silent, and
#3697-9j pins the exemption's scope in both directions. Reversion control:
restoring the blanket exemption fails both.

* fix(#3697): warn on zero selection, as the acceptance criterion asks

`Deferred`, `N/A`, `Pending`, `TBA` and `-` selected no REQ-IDs and stayed
SILENT, while three shipped artifacts said they warned: `docs/CLI-TOOLS.md`,
the `placeholderLed` census comment, and the advice string the command emits
to the user. The asymmetry was the tell — `Deferred (see ADR-7)` warned,
because the citation supplied the ID-shaped residue R3 required, while bare
`Deferred` did not. The claim was written into three places and never
executed once.

This is also #3697's AC-1b/AC-4 verbatim: "warn when `citedReqIds.length ===
0` while the raw capture is non-empty and not `TBD`".

R3b keys on the SELECTION being empty, never on what the prose means, so it
adds no free-text heuristic. It is deliberately not gated on ID-shaped
residue the way R3 is, and the negative space is what settles that: all
fifteen #2334/#2339 fixtures are held silent by non-zero selection or by
`placeholderLed`, and not one of them by the ID-shape gate — measured, not
argued. The gate was buying no negative space while costing the acceptance
criterion.

`tokens.length > 0` keeps an empty line and a comment-only line silent: the
tokenizer strips `<!-- ... -->` before splitting, so the shipped template's
own comment cannot reach the rule.

Selection behavior is unchanged. This warns; it never invents an ID.

Also extracts `warn` to a named const (round 3 review Minor 3) — this commit
adds a disjunct to exactly that predicate, and in the return literal a later
reordering would be a TDZ ReferenceError rather than a reader-visible error.

Tests: #3697-14 (six zero-selection lines warn and tick nothing, and the
warning names the TBD/None escape), #3697-14b (five placeholder spellings
stay whole-channel silent), #3697-14c (comment-only line stays silent).
Fail-first controls: all six #3697-14 cases fail against the pre-fix tree;
-14b and -14c pass at both ends, which is correct — they pin silence the
widening must preserve.

* fix(#3697): name the REQ-ID a glued delimiter dropped

`RANGE-01; RANGE-02` selects only RANGE-02 and marks only RANGE-02, with
`requirements_updated: true` — #3697's own half-success failure mode, reached
by one wrong delimiter, and silent before this rule. It is the issue's AC-1a
("a warning whenever the line contains ID-shaped content that the tokenizer
did NOT select") at the shape most likely to be typed by accident.

Round 4 review rated this Major rather than Blocker on the ground that the
case is indistinguishable from a parenthesised citation, since `(ADR-7)` also
shaves down to a bare ID. At the RAW token level it is distinguishable, and
that is what makes the rule shippable: `REQ-01;` is shaved of a trailing
DELIMITER, `ADR-7)` of a citation wrapper. R4 keys on that shave class and
requires the token to sit outside any parenthetical.

Measured before implementing: 0 false positives and 0 false negatives across
21 probes, including all fifteen #2334/#2339 negative-space fixtures. A first
cut without the parenthetical test scored 3 false positives — every one of
them a colon inside a citation (`(see ADR-7: section 3)`) — which is why that
test is the rule's boundary rather than an optimisation.

Adds the delimiter census the module did not have. The range-operator domain
was already censused; the comma-substitute domain was not. Swept 26
spellings: exactly two produce a silent under-selection, `; ` and `: `. Every
other spelling either selects both IDs or selects none and already warns. The
review hand-listed the semicolon; the colon is the sibling that sweep found,
and it fails identically.

`rangeReadingOnly` now excludes an R4 hit — the ambiguous voice claims nothing
was dropped, and must not speak for a line where something demonstrably was.

Tests: #3697-15 (four delimiter shapes warn, name EVERY dropped ID, and tick
exactly the unchanged selection), #3697-15b (three citation forms stay
whole-channel silent). Fail-first control: all four #3697-15 cases fail
against the previous commit's tree; -15b passes at both ends, pinning the
boundary the widening must not cross.

* fix(#3697): give the Requirements-line warning a stable machine kind

The warning's kind existed only in the prose of its message, so every consumer
and every test had to regex an English sentence — and rewording a message
silently un-asserted the tests that pinned it. Round 4 review Major 3.

The repo already had the settled seam for exactly these semantics.
`CONTEXT.md` records `diffLiveConfig` emitting `kind:'unverified'` for a
truncated scan, which is precisely this module's third voice; and
`WAVE_CLEANUP_WARNING` in `src/worktree-safety.cts` carries codes for the same
reason. ADR-3473 Decision 3 ("failure is a value") points the same way.

`formatRequirementsLineWarning` now returns `{ code, message }` instead of a
bare string, which also settles round 4 Nit 3 — `null` still means CLEAN, a
legitimate value, but the success arm is no longer a naked string one field
away from the shape the ADR standardises on.

The kind is carried ALONGSIDE the prose, never instead of it. `warnings[]` is
a documented `string[]` in `phase complete`'s JSON output, rendered by
execute-phase.md's "If has_warnings is true" step, so re-typing its elements
would be a breaking output-contract change for a shipped command. The code is
emitted as its own additive `requirements_line_warning` field, absent
entirely when the line is clean.

Vocabulary, exported so tests key on it rather than on string literals:
`req-line-misparse`, `req-line-range-reading`, `req-line-unverified`.

Tests: channel ROUTING in #3697-9/-9b/-9c/-9d/-9e/-9g/-10 now asserts the code;
message-content assertions stay where the user-visible wording is itself under
test. #3697-16 pins the code end-to-end through the CLI's JSON for four line
shapes and asserts warnings[] is still a string[]; #3697-16b pins that a clean
line emits no kind at all, because a field present on every run carries no
information. #3697-P4 holds kind-and-message-appear-together and
kind-is-in-the-declared-vocabulary over arbitrary input, so a channel added
later cannot ship without one.

* test(#3697): pin the divergence against the second parser of the same line

CLAUDE.md, KNOWN DEFECTS & ANTI-PATTERNS: "Generative Fix Divergence: when
sharing constants/arrays/parsers between parallel surfaces, add a parity
assertion test that fails if they diverge." Round 4 review Major 2.

`normalizePhaseReqIds` (src/gap-checker.cts) parses the SAME ROADMAP
`**Requirements:**` value — its own docblock says callers "may pass the
roadmap value through verbatim" — and diverges on four axes. Measured, not
inferred:

  line                                    phase complete      gap-checker
  RANGE-01..RANGE-05                      []                  5 IDs
  None (per ADR-7)                        []                  ["ADR-7"]
  (REQ-02)                                []                  ["REQ-02"]
  REQ-01a                                 []                  ["REQ-01a"]
  REQ-01, REQ-02                          both                both

This pins the divergence rather than removing it, which is the review's
second option and the correct one here: unifying the two would change what
`phase complete` MARKS, and "the ledger-writing set is byte-identical to base"
is the one invariant this PR holds fixed. Every axis is now asserted in BOTH
directions, so drift on either side fails here instead of widening silently.

The range axis is a DELIBERATE disagreement and is labelled as such — #3697
declines range expansion in terms ("I am not asking for range syntax to be
supported") while gap analysis adopted it under #1269.

The placeholder axis is the one worth reading twice: `None (per ADR-7)` is a
declared-empty line to `phase complete`, which reads the lead token, and a
one-requirement line to gap-checker, which strips parentheses first so the
citation survives its ID-shape filter. That is a citation being reported as a
requirement.

#3697-17b states the cost concretely: one line, five requirements in scope to
gap analysis and zero to phase complete. This PR is what makes that
contradiction visible, by finally giving the silent side a voice.

* fix(#3697): stop the skipped-text rider reporting a date, and close the 4b channel gap

Two round 4 review minors, both about a warning saying something it cannot
support.

MINOR 2 — false rider content. `REQ_ID_SUBSTRING_RE` is unanchored, so
`FY-2026-08` matches as `FY-2026` and lands in `unselectedIdShaped`.
`REQ_RANGE_TOKEN_RE`'s entire strict-dash arm exists to keep that shape
silent, and #3697-4 pins `RANGE-01 (target FY-2026-08)` as producing no
warning at all — but whenever some OTHER rule fired on a line that also
carried a date annotation, the rider told the author to "check whether any of
it is a requirement" about a date. Not a false warning, since the line was
warning anyway; false CONTENT, in the #2334 voice, through the side door.

Filtered at the MESSAGE rather than in the analysis: `unselectedIdShaped`
stays a faithful record of what the selector skipped — it is documented as a
fact that never routes — while the user-facing clause declines to assert
requirement-ness about a shape the design already ruled unadjudicable.
#3697-18b is the other half, so the filter cannot become a silencer: a
genuinely dropped REQ-ID is still named.

MINOR 1 — `#3697-4b` asserted only that the ASSERTIVE channel stayed silent,
so a regression routing `RANGE-01, TORANGE-05` into the AMBIGUOUS channel
would have passed. Whole-channel silence is not available on that fixture (the
pre-existing ghost-ID warning legitimately fires on the unregistered
`TORANGE-05`), so the precise assertion is that no Requirements-line warning
of ANY kind was emitted. The machine code added earlier in this round is what
makes that statable; before it, "both channels" could only have meant a second
prose regex.

* docs(#3697): record the Requirements-line seam in CONTEXT.md

CLAUDE.md names the CONTEXT.md glossary as a PR gate, and
`get_cochange_context(src/phase.cts, 45d)` ranks CONTEXT.md 4th at 25
co-changes — above src/init.cts and src/roadmap.cts. This PR introduced a
named seam, three warning kinds, a bound, a rule taxonomy and a deliberate
cross-parser divergence, and recorded none of it. Round 4 review Major 4.

The precedent is explicit rather than inferred: the directly analogous seam
is already there as `LIVE-CONFIG.GUARD.SEAM.truncation`, including its bound
and its boundary obligation — and that entry is the one this module's third
voice was modelled on.

Eight predicates, in the machine-oriented section beside it:

  .module                    the two exported functions and the code vocabulary
  .selector-identity         citedReqIds is byte-identical to base and is the
                             only thing reaching the ledger — a change to what
                             phase.complete MARKS is outside this contract
  .rules                     R1 / R2 / R2' / R3 / R3b / R4 / over-cap
  .kinds                     the three codes, and why they ride beside
                             warnings[] rather than inside it
  .cap                       2048, neighbours included, and the boundary rule
  .placeholder               the gate that actually holds the negative space
  .census-domains            both open domains with their NOT-reached members
  .gap-checker-divergence    the four axes, pinned not unified

The changeset type is `Fixed`, which exempts this PR from the docs/
co-change requirement — but the glossary gate is separate from that
exemption, and the 2048 cap in particular is a machine-canon-shaped fact
that until now existed only inside a source comment.

`docs/CONTEXT-INDEX.json` regenerated (269 predicates); lint:generated-sync
confirms all six targets in sync.

* docs(#3697): document what the command now does, in one changeset sentence

DOCS. The grammar section predated this round's two new rules, so it
under-described the behaviour it exists to make lookup-able:

- The placeholder paragraph enumerated three words; the rule is a DEFAULT.
  Any wording that selects no REQ-IDs warns, and the placeholders are matched
  as the LEAD token, so `None (per ADR-7)` and `**None**` are declared-empty
  too. The comment-only line is called out, because "any other wording" would
  otherwise read as covering the shipped template's own `<!-- ... -->`.
- The comma rule was implicit. `REQ-01; REQ-02` marks only REQ-02, and it is
  the quietest way to lose a requirement on this line — `requirements_updated`
  reads `true` either way — so it gets its own paragraph, with the
  parenthetical exemption stated beside it.
- The machine kind is documented where a consumer would look for it, with the
  instruction to key on the kind rather than the wording.
- The skipped-text note no longer implies it reports date shapes; it
  deliberately does not, and silently omitting that left the doc promising the
  behaviour this round removed.

CHANGESET (round 4 review Minor 4). CONTRIBUTING.md's format is
`**<Bold user-visible change>** — <symptom-led explanation>.` and both
canonical examples are one sentence; this fragment ran three. Now one, and
covering what the round actually delivers rather than only the range shape it
started from.

* test(#3697): keep phase.test.cjs off the docs-guard exemption fingerprint

A comment added earlier in this round named `docs/CLI-TOOLS.md` by path. The
docs-guard exemption ratchet (#3753 FIX 3) fingerprints literal `docs/`
references in exempt test files and fails when a new one appears, so that
comment turned four green gates red — `ci-docs-guard-registry` and the
registration lint — for a file that reads no documentation at all.

Caught by diffing the full suite's failing-name set against the same suite run
at `upstream/next` in a probe worktree: 33 of 37 failures reproduce at base
(install / config-home / shadowing tests under the sandbox HOME), and exactly
these 4 did not.

Rephrased rather than baselined. Adding the path to
DOCS_GUARD_EXEMPT_DOCS_PATHS is the sanctioned response when a test genuinely
starts READING a new docs path — the violation text asks the author to
re-confirm the exemption still holds. Nothing here reads documentation; the
guard matched prose. Baselining would have recorded a coupling that does not
exist and made the next reader wonder what phase.test.cjs does with
CLI-TOOLS.md. The comment still names where the contract is written, just
without planting a path string.

* fix(#3697): close four defects this round's own pre-push review drove

An adversarial cross-AI review of this round, run before the push, returned 6
CONFIRMED and 4 REFUTED. Every refutation was driven against the built tree,
and every one was a shape the author had not probed — the rules were correct
across the probe set and wrong just outside it.

(1) R4 FALSE POSITIVE, and it is the #2334 over-warning class arriving through
the rule added to close a different hole. `REQ-01, see ADR-7: section 3` fired:
`ADR-7:` is the same shave class as `REQ-01;`, and the parenthetical test does
not reach a BARE citation. The FP probe that scored this rule 0/0 only ever
tested the parenthesised form.

Fixed by requiring the dropped id's prefix to agree with a SELECTED id — the
module's own idiom, not a new heuristic: `reqEndpointsImplyInterior` already
demands an agreeing prefix for the same reason. Cost, stated in the census: a
dropped id whose prefix is on no selected id (`REQ-01, FOO-02: x`) stays
silent. Same trade the strict-dash rule takes — under-report a rare shape
rather than over-report a common one. Pinned as a declared blind spot by

(2) R4 FALSE NEGATIVE, on the DOCUMENTED form. `[REQ-01; REQ-02]` dropped
REQ-01 silently: the selector strips square brackets and R4's raw scanner did
not. The bracket spelling the shipped template recommends was the one shape the
rule could not see.

(3) The rider filter suppressed a REGISTERED requirement. `API-2-01` is a legal
requirement id — gap-checker's `parseRequirements` accepts it from
REQUIREMENTS.md — so a `\d+-\d+` filter hid a genuinely dropped requirement
behind a rule meant only to hide dates. Narrowed to a four-digit year segment.
The earlier #3697-18 case asserting `API-2-01` should be suppressed is REMOVED,
and the removal is recorded in place: its premise was refuted, it was not
inconvenient.

(4) An INVISIBLE line warned. A lone U+200B carried a token to the parser while
reading as empty to the author, so R3b fired with nothing on screen to explain
it. Zero-width and format characters are now stripped — stripped rather than
treated as delimiters, because splitting on one would fabricate two fragments
out of one ID.

Also corrects the documentation the same review found overstated: the line is
split on commas AND whitespace, and the ID shape is matched case-insensitively,
so `REQ-01 REQ-02` and `req-01, req-02` both select and neither warns. That was
pre-existing selector behaviour; this round is the one that asserted the docs
were true of it.

Tests: #3697-19 (four invisible-only shapes), -19b (embedded zero-width is
stripped, not split on), -19c (three citation forms), -19d (both bracket
spellings), -19e (both halves of the rider boundary), -19f (the declared blind
spot), -19g (the two documented tolerances). 500 tests in phase.test.cjs, 0
failures; lint:ci clean.

* fix(#3697): generalise the drop rule, and stop the invisible fix hiding a drop

The pre-push review's continuation refuted six of seven follow-up claims. The
first one is the one that mattered: the invisible-character fix committed in
c7dce173a INTRODUCED #3697's own defect. Stripping zero-width characters from
the detector wholesale made `REQ-01<ZWSP>, REQ-02` go SILENT — the selector
really does drop REQ-01, and the strip removed the only evidence of it. The
test written alongside asserted the tokens and the empty R4 result and never
asserted `warn`, so it DOCUMENTED the bug rather than catching it; that
omission was the reviewer's own MISSED finding.

An invisible is two different questions about one character, and the fix is to
stop conflating them: absence-of-content for the empty test, DECORATION on a
token for the drop rule. Neither is a reason to delete it from the line.

R4 is generalised accordingly, because the continuation drove four more shapes
a trailing-delimiter-only regex could not see — `REQ-01 ;REQ-02`,
`REQ-01 :REQ-02`, `**REQ-01;** REQ-02`, the backticked form — plus
`**REQ-01**, REQ-02`, where emphasis alone defeats the selector. These are one
class: decoration on a token the selector then cannot take. One rule, not four
patches; patching them individually is how a list stays short and wrong.

PARENTHESES ARE NOT DECORATION, and the suite caught me learning that: shaving
them made `REQ-01, (REQ-02), REQ-03 — REQ-05` report a glued delimiter that was
never there and broke #3697-9d's channel routing with it. A parenthesis is this
rule's citation marker.

The rider stops adjudicating an undecidable shape. `API-2-01` is a legal
requirement id and `API-2026-08` is too, while `FY-26-08` and `FY-2026-08-15`
are dates — no regex separates them, and both filters this round tried scored a
miss in each direction. It now NAMES the token and states the ambiguity, which
is the same thing the two warning voices already do about a range separator.
Filtering hides a real dropped requirement; reporting it bare asks the author
whether a date is a requirement; saying "this may equally be a date" does
neither.

The census and the docs are corrected to what the code does, including the part
that is NOT complete: the prefix gate does not stop a citation that SHARES a
selected prefix (`ADR-01, see ADR-7: sec 3` fires), and nothing at token level
separates that from a real drop. A prose heuristic on "see" is the free-text
detector this module exists to avoid, so the honest move is to say so.

CLAIM 17 — the invariant that actually matters — came back CONFIRMED on a
20,000-run fast-check property over arbitrary Unicode: `citedReqIds` is
identical to upstream/next's for every input, and marking is untouched.

513 tests in phase.test.cjs, 0 failures; lint:ci clean.

* fix(#3697): gate the drop rule on evidence, and stop an unmatched paren swallowing the line

Third pass of the round's own pre-push review, scoped to regression-hunting
rather than further polish. Three findings, all driven, all mine.

R4 OVER-WARNED on markdown styling. `REQ-01, see **REQ-7** for context` claimed
a dropped requirement: the previous cut treated any shaved decoration as
evidence, and emphasis is not evidence. Nothing separates that line from
`**REQ-01**, REQ-02` meaning to list one, so the rule now requires a positive
signal — a glued `;`/`:` (a list separator was INTENDED) or an invisible (the
token is CORRUPTED; nobody types one on purpose). Emphasis alone falls back to
the skipped-text rider, which names the id without asserting a drop, exactly as
`(REQ-02)` is handled. That is the #2334 class caught one cut before shipping.

R4 UNDER-WARNED on `**REQ-01**; REQ-02` — one shave pass cannot reach a wrapper
sitting behind a delimiter. Shaves to a stable point now.

The range OPERATOR lost its invisibles handling. `REQ-01 <ZWSP>..<ZWSP> REQ-05`
went silent, because the previous commit removed the invisible strip from BOTH
the tokenizer and R4 when only R4's was wrong. An invisible is two questions
about one character: for the classification rules it is noise and is stripped
from the token; for the drop rule it is the evidence and must survive on the
raw line. Stripping in both places hid a dropped id; stripping in neither hid a
range. The reviewer's MISSED finding named the missing control — regression
tests covered invisibles inside ids and not beside operators — and #3697-19i is
that control.

UNBALANCED PARENTHESES swallowed the line. `REQ-01, (note REQ-02; REQ-03`
reported nothing: a running-depth counter left the unclosed `(` open through
end-of-line, so every genuine drop after it inherited citation immunity. A
parenthesis confers that immunity only as part of a MATCHED span now — an
unmatched one is a typo, not a citation.

CLAIM 22 re-confirmed on a fresh 20,000-run property over arbitrary Unicode:
`citedReqIds` identical to upstream/next, marking untouched, warnings appended.

522 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite carries zero
head-only failures against a probe worktree at upstream/next.

* fix(#3697): make delimiter ADJACENCY the rule, and delete matched citations outright

Fourth and final pass of the round's own pre-push review. Three findings, and
they shared one root cause, so this is a narrower rule rather than a longer list
of shapes.

TOKEN-WIDE PAREN IMMUNITY LEAKED. `REQ-01, REQ-02;(note) REQ-03` is a single
whitespace token, so a matched parenthetical inside it conferred immunity on the
`REQ-02;` sitting OUTSIDE the parens, and the drop went silent. Matched spans
are now deleted from the line outright — which states what is actually meant,
that for this rule a citation is not on the line — and an UNMATCHED paren is a
typo that confers nothing. That also retires the running-depth counter whose
previous bug was the mirror image: an unclosed `(` swallowing the rest of the
line.

DECORATION WAS TESTED TOKEN-WIDE, so `REQ-01, see **REQ-7**; next topic` was
reported as a dropped requirement. It is a citation with sentence punctuation.
The rule is now ADJACENCY: styling is stripped, then the `;`/`:` must be
touching the id. `REQ-01;`, `;REQ-02` and `**REQ-01;**` qualify;
`**REQ-01**;` does not, because outside the styling that character is
punctuation. An invisible needs no adjacency test — nobody types one on
purpose, so anywhere in the token it is corruption rather than intent.

`**REQ-01**; REQ-02` therefore goes silent, and the test row asserting
otherwise is inverted rather than deleted quietly: it was added one commit ago
on the reasoning this pass refuted, and nothing distinguishes it from
`see **REQ-7**; next topic`.

Worth recording plainly: three successive cuts of this rule fired on a
citation, and each fix was a narrower definition of EVIDENCE, never a longer
list of shapes. The list-lengthening instinct is what produced the bug each
time.

The review's last MISSED finding named the missing control — the paren tests
all surrounded matched spans with whitespace, so none covered a span sharing a
token with an id outside it. #3697-19j carries both directions now.

526 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite zero
head-only failures against a probe at upstream/next; the invariant that
`citedReqIds` is identical to upstream re-confirmed on 20,000 arbitrary
Unicode inputs.

* fix(#3697): state R4's real boundary, and stop the over-cap voice masking a drop

Two defects, both found by this round's own pre-publication body claim-audit.

1. A DEMONSTRATED drop was discarded by the unverified voice. On
   `REQ-01, REQ-02: <2049 chars>` the analyzer names REQ-02 in
   delimiterDroppedIds and the formatter then reported `req-line-unverified`,
   whose message never mentions it — the one actionable finding masked by the
   token beside it. The over-cap channel now excludes a line carrying an R4
   hit, exactly as rangeReadingOnly already did and for the same reason: that
   voice's whole claim is that nothing could be checked, and R4 has already
   checked something. The assertive channel still carries the over-cap rider,
   so nothing about the cap is traded away. Pinned by #3697-19l, which fails
   against the pre-fix build and nothing else does.

2. Three shipped artifacts asserted behaviour the code does not have — the
   same class as this PR's round-4 blocker, re-committed. CONTEXT.md's rules
   predicate, the CLI tools reference, and the warning's own advice string all
   listed markdown emphasis as an R4 trigger. It is not: styling is shaved
   BEFORE the test and tolerated around an id, never a trigger on its own, so
   `**REQ-01**, REQ-02` and `**REQ-01**; REQ-02` are both silent. The trigger
   is exactly a glued `;`/`:` or an embedded invisible.

   The census predicate was wrong in a second way. Its 26-spelling separator
   sweep found only `;` and `:` because the sweep was SYMMETRIC-ONLY and
   therefore biased: one-sided attachment drops silently for every punctuation
   outside the set — `/ | & + . > \` and the full-width and non-ASCII forms
   `; , ؛` all measured silent. The domain is wide open and R4 covers two
   characters of it. Said plainly in all three places rather than widened
   here: every previous widening of this rule first fired on a citation, so it
   is not done blind at the end of a round.

Both blind spots are now PINNED as tests (#3697-19m styling-only, #3697-19n
one-sided separators) so the documents and the code cannot drift apart again —
which is what the round-4 blocker asked for.

tests/phase.test.cjs: 542 tests, 542 pass, 0 fail, 0 skip. lint:ci rc=0.
Both CONTEXT-INDEX consumers regenerated.

* fix(#3697): re-sweep the separator census properly, and say what it really found

The round-4 census in src/phase.cts concluded "exactly two — `; ` and `: `"
from a 26-spelling sweep. That conclusion was forced by how the sweep was
built, not by the code: it swept the ONE-SIDED form (`REQ-01; REQ-02`) for the
semicolon and colon, and only the BARE and SYMMETRIC forms (`|`, ` | `) for
every other separator. Different members of the domain were tested in
different shapes, so no other answer was reachable. Caught by this round's
pre-publication claim-audit of the response comment, reading the census
comment against its own swept list.

Re-swept fully crossed and driven through the built artifact: 21 separators x
{bare, trailing-space, leading-space, both-spaces} = 84 combinations. 26 select
both ids, 24 under-select and already warn, and 34 UNDER-SELECT SILENTLY. All
34 are one shape — a separator glued to exactly one of the two ids, e.g.
`REQ-01/ REQ-02` or `REQ-01 /REQ-02` — for every punctuation except `,` and
the `;`/`:` that R4 covers.

So R4 covers TWO CHARACTERS of a wide-open domain. That is now what the census
comment, the CONTEXT.md census-domains predicate and the CLI tools reference
all say. The set is deliberately not widened here: three successive cuts of
this rule fired on a citation, and a fourth at the end of a round with no
adversarial pass is how each of those got in.

Second false passage in the same block: styling-only decoration was described
as "left to the skipped-text rider, which names the id". A rider only exists
inside a message, and a message only exists once some rule sets `warn` — so on
a line where nothing else fires, `REQ-01, **REQ-02**` is wholly silent.
Describing it as handled reads as coverage. #3697-19m already pins the silence.

so the test matches the documented claim.

tests/phase.test.cjs: 552 tests, 552 pass, 0 fail, 0 skip. lint:ci rc=0.
Both CONTEXT-INDEX consumers regenerated.

* chore(#3697): regenerate both CONTEXT-INDEX.json after rebasing onto next

Rebased onto next @ f4fefb0be. The two generated indexes conflicted on the
replay and were resolved by regenerating, not by hand-merging:
`npm run gen:context-index` for docs/CONTEXT-INDEX.json and
`node examples/dynamic-context-management/gen-context-index.cjs --write` for
the example's copy. Against next, each now differs only in the eight
PHASE.REQ-LINE.SEAM.* predicates this PR adds (plus the PHASE class and the
count); every base-side change (the SEAM.* predicates, the ADR-3942 value
rewrites) is carried. `lint-example-parser-parity` and
`gen-context-index --check` both pass.

* fix(#3697): an over-cap token outranks the ambiguous range voice

Round 7 review, Minor 1. `rangeReadingOnly`'s guard conjunction checked
`hasGluedRangeFragment`, `inertIdShaped` and `delimiterDroppedIds` but not
`oversizedTokens`, so a line carrying a clean, fully-selected spaced range
*and* an unrelated token past the 2048-char scan cap was coded
`req-line-range-reading` — a code CONTEXT.md's PHASE.REQ-LINE.SEAM.kinds
predicate documents as "nothing was dropped" — over a token no rule (R1-R4
all skip over-cap tokens) had ever examined. Since the PR tells machine
consumers to branch on `.code` rather than parse prose, that is a false
"nothing to verify further" signal.

The review's suggested fix was to add `oversizedTokens.length === 0` to the
conjunction. That clause is right and is here, but on its own it routes the
line to the ASSERTIVE channel: `req-line-misparse`, whose message says the
line "could not be parsed as a comma-separated REQ-ID list" on a line where
every ID present was in fact selected. That is the #2334 over-warning class
this module's own channel-selection docblock exists to prevent — a false-clean
code traded for a false-assertion one. So the correct destination is
`req-line-unverified`: the line was not CHECKED, which is not the same as
clean, and nothing on it demonstrably failed to parse either.

Both non-assertive voices were carrying their own inline copy of the same
"nothing was demonstrably dropped" conjunction, and that duplication is what
let them drift: `rangeReadingOnly` omitted the cap, while the over-cap channel
excluded a spaced range wholesale via `!hasSpacedRange`. Extract it once as
`nothingDemonstrablyDropped` and have both read it. The new predicate is a
strict superset of the old `!hasSpacedRange` guard — `spacedRangePairs.every()`
is vacuously true when no spaced range fired — so the over-cap channel's
behaviour on every line without a spaced range is unchanged, and R2 firing on
an endpoint the selector did not take still routes to the assertive channel.

Exactly one input class changes routing: a clean fully-selected spaced range
beside an unexamined over-cap token, which moves from `req-line-range-reading`
to `req-line-unverified`. Pinned by `#3697-19n` (the twin of `#3697-19l` on
the other side of the boundary — there a demonstrated drop outranks the
unverified voice, here the cap outranks the ambiguous one), with `#3697-19o`
as the negative control asserting a clean range with nothing over the cap is
still a range reading.

* docs(#3697): state the warning-code precedence, and what the cap condition actually is

Two surfaces, one point. The round 7 finding cited CONTEXT.md's
PHASE.REQ-LINE.SEAM.kinds predicate as the documentation of what
`req-line-range-reading` claims, and it was right to: that predicate said "R2
alone fired on selected endpoints; nothing was dropped" with no mention of the
scan cap, describing a line the code could not distinguish from one carrying a
token it never examined.

The range reading is the weakest of the three claims and yields to the other
two: a demonstrated drop makes the line a misparse, and a token the cap left
unclassified makes it unverified. `docs/CLI-TOOLS.md` gains that sentence; the
CONTEXT.md predicate gains it plus two precision points that this round's own
pre-push adversarial review extracted over three passes, each with a driven
counterexample I reproduced before acting on it:

  * "nothing was dropped" overstates the rule. The discriminator is RULE-scoped
    by design (round 3, Major 3), so that a parenthesised ID-shaped token does
    not re-open the #2334 over-warning class — `(REQ-02)` is indistinguishable
    from `(ADR-7)` at token level and is carried by the skipped-text rider, never
    by this code. `REQ-01, (REQ-02), REQ-03 - REQ-05` selects three, names REQ-02
    as skipped, and is still a range reading. The predicate now says "no rule
    named a dropped ID", which is what the code tests.

  * The deferral condition is `oversizedTokens` being non-empty, NOT the presence
    of a token past the cap. Those differ: a long token the selector itself took
    can be exempt, because selection is uncapped and anchored and such a token
    was therefore examined. `RANGE-01 - RANGE-05, R-<2047 sevens>` yields
    `oversizedTokens=[]` and stays a range reading. Two earlier attempts to
    characterise WHEN the exemption applies were both refuted — the neighbour
    test admits any non-short neighbour, not just a range operator — so the
    predicate now states the condition and defers the exemption's own rule to
    SEAM.cap rather than paraphrasing it a third time.

Verified after the edit: deferral holds if and only if the cap left a token
unclassified, across all three counterexamples plus controls, and an exempt
selector-taken long token is exhibited. No behaviour change — predicate text,
one CLI-reference paragraph, and the two regenerated indexes. Predicate count
unchanged at 285 across 22 classes, 0 duplicate ids; lint-example-parser-parity
and both `--check` generators pass.

* test(#3697): rename the round 7 regression pair — 19n was already taken

Self-found immediately after the push, before anything was published to the
review thread. The two tests added this round were named `#3697-19n` and
`#3697-19o`, and `#3697-19n` was already in use: it is the declared-blind-spot
case for one-sided separators in both attachment directions, generated inside a
loop with a computed label, which is why a grep for a literal `test('#3697-19n'`
did not find it. The PR body's own *Declared blind spots* list already refers to
`#3697-19n` with that meaning.

Nothing failed. Duplicate test names do not error, and the suite stayed green at
567/567 — which is the argument for fixing it rather than against. Two concrete
costs: any TAP name-set differential collapses same-named tests under `sort -u`,
so one of the two becomes invisible to exactly the did-not-run and pass->fail
checks that name-set comparison exists to perform; and a reviewer reading
"#3697-19n" in the body now gets a different test than the one the body means.

Renamed to `#3697-19p` (over-cap beside a clean range) and `#3697-19q` (its
negative control), the next free ids in the series; the pre-existing `#3697-19n`
is untouched at its original 4 occurrences. Cross-references inside the renamed
block were updated with them. Negative control re-run under the new ids: 19p
still fails against pre-fix source and passes after, 19q passes at both ends.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-01 14:36:57 -04:00
Cody Anderson
8c9265d4e5 fix(#3724): warning-only Dimension 3b findings no longer force the revision loop (#3758)
* fix(#3724): stop advisory Dimension 3b findings from forcing the revision loop

Dimension 3b (undeclared/temporal coupling, #1954) is spec'd "never a
blocker" but tagged severity: warning — the tier plan-phase's revision
loop counts as must-fix — and the planner is never taught the rule, so
every multi-wave phase touching shared mutable state replans at least
once, and intentionally coupled plans re-flag identically every
iteration to the stall prompt.

Three coordinated changes:
- gsd-plan-checker: retag 3b to severity: info, the tier
  references/revision-loop.md already exempts by design; recognize a
  coupling_justified frontmatter declaration in the Do-NOT-flag list so
  deliberate pairs converge. Additions are offset by trimming 3b
  motivation prose — the checker sits 45 bytes under its LARGE hard cap.
- plan-phase step 12: INFO-only accept — an issues block with zero
  BLOCKER/WARNING entries accepts the plan and surfaces the advisories
  instead of re-entering the revision loop. Real blockers and warnings
  still gate unconditionally.
- gsd-planner: slim pointer in assign_waves to the new
  progressive-disclosure reference gsd-core/references/planner-coupling.md
  (the planner sits 19 chars under its own cap), which carries the
  shared-mutable-state rule and the coupling_justified escape hatch so
  first-pass plans avoid the finding when the coupling is unintentional.

Documented the coupling_justified field in docs/reference/plan-md.md.
Growth acks per #2914; inventory manifest and install-tree fixtures
regenerated for the new reference file.

Closes #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin Dimension 3b at severity: info

The severity retag makes the old assertion (severity: warning) stale;
lock the advisory tier from both directions — info must be present,
warning must not — so a future edit cannot silently re-arm the
revision-loop trigger.

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* chore(#3724): changeset fragment for PR #3758

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* docs(#3724): roster planner-coupling.md in docs/INVENTORY.md

The new reference was enumerated in the manifest and all 19 install-tree
fixtures but missing its row in the Modular Planner Decomposition table —
the roster half the manifest-sync test cannot check. (Review Blocker.)

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): cover all four acceptance criteria (review round 1)

- plan-checker-coupling: the 3b severity assertion is now a PARITY check
  deriving the exempt tier from revision-loop.md's flow instead of
  hardcoding info — editing either side alone reds the suite. New
  describe pins the other three criteria: plan-phase's INFO-only accept
  clause (proven failing-first), the BLOCKER + WARNING count staying
  intact, the coupling_justified Do-NOT-flag exemption + fix_hint, and
  the planner pointer + planner-coupling.md content.
- ack fragment: $comment's plan-phase figure corrected to +79B; the 2775
  pin note carried forward into the gsd-planner.md entry, updated for
  upstream's #3761/#3764 Rule-paragraph anchor (which this diff leaves
  verbatim).

The parallel-dependent-plans re-anchor this commit originally carried was
superseded by upstream #3764 during review; this branch no longer touches
that file.

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 2 — align the stance enumeration, complete the template contract

MAJOR: <adversarial_stance>'s severity enumeration gains the INFO bullet so it
agrees with Dimension 3b's 'ALWAYS INFO' mandate instead of contradicting it.
Funded by extracting the inline <examples> block to the new progressive-
disclosure reference gsd-core/references/plan-checker-examples.md (@-inlined
from the same spot; #1949 precedent), which also restores the 3b motivation
clause round 1 traded away (Nit 4) and nets the agent file SMALLER than base
(49107 -> 48486) — the extraction the byte pressure was owed.

MINOR: gsd-core/templates/phase-prompt.md now carries coupling_justified, and
the field's shape becomes one 'plan-id: reason' string per coupled peer so a
plan justified against two peers can express it; docs/reference/plan-md.md's
Type column names the shape.

NIT: the 3409 ack's plan-phase entry no longer calls the #1168 workflow
ratchet an 'XL tier'.

Acks and derived artifacts updated accordingly (checker entry removed — a
shrink needs no ack; INVENTORY roster row + regen:derived for the new file).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): derive the 3b negative severity assertion (review round 2)

Every severity token in the 3b span must BE the tier revision-loop.md exempts,
replacing the hardcoded severity:warning negative — if the loop's exemption
ever moves, the failure names the real conflict instead of blaming the agent
file with a mutually-unsatisfiable pair.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): refit the planner coupling pointer under the char cap

Upstream #3299 (PR #3390) grew agents/gsd-planner.md to 49146 chars at the
base, leaving 5 chars of headroom where the +16-char pointer was measured
against 13 more. The pointer prose shortens to 'Non-file coupling:' —
49150 chars, back under the strict 49152-char cap — and the ack figures
follow. The @-path the tests pin is unchanged.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): re-home the plan-phase ack after the #3823 spent-fragment sweep

Upstream #3078/#3823 deleted all fully-spent ack fragments, including
3409-unreachable-guard-arms.json, which carried this PR's plan-phase.md
+79B append. Per the collision remedy that sweep added: take the deletion
and home the still-live entry in this PR's own fragment. Figures
re-measured at this merge base (90871 -> 90950 LF bytes).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): absorb the spent #3172 plan-phase fragment into this PR's ack

Upstream #3825 shipped 3172-stated-failing-direction.json naming only
plan-phase.md, now spent at the base — colliding with this PR's live
plan-phase entry. Per the #3003 pattern the fully-spent single-path
fragment is deleted and this fragment stays the path's one source;
figures re-measured at this base (93073 -> 93152 LF bytes).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 3 — true up the ack figures, restore the wave comment

The fragment's absolute sizes are re-measured and anchored to base
e40e9670 (planner 47259 -> 47330 chars, checker 45537 -> 44916 B,
plan-phase 91186 -> 91265 LF bytes), with a note that absolutes rot as
next moves — the deltas are the durable claims. The round-1 removal of
the '# Implicit dependency: files_modified overlap forces a later wave.'
pseudocode comment offset headroom base drift had already returned, so
it is restored (findings 2-3). Changeset gains the (#3724) backlink
(finding 4).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 4 — close the verify-work surface, harden the boundaries

BLOCKER: verify-work.md's verify_gap_plans is the second multi-plan
consumer of the checker's sentinels, and its ISSUES FOUND handler entered
revision_loop with zero severity parsing — the guaranteed replan #3724
fixed in plan-phase, alive on the gap-closure surface. The handler now
counts BLOCKER + WARNING and accepts INFO-only returns with advisories
displayed. The checker's INFO stance bullet is reworded to the claim that
is true everywhere ('revision gates count only BLOCKER + WARNING').

Minor 1: plan-phase's iteration_count >= 3 arm recounts severities, so an
INFO-only third check accepts instead of halting on a '0 issues remain'
user gate. Minor 2: the coupling_justified exemption now requires the
entry to NAME the other plan, closing the blanket-suppression reading.
Nit 1: INVENTORY row states the extraction buys cap headroom, not context.
Nit 2: the advisory display gains a concrete format on both surfaces.

Ack fragment re-anchored at base ddde001a: verify-work.md +264B (new
entry), plan-phase.md +395B, checker still net negative (-512B).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin the verify-work accept and the iteration-cap boundary (review round 4)

Two wiring assertions: verify_gap_plans' ISSUES FOUND handler gates on
BLOCKER + WARNING and accepts INFO-only blocks, and plan-phase's
iteration_count >= 3 arm recounts severities instead of gating advisories
— the limit+1 boundary of the gate this PR fixes.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 5 — fail closed at the gates, surface the advisory

Blocker 1: the checker's step-10 status rule routes an INFO-only result to
## ISSUES FOUND (with a new ### Advisories (info) template section and a
severity-aware recommendation) so the orchestrator receives the block and
displays the advisory instead of silently accepting a bare PASSED.

Blockers 2+3: all three gate surfaces (plan-phase step 12 both arms,
verify-work verify_gap_plans) carry one canonical clause verbatim — an entry
whose severity is missing or unrecognized counts as a BLOCKER (fail closed) —
making the accept condition an explicit-INFO whitelist while keeping
issue_count coherent for stall math.

Major 1: the INFO stance bullet scopes its claim to the plan-phase and
verify-work gates (quick mode's loop still revises on any ISSUES FOUND).
Major 2: INVENTORY row and ack $comment state the extraction's real trade
(readability, +0.6 KB eager runtime context), not a cap remedy.
Minor 1: plan-md.md marks coupling_justified as prompt convention, unvalidated.
Nit 1: ack absolutes re-anchored at base 1e67ec97; checker now +120B and acked.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin the round-5 contract — fail-closed parity, INFO-only return shape

New: three-surface verbatim parity test for the fail-closed clause (Blockers
2+3); checker return-contract test for the INFO-only ## ISSUES FOUND route and
advisories section (Blocker 1). All seven newly pinned tokens are absent at
f3a5682d, so each new assertion fails pre-fix.

Updated: accept-clause regexes track the explicit-INFO whitelist wording;
the severity sweep scopes to the span's fenced yaml examples via
yamlSeverityTiers (round-5 Minor 3, applied to the blocker negative too);
the iteration-cap comment states it is a prose pin, not an executed boundary
check (Minor 4); splitLines call sites document the line-pin coupling (Nit 2).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): adopt next's line wrap in the 3b motivation clause — drops a wrap-only hunk from the diff

Byte-identical content; the wrap difference was an artifact of the round-1
base adaptation predating upstream's #3003 landing.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 7 — gate every checker consumer, not just the two audited ones

Blocker: quick/steps/plan-checker-loop.md (issue-named in #3724) gets the
same canonical fail-closed clause and explicit-INFO whitelist accept as
plan-phase/verify-work — an INFO-only result proceeds instead of entering
quick mode's revision loop.
Major: import.md plan_validate handles the checker return by severity
(INFO-only never blocks an import) and is added to agent-contracts.md's
consumer enumeration, which had omitted it.
The checker's INFO stance bullet drops the quick-mode carve-out — the claim
is universally true again now that every consuming gate is severity-aware.
Minor: an applied coupling_justified exemption is surfaced as its own info
advisory so a stale one-sided declaration stays observable.
Nit: plan-phase's revision-iteration Display line is explicitly conditioned
on not having already proceeded to step 13.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM
Emitted-Drift-Ack-Growth: import.md — #3724 round 7: the plan_validate step's checker-return handler becomes severity-aware — counts BLOCKER + WARNING failing closed and accepts an explicitly-INFO-only return with advisories displayed instead of blocking the import

* test(#3724): pin the round-7 surfaces — five-gate parity, quick/import accepts, exemption visibility

The verbatim fail-closed parity test extends to quick/steps/plan-checker-loop.md
and import.md plan_validate; new assertions pin quick mode's INFO-only proceed,
import's never-blocks accept, import.md's presence in agent-contracts.md's
consumer row, and the surfaced coupling_justified exemption advisory. All four
newly pinned token families are absent at the pre-fix head, so each new
assertion fails first.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-31 19:08:11 -04:00
BeeHiggs
1a358ce0fd feat(#2761): bracket-tolerant read path — roadmap/validate/verify/state recognize bracket ids (epic #612 PR-2) (#2867)
* feat(#2761): gated heading-intro selection + one bracket identity grammar

Foundation. Two owner-level changes plus a federated convention resolver; no
reader consumes them yet.

1. GATED SELECTION, not an ungated widening.

   Widening every heading matcher requires the claim "no legacy ROADMAP contains
   a `[CODE.MM]` bracket followed by a digit", and that is false:
   `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and
   `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each
   as a phase — moving phase_count and total_phases and adding W006 on projects
   that never opted in. No narrowing rescues it: the premise is about documents
   we do not control.

   `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the
   pattern SOURCE at construction time. A project whose resolved
   `phase_id_convention` is not exactly 'bracket' compiles the same source string
   it compiled before. `baseline` is explicit because whether a site spells the
   any-bracket prefix or a bare `Phase\s+` is a fact about that site's history:
   handing the wider grammar to a bare site retro-grants tolerance it never had,
   in both directions — warnings appear, and a warning that fires today vanishes.

   Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through
   the base alternative, which captures nothing, so a reader saw no bracket, fell
   back to the legacy token rule, and counted a labeled icebox heading while
   excluding the label-less one beside it — two derivations of one ROADMAP
   disagreeing.

2. ONE bracket identity grammar, one width rule.

   The milestone width is reconciled with the emit validator: pad2 output, so
   two digits or 3+ with no leading zero. Earlier spellings diverged in both
   directions — admitting `002`, which the validator rejects, and a bare `0` pad2
   never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED
   a milestone no phase heading could then resolve into, recreating the
   on-disk-count fallback this epic removes. An unpadded bracket is now uniformly
   malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its
   directories is the surfacing signal.

   The milestone field is boundary-anchored, so a malformed run cannot match by
   its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays
   case-insensitive because readers compile `/i`, but identity helpers match
   `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:`
   failed every sentinel test. The qualified key shares the width, the `(?=-|$)`
   boundary and the single-sub-phase shape of the directory token, because
   phaseTokenMatches returns unconditionally on a qualified hit: a key matching a
   directory isPhaseDirName rejects would be a final wrong answer.

3. resolvePhaseIdConvention federates workstream -> root exactly as
   config-loader does — including that root is a fallback only when a WORKSTREAM
   is active, so a project-scoped directory stands alone. loadConfig cannot serve
   this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and
   this key is not among them. It governs the bracket-selection reads ONLY.

PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing
consumes it, and it is superseded rather than redefined.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#2761): roadmap.cts selects its heading grammar from the convention

Six matchers build their intro through the gated selector, and
cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each
resolve the convention ONCE per command and thread it down.

Three sites take the any-bracket baseline (they already tolerated
`[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`).
Handing the wider grammar to a label-only site retro-grants tolerance it never
had — and not only by adding matches: on a legacy repo an unchecked
`- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that
fires today.

Sentinel handling under bracket ADDS a rule rather than replacing one: a
bracketed heading is a sentinel when its bracket milestone is reserved
(`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog
convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a
mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly
the content this epic targets — add entries to the progress denominator. The
captured id is folded before the identity test, so a lowercase
`### [gsd.999] 07:` is excluded too.

The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention
once and hands it to all four of its heading/checklist patterns, but the single
`phaseTokenMatches` call that decides `disk_status`, `plan_count`,
`summary_count`, `has_context` and `has_research` was left two-argument — so
every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with
zero counts, on the PR's own headline verb, while the SAME build resolved those
same directories correctly in three other places on the same repo (W006/W007 via
phaseTokenFromDir, `state json` via the milestone filter, and the W021
milestone-complete read through this very helper's three-argument form). It
failed ONLY for the directory shape the convention exists to name: a
mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is
why nothing caught it. Measured, bracket vs its flat-legacy twin:
`[["01","no_directory",0,0],["02","no_directory",0,0]]` against
`[["01","complete",1,1],["02","planned",1,0]]`.

The oracle is the twin, computed in the same test run, plus exact literals —
`grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix
nor a future regression had any gate at all.

Disclosed: a ROADMAP written in bracket form before config.json is switched
reads as empty rather than mis-counted. Silent invisibility during the migration
window is the deliberate trade against claiming phases on projects that never
opted in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#2761): validate.cts selects its grammar; gated directory recognition

The W006/W007 feeders take the resolved convention as a threaded parameter.
These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them
where an ungated widening does the most damage: `### [RFC.2119] 5:` enters
roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on
disk" on a project that never opted in.

buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket
headings. Surfaced rather than filtered in place because roadmapPhases feeds both
a membership check and a missing-directory warning, and only the latter should
ignore an icebox item.

That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is
a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying
suppression on the token alone let an icebox heading silence a REAL phase that
happens to share its number — a false negative strictly worse than the warning it
removed. A token is suppressed only when no non-sentinel heading bears it.

Directory recognition is added as gated FUNCTIONS beside the exported RegExp
constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is
string-indistinguishable from the letter-prefixed-decimal family this repo
documents as ambiguous, and folding a branch in changes those constants' answers
on exactly that family. A RegExp constant has nowhere to attach a gate.

The recognizer mirrors the emit grammar and delegates the token to the canonical
owner, so recognizer and resolver agree on rejected input as well as accepted.
Both functions throw on a non-string, matching the call pattern they replace.

buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin
and like the sibling checklist scan in roadmap.cts, and for the reason that one
states: the bracket id has to ride along or the sentinel filter is blind to
`- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every
checklist token REAL, and the occurrence-aware un-suppression loop then deleted
the icebox token the HEADING scan had correctly marked sentinel — so `validate
consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE
ROADMAP shape where an icebox appears as both a bold bullet and a detail
heading. `validate health` stayed silent on that same repo, so the two verbs
disagreed — which is the disagreement `sentinelPhases` exists to close.

Both directions are pinned, because the failure mode of a careless fix here is
the opposite one: a real phase sharing a sentinel's token must still warn. It
does, in all four shapes that attack it (sentinel heading + real bullet,
lowercase sentinel, sentinel after the real heading, colon-less bullet).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2761): count bracket headings, and retire them, in both derivations

Both `total_phases` derivations select their grammar from the resolved
convention, in one commit — cmdStateSync already carries the comment that it
mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)",
so teaching one and not the other ships that divergence.

The #1514 retirement filter widens WITH the counter it protects. The canonical
gesture strikes the checklist BULLET and leaves the detail heading intact, so a
bracket-form retirement went undetected and the phase stayed in the denominator
forever. That is half a fix alone: the retired key is compared against
phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves
land here.

Under bracket the sentinel token rule composes as the full engine set {0, 999},
so this counter agrees with `roadmap analyze`, which has always excluded both —
otherwise the two derivations report different numbers for one ROADMAP and the
changeset's "excluded from every count" is false as written. The LEGACY path
keeps its pre-existing 999-only rule: widening it there would move legacy totals,
so the two stay split off the bracket path exactly as they are today.

The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not
the frontmatter total_phases. Sync's own counter never reaches that field — the
read derivation writes it — so asserting the frontmatter after a sync measures the
read path twice and lets a mutation to the write-path guard survive.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read

The shipped milestone-prefixed W021 gate keeps its ROOT-only config read,
verbatim base semantics. Federating it silently moved a legacy convention's
answer in BOTH directions on workstream repos — a W021 that fires at base
vanishing, and one that is silent at base firing. resolvePhaseIdConvention
governs the new bracket-selection reads only.

B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it
with an empty config) but selects its grammar from the convention. Inferring
'bracket' from the shape of a matched bracket ran a repo-failing check against a
legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution
widens with the heading read, so a bracket repo whose phases are on disk stays
silent, and a bracket sentinel is not reported as unstarted.

checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so
fenced examples cannot warn and heading level is structural. Its scope rules each
close a way it silently did nothing or fired wrongly: only a genuine MILESTONE
heading opens or closes a section (a `### Notes` used to reset scope and disable
both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed
phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a
phase; the full h2-h6 range is processed. Its section recognizer shares the one
milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the
id grammar and a section to the section grammar at once, silently re-scoping
every warning after it.

validate consistency suppresses bracket sentinels in its missing-directory
warning — the two verbs disagreed, health suppressing via notStartedPhases while
consistency did not. The legacy reading is untouched, including its pre-existing
wart that `### Phase 999:` still warns there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2761): scope the milestone by its bracket; select the disk-side filter

Two roadmap-parser reads, both of which made a bracket project's totals track
the disk instead of the ROADMAP.

The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name,
no version — but scoping matched STATE's `milestone: v2.0` STRING against a
heading, so the canonical form matched nothing and total_phases fell back to the
directory count. The rule was re-derived in THREE places: extractCurrentMilestone
plus two `milestoneBounded` guards; fixing one left the others falling back
regardless, so they are now one gated helper. It matches the CANONICAL padded
spelling only — accepting `0*N` bounded a milestone whose phases were invisible,
which un-suppressed a progress percent computed off an unscoped disk count.

getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a
bracket ROADMAP it collected nothing, so the filter degraded to pass-all and
buildStateFrontmatter counted every other milestone's directories — making the
bracket convention strictly worse than the M-NN one it supersedes on the property
that matters most: totals must track the ROADMAP, not the disk.

The DIRECTORY side of that same filter is selected with it. Teaching only the
heading scan was half a fix and a worse one: `milestonePhaseNums` became
non-empty, so the pass-all degrade stopped firing, but no bracket directory could
satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the
custom-id match captures the project code `GSD`, and stripProjectCodePrefix does
not strip a dotted prefix). Every bracket directory was rejected, and
completed_phases / total_plans / completed_plans / percent all collapsed to 0
while `state sync` went on writing a percent off the unfiltered disk — `state
json` reporting 0% on the same repo, in the same second, that STATE.md's body
called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and
total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)`
floors it at the ROADMAP count no matter how many directories are rejected.

The dir side matches on the milestone-QUALIFIED id, delegated to the owner's
gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B
puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one`
share the token `01` and only the qualified key separates them. The qualified ids
are kept in their own set — a hyphen in `milestonePhaseNums` would flip
`roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo
— and the branch is ADDITIVE: on a miss it falls through to the three legacy
checks, so a bracket project carrying legacy-shaped directories reads unchanged.

Both are resolved lazily and gated, so the legacy path pays neither a config read
nor a second scan and cannot change answer. The scoping call is also GUARDED:
resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a
GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only
planningDir call in extractCurrentMilestone sits inside the STATE-read try, so
the function returned normally on such an environment; an unguarded one here let
that escape and broke the never-throws invariant that getRoadmapPhaseInternal and
getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about.
Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the
workstream-name policy and GSD_PROJECT throws identically at base — but reachable
by any in-process embedder, which is precisely who that invariant is for. The
filter's own resolve call was already inside its try and is unaffected.

The milestone-qualified key is formed only for a token that is itself a bracket
phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration
heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to
`GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02:
the `-01` truncated, both such headings collapsing to one key, and the heading
claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting
`GSD.02-01-one`, the one it does. The guard drops those headings back to the
unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned
against the milestone-prefixed reading of the same ROADMAP, which is
base-identical on this shape.

Scoped precisely, because the fixture moves one number that the guard does not
touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket
heading COUNT this PR exists to add, not the splice — measured identical with and
without the guard, and identical to what the canonical `### [GSD.02] 01:`
spelling does on the same fixture (both read 2 with zero directories on disk,
where base reads 0). The claim is base-equivalent ACCEPTANCE, not a
base-equivalent reading.

One consequence is stated rather than fixed: a heading whose token carries a
hyphen still puts that hyphen into milestonePhaseNums and so still flips
`roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it
is what keeps the shape base-equivalent; excluding the token would have moved
answers versus base on malformed input. The comment at the qualified-set
declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out
of that flag's input, not hyphens in general.

The oracles ship with it, and they are the five numbers, not the one: the parity
gate now asserts total_phases, completed_phases, total_plans, completed_plans AND
percent, on both derivations, on two fixture shapes (one milestone; two
milestones with stale prior-milestone directories on disk). The oracle is the
flat-legacy twin, built in the same test run and compared number for number,
plus exact literals so a shared wrong answer cannot pass.

The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could
not serve, because buildStateFrontmatter's #2445 de-dup key captures only a
directory's leading integer and collapses `02-01-one` / `02-02-two` /
`02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's
[3,2,3,2,67], identically at base and before this fix, and structurally
unreachable from the bracket key space. That reasoning is only sound while it
stays true, so a characterization test holds the M-NN reading down on the two
numbers that do not depend on which directory wins the mtime race. Widen the
de-dup key and it fails, instead of quietly invalidating the changeset's
disclosure.

Also adds the call-site pin. The structural table pins transcription against the
selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping
verify.cts's milestone-complete site to the wider baseline grants a
fires-on-every-repo check tolerance it has never had, and every behavioural test
still passed. The pin reads the shipped sources and asserts the mode at each of
the 14 sites, count-exact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): pin the bracket read surfaces in the parity gate

This gate exists because #2043 fixed one bug across five hand-edited copies of a
rule and #2232 was the residual that survived, because a later reader could not
tell the copies were one rule. PR-2 adds two consumers, so they belong here.

Surface 7 — the heading read and the directory read must agree about WHICH phase
a `MM-<seg>` pair names, across the shared width corpus, and the bracket and
legacy spellings of one heading must yield the same token.

Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on
ACCEPTED input was already pinned; agreement on REJECTED input is where they
actually diverged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2761): changeset

Disclosures for the PR body (deliberate, not defects):

- phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and
  cannot serve as the convention resolver however the file is federated. This PR
  ships its own workstream->root resolver; adding the key and its value enum is
  later-slice work.
- Convention matching is strictly === 'bracket'. A misspelled value reads as
  not-configured and the project keeps legacy behaviour silently.
- An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing,
  bounds nothing, sections nothing, and is not a phase id. W005 on its
  directories is the surfacing signal.
- WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical
  inputs, none of which toDir can emit and none of which had a bracket caller at
  base:
    isSentinelPhaseId('GSD.0-01',    'bracket')  true  -> false
    isSentinelPhaseId('GSD.0999-01', 'bracket')  true  -> false
    getMilestoneFromPhaseId('GSD.2-01',   'bracket')  'v2.0' -> null
    getMilestoneFromPhaseId('GSD.002-01', 'bracket')  'v2.0' -> null
  The canonical pad2 sentinel spelling `[GSD.00]` still tests true.
- FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x ->
  milestone null) is preserved". After the unification that holds for the
  canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours;
  flagging the tension rather than editing it.
- The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is
  a sentinel when its bracket milestone OR its token is reserved. Under bracket
  the state-side token rule is the full {0, 999} set so both derivations agree;
  the LEGACY path keeps its pre-existing 999-only rule, unchanged.
- validate consistency's legacy reading is untouched, including the pre-existing
  wart that `### Phase 999:` warns there while validate health suppresses it.
- find-phase still cannot resolve a bracket phase directory. phase-locator.cts is
  outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls
  phaseTokenMatches without a convention, so whichever slice lands second must
  thread it through.
- Four of the five bracket readers scan raw ROADMAP content, so a bracket heading
  inside a fenced code block is read as a phase. Pre-existing for the legacy
  spelling; parity, not a new class.
- roadmapPhaseLookupSources gained no bracket source: nothing emits a
  milestone-qualified query into it yet.
- roadmap validate remains a separate, unfederated convention reader.
  Pre-existing and base-identical, but two verbs can disagree about the active
  convention on one project.
- _diskScanCache keys on cwd while the values it caches are now
  convention-dependent. Not reproducible through the CLI; pre-existing for the
  workstream dimension, widened here. Stated as inconclusive.
- A ROADMAP written in bracket form before config.json is switched reads as empty
  rather than mis-counted — the deliberate migration-window trade.
- THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that
  divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter
  applies the milestone filter; cmdStateSync does its own fs.readdirSync and never
  calls it, so on a repo carrying prior-milestone directories the read path
  reports the SCOPED percent and the sync body reports the WHOLE-DISK one.
  Measured on the true base build (d04592de), flat-legacy spelling, 3 in-scope
  phases with 1 complete plus 2 stale prior-milestone dirs: `state json`
  [3,1,3,1,33], sync body 60%. The bracket twin of that repo now reads the same
  two numbers — 33 and 60. Scoping the sync counter would move every legacy
  repo's percent, which a bracket read-path PR must not do. The gate pins both
  sides, so the mirror cannot silently become a one-sided fix.
- THE PARITY ORACLE IS THE FLAT-LEGACY TWIN, NOT THE M-NN ONE, and that is a
  measurement finding rather than a preference. buildStateFrontmatter's #2445
  de-dup key captures only a directory's LEADING integer, so the M-NN dirs
  `02-01-one` / `02-02-two` / `02-03-three` all key to `2` and two of the three
  are dropped before they are ever counted: base reads [3,0,1,0,0] where the
  flat-legacy twin of the same repo reads [3,2,3,2,67]. Present identically at
  base and at HEAD, untouched here, and structurally unreachable from the bracket
  key space — `GSD.02-01-one` does not match that pattern at all, so every bracket
  directory keys to its own name. The source line already carries a
  `phase-id-owner:` sanction recording the divergence. Mirroring it under bracket
  would mean manufacturing a collision that cannot occur, so the gate compares
  against the flat-legacy spelling, which is uncontaminated. This paragraph is
  itself pinned: a characterization test holds the M-NN reading on the two
  numbers that do not depend on which directory wins the mtime race, so widening
  the de-dup key in a later slice fails the suite rather than silently making
  this disclosure false.
- THE `phaseTokenMatches` CALL-SITE CENSUS, stated so the remaining gaps are
  auditable rather than implied. 13 call sites outside the owner (phase-id.cts).
  THREE are three-argument: verify.cts:2229 (the W021 milestone-complete read,
  already was), roadmap.cts:436 (`roadmap analyze`'s directory lookup, threaded
  by this PR) and roadmap-parser.cts:792 (the disk-side milestone filter, added
  by this PR). The other TEN are two-argument and stay that way — phase.cts ×5
  (220, 277, 444, 585, 1547), phase-locator.cts:62, smart-entry.cts:243,
  init.cts:1414, milestone.cts:551 and verify.cts:2467. All ten are untouched by
  this PR and base-identical.
  One of them sits in a file this PR DOES edit, so it is named rather than left
  to a reader's grep: verify.cts:2467, `verify schema-drift <phase>`. Measured on
  a bracket repo across base / pre-fix branch / this HEAD, all three agree on all
  three argument forms — `verify schema-drift GSD.02-01` and `… 01` both report
  "Phase directory not found" on every build, and `… GSD.02-01-one` resolves on
  every build through the exact-directory-name fallback. So the user-visible
  shape of what stays broken is: a bracket phase is addressable there by full
  directory name only, exactly as at base. Threading the convention into a
  function this PR never touched, in the last round before ship, is the wrong
  trade; it is where the same one-argument fix goes next, alongside
  milestone.cts:551 and init.cts:1414.
- A BRACKET HEADING WHOSE TOKEN CARRIES A HYPHEN (`### [GSD.02] Phase 02-01:`,
  a mid-migration spelling) forms NO milestone-qualified key, and therefore
  scopes through the unqualified legacy path — base-equivalent ACCEPTANCE, which
  is the claim, and not a base-equivalent reading: `total_phases` on that shape
  moves 1 -> 2 for the same reason it moves on the canonical `### [GSD.02] 01:`
  spelling, because counting bracket headings is what this PR does. Such a token
  still flips `roadmapUsesHyphenedIds`, as it also does at base. The comment at
  the qualified-set declaration now claims only that narrower, true thing.

The `pr:` field carries the sub-issue number as a placeholder — it must be
updated to the real PR number when the PR is opened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2761): point the changeset at PR #2867

* test(#2761): fast-check properties for the convention-selection layer

CONTRIBUTING.md mandates a generative property test for parser/bijective-contract
changes; PR-2 shipped six example-based files and none. This adds the missing
layer, scoped to what PR-2 actually contracts — WHICH pattern each reader
compiles, decided by the resolved `phase_id_convention` — rather than restating
PR-1's grammar round-trip properties, which already live in
tests/adr-612-bracket-grammar.test.cjs.

Four properties: P1 an opted-in repo reads the ADR-canonical label-less bracket
heading/dir and a non-opted-in repo is byte-blind to the identical input; P2
every non-bracket convention agrees with the hand-transcribed BASE source over
generated content, including bracket-DOTTED legacy prose (`[RFC.2119] 5:`) that
must never be claimed as a phase; P3 nine per-field mutations are rejected and
the one case variation folds instead; P4 both sides of a phase comparison derive
the same key under the same convention.

Generators template every input from raw primitives — nothing is seeded through
renderPhaseId/toDir, the p2() tautology that made #2258 round 1's property test
structurally unable to find B1. Domain reaches past 99 into the 3+-digit branch
(round 2's numArb-capped-at-99 miss), forces sub-phases in at weight, and pins
both sentinel milestones.

Falsified against the COMPILED lib, not the source: five deliberate mutants
(gate never fires; gate always fires; milestone width widened to \d+; the #612
convention forwarding dropped from phaseKeyFromDir; extractPhaseToken's bracket
branch ungated) each fail the specific property that should catch them —
16/2, 16/2, 15/3, 17/1, 17/1 pass/fail — and the lib restores byte-identical.

An earlier draft of P2 held vacuously: its base regex omitted the markdown
furniture the selected one carried, so every realistic `### Phase NN:` line
matched neither side. The gate-always-fires mutant did not kill it. Both are now
compiled through one function, and that mutant kills P2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): adversarial malformed bracket tokens across the tolerant readers

The existing boundary coverage stopped at shapes the emit grammar rejects
(unpadded `[GSD.3]`, wrong-case, `12A`). It never exercised a STRUCTURALLY
broken token — a non-numeric milestone, a bracket that never closes, a bracket
nested in another — which is the input a tolerant reader is most likely to
half-read, and the one the PR's own regex commentary is explicit about.

read-tolerance (roadmap heading scan + validate's dir and variant builders):
ten malformed headings, each asserted to be read as a phase by NO convention and
to give the opted-in repo the same answer as the legacy one; the corpus driven
through `roadmap analyze` end to end; malformed DIRECTORY names asserted
unrecognized and non-throwing on all four conventions; and the two variant
builders asserted to agree, since a widening that reaches only one splits
`validate consistency` from `validate health` (the #3242 Bug B shape).

coherence (verify.cts W021): the same six broken shapes asserted to raise no
W021 of their own AND not to re-scope the W021 that follows them — the G2
failure mode reached from a different shape, where a heading that is not a phase
but IS read as a section silently moves later warnings onto the wrong milestone.

Both files gain a pathological-input time bound. Nested quantifiers over a long
unclosed bracket are the classic ReDoS shape and two commits on next (#2828,
#2944) were CodeQL-flagged for exactly that, so the bound is asserted rather
than argued from reading the pattern. The probes themselves parse no regex.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2761): document "bracket" as a phase_id_convention value

The row listed only `"milestone-prefixed"` and `null`, so after two shipped
slices (#2258 grammar, this PR's read path) the convention had no documented
enum value. CONFIGURATION.md is also a top-10 historical co-changer of both
src/verify.cts and src/state.cts and was absent from this PR.

The row states the boundary rather than the ambition: `"bracket"` changes the
READ path only, there is no migrator and no emit yet, and a project on any other
value compiles the patterns it compiled before. That keeps the docs honest for
the two releases before PR-3 and PR-4 land, instead of describing a convention a
user cannot yet migrate to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2761): retype the changeset Added, drop the docs-exempt marker

`Fixed` was wrong by CONTRIBUTING.md's own definition — a fix restores
documented behavior, and bracket read tolerance is the second slice of a
capability that did not exist before #2258. The type also carried a
`docs-exempt` marker, and `Fixed`/`Security` are exempt from the docs-required
lint, so the typing had the effect of routing around a gate this change should
pass. It now passes it: `lint-docs-required` returns ok_docs_updated on the
CONFIGURATION.md row added in the previous commit.

Body gains one sentence pointing at that row and restating that `"bracket"` is a
read-path opt-in until the migrator and write path land.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): pin the version-less bracket milestone scoping gap; narrow the claim

The changeset asserted that milestone scoping "recognises the ADR-canonical
`## [GSD.02] Foundation` heading and applies to the phase DIRECTORIES too." A
CLI probe on that exact heading form falsifies the second half: with no `vN.N`
in the milestone heading the directory side does not scope, and directories from
BOTH the prior and the later milestone are admitted. Measured 4 dirs counted
where the milestone declares 2.

Every bracket fixture in the suite writes `## [GSD.02] v2.0: …`, so nothing
covered the form the ADR actually specifies — and the state.cts doc comment
calls that version-less form canonical.

Mechanism, in extractCurrentMilestone: the bracket scope branch selects the
right currentSection, but `preambleCutoff` keys off a pattern requiring a
version or status emoji, so a version-less roadmap falls back to the current
milestone's own offset and every PRIOR milestone lands in the preamble — whose
phase-stripping regex only strips `Phase N:`-labelled headings, so bracket phase
headings survive it. Independently, `computeSectionEnd` accepts a boundary only
on a version/emoji heading, so the section runs to EOF and every LATER milestone
is swept in. Two sites, bidirectional.

Not fixed here: it changes milestone scoping, which is shared with the legacy
path. Five characterization tests pin today's reading plus a versioned CONTROL
proving the version string is the only difference, and the changeset sentence is
narrowed to what the code does. The DEFECT assertions are written to be
INVERTED by the fix, not deleted — that inversion is its regression proof.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2761): re-anchor the branch's own cross-file line citations after the rebase

Three of this branch's code comments cite sibling call sites by line number, and
the rebase onto 178ec000 moved two of the three targets:

  roadmap-parser.cts  validate.cts:210 -> :218   (const g = capturing ? 1 : 0)
                      state.cts:1715   -> :1752  (const bg = … 'bracket' ? 1 : 0)
  roadmap.cts         verify.cts:2229  -> :2355  (phaseTokenMatches 3-arg form)

`state.cts:1715` had drifted 37 lines and now lands on the retirement skip, not
the capture-offset idiom the sentence is about — the citation read as evidence
for a claim the cited line does not support.

planning-workspace.cts's `config-loader.cts:618/:649` was checked and is still
correct; left alone.

Comment-only. Build, drift guard and the bracket suites re-run unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2761): scope the version-less bracket milestone heading too (B1)

computeSectionEnd and the preambleCutoff scan in extractCurrentMilestone
(roadmap-parser.cts) only recognized a milestone boundary heading that
carried a vN.N token or a status emoji. The ADR-canonical bracket
heading (## [GSD.02] Foundation) carries neither, so on that shape
computeSectionEnd fell through to content.length (sweeping every LATER
milestone into scope) and preambleCutoff fell back to the current
milestone's own offset (leaking every PRIOR milestone's bracket phases
into the preamble, whose Phase-N: strip regex never matches them).

Under the bracket scope branch, both sites now also accept a
`#{1,2}\s+\[CODE.MM\]` boundary, built from phase-id.cts's BRACKET_ID_SRC
(single owner of the bracket-id grammar) rather than a re-typed literal.
`#{1,2}` is the deliberate discriminator: a bracket PHASE heading is
level 3 and shares the same `[CODE.MM]` prefix, so a `#{1,3}` boundary
would swallow it too. Reachable only when bracketScopeConvention ===
'bracket' was already resolved (i.e. the bracket scope branch actually
fired), so version-bearing/emoji headings and non-bracket conventions
take the exact pre-existing code path byte-identically — confirmed by
the full adr-612 suite staying green.

Inverts the four DEFECT assertions in the
"#612 PR-2 CHARACTERIZATION: a version-less bracket milestone does not
scope" describe block (tests/adr-612-bracket-phase-counting.test.cjs)
into their regression-proof form, per the block's own doc comment, and
reframes the describe title/comments accordingly. Corrects the
.changeset/2761-bracket-read-tolerance.md fragment, which described the
directory-side version-less gap as an open, un-closed bound.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): thread sentinelPhases into validate health's W006 loop (B2)

cmdValidateHealth's W006 loop (src/verify.cts) destructured only
roadmapPhases from buildRoadmapPhaseVariants, not sentinelPhases —
unlike cmdValidateConsistency, which already skips sentinelPhases with
the identical guard a few hundred lines up. A heading-only bracket
icebox/pre-milestone entry ([GSD.999] / [GSD.00]) therefore gained a
false W006 "no directory on disk" from validate health while validate
consistency correctly stayed silent on the very same ROADMAP — the two
validators contradicting each other.

Threads sentinelPhases through and skips it before the existsOnDisk
check, mirroring the consistency guard exactly. Gated the same way
sentinelPhases already is (empty unless phase_id_convention is
'bracket'), so a legacy repo's W006 reading — including its own
pre-existing wart where a legacy `### Phase 999:` still warns on both
verbs — is untouched; confirmed by the existing "INHERITED WART,
unchanged" test staying green.

Adds the paired-agreement regression test (#612 PR-2 B2 describe block
in tests/adr-612-bracket-read-tolerance.test.cjs): a sentinel-only
bracket roadmap must produce no missing-directory warning from EITHER
validator, plus a CONTROL proving a real phase with no directory still
warns on both. Confirmed red (health false-W006) against the pre-fix
code before applying the fix.

Corrects the .changeset/2761-bracket-read-tolerance.md fragment, which
described the asymmetry as already closed and in the wrong direction
(it credited validate health with already staying silent, when health
was the one falsely warning).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#2761): pin mixed-shape preamble cutoff and boundary heading levels

Closes two self-flagged coverage gaps in the B1 fix (commit 08d5b0c4)
ahead of adversarial review. No src change — all three new tests are
green against the code as committed.

1. earliest-of-either preambleCutoff comparison: only exercised where
   the version/emoji match and the bracket match happen to land on the
   same heading. Adds the mid-migration mixed shape (version-bearing
   PRIOR + version-less CURRENT) and asserts scoping outcomes (accepts
   booleans + total_phases), not internals.

2. `h.level <= 2` conjunct in computeSectionEnd: provably redundant
   whenever the selected milestone heading is level 2 (every existing
   fixture), since `h.level > level` alone already implies it there —
   a mutant deleting the conjunct would have survived every prior test
   in this file. Adds a level-3 CURRENT-heading fixture (with a real
   PRIOR milestone so the preamble side-channel can't independently
   rescue the truncated phases) that makes the conjunct's deletion
   test-visible, confirmed by hand-mutating a throwaway copy of the
   compiled output (never touching tracked src or the real build) and
   observing the assertion flip. Also pins a level-1 companion case
   (#{1,2} tolerance, not just level 2).

NOT included here: the other mixed-shape direction (version-less PRIOR
+ version-bearing CURRENT) turned out to be a genuine, currently-unfixed
gap — reported separately rather than silently patched or weakened, per
instruction not to touch src while a probe run is in flight.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): engage bracket boundaries when the current milestone heading is version-bearing (B3, self-caught)

Found during round-2 self-verification of B1 (commit 08d5b0c4), while
closing the mixed-heading-shape coverage gaps flagged in my own review
notes. The B1 fix resolved `bracketScopeConvention` only inside the
`if (headingMatches.length === 0)` gate that also drives SELECTION's
own bracket fallback (which heading counts as "current"). That gate is
correct for selection, but `bracketScopeConvention` also feeds
computeSectionEnd's and preambleCutoff's boundary detection further
down — which accidentally inherited selection's gate instead of having
its own.

Trigger shape: the CURRENT milestone heading is itself version-bearing
(`## [GSD.02] v2.0: Current Milestone`), so the primary version-string
match succeeds immediately — headingMatches.length !== 0 from the very
first check — and the entire bracket-resolution branch was skipped. A
sibling milestone (PRIOR or LATER) that is version-less then got
neither the version/emoji boundary rule (it has none) nor the bracket
boundary rule (never resolved), reproducing the original #612 defect
(total_phases falling back to the whole-disk count) through a
structural shape B1's own fixtures never exercised — every one of them
is uniformly version-bearing or uniformly version-less across all
three milestones, never mixed with CURRENT specifically being the
version-bearing one.

Fix: resolve `bracketScopeConvention` unconditionally, decoupled from
`headingMatches.length`. SELECTION is deliberately left untouched — the
`if (headingMatches.length === 0 && bracketScopeConvention === 'bracket')`
fallback that picks which heading is "current" keeps its original gate
byte-for-byte (confirmed by diff: that line is unmodified). Only the
convention *resolution* moved out from behind it, so boundary detection
can consult it regardless of which branch selected the heading. The
extra `resolvePhaseIdConvention` call this now costs on every
invocation (previously paid only when the version match found nothing)
is the accepted cost: a non-bracket repo still resolves to something
other than 'bracket' (or null on a poisoned env, caught exactly as
before), so `bracketMilestoneHeadingRe` stays null and every downstream
branch is byte-identical to today — confirmed by the full adr-612 +
roadmap-parser + state + verify + health-validation suite staying green
(1260/1260) and the all-version-bearing/legacy fixtures showing no
behavior change.

TDD: tests/adr-612-bracket-phase-counting.test.cjs describe block
"#612 PR-2 B3: bracket boundaries engage even when CURRENT is
version-bearing but a sibling is not" — 4 tests, confirmed red against
pre-fix code (leak-in booleans true/true, total_phases 4) before this
change, green after.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): reject same-milestone continuation headings as boundaries (B1)

Gate-2 adversarial review Blocker 1: the B1/B3 boundary fired on ANY
as the one currently selected — a version-less checklist/detail split
(`## [GSD.02] Foundation (Phase Details)`, or an ad-hoc continuation
heading) truncated the current milestone's own section instead of
being recognised as a continuation of it. The `(Phase Details)`
re-append only searches VERSION-STRING matches, so a version-less
continuation heading was cut out and never re-appended — a confidently
wrong, non-degraded phase count for a still-incomplete milestone
(repro8 case 1: 1/1/100 instead of 2/1/50; repro5: same, on a fully
version-less roadmap with no sibling milestones at all).

Introduces one shared helper, isBracketMilestoneBoundary(headingText,
level, selectedBracketId), used by both computeSectionEnd and the
preambleCutoff bracket scan, replacing the ungated `h.level <= 2 &&
bracketMilestoneHeadingRe.test(...)` inline check. `selectedBracketId`
(case-folded via phase-id.cts's foldBracketId, matching the branch's
own fold-before-identity convention) is derived from `selected[0]`,
which is the full matched heading line on BOTH selection paths
(version-string and bracket-fallback), so one extraction covers both.

Level cap stays at `level > 2` for now (temporary — ADR-612's content
discriminator replaces it in the next commit); same-milestone rejection
is the change this commit is scoped to.

DEVIATION from the reviewed plan, caught empirically: applying the
same-milestone rejection at the preambleCutoff site (as literally
specified) regressed an existing pin ("boundary heading level: a
level-1 CURRENT milestone heading also scopes correctly") and a
fenced-heading case (repro10 A3) — because preambleCutoff's job is
"where does the earliest milestone-shaped heading sit, scanning from
the TOP of the document," and the selected heading's own occurrence is
always a correct answer to that question regardless of same-id-ness;
rejecting it let the earliest-of-either comparison fall through to a
stray LATER heading instead. `selectedBracketId` is threaded through as
`null` at the preambleCutoff call site for this reason — bracket-shaped
(and, from the next commit, phase-tail) discrimination still applies
uniformly at both sites; only the same-milestone component is
call-site-specific, since it encodes a "keep scanning past this
heading" instruction with no counterpart in a top-of-document search.

Tests: new describe block "#612 PR-2 B1 round-2: a same-milestone
continuation heading is not a boundary" — RED-turned-GREEN fixtures for
repro8 case 1 and repro5, plus PINs for repro8 case 3 (trailing
different-id icebox still terminates) and repro10 A1 (all-version-
bearing + icebox + Phase Details stays exactly 2/1/50 — no double-count
from the same-milestone exclusion interacting with the pre-existing
detailsMatch re-append). syncedTotal()/syncedPercent() assertions
omitted from the repro10 A1 pin: that fixture carries dirs outside the
current milestone, which exposes the SEPARATE Major 1 defect
(cmdStateSync's body percent from an unfiltered disk scan) — asserted
once Major 1 is fixed, not here.

Full suite green (796/796 across the targeted adr-612 + roadmap-parser
+ state files); node scripts/lint-phase-id-drift.cjs clean; eslint
clean on both changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): bracket boundary discriminates by content, not heading level (B2)

Gate-2 adversarial review Blocker 2: three sites disagreed about which
heading levels are a bracket milestone. The selector
(roadmap-parser.cts's bracket-fallback SELECTION branch,
`^#{1,3}\s+\[CODE.MM\]`) and `isMilestoneBounded` (state.cts) both
admit level 1-3, but isBracketMilestoneBoundary's level cap only
admitted level 1-2 (`h.level <= 2`, from the B1 commit). A `###`-level
bracket milestone heading was therefore SELECTED and BOUNDED but never
TERMINATED: computeSectionEnd ran with level=3, a level-3 SIBLING
milestone survived the pre-existing `h.level > level` (not-deeper)
filter, failed the version/emoji test (version-less), then failed
`h.level <= 2` — falling through to `return content.length` and
sweeping the sibling milestone's own phases into the current one.
Reproduces trek-e's original #612 defect verbatim ("a safe degrade
became a confidently-wrong persisted number") on a heading level the
selector and bounding predicate both already admit (repro2 case C:
4/75% instead of 2/100%; mechanism confirmed directly via repro7 —
extractCurrentMilestone returned the whole 214-byte document).

ADR-612 Decision 1 (docs/adr/612-bracket-phase-id-convention.md:56)
specifies the discriminator as CONTENT, not level: "a phase heading is
a bracket followed by a digit-then-colon ([GSD.02] 05:); a milestone
heading is a bracket followed by a name." Replaces the `level > 2`
rejection with BRACKET_PHASE_TAIL_RE — built by interpolating
phase-id.cts's single-owner phaseHeadingPrefixSrcFor(ANY_BRACKET,
'bracket', false) plus the digit + optional-tag + colon tail every
phase-heading counter in this file already spells, not a re-typed
grammar — and widens the level check to a depth-sanity cap of 3
(mirroring the selector's own `#{1,3}` ceiling; NOT itself a
phase/milestone discriminator). Covers the dotted sub-phase heading
form (`[GSD.02] 05.03:`) via the same `[\w][\w.-]*` token, pinned by a
new fixture — the shape where a regex slip in the tail grammar would
hide.

preambleCutoff's own raw-scan regex is widened from `^(#{1,2})` to
`^(#{1,3})` in lockstep: the outer pattern's level ceiling must track
the helper's cap, or a level-3 PRIOR milestone heading is invisible to
that scan and its own phase heading leaks into the preamble
un-stripped (a real double-count this widening closes, verified
against repro2 case C directly).

The existing "boundary heading level: a level-3 CURRENT milestone
heading still scopes correctly" pin (3e562f12) now passes via a
DIFFERENT mechanism than before — its own neighbours are version-
bearing, so it previously passed via the version/emoji rule (the level
cap was never actually exercised by that fixture, per the round-2
review's own finding); with the content discriminator, the SAME
fixture's level-3 phase headings are now correctly excluded because
they are phase-tail-shaped, not because they are too deep. A
deliberate mechanism change, confirmed by re-running that test green
after this commit.

Also updates the "every selector call site declares the right
baseline" governance pin (adr-612-bracket-heading-selection.test.cjs):
BRACKET_PHASE_TAIL_RE is a new, legitimate ANY_BRACKET call site in
roadmap-parser.cts (always passing the literal 'bracket' convention,
since its only caller is already gated on bracketBoundaryActive) —
EXPECTED count bumped 1->2, with a matching BASE_SITES transcription
entry (identical src to every other ANY_BRACKET site, since the
function is pure).

Tests: new describe block "#612 PR-2 B2 round-2: the bracket boundary
is a CONTENT discriminator, not a level cap" — RED-turned-GREEN for
repro2 case C (exact total AND truthful percent, since
isMilestoneBounded already returns true at #{1,3}) and repro7's
mechanism, a PIN for the dotted sub-phase form, and a re-pin of repro8
case 3 (icebox) under the new mechanism.

Full suite green (907/907 across the targeted adr-612 + roadmap-parser
+ state + phase-id files); node scripts/lint-phase-id-drift.cjs clean;
eslint clean on all changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): fence-aware preamble cutoff on the bracket branch (Blocker 3)

Gate-2 adversarial review Blocker 3: preambleCutoff's bracket scan used
a raw content.match/matchAll — blind to fenced code blocks — while its
sibling computeSectionEnd (a few lines above it) already consumed
tokenizeHeadings(content), which strips fences. The two halves of one
boundary semantic disagreed about what a heading is.

A fenced markdown example in the preamble containing a bracket heading
(ADR-612's own docs do exactly this) was textually the earliest
`#{1,3} [CODE.MM]` match: preambleCutoff landed INSIDE the fence,
`preamble = content.slice(0, preambleCutoff)` ended with an unclosed
opener, and the unbalanced fence then blinded
getMilestonePhaseFilter's own tokenizeHeadings(scope) call — every
heading in the returned scope vanished, phaseCount degraded to 0, and
the pass-all filter admitted every directory on disk (repro11's
mechanism, confirmed directly: fence count 1/odd, tokenizeHeadings(scope)
-> only "Roadmap"). Regression vs round-1, which had no bracket pattern
to blind and so fell back to the correct heading (repro12 bracket row:
2/1/50 at round-1, 4/3/75 at HEAD).

Fixed by hoisting one tokenizeHeadings(content) call
(currentMilestoneHeadings) shared by computeSectionEnd and the
preambleCutoff scan, which now iterates that same fence-aware token
list instead of a raw regex. HeadingToken.text is already hash-stripped
and trimmed, so isBracketMilestoneBoundary needs no `^#{1,3}\s+`
re-derivation at this site (that spelling would not match h.text — a
note the round-2 review called out explicitly, confirmed while
porting). selectedBracketId stays `null` here, unchanged from the B1
commit's same-milestone-exclusion reasoning.

DISCLOSED, not fixed (explicitly out of scope per the round-2 review's
own minimal-fix note): the LEGACY (non-bracket) anyMilestonePattern
raw-match path shares the identical fence-blindness hazard and stays
byte-identical — a bracket repo whose preamble has a fenced
VERSION-BEARING heading still has the legacy raw-match win the
earliest-of-either min() (repro12's LEGACY control: 4/3/75, unchanged
across base/round-1/HEAD/this commit). Pinned here so a future reviewer
files this as a known, pre-existing gap rather than a new regression.

Tests: new describe block "#612 PR-2 Blocker 3 round-2: preambleCutoff
is fence-aware (bracket branch only)" — RED-turned-GREEN for repro12's
bracket row and repro11's mechanism (fence balance + non-degraded
phaseCount + correct per-directory admission), a PIN for repro12's
LEGACY control (the disclosed gap, explicitly unchanged), and a PIN for
repro10 A3 (a fenced heading INSIDE the current section must still not
terminate it).

Full suite green (1072/1072 across the targeted adr-612 + roadmap-
parser + state + phase-id + markdown-sectionizer files); node
scripts/lint-phase-id-drift.cjs clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): scope cmdStateSync's disk scan by milestone under bracket (Major 1)

Gate-2 adversarial review Major 1: `state sync` wrote a Progress
PERCENT computed from an UNFILTERED whole-disk scan, beside the
milestone-scoped total_phases/completed_phases it writes into the same
STATE.md via the refreshed frontmatter (syncStateFrontmatter ->
buildStateFrontmatter, which has always applied getMilestonePhaseFilter
for the READ path). cmdStateSync's own `fs.readdirSync` chain (the
WRITE-path scan) never called the milestone filter at all, unlike
buildStateFrontmatter's identical-purpose scan. One command therefore
wrote two contradictory numbers into one file: on the ADR-canonical
version-less bracket fixture (4 dirs, 3 complete; asserted milestone =
2 phases, both complete), base wrote total_phases:2/completed_phases:2
(correct, from the READ derivation) alongside body Progress 75% (wrong
— from the unfiltered WRITE derivation; repro3).

Fixed by threading `getMilestonePhaseFilter(cwd)` through the same
`.filter()` chain buildStateFrontmatter already applies, gated on
`syncConvention === 'bracket'` (falling back to a pass-all predicate
otherwise) — so totalDiskPlans/totalDiskSummaries/diskCompletedPhases/
syncTotalPhases become milestone-scoped under bracket, byte-identical
under legacy.

DEVIATION (approved, stated plainly): an earlier phrasing of this fix
called for mirroring buildStateFrontmatter's filter UNCONDITIONALLY.
Implemented GATED instead — an unconditional filter would ALSO move
every LEGACY repo's persisted percent, since the milestone-scoping-vs-
whole-disk divergence this closes is engine-wide, not bracket-specific.
The gate keeps legacy byte-identical, which is the binding constraint:
this is a bracket read-path PR, not a legacy behavior change.

Nit 2 (informational, no code change): 10 calls to
extractCurrentMilestone on a legacy repo cost 10 config.json
existsSync + 10 readFileSync (0 before B3); accepted, unmemoized cost,
unaffected by this commit.

Also folds in two minors from the round-2 review:
- Corrects .changeset/2761-bracket-read-tolerance.md: the sibling-
  exclusion sentence now states it holds at any heading level 1-3 and
  across a milestone split over two headings (true again now that
  Blockers 1 and 2 are fixed); the percent sentence states plainly that
  `state sync`'s body percent is now milestone-scoped under bracket,
  and unaffected under legacy.
- Records the read/write scoping divergence at currentMilestoneRawRanges
  (src/roadmap-parser.cts) in a comment: it did not receive B1/B2's
  bracket boundary fixes, currently harmless (its only consumer falls
  back to whole-content mutation, and every mutation there is still
  Phase-labelled-only, not bracket-widened), but live the moment the
  write path is bracket-widened — flagged so a future PR closes it in
  lockstep with that work, not after.

Tests: 6 pre-existing tests in tests/adr-612-bracket-phase-counting.test.cjs
needed fixture updates, not logic changes — they used the default
single directory (`GSD.02-01-setup`, phase "01"), which the SENTINEL/
retirement/mixed-heading fixtures in those tests never declare as a
real phase (only 04/05/06/999/etc are declared). Before this fix,
cmdStateSync's unfiltered scan counted that off-roadmap directory
anyway; after this fix the milestone filter correctly excludes it,
which for several of these fixtures made `state sync` a no-op (the
computed 0% coincided with STATE.md's initial template default) and
broke `syncedTotal()`/`syncedPercent()`'s ability to observe anything.
Updated each to pass an EXPLICIT directory naming one of the fixture's
REAL declared phases, preserving each test's original numerator/
denominator intent. One test — "shape 2 WRITE" — was substantively
rewritten: it was a CHARACTERIZATION of the Major 1 bug itself ("the
DISCLOSED legacy gap, mirrored — not closed"), and now correctly pins
bracket closing to 33% (agrees with the read path) while legacy stays
at the disclosed 60% (unchanged, deliberately, per the gating decision
above).

Full suite green: `npm test` 1449/1449 (0 fail, 0 skipped, 0 todo,
single-shard "all" run — includes issue-2765-brace-expansion-lockfile
passing); `npm run lint:ci` clean (0 errors; 2 pre-existing timing-
assertion warnings in files this PR does not touch); node
scripts/lint-phase-id-drift.cjs clean; node scripts/changeset/lint.cjs
ok.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): preambleCutoff identity is offset- and child-aware (round-3 Blocker 1)

Gate-2 round-3 re-verify Blocker 1 (NEW): the round-2 B1 deviation
(39c42a89) threaded `selectedBracketId` as the real value at
computeSectionEnd but as bare `null` at the preambleCutoff scan. The
deviation's rationale — "the selected heading's own occurrence is
always a correct earliest answer" — was right, but `null` disables the
same-milestone check for EVERY candidate, not just the selected one.
Any bracket-shaped heading earlier than the selected milestone was
accepted as a boundary regardless of identity: a same-id checklist/
overview heading preceding the version-bearing selected heading (cases
A, B — the version lands on the LATER half of a split, or a plain
overview heading with no "(Phase Details)" spelling), or a DIFFERENT-id
bracket-shaped PROSE heading with no children of its own sitting above
the current milestone's content (case D — `## [ADR.612] Heading
convention used by this roadmap`). In every case the region between
that false boundary and the real sectionStart was silently dropped —
a completed phase vanished and `state sync` persisted a confident 0%
where base and round-1 both correctly wrote 50%. Regression vs base
AND round-1 (not merely "under-fixed", per the round-3 review's own
severity note).

Fixed with two changes, both scoped to the preambleCutoff scan only
(computeSectionEnd already threads the real `selectedBracketId` and is
untouched):

(a) `h.offset === sectionStart` now bypasses BOTH the same-milestone
    check inside isBracketMilestoneBoundary (passing the REAL
    `selectedBracketId` for every other candidate) and the new child
    rule below — the selected heading's own position is definitionally
    the correct answer, so neither discriminator should run against it
    (rejecting it would mean rejecting the heading against ITSELF).
    Closes cases A and B — verified by the reviewer's own one-liner,
    reproduced here.

(b) New `bracketHeadingHasMatchingChild`: an otherwise-accepted
    candidate (bracket-shaped, not phase-tail-shaped, not the same id
    as the selected milestone) must ALSO have a next-strictly-deeper
    heading carrying its OWN bracket id to count as a boundary. This is
    what a genuine sibling milestone has (its own phase children share
    its bracket id — `## [GSD.01] Setup` / `### [GSD.01] 01: …`) and an
    unrelated bracket-shaped prose heading does not. A candidate with
    no such child at all (childless — e.g. an empty prior milestone, or
    one immediately followed by a same-or-shallower heading) degrades
    to NOT a boundary — over-inclusive, the safe direction: its own
    heading text stays in the preamble, contributing nothing to any
    phase count (not phase-shaped). Closes case D, which (a) alone does
    not — verified: without this rule, `[ADR.612]`'s prose heading is
    indistinguishable from a genuine prior sibling at this site.

As a side effect, also neutralizes Nit 2 (a colon-less `[GSD.02] 05`
heading spuriously terminating the preamble): a colon-less bracket
heading is not phase-tail-shaped so isBracketMilestoneBoundary alone
would accept it, but it is — precisely because it is malformed/
incomplete rather than a real milestone — childless, so the child rule
rejects it too. Pinned.

Known interaction with the fence-blind SELECTION path (disclosed by
the reviewer, not introduced here, tracked for the next commit): when
`sectionPattern` selects a FENCED version-bearing heading (an
extremely pathological shape — a fenced example whose text happens to
match STATE's asserted version), no token exists at `sectionStart`, so
the `h.offset === sectionStart` bypass never fires and the loop falls
through to the ordinary same-id / child-rule checks. This composes
with the round-3 Major 1 fix (next commit) rather than introducing a
new defect — SELECTION itself is untouched by any of this — but is
worth stating plainly rather than rediscovering.

Tests: new describe block "#612 PR-2 Blocker 1 round-3: preambleCutoff
identity is offset- and child-aware" — RED-turned-GREEN for cases A, B
(rv-attack1) and D (rv-attack1b) with syncedTotal()/syncedPercent()
assertions (the persisted 0% is the point), a PIN for a genuine prior
sibling with real children (still excluded), a PIN for a childless
prior sibling (degrades to not-cutting, over-inclusive/safe), a PIN
for the colon-less Nit 2 shape, and the reviewer's rv-mech1 mechanism
re-run as a proper test (scope now equals the full input document,
phaseCount 2, both dirs accepted).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 937/937 pass (930 baseline + 7 new). node
scripts/lint-phase-id-drift.cjs clean; eslint clean on both changed
files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): fence-aware version/emoji half of preambleCutoff on the bracket branch (round-3 Major 1)

Gate-2 round-3 re-verify Major 1 (NEW): ff6bf0a8 (round-2 Blocker 3)
made the BRACKET half of preambleCutoff's "earliest milestone-shaped
heading" search fence-aware, but left the VERSION/emoji half a raw
`content.match` even on the bracket branch. A fenced VERSION-BEARING
example heading in a bracket repo's preamble (ADR-612's own docs
illustrate the LEGACY heading shape exactly this way, inside a fenced
authoring-guide block) was still textually the earliest match for that
raw regex, winning the min() and un-suppressing a wrong persisted 75%
that base correctly suppressed (rv-attack3c fixture C1: base
suppressed the percent entirely — `isMilestoneBounded` false — HEAD
wrote 75% where truth is 50%).

Fixed by deriving the version/emoji half from the SAME fence-aware
`currentMilestoneHeadings` token list as the bracket half, on the
bracket branch only — the exact `/^Phase\s+\S/i` / `/v\d+\.\d+|✅|📋|🚧/i`
pair `computeSectionEnd` already uses against `h.text`. The non-bracket
(legacy) path is untouched: it keeps the raw `content.match`, byte-
identical to before, including its own fence-blindness (repro12's
LEGACY control, pinned unchanged in the round-2 Blocker-3 test block —
not re-pinned here to avoid duplicating an already-covered assertion).

Not rated Blocker (per the review) because it is not a regression vs
round-1 and the fixture (a version-BEARING fenced example in a bracket
repo) is rarer than the already-fixed bracket-heading case; still
fixed now rather than disclosed, per this arc's own precedent (every
prior "disclose instead of fix" call in this PR has been overturned on
re-review).

Tests: new describe block "#612 PR-2 Major 1 round-3: preambleCutoff's
version/emoji half is fence-aware on the bracket branch" — RED-turned-
GREEN for case C1 (readTotal + syncedPercent, so the persisted 75% is
directly observed, not just the read-path total), PIN for case C2 (the
already-fixed fenced-bracket-heading shape, unchanged).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 939/939 pass (937 + 2 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean on both changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): correct changeset claims + mark runtime-gated BASE_SITES row (round-3 minors)

Gate-2 round-3 re-verify Minor 2: three changeset sentences in
.changeset/2761-bracket-read-tolerance.md were overstated in a new
direction after round-2:

- The split-milestone claim ("across a milestone split over two
  headings") was true only when the version-bearing heading came
  FIRST (repro8 case 1); false when it came LATER (round-3 Blocker 1
  case A). Now restated to say plainly "with the version-bearing
  heading in EITHER position" — true again now that round-3's Blocker
  1 fix lands earlier in this range.
- The "counted from the phases... rather than from every directory on
  disk" claim was false on cases A/B/D (a strict subset of the
  milestone's own phases). Restated as "ALL of the phases... not a
  subset", and extended to state that an unrelated bracket-shaped
  heading with no phase children of its own (case D's `[ADR.612]`
  shape) does not truncate the milestone either — true now, not before.
- "Each widened read is SELECTED by the project's phase_id_convention"
  was literally false for BRACKET_PHASE_TAIL_RE, which is RUNTIME-gated
  (via its only caller, isBracketMilestoneBoundary, itself only
  consulted when bracketBoundaryActive) rather than selector-gated.
  Restated behaviourally: "every widened read ENGAGES only when the
  project's resolved phase_id_convention is bracket" — true for both
  gating mechanisms, so it no longer implies a selector call this site
  does not make.

Minor 1: the STRUCTURAL IDENTITY test's BASE_SITES row for
BRACKET_PHASE_TAIL_RE (added in the B2 commit) asserts a property of
`phaseHeadingPrefixSrcFor` — the function — not of the call site; it
would pass unchanged even if the site were deleted. Safety at that
specific site rests entirely on a runtime gate the test cannot see.
Added `runtimeGated: true` to the row and threaded it into the
generated test's own title (`… [runtime-gated, not selector-covered]`),
so the gating mechanism is visible in test OUTPUT, not only in a source
comment that could drift silently.

No production code changed. Targeted suite green (48/48 in the
affected file); `node scripts/changeset/lint.cjs` ok; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): harden round-3 preamble cutoff — subtree child scan + level cap

Team-lead review of f87bba0e found two edges in the round-3 preamble-
cutoff code, both hardened here.

AMENDMENT 1 — bracketHeadingHasMatchingChild (2e06aef5) checked only
`headings[index + 1]`, the IMMEDIATE next heading, not the candidate's
whole subtree. A genuine prior sibling milestone whose section opens
with a non-bracket subsection before its first phase heading
(`## [GSD.01] Setup` / `### Notes` / `### [GSD.01] 01: Old`) was
therefore wrongly rejected as a boundary — its real phase heading sits
TWO headings deep, not one — leaking its entire section into the
preamble unstripped.

CONFIRMED RED, not merely theoretical (built and ran the fixture
against f87bba0e before touching the fix, per instruction): scope
membership DOES drive the disk-side filter on this shape.
`GSD.01-01-old`'s directory was wrongly admitted into the CURRENT
milestone's filter via the leaked heading's qualified key
(`GSD.01-01`) — 3/2/67% where truth is 2/1/50%.

Fixed by scanning the candidate's full SUBTREE: continue past a
non-matching deeper heading instead of returning false on the first
one; only a same-or-shallower heading actually closes the subtree and
yields "no match found". A candidate whose entire subtree closes with
no same-id hit (including a genuinely childless one) still degrades to
`false` — over-inclusive, safe, unchanged from before.

AMENDMENT 2 — c483552a ported the version/emoji half of preambleCutoff
to the token-based scan with no level cap; the raw
`content.match(anyMilestonePattern)` it replaced was anchored
`^#{1,3}\s+`. A level-4+ version-bearing heading in the preamble
(`#### v2.0 notes`) therefore won the scan on the bracket branch where
the raw pattern — and the legacy path, unaffected — ignores it
outright. Fixed with `if (h.level > 3) continue;`, mirroring the
depth-sanity cap isBracketMilestoneBoundary already applies to the
bracket half of this same scan.

Tests: new describe block "#612 PR-2 round-3 hardening: subtree child
scan + level cap on preambleCutoff" —
- RED-turned-GREEN for the Notes-intervening fixture: exact 2/1/50 (was
  3/2/67), plus the disk-filter observable (`GSD.01-01-old` now
  correctly excluded).
- PIN for the level-4 preamble heading: the scope now PRESERVES the
  heading's text (was silently dropped before this fix — harmless in
  this minimal fixture's total_phases specifically, since the dropped
  text carries no phase-shaped content, but a real correctness gap
  against the raw pattern's own ceiling) — asserted via scope content,
  not total_phases, since that number is invariant here either way.
- PIN for the LEGACY control on the same level-4 shape — unchanged,
  confirming the raw content.match path is untouched.

Re-verified the existing genuine-prior-sibling and childless-sibling
pins (round-3 Blocker 1 commit) still pass under the subtree scan —
both fixtures' outcomes are unchanged since their same-id hit (or its
absence) was already at the first deeper heading.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 942/942 pass (939 + 3 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean on both changed files. Full `npm test` + `npm run
lint:ci` deferred to the team lead's own run per instruction.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): bracketHeadingHasMatchingChild requires a same-id PHASE child (round-4 Blocker 1)

Gate-2 round-4 re-verify Blocker 1 (NEW): the round-3 hardening's
subtree scan (fbfd0fca) proved SAME-ID-NESS but never asked whether
the matching child was PHASE-shaped. Case F1 re-opens round-3's case D
one heading later: `## [ADR.612] Heading convention` is followed by
its OWN sub-heading `### [ADR.612] Examples` — same bracket id as the
candidate, but MILESTONE-shaped (a name, no digit-then-colon), not a
phase. Same-id-ness alone satisfied the subtree scan and re-cut the
preamble at exactly the shape the round-3 hardening was written to
close.

Failing input: `## [ADR.612] Heading convention` / `### [ADR.612]
Examples` (prose) / `### [GSD.02] 01: One` (the current milestone's
own first phase, now unreachable) / `## [GSD.02] v2.0: Foundation` /
`### [GSD.02] 02: Two`. Truth 2/1/50. HEAD read 1/0/0, and `state sync`
reported "nothing to do" (exit 0, `{synced:true,changes:[]}`) because
its wrong 0% happened to equal the STATE.md seed — a half-done
milestone read as untouched with no write-path signal at all.

Fixed with the reviewer's one-conjunct addition: a same-id child only
counts if it is ALSO phase-tail-shaped (`BRACKET_PHASE_TAIL_RE`) — the
same single-owner discriminator `isBracketMilestoneBoundary` already
uses one level up for the identical distinction (phase vs milestone),
reused here rather than re-derived. This is exactly what the
changeset's own wording already claimed ("no phase children of its
own") — the code now matches the sentence rather than the other way
around.

Docstring updated at the function itself: the rule is "same-id PHASE
child", not "same-id child".

Tests: new describe block "#612 PR-2 round-4 Blocker 1: the same-id
child must be PHASE-shaped" — RED-turned-GREEN for F1 with
syncedTotal()/syncedPercent() (the persisted 0% — and the
report-nothing-to-do write-path silence — is the point), PINs for F11
(colon-less same-id child) and F11b (bullet-only phase list): both
correctly stay excluded either way, and the leak the phase-shape
requirement newly creates for these two shapes is INERT — a colon-less
heading forms no qualified key (getMilestonePhaseFilter's own
phase-heading pattern requires the colon too) and a bracket bullet
never matches the legacy-only BULLET_PHASE_LINE_PATTERN — confirmed
directly via getMilestonePhaseFilter, not merely inferred. Re-verified
the four existing child-rule pins (F2 subtree-closure, F3 deep-nested
same-id, F4 level-4 same-id, F5 childless-at-EOF) are unaffected by the
phase-shape requirement, since every one of them already used a
colon-bearing same-id child.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 946/946 pass (942 + 4 new). Zero drift measured across the full F1-F12
corpus except F1 itself (F9/F10/F12 remain red, deferred to the
separate Major 1 fix). node scripts/lint-phase-id-drift.cjs clean;
eslint clean on both changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): fence-aware phase counting, milestone bounding, and bracket-fallback selection (round-4 Major 1)

Gate-2 round-4 re-verify Major 1 (NEW): roadmapPhaseCount is a
fence-blind raw `.exec()` over the scope string, duplicated in TWO
independent copies (buildStateFrontmatter's read path, cmdStateSync's
write path). With the bracket alternative now compiled into it (#612),
a fenced EXAMPLE phase heading in the preamble inflates total_phases
and persists a wrong percent that base got right. Two further
fence-blind sites participate: isMilestoneBounded (a raw
`.test(roadmapRaw)`) and the bracket-fallback SELECTOR inside
extractCurrentMilestone (a raw `content.matchAll`, only reachable when
version-string selection finds nothing).

Failing inputs:
- F10 (clean isolate, version-bearing selection): a fenced
  `### [GSD.02] 05: Example phase` in the preamble inflates
  total_phases 2->3, persisting 33% where truth is 50%. LEGACY control
  on the same shape is correct on every build — not a pre-existing
  hazard being inherited, bracket-only.
- F9 (version-less selection): a fenced example carrying the project's
  OWN milestone id additionally confuses the bracket-fallback selector
  (the fenced heading gets SELECTED), compounding with the same
  fence-blind counter. base suppressed the percent; round-1 and HEAD
  both wrote 33%.
- F12 (isMilestoneBounded isolate): the ONLY `[GSD.02]` heading in the
  document is inside a fence, and the asserted milestone genuinely has
  no section at all — HEAD persisted 67% where base correctly
  suppressed the percent (the milestone is absent from the roadmap).

Fixed at the CONSUMER level, not the producer — extractCurrentMilestone's
returned scope string is deliberately UNCHANGED, since every other
consumer of that string needs its full content fidelity and legacy
identity forbids touching the shared string (this branch's own
precedent, ff6bf0a8/c483552a, was producer-level; here the ruling is
consumer-level because the string is shared far more broadly than the
two round-3 fixes' narrower producer edits):

(a) New `countRoadmapPhaseHeadings` (src/state.cts, immediately above
    extractRetiredPhaseNumbers) — ONE shared implementation for both
    call sites, replacing two independently-maintained copies. BRACKET
    convention counts via `tokenizeHeadings(scope)` at levels 2-4,
    testing each heading's hash-stripped text directly — fence-aware by
    construction, since tokenizeHeadings never produces a token for a
    fenced line. LEGACY convention keeps the exact pre-existing raw
    `.exec()` loop, byte-for-byte. A pre-existing, deliberately
    PRESERVED asymmetry between the two original call sites — the read
    path always excluded a bare `/^999\b/` token, the write path never
    did — is threaded through as an explicit
    `includeUnconditional999Check` parameter per call site, so sharing
    the implementation does not silently unify (and thereby move)
    either total.
(b) isMilestoneBounded's bracket branch now scans
    `tokenizeHeadings(roadmapRaw)` for a matching heading (level <= 3)
    instead of a raw regex test. Legacy version-string branch untouched.
(c) The bracket-fallback SELECTOR now builds its candidate set from
    `tokenizeHeadings(content)` instead of `content.matchAll`,
    reconstructing a match-shaped array so every downstream consumer of
    `headingMatches` sees the identical shape the raw-regex path always
    produced. This is the ONE site in this entire arc where SELECTION
    itself changes — selection SEMANTICS are otherwise unchanged (same
    pattern, same first-match-wins by document order); only the
    candidate set is now fence-aware. Pinned that unfenced selection is
    byte-identical.

Zero drift measured across the full historical corpus (repro2-13,
rv-attack1/1b/3c, rv-mech1, rv2-amend1/2, and F1-F11b) except the three
target fixtures.

Tests: new describe block "#612 PR-2 round-4 Major 1: four fence-blind
sites on the bracket path" — RED-turned-GREEN for F10 (with
syncedTotal()/syncedPercent()) and F9 (both layers), PINs for F10's
LEGACY control and F10c (non-phase-shaped fence, unaffected either
way), RED-turned-GREEN for F12 (asserts the percent KEY is absent from
`state json`'s output and that `state sync`'s body stays at its
unmodified seed — the persisted-suppression signal, not merely a
total_phases number), and an explicit PIN that unfenced bracket-fallback
selection (first real milestone-shaped heading wins, no fences
involved) is unaffected.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 952/952 pass (946 + 6 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean on all three changed files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): correct docstring overstatement + stale consumer-count sentence (round-4 minors)

Gate-2 round-4 re-verify Minor 1 + Nit 1. No production code changed.

Minor 1: bracketHeadingHasMatchingChild's own docstring said a
rejected (no-same-id-PHASE-child) candidate's degrade "contributes
nothing to any phase count" — true of the candidate's OWN heading
text, but not of its SUBTREE, which is what actually stays in the
preamble. F7 (`## [GSD.01] Setup` / `### [GSD.07] 01: Foreign`) shows
a DIFFERENT-id bracket PHASE heading inside a rejected candidate's
subtree DOES form a qualified key and CAN admit a foreign directory —
3/2/67%, stable across base, round-1 and HEAD (base via its own
pass-all degrade). Not a regression, still the declared over-inclusive
/ never-under-inclusive safe direction — the comment now says that,
with F7's numbers cited, at the call site that actually decides
`isBoundary` (roadmap-parser.cts's preambleCutoff loop) rather than
only at the helper's own definition.

Minor 2 (changeset) — VERIFIED, no wording change needed: re-ran F1,
F9, F10, F12 at this HEAD. The "no phase children of its own... does
not truncate it either" sentence (naming the `[ADR.612]` shape
directly) is now literally true — F1 reads 2/1/50. The "counted from
ALL of the phases... not a subset" sentence is now true on every
measured shape — F1/F9/F10 all read 2/1/50, F12 correctly suppresses
the percent. No carve-out for F9/F10 is needed since round-4 Major 1
(3be5c412) closes both; per the fix-round instruction to "only carve
out anything genuinely left," nothing is.

Nit 1: tests/adr-612-bracket-heading-selection.test.cjs's runtimeGated
row claimed `BRACKET_HEADING_INTRO_RE` has "no other consumers" — true
when round-3's f87bba0e wrote it, stale since 2e06aef5 (round-3's own
earlier commit) had already added two more uses inside
bracketHeadingHasMatchingChild. Corrected to state the true count
(three consumers) and re-confirm the conclusion is unaffected: all
three are still nested inside the same bracketBoundaryActive runtime
gate, and BRACKET_HEADING_INTRO_RE is built from BRACKET_ID_SRC, not
phaseHeadingPrefixSrcFor, so it was never a selector site regardless of
consumer count.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 952/952 pass (comment-only changes, no count movement). eslint
clean; node scripts/changeset/lint.cjs ok.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): restore the bracketId guard in countRoadmapPhaseHeadings (round-5 Blocker 1)

Gate-2 round-5 re-verify Blocker 1 (NEW, introduced by 3be5c412): the
merge that created the shared countRoadmapPhaseHeadings helper dropped
the `bracketId && ` guard both original inline loops carried before
calling isSentinelPhaseId. Every other isSentinelPhaseId call site in
src/ (roadmap-parser.cts, roadmap.cts, validate.cts x2, verify.cts)
keeps the guard; state.cts's shared counter was the only one of seven
without it.

When the phase-heading-intro grammar's LEGACY alternative matches (a
`### Phase 00:` heading in a `phase_id_convention: "bracket"` repo —
the mid-migration shape this PR exists for), the bracket capture group
is `undefined`, so the unguarded call became
`isSentinelPhaseId("undefined-00", 'bracket')` — measured TRUE, so
phase 00 (and 000, 0a, 0.5, 999.1 — any token whose splice with the
literal string "undefined" happens to fall in a sentinel range) was
silently dropped from the denominator. `getMilestonePhaseFilter` (which
still carries its own guard) counts the phase and admits its directory
regardless, so the filter and the counter disagree — a half-done
milestone reads as 100% complete, persisted.

One-line fix, restoring the guard every sibling call site already has:

    if (bracketId && isSentinelPhaseId(`${bracketId}-${token}`, 'bracket')) continue;

Line count: `git diff --stat src/state.cts` -> 1 file changed, 1
insertion(+), 1 deletion(-).

Tests: new describe block "#612 PR-2 round-5 Blocker 1:
countRoadmapPhaseHeadings restores the bracketId guard" — RED-turned-
GREEN for G3 (3/2/67, was 2/2/100) and G3d (the mixed bracket+legacy
mid-migration shape, same numbers) with syncedTotal()/syncedPercent(),
PIN for G3's legacy control (unaffected), PIN for G3b (isolates the
counter with no directory to admit), PIN for G3c (legacy 01/02 only,
no sentinel-shaped token present).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 957/957 pass (952 + 5 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): extractRetiredPhaseNumbers is fence-aware on the bracket path (round-5 Major 1)

Gate-2 round-5 re-verify Major 1 (NEW) — a FIFTH fence-blind site on
the bracket path, missed by 3be5c412's own enumeration of "four".
extractRetiredPhaseNumbers' line scan (`scope.split(/\r?\n/)`) has no
fence awareness. This PR compiles the bracket alternative into
`introSrc` ("the retirement filter has to widen with the counter it
protects" — the function's own pre-existing comment), so a FENCED
authoring EXAMPLE showing the #1514 retirement gesture in bracket
spelling is now indistinguishable from a real one: it retires a
genuine phase, shrinking the denominator and persisting a confident
100% where base correctly read 50%.

Fixed with the SAME consumer-level ruling this arc has used at every
other fence-blind site, reusing markdown-sectionizer's existing
exported `stripFencedCode` rather than hand-rolling a second fence
parser (single-owner rule) — retirement lines are BULLETS, not
headings, so `tokenizeHeadings` doesn't serve here; `stripFencedCode`
is the general-purpose fence stripper the tokenizer itself is built on.
Gated on `convention === 'bracket'`; the LEGACY line scan stays the raw
`scope` string, byte-identical — its own fenced-example hazard is
pre-existing (wrong at base too) and out of scope.

Line count: `git diff --stat src/state.cts` -> 1 file changed, 9
insertions(+), 2 deletions(-) — one import added, four lines inside
the function (a comment + the `scanScope` computation + the changed
`.split()` call).

Tests: new describe block "#612 PR-2 round-5 Major 1:
extractRetiredPhaseNumbers is fence-aware on the bracket path" —
RED-turned-GREEN for G2 (fenced example in the preamble) and G2b (the
same example placed INSIDE the milestone section, ruling out a
preamble-scoping artifact — the site itself was fence-blind wherever
the fence sits), PIN for the LEGACY control (unchanged, pre-existing,
out of scope — base is wrong on this shape too).

Also folds in the "four fence-blind sites" correction: 3be5c412's
commit message and any restatement of it should read FIVE going
forward; this commit's own message states the count correctly.

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 960/960 pass (957 + 3 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): state sync's counter excludes the bracket 999 icebox token (round-5 Major 2)

Gate-2 round-5 re-verify Major 2 (NEW) — `includeUnconditional999Check`
left `state json` and `state sync` reporting different totals for one
bracket repo. Under bracket, READING-B puts the sentinel in the
bracket, so `isSentinelPhaseId("GSD.02-999", 'bracket')` is false —
the `/^999\b/` TOKEN rule is the only thing excluding a
`### [GSD.02] 999:` icebox heading, and it ran on the read path
(buildStateFrontmatter, `true`) and on getMilestonePhaseFilter
(unconditional), but not on cmdStateSync's own counter (`false`). One
`state sync` call could leave a single STATE.md with its own
frontmatter (percent 50, from the read-path re-sync inside
writeStateMd) and body (percent 33, from the write-path counter that
alone still counted the icebox heading) disagreeing — falsifying this
PR's own stated invariant that sharing countRoadmapPhaseHeadings made
"the two counters must see the same phases" structural.

Functional change is one argument, exactly as specified: the write
call site now passes `syncConvention === 'bracket'` instead of the
literal `false`. `syncConvention === 'bracket'` is `false` for every
non-bracket value, so the LEGACY path resolves to the exact same
`false` it always did — this file's own pre-existing, deliberately-
unchanged read/write divergence on that path is untouched. The READ
site (`:1860`) is NOT touched — its historical behaviour applied
`/^999\b/` to legacy and bracket alike, so changing it would move
legacy READ totals, exactly the class of mistake this arc's own
Blocker 1 (this round) was.

Line count: the functional change is ONE argument
(`false` -> `syncConvention === 'bracket'`); the surrounding comment
was rewritten because the previous one asserted the now-superseded
behaviour ("preserving this file's pre-existing... divergence... a
bare 999 token is not excluded here") and leaving it would mislead the
next reader — not a structural change.

Tests: new describe block "#612 PR-2 round-5 Major 2: state sync
excludes the bracket 999 icebox token like the read path" —
RED-turned-GREEN for G1, asserting `state sync`'s body percent equals
`state json`'s own percent (both 50, not 33 vs 50), PIN for G1's
legacy control (33 vs 50 unchanged, deliberately).

Targeted suite green: `node --test tests/adr-612-*.test.cjs
tests/roadmap-parser.test.cjs tests/state.test.cjs tests/verify.test.cjs`
-> 962/962 pass (960 + 2 new). node scripts/lint-phase-id-drift.cjs
clean; eslint clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): restore indent parity at the two line-start-anchored bracket-heading reconstructions (round-5 Minor 1)

`HeadingToken.offset` (tokenizeHeadings) is the LINE-START character offset,
not necessarily the `#` character's own offset — a ≤3-space-indented ATX
heading has both. Two round-4 reconstructions built on tokenizeHeadings
inherited this gap and accepted indented headings their raw, line-start-
anchored predecessors (`^#{1,3}\s+\[...`) never matched:

  1. roadmap-parser.cts's bracket-fallback SELECTOR (extractCurrentMilestone,
     ~line 377) — an indented, version-less `[GSD.02]` milestone heading
     could be reconstructed into headingMatches, then mis-parsed downstream
     (selectedBracketId null, level fallback to 1), leaking a SIBLING
     milestone's phases into the counted scope (G6: HEAD read 3/2/67 instead
     of 2/1/50 — a real phase heading's own directory belonging to the NEXT
     milestone got counted).

  2. state.cts's isMilestoneBounded (~line 1571) — the same gap let an
     indented-only `[GSD.02]`-shaped heading wrongly bound a milestone
     absent from the roadmap, un-suppressing a percent that should stay
     suppressed (mirrors round-4's F12 fenced-only case, but via indentation
     instead of a fence).

Fix: one added conjunct per site — `content[h.offset] === '#'` (source-named
`roadmapRaw` in state.cts) — filtering to tokens whose LINE-START offset IS
the `#` character, i.e. exactly the set the raw line-start-anchored regex
would ever have matched. Restores byte-for-byte raw parity; no other logic
in either function changes. computeSectionEnd and the preamble version/
emoji-token scan are untouched, as instructed — they consumed tokenizeHeadings
output before this arc and are out of scope here.

Also corrects roadmap-parser.cts's now-provably-false docstring claim that
`h.offset` is unconditionally "the same `#`-character coordinate space
`content.match().index` used" — true only for the survivors of the new
filter, not for every token tokenizeHeadings produces.

Line count: the FUNCTIONAL change is exactly 2 lines (one added `&&` conjunct
per call site — `git diff --stat` on the two source files shows 20
insertions/7 deletions, but only those 2 lines change behavior; the rest is
docstring/comment rationale, per this round's "state the line count" ask).

TDD: both fixtures verified RED at HEAD before this commit, GREEN after,
via /tmp/pr612rev/rv5-attack.cjs G6 and a locally-authored isMilestoneBounded-
isolating probe (G6 alone doesn't distinguish the two sites — its unindented
phase headings already satisfy isMilestoneBounded's loose prefix regex
either way, so a second, indentation-only fixture was needed to prove that
site's fix is not a no-op; verified by temporarily reverting just that one
conjunct, confirming 100%-wrongly-bounded RED, then restoring it, confirming
suppressed-percent GREEN).

New tests (adr-612-bracket-phase-counting.test.cjs):
  - RED (G6): indented version-less bracket milestone heading — 2/1/50, not
    the pinned-before-fix 3/2/67.
  - PIN (G6c unindented control): identical document, no indent — 2/1/50
    unaffected on every build.
  - RED (isMilestoneBounded site, indented-ONLY): mirrors round-4's F12
    shape (fenced-ONLY → indented-ONLY) — percent stays suppressed instead
    of the pinned-before-fix wrongly-bounded 100%.

Full G1-G10 (rv5-attack.cjs) + G2b/G3c/G3d/G6c (rv5b.cjs) re-verified
zero-drift against TRUTH after this change. Targeted suite (adr-612-*,
roadmap-parser, state, verify): 962 -> 965 (+3), 0 fail.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): changeset discloses counting-set narrowing + correct stale consumer-count sentence (round-5 minors)

FIX 5 (Minor 2): .changeset/2761-bracket-read-tolerance.md did not disclose
that bracket phase-heading counting is narrower than the raw-regex
predecessor in two ways the review's G4/G5 fixtures surfaced: a level-5
heading (`##### [GSD.02] 05: ...`) is no longer counted (the counter's
tokenizeHeadings scan caps at level 4, matching the selector/isMilestoneBounded
ceiling), and a space-less heading (`###[GSD.02] 05: ...`) is no longer
counted (CommonMark requires ≥1 space/tab after the hashes, which
tokenizeHeadings correctly enforces and the old raw regex did not). Both
G4 and G5 moved from round-1's wrong (inflated) values back to base's
original values as an incidental side effect of routing through
tokenizeHeadings — never a deliberate feature of this PR, and previously
undocumented. One clause added to the existing run-on paragraph; no other
wording in the changeset touched.

FIX 6 (Nit 1): tests/adr-612-bracket-heading-selection.test.cjs:76 —
`BRACKET_PHASE_TAIL_RE` has TWO consumers as of round-4's 65d257ce
(isBracketMilestoneBoundary's own use, plus bracketHeadingHasMatchingChild's
same-id-PHASE-child conjunct), not the "no other consumers" the comment
claimed. Same correction pattern round-4 already applied to this row's
BRACKET_HEADING_INTRO_RE neighbor (that sentence's own staleness was fixed
in 4d7184b8): note the true consumer count, confirm both stay nested inside
the same bracketBoundaryActive runtime gate (verified at
src/roadmap-parser.cts:178 and :250, both reached only through the
`if (bracketBoundaryActive)` block starting at :529), and record which
commit and which fix introduced the drift. Comment-only; no assertion
logic changed.

Line count: 2 files, 8 insertions / 2 deletions total — one added clause
in the changeset (1 line changed) and one comment block replacing the
single stale line in the test file (6 comment lines replacing 1).

Verified: targeted suite (adr-612-*, roadmap-parser, state, verify)
unchanged at 965/965 pass (comment/prose-only diffs, no test count change).
`node scripts/changeset/lint.cjs` and `npx eslint` on the touched files both
clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): restore the bracketId guard on the bracket-only /^0\b/ sibling rule (round-6 Blocker 1)

Round-5's Blocker 1 was `isSentinelPhaseId("undefined-00")` — the merge
that introduced the shared counter dropped the `bracketId &&` guard on the
sentinel check. This is the identical failure one line further down, in the
sibling rule this PR itself added: when the phase-heading grammar's LEGACY
alternative matches (`### Phase 0:` in a `phase_id_convention: "bracket"`
repo — the mid-migration shape this PR exists for), `bracketId` is
`undefined`, and the unguarded `/^0\b/` fires on the bare token anyway.
Neither the LEGACY branch of this same function nor
`getMilestonePhaseFilter` has a `/^0\b/` rule at all, so the filter counts
the phase and admits its completed directory while the counter refuses to
count its heading — a milestone with an unstarted phase 02 persists as a
confident 100%.

`/^0\b/` matches `0` and `0.5` (word boundary before the `.`) but not `00`
(no boundary between the two zeros), which is exactly why round-5's G3/G3d
fixtures (`### Phase 00:`) never tripped this one — same defect class,
different token spelling.

Fix: one word, mirroring the guard round 5 restored two lines above —
`if (/^0\b/.test(token)) continue;` -> `if (bracketId && /^0\b/.test(token)) continue;`.

Expected and intentional side effect: `roadmap analyze` and `state json`
now disagree again on this bracket-repo shape (analyze phase_count=2, json
total_phases=3) — exactly as they already do under the legacy convention
today (verified via /tmp/pr612rev/rv6c.cjs on both conventions). That is
the counter regaining agreement with `getMilestonePhaseFilter` (the tighter
constraint — it is what actually decides `completed_phases`), not a new
break; the counter/filter disagreement is what was wrong.

Out of scope, deliberately NOT fixed here (Minor 1, disclosed via a PIN
test only): the bracket-SPELLED `### [GSD.02] 0:` shape has the same
counter/filter disagreement, but reads 2/2/100 on base too — never closed
by any build in this arc, so it is a pre-existing gap rather than a
regression this commit could introduce.

Line count: src/state.cts is exactly 1 insertion / 1 deletion (one word,
`bracketId && ` prepended to the existing condition).

TDD: T0 (`Phase 0:`) and T05 (`Phase 0.5:`) verified RED at HEAD before this
commit (json 2/2/100, sync body 100%) via /tmp/pr612rev/rv6b.cjs, GREEN
after (3/2/67 on both derivations, matching TRUTH). Zero-drift verified by
diffing the FULL corpus (rv5-attack, rv5b, rv4-attack, rv-attack1,
rv-attack1b, rv-attack3c, rv-mech1, rv2-amend1, rv2-amend2, rv6-attack
H1-H19, rv6b) between a pre-fix and post-fix build: the only differing
lines in the entire diff are T0, T05, and H14 — the three target fixtures.

New tests (adr-612-bracket-phase-counting.test.cjs):
  - RED (T0) + PIN (T0L legacy control)
  - RED (T05) + PIN (T05L legacy control)
  - PIN (B0, bracket-spelled `0` — base-parity characterization, Minor 1,
    not fixed this round)
  Round-5's own G3 test (`### Phase 00:`) already serves as the T00 pin —
  `/^0\b/` never matched `00`, so it is unaffected and untouched.

Targeted suite (adr-612-*, roadmap-parser, state, verify): 965 -> 970 (+5),
0 fail. `lint-phase-id-drift.cjs` and `npx eslint` on touched files clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): changeset re-attributes the phase-count level cap + discloses the bare-0 carve-out; correct two stale test comments (round-6 Minor 1 + Nits)

FIX 2 (Minor 1 — disclosure only, NOT a code fix): the bracket-SPELLED
`### [GSD.02] 0:` / `0.5:` shape has the same counter/filter disagreement
round-6's Blocker fixed for the legacy spelling, but it reads 2/2/100 on
BASE too — never closed by any build in this arc, so it is a pre-existing
gap rather than a regression this round could introduce. Two changes:

  - Pinned as a base-parity characterization: new PIN test (B0) in
    04b5a95f's describe block asserting the exact unchanged value with a
    comment stating why it's deliberately not touched.
  - .changeset/2761-bracket-read-tolerance.md: one qualifying clause added
    to the "counted from ALL of the phases … not a subset" sentence — a
    bare `0`/`0.x` phase token (however spelled) keeps `roadmap analyze`'s
    own pre-existing sentinel reading under bracket too, so it stays
    excluded from these counts. Carried over, not newly introduced by this
    PR — `roadmap analyze` has always read it this way.

FIX 3(a) — same changeset paragraph misattributed the phase counter's
2-4 level cap to "the milestone-boundary machinery in this PR" (the
selector, `isMilestoneBounded`, the preamble scan) — that machinery caps
at `###` (level ≤3), not 2-4. The 2-4 cap belongs to
`getMilestonePhaseFilter`'s own phase scan. Re-pointed the attribution;
the paragraph's earlier, correct ≤3 claim (selector/isMilestoneBounded)
is untouched.

FIX 3(b) — tests/adr-612-bracket-heading-selection.test.cjs:76-82 (added by
1395bd89, round-5's own Nit-1 correction) claimed both
`BRACKET_PHASE_TAIL_RE` consumers are reached "only through the
`if (bracketBoundaryActive)` block starting at :529". Verified at HEAD:
`bracketHeadingHasMatchingChild` is (only caller :593, inside that block).
`isBracketMilestoneBoundary` is not — it has two callers, an inline
`bracketBoundaryActive &&` conjunct at `:473` (inside `computeSectionEnd`)
and the `:529` block at `:567` — the very fact this row's own earlier
"exactly two callers" paragraph already stated correctly. Corrected the
mechanism claim; the CONCLUSION (every consumer is still gated on the same
flag) is unchanged, exactly as round-5's own Nit-1 fix left round-4's
conclusion unchanged when it corrected the consumer count.

FIX 3(c) — tests/adr-612-bracket-phase-counting.test.cjs:2311,2340 still
said "four fence-blind sites" after round-5 (357ba671) added a fifth
(the retirement scan) in its own block below. Reworded both the section
comment and the `describe()` label to read as historical scoping of
round-4's own fix ("the four sites known at round 4 … a fifth was found at
round 5, see its own block below") rather than a live exhaustive claim.

Line count: 3 files, changeset 1/1, heading-selection test 17/8,
phase-counting test 6/2 (comment/prose-only; B0's own PIN test landed in
04b5a95f alongside T0/T05 since it was investigated as part of that same
Blocker's defect class, not in this commit — noted here for the record).

Verified: targeted suite (adr-612-*, roadmap-parser, state, verify)
unchanged at 970/970 (comment/prose-only diffs, no test count change).
`npx eslint` and `node scripts/changeset/lint.cjs` both clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): W021's bracket remediation hint stops pointing at a command that hard-errors (round-7 Minor 2)

`checkBracketCoherence`'s W021 (added by 94abf5df) attached the fix hint
`Run \`gsd-tools roadmap upgrade --convention bracket\` to migrate
(dry-run by default)` to every bracket-convention W021 — both sub-checks
(`missing-bracket` and `mismatch`) share the single `addIssue` call at
src/verify.cts:2255-2256. But `roadmap-command-router.cts:204` throws
unconditionally for any `--convention` value other than
`milestone-prefixed`:

  $ gsd-tools roadmap upgrade --convention bracket
  Error: Only --convention milestone-prefixed is supported

This contradicts the PR's own two disclosures: the changeset ("`bracket`
is a READ-path opt-in until the migrator and write path land") and
docs/CONFIGURATION.md:188 ("There is no bracket migrator and no bracket
emit yet"). Reachability is the exact mid-migration repo this PR targets —
any bracket project with one un-migrated heading gets an unfollowable
instruction on every `validate health`.

Fix: one string. Replaced the hint with what a user can actually do today —
manually align the heading's bracket milestone to its section — and named
the tracked future landing (#612 PR-3) instead of a command that errors.
The milestone-prefixed sibling hint at src/verify.cts:2240 (a different,
already-functional convention/command pair — verified against
roadmap-command-router.cts:65) is untouched.

Line count: src/verify.cts is exactly 1 insertion / 1 deletion (one string
literal).

No test in the suite previously asserted this string's content
(`grep -rn "upgrade --convention bracket" tests/` was empty), so the
unfollowable hint shipped unpinned — `lint-fix-has-regression-test.cjs`
would not have caught a string-only change without a new test. Added one:
asserts the fix string both (a) does not match `--convention bracket`
(the specific pinned-before-this-fix hazard) and (b) equals the new string
exactly, covering the invariant a future edit must not re-break: no
unsupported `--convention` value named in remediation text users are
expected to run verbatim.

Verified end-to-end (not just unit-level) via /tmp/pr612rev/rv7f.cjs: the
W021 issue's `fix` field now reads the new string; `gsd-tools roadmap
upgrade --convention bracket` (and its --dry-run variant) still correctly
hard-error — that command remains unsupported, only the hint text changed.

Targeted suite (adr-612-*, roadmap-parser, state, verify): 970 -> 971 (+1),
0 fail. `lint-phase-id-drift.cjs` clean; `npx eslint` on touched files:
0 errors (1 pre-existing unrelated no-elapsed-assertion warning, same file,
same line this arc's prior rounds already disclosed).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): correct the bare-0 changeset disclosure + disclose W021's second sub-check (round-7 Minor 1 + Nit 1)

FIX 1 (Minor 1) — round-6's bare-0 changeset clause (1395bd89) was wrong in
three falsifiable ways:

  (a) Its own illustration (`### Phase 0:`) is exactly the LEGACY spelling
      04b5a95f made COUNTED. The residual exclusion after that fix applies
      only to the BRACKET-spelled token (`### [GSD.02] 0:`) — the clause
      named the wrong shape as its example.
  (b) "excluded from these counts" over-scoped the carve-out. The phase-0
      directory is admitted into `completed_phases`/`total_plans` in BOTH
      spellings (measured: B0 at HEAD has `completed_phases=2`,
      `accepts {"GSD.02-0-bootstrap":true}`). The exclusion that survives
      lives in `total_phases`/`phase_count` (heading counting) only.
  (c) The pre-existing "so `roadmap analyze` and `state json` report the
      same number" sentence (present since before this arc's bracket work)
      is now false for the bare-0 LEGACY-spelled shape: 04b5a95f's own
      commit message discloses this exact re-divergence as expected —
      `roadmap analyze` phase_count=2 vs `state json` total_phases=3 on T0
      — matching the disagreement legacy already carries today. The
      changeset still asserted unconditional agreement.

Rewrote both sentences: dropped the `### Phase 0:` example, scoped the
carve-out to `total_phases`/`phase_count`, named the bracket-spelled token
as the one that keeps analyze's sentinel reading, stated plainly that the
legacy-spelled form in a bracket repo IS counted (the mid-migration guard),
and qualified the "report the same number" claim to the `999` token, with
the bare-0 legacy-spelled shape named as the one exception and why.

FIX 3 (Nit 1) — `checkBracketCoherence` has always had two sub-checks
(its own docstring: "Two sub-checks, both surfaced as W021") but both the
changeset and docs/CONFIGURATION.md:188 described only the `mismatch`
sub-check. The `missing-bracket` sub-check — which fires on every
legacy-spelled heading in a bracket repo, the noisier of the two on a
mid-migration project — was undisclosed in both places. One clause added
to each.

Line count: 2 files, 1 line changed each (both are single-paragraph/
single-row files; git diff --stat reports 1/1 per file though several
distinct clauses were edited within that one line each).

Verified: `node scripts/changeset/lint.cjs` clean. Targeted suite
(adr-612-*, roadmap-parser, state, verify) unchanged at 971/971
(prose-only diffs, no test count change, no code touched).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): resolvePhaseIdConvention's docstring no longer claims loadConfig drops the key

Upstream #2997 (aa7697fe, landed in `next` during this PR's final
verification) added `phase_id_convention: get('phase_id_convention') ?? null`
to `_baseConfig` in config-loader.cts — `loadConfig` now surfaces
`phase_id_convention` in its resolved config. This function's docstring
gave "loadConfig merges against CONFIG_DEFAULTS and drops keys it does not
know, and `phase_id_convention` is not among them" as the reason for
reading config.json directly instead of calling `loadConfig(cwd)`. That
rationale is now stale against live `next`.

Corrected the comment to state the two reasons that actually survive #2997:

  1. The workstream->root federation this function performs is a standalone
     resolution run against a GIVEN cwd, not necessarily the same base a
     `loadConfig(cwd)` call elsewhere in the codebase would federate from.
  2. Convention-ENUM validation is still #612 PR-4 work — this function
     returns the raw string unvalidated, exactly as the now-surfaced
     resolved key would.

Noted that #2997 surfacing the key makes consuming it from resolved config
(instead of re-reading config.json here) a natural PR-4 consolidation —
not this PR's scope. The "cycles were never the obstacle" close and every
other paragraph in the docstring (federation rationale, SCOPE note) are
untouched; they still hold.

Comment-only — no code behavior changed. This worktree's own history does
not contain aa7697fe (git merge-base --is-ancestor confirms neither branch
is an ancestor of the other; not rebasing per instruction), so this is a
textual correction against a documented external fact, not a functional
sync with upstream.

git diff --stat:
 src/planning-workspace.cts | 17 ++++++++++++-----
 1 file changed, 12 insertions(+), 5 deletions(-)

Verified: `npm run build:lib` clean, `node scripts/lint-phase-id-drift.cjs`
clean, `npm run lint:ci` exit 0 (same 2 pre-existing unrelated eslint
warnings as every prior round this arc, 0 errors; lint-fix-has-regression-test
PASS). Full suite not re-run per instruction (comment-only diff).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2761): resolve phase_id_convention against the caller's workstream

`resolvePhaseIdConvention` took no workstream, so it resolved from
`planningDir(cwd)` — which falls back to `GSD_WORKSTREAM` only when its `ws`
argument is `undefined`. Every caller that passes a workstream by ARGUMENT
(they cannot set the env var per iteration) therefore read the convention from
the ROOT config while reading that workstream's ROADMAP.

Two reproduced consequences:

- A workstream that explicitly declares its own `phase_id_convention` had it
  ignored. Flipping ONLY the root config between bracket and milestone-prefixed
  changed which milestone that workstream extracted.
- `--workstream foo` and `GSD_WORKSTREAM=foo` disagreed on the same repo: the
  arg form fell through to the root config, the env form did not.

`resolvePhaseIdConvention(cwd, ws?)` now forwards `ws` to `planningDir`, and
both roadmap-parser call sites pass theirs — `extractCurrentMilestoneScoped`
(which already reads STATE from `planningDir(cwd, ws)`) and
`getMilestonePhaseFilter`'s lazy branch. `undefined` keeps the env fallback, so
convention-less call sites are byte-identical; `null` still means "explicitly no
workstream". The `undefined` vs `null` discriminator on `phaseIdConvention` is
untouched — only the base the `undefined` branch resolves FROM moves.

workstream-inventory's two sites (`countRoadmapPhases`, `inspectWorkstream`)
passed a literal `null` where they meant `undefined`, pinning every workstream
to the legacy grammar. On a bracket workstream whose milestone declares 3
phases that returned phaseCount 0 and fell back to the on-disk directory count;
it now returns 3. A non-bracket workstream is unchanged (legacy control pinned).

The workstream -> root federation is preserved: a workstream that declares no
convention still inherits the root, as config-loader does. That inheritance is
pinned so the fix cannot be over-applied into isolation.

* fix(#2761): classify missing phase details per occurrence, not per token

Under READING-B a phase's sentinel status lives in the BRACKET, so
`[GSD.999] 01` and `[GSD.02] 01` share a token and are not the same phase.
Both sides of the missing-detail check were keyed by the bare token anyway, and
both produced false negatives in `missing_phase_details`:

- The checklist scan built a token -> bracket-id Map, FIRST-WINS. Of two entries
  sharing a token, whichever the author wrote first classified both. With
  `- [ ] **[GSD.999] 01: Icebox**` above `- [ ] **[GSD.02] 01: ...**` the real
  phase inherited the icebox's sentinel verdict and vanished from the report;
  swapping the two bullets — same document, same phases — reported it. Pinned
  with a test asserting BOTH orders.

- The detail set was `new Set(phases.map(p => p.number))`, also token-keyed, so
  `[GSD.02] 01`'s heading marked token `01` present and satisfied
  `[GSD.03] 01`, which has no heading anywhere. Order-independent, same class.

Both sides now key on an occurrence key: the owner's `bracketQualifiedKey`
(fold- and padding-insensitive, so `[gsd.2] 01` and `[GSD.02] 01` are one
phase), falling back to a fold-normalized composite for the two shapes that
grammar refuses — a token carrying its own hyphen, which splices to an id whose
trailing segment the qualified-key grammar truncates (the hazard
`getMilestonePhaseFilter` guards with its own `!token.includes('-')`), and any
id it does not accept. With no bracket id the key IS the bare token, so the
legacy path keeps its exact keys and dedupe order.

The emitted value is unchanged — `missing_phase_details` stays an array of bare
tokens, matching `phases[].number`. Only the classification moved to the
qualified key, so two brackets' `01` both missing report `01` once instead of
one silently covering for the other.

* test(#2761): replace the two wall-clock ReDoS assertions with algorithmic bounds

Both guards asserted elapsed wall-clock time — `Date.now()` against a 20s
ceiling in the coherence suite, `process.hrtime.bigint()` against 1s in the
read-tolerance suite. Those measure the host machine rather than the SUT and
flake on a loaded CI runner (RULESET.TESTS.no-timing-assertion). Per
RULESET.TESTS.delete-bad-tests they are REPLACED, not skipped, and the
behavioral property each one guarded is preserved.

The property is "the widened bracket patterns do not backtrack
catastrophically", which is a claim about growth, so it is now stated by
scaling the input instead of by timing it:

- validate health runs the pathological unclosed bracket at 4,000 and 16,000
  characters and must return the same correct result (no W021) at both.
- The four reader entry points run five ReDoS shapes at 5,000 and 20,000 and
  must name exactly the phases each shape should name at each size, with the
  variant cardinality unchanged across the two.

Catastrophic backtracking is superlinear, so a regression cannot complete the
4x leg under any ceiling, while a bounded matcher is indifferent to the
scaling. The one attack that is well-formed-but-oversized now pins its reading
precisely (`['1'.repeat(n)]`) rather than being lumped in with the malformed
ones. A positive control asserts the readers still extract a well-formed
bracket heading, so "names no phase" cannot pass by the readers being inert.

`{ timeout }` is a hang backstop, not an assertion: it turns a runaway into a
deterministic failure instead of a suite that never returns.

* fix(#2761): give the bracket grammar one owner and teach the drift guard to see it

The bracket milestone-intro grammar was re-typed verbatim in three readers —
roadmap-parser's bracket-fallback selector, state's `isMilestoneBounded` and
verify's `checkBracketCoherence` — which is exactly what #2761's own gate
forbids ("no token literal outside src/phase-id.cts"). `check:phase-id-drift`
reported clean the whole time: its detector only ever knew the phase-NUMBER
token grammar, so the bracket class `[A-Z][A-Z0-9_]*` was invisible to it.

Ownership. `src/phase-id.cts` now exports the class as
`BRACKET_PROJECT_CODE_SRC` and the intro in the two shapes its readers need:
`bracketMilestoneIntroSrcFor(milestone)` (pinned to one milestone) and
`BRACKET_MILESTONE_INTRO_CAPTURING_SRC` (milestone captured). The pinned
builder owns the pad2 spelling rule too — "canonical spelling only, not `0*N`"
was previously restated in prose beside each copy, a convention two files had
to keep agreeing on by hand. `BRACKET_ID_SRC`, `BRACKET_ID_PREFIX_RE` and
`BRACKET_QUALIFIED_KEY_RE` now compose the class rather than re-spelling it.
All three call sites consume the owner; the regex sources are byte-identical to
what they spelled, asserted against hand transcriptions of the pre-fix lines.

Guard. `scripts/lint-phase-id-drift.cjs` gains a bracket rule
(`findBracketGrammarDrift`), wired into `scanRepo` and tagged `kind`. It
deliberately does NOT copy the token rule's `line.includes(CANON_REF)` escape:
that escape is line-level, and verify's copy referenced the owner for the
MILESTONE field on the same line as the re-typed PROJECT-CODE class — so a
bracket rule with that escape would have kept passing on the very site under
review. Partial ownership is the drift; only a dedicated `// phase-id-owner:`
comment suppresses it.

Proof, end to end: planting the shipped verify.cts literal back into src/ makes
`npm run check:phase-id-drift` exit 1 naming `[bracket] src/verify.cts:1475`;
restoring it returns the gate to ok.

Tests. phase-id-drift-guard carries all three shipped literals as negative
fixtures, the same-line-owner-reference case, the case-widened evasion variant,
the sanction rules, and a temp-tree scan proving the rule is wired into
scanRepo rather than merely exported. continuation-grammar-parity drives the
pinned and capturing shapes over a 12-entry corpus and requires the same
verdict from both plus the same captured milestone — widening either alone
fails there.

* test(#2761): drop the out-of-scope source-grep exemption from the selector pin

The baseline-selector pin in tests/adr-612-bracket-heading-selection.test.cjs
read `src/*.cts` with readFileSync and regex, claiming the no-source-grep
escape with a source-text-is-the-product reason. CONTEXT.md's documented scope
(RULESET.TESTS.no-source-grep.exemption) reserves that escape for tests whose
subject is a runtime CONTRACT FILE — STATE.md, config.toml, hooks.json, agent
.md — and `src/*.cts` is none of those.

It was also broader than it looked: eslint-rules/no-source-grep.cjs matches the
marker in ANY comment in the file, so one block's claim disarmed the rule for
the whole ~700-line suite.

The pin itself is worth keeping — the BASELINE ARGUMENT at each call site is a
fact no behavioural test can recover (flipping verify's milestone-complete site
from LABEL_ONLY to ANY_BRACKET grants a tolerance it has never had, and every
behavioural test still passes), so pinning it does require reading the authored
source. That reading moved to `scripts/lint-phase-id-drift.cjs` — the seam's
own guard, where source scanning is sanctioned (`warn` scope) and already
happens for the grammar rules — as `countSelectorBaselines` /
`scanSelectorBaselines`, which return a structured census. The test asserts on
the returned data and touches no file text.

CONTEXT.md is unchanged: the exemption scope was not widened to fit the test.
No marker remains in the suite, so the rule is live across all of it again, and
eslint passes with the escape removed rather than relocated. A floor assertion
pins that the census actually found the five consumers, so the "no other file
consumes the selector unpinned" check cannot pass on an empty scan.

* docs(#2761): rewrite the changeset as a lead + deltas instead of one paragraph

The fragment was a single ~6,400-character paragraph that opened on internal
mechanics, buried the user-visible change, and named an internal test path
(tests/adr-612-bracket-phase-counting.test.cjs) that means nothing to a reader
of the CHANGELOG.

It now leads with what a user sees — bracket-style phase IDs are recognized on
the read path across roadmap, validate, verify and state — followed by compact
bullets for the behavioral deltas, the opt-in caveat, and the upstream
consequences. Both merge-added disclosures are kept: the #3185 legacy-sentinel
Phase-0 delta with the reason the bracket counter keeps the narrower rule, and
the enumerator's convention-argument default flip with the four newly-scoped
read surfaces and the note that archival and milestone-completion paths are
unchanged. Two deltas from this review round are stated as their own bullets:
per-occurrence classification in `missing_phase_details`, and workstream-scoped
convention resolution including the workstream-rollup count change.

The internal test path is gone and the body is down to ~3,850 characters. The
bullets render as sub-bullets under the CHANGELOG entry and the `(#NNNN)`
suffix still lands on a trailing paragraph rather than mid-list.
`npm run lint:changeset` passes.

* test(#2761): use helpers.cleanup in the drift-guard temp-tree test

`local/no-raw-rmsync-in-tests` rejects a bare `fs.rmSync` in tests — the shared
helper carries the Windows-EBUSY retry budget. The temp-tree scan added with
M3's guard coverage test used the raw call.

* docs(#2761): correct the occurrenceKey comment on qualified-key normalization

The comment claimed `bracketQualifiedKey` is "fold- and padding-insensitive, so
`[gsd.2] 01` and `[GSD.02] 01` are one phase". The fold half is right; the
padding half is not. `BRACKET_MILESTONE_NUMERIC_SRC` is `(?:[1-9]\d{2,}|\d{2})`,
so `bracketQualifiedKey('GSD.2-01', 'bracket')` returns null and that input
takes the composite fallback instead.

No behavior change — the two key spaces use different separators and cannot
collide — but the claim was checkable and wrong. Restated: the owner case-folds,
and padding-tolerance is not a property it has or needs, because each accepted
milestone has exactly one canonical spelling and `[GSD.2]` is malformed rather
than an alternate spelling of `[GSD.02]`.

* ci(#2761): give the coverage-gate job the heap floor the shard jobs already have

The gate's report steps parse ~2.5GB of merged raw V8 dumps in one
process at the runner's implicit ~4GB ceiling, so pass/fail comes down
to GC timing (run 31338081337 OOM'd; next passes the same volume at
28s). Deterministic locally: crashes at --max-old-space-size=4096,
completes in 11s at 8192. PR shard data is within 0.15% of green next
runs — the load is pre-existing, only the ceiling was missing. #2952
set 6144 on the shard-collection step; the gate job was missed.

* Revert "ci(#2761): give the coverage-gate job the heap floor the shard jobs already have"

This reverts commit 03c74a97c911e9ddc53675aded9c9bbb89f14cd1.

* fix(#2761): thread the resolved convention into cmdStateValidate's phase-directory lookup

#3208 replaced cmdStateValidate's `startsWith` prefix test with the canonical
`phaseKeyFromDir(...) === selectedPhaseKey` comparison. That is the right
surface, and it is why the lookup now needs the resolved `phase_id_convention`,
which the rewrite does not pass.

`phaseKeyFromDir` deliberately refuses to read a bracket directory without an
explicit signal (ADR-2121: a bracket dir is string-indistinguishable from the
legacy letter-prefixed-decimal family), so un-threaded it returns the whole dir
name as the key — `GSD.02-05-real-work` -> `GSD.02-5-REAL-WORK` — while the
STATE side is the bare `05` that `parsePhaseFromProse` yields. Both sides of one
comparison derived under different conventions is #2562's defect class, and this
file's other three `phaseKeyFromDir` call sites already thread against it.

Observable: a bracket repo whose phase directory plainly exists reported
`valid: false` and "no phase directory matches phase 05", and the drift scan
(plan-count mismatch, verification status) never ran. The pre-#3208 `startsWith`
missed the same directory but skipped silently, so this is a visible-failure
regression on bracket repos, not a new miss.

Non-bracket conventions are byte-identical by construction: `extractPhaseToken`
branches only on `=== 'bracket'`, so null / 'milestone-prefixed' / unresolvable
compile the same path as the un-threaded call. The flat-legacy twin assertion
pins that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2761): note the cmdStateValidate convention threading in the changeset fragment

The fragment described the bracket read path as of the pre-merge branch. 504c64ff
added a shipped behaviour change — `state validate` now resolves bracket phase
directories — that the fragment did not mention, so the rendered changelog would
have under-described what ships.

Body text only; `type` and `pr` are untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): invert the inherited-wart characterization — #3225 fixed it upstream

Required by the merge of next @ 86101ee6; test-only, no src delta.

This branch disclosed rather than fixed a pre-existing upstream wart: on a
LEGACY repo, `### Phase 999:` warned from `validate consistency` while
`validate health` suppressed it — the two verbs contradicting each other.
The case was pinned as a characterization test ("INHERITED WART, unchanged")
precisely so it would INVERT if upstream ever closed it, rather than rot
silently.

ae7dc529 (#3225) closed it, by adding the `isSentinelPhaseId` guard to this
very loop. So the assertion inverted on the merge — as designed. Flipped to
assert the FIXED behaviour rather than deleted: it is the negative-space
proof that this branch's `sentinelPhases` guard never had to grow a legacy
reading of its own, and it reds if a future conflict resolution keeps our
guard while dropping upstream's.

Added a scope control alongside it (a legacy NON-sentinel `### Phase 09:`
with no directory still warns), so deleting the loop outright cannot pass.

Union proven load-bearing in BOTH directions against the merged tree —
neither guard subsumes the other:
  - drop `isSentinelPhaseId(p)` (ours only) -> 1 red, exactly this case.
    Note upstream's own #3225 tests stay GREEN there: they cover the
    disk-side loops (sentinel dir on disk) and the gap-numbering filter,
    not the ROADMAP-side loop. This case is that site's only coverage,
    which is the second reason to keep it rather than delete it.
  - drop `sentinelPhases.has(p)` (upstream only) -> 4 red bracket suites
    (icebox-not-missing, health/consistency agreement, occurrence-aware
    suppression, checklist-index suppression).

tests/adr-612-bracket-read-tolerance.test.cjs 79 -> 80, all green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2761): pin the three re-homed W006/W007 reads the merge left unfalsifiable

#3309/#3310 deleted the helpers this PR threaded (collectDiskPhases,
collectDiskPhaseEntries, collectArchivedPhaseDirNames,
forEachArchivedPhaseToken) and rebuilt their reads inside
src/planning-snapshot.cts + src/health-diagnostic-rules/*.cts. The convention
threading moved with them in the merge commit — but a mutation sweep over the
re-homed sites found three where reverting the convention argument changed real
CLI output and NOT ONE existing test went red. The pre-migration sites were
covered indirectly, through helpers that no longer exist, so the coverage did
not survive the relocation even though the behaviour did.

Each case is pinned at the CLI with its flat-legacy twin as the byte-identity
control, plus a non-vacuity control in the opposite direction:

  archivedPhaseTokens      revert -> "Phase 05 in ROADMAP.md but no directory on
  (planning-snapshot.cts)            disk" for a phase whose only directory is
                                     under .planning/milestones/v1.0-phases/
  roadmapPhaseCheckboxes   revert -> "Phase 09 in ROADMAP.md but no directory on
  (planning-snapshot.cts)            disk" for an unstarted `- [ ]` phase
  W007's extractPhaseToken revert -> "Phase GSD.02-77-orphan exists on disk"
  (roadmap-disk-consistency)         instead of "Phase 77 exists on disk"

Red/green: each single-line revert reds exactly its own case and nothing else;
every legacy control stays green in all three runs.

NOT pinned, and disclosed as droppable rather than given an unfalsifiable test:
buildValidPhaseSet's extractPhaseToken(dir, convention) in W002. That rule
unions disk tokens with roadmapDeclaredPhases and archivedPhaseTokens, and any
STATE.md reference a bracket disk token would rescue is already rescued by the
ROADMAP half — probed directly, the argument makes no observable difference. It
is threaded because it restores collectDiskPhases(planBase, convention)'s
derivation exactly, not because a test needs it.

Test-only; revertable independently of the merge commit.

* test(#2761): pin the two reads the #3165 extraction left unfalsifiable

DROPPABLE, offered as such. Test-only; the merge commit is correct without it.

#3428's extraction gave `collectAnalyzePhases` TWO call sites — the scoped
milestone window and the truncated-window recovery path. The merge threads
`convention` into both and swaps `detailKeys` with `phases` across the recovery.
Neither of those was falsifiable by the suite as it stood:

  - passing a NULL convention at the FALLBACK site, with the scoped site still
    threaded, is green across all 518 tests of the bracket, roadmap and
    milestone-window files. Every pre-existing bracket assertion reaches the
    enrichment through the scoped window, so none of them can observe the
    fallback's reading at all;
  - dropping `detailKeys = fallbackScan.detailKeys;` — the line this merge
    authored — is likewise green, because no fixture that reaches the recovery
    path carries a checklist bullet, so `missing_phase_details` is `null`
    either way.

That is the same hole class lap 3 closed in 9914359c: an argument whose revert
changes real CLI output with zero reds.

Pinned at the CLI on the shape that reaches the fallback on a bracket repo — a
MID-MIGRATION ROADMAP: bracketed ACTIVE milestone, its phase-detail sections
and checklist bullets sitting after an intervening CLOSED legacy milestone (so
the window closes over prose only), plus one legacy `### Phase N:` section of
its own. Six tests:

  - a NON-VACUITY control that removes the phase directories, so #3428's own
    precondition fails and the result is the empty one the recovery exists to
    replace — this is what proves the rest read the FALLBACK's output rather
    than the scoped scan's;
  - the directory assertion (`disk_status`/`plan_count`/`summary_count` for
    canonical `{CODE}.{MM}-{PP}-slug` dirs) + a flat-legacy twin asserting
    byte-equal shape, and that the twin is the right answer rather than a
    shared wrong one;
  - `missing_phase_details: null` for phases the recovery just found, and its
    companion direction — a bullet with no heading anywhere is STILL reported,
    so a fix that merely suppressed the report does not pass;
  - a non-bracket repo taking the same path unaffected.

Red/green, each mutation reverted afterwards:
  - fallback call site -> `null` convention: exactly 2 reds, both directory
    assertions in this block; 307 other tests green, including upstream's own
    #3428 tests in tests/milestone-window-single-owner.test.cjs.
  - `detailKeys` swap deleted: exactly 2 reds, both `missing_phase_details`
    assertions; 414 other tests green.
  - `matchPhaseDirs` 3rd argument dropped (the lap-2 site): 4 reds — the 2
    already on record plus this block's 2, which is the point: one seam, now
    reached by two paths.

DISCLOSED IN THE TEST BODY WITH MEASURED OUTPUT, not fixed: `hasPhaseEntries`
(src/roadmap-parser.cts) is convention-blind, so on a PURE-bracket ROADMAP the
window classifies COMPLETE and #3428's recovery is gated off entirely. That
document returns `{"scope":"complete","phase_count":0,"next_phase":null,
"phases":[]}` — the "genuinely empty milestone" answer, the indistinguishability
#3184 introduced `scope` to remove — where the flat-legacy twin returns
`{"scope":"truncated","phase_count":2,...}`. Widening it changes the value of an
upstream-owned output field on bracket repos, so it is a behaviour slice (the
PR-2.5/PR-4 convention-less-readers question), not a merge-round change. The
mid-migration shape pinned here is the reachable half.

* fix(#2761): mirror the bracket terminator in the milestone-scope write guard

parser's terminator vocabulary — "a level 1-3 heading that is not a Phase
heading and carries a milestone signal". On this branch that vocabulary is
convention-SELECTED: `computeBracketSectionEnd` adds `isBracketMilestoneBoundary`
as a terminator arm, and the ADR-canonical `## [GSD.09] Hidden` carries NO
vN.N token, NO ✅/📋/🚧/🔄 marker, and not the word "Milestone". Left
unmirrored, the guard accepted exactly the description it exists to reject.

NOT a defect on clean next — a bracket heading terminates nothing there. The
branch widens the terminator set, so the branch owns the mirror.

MEASURED at the CLI seam before the fix (bracket fixture, one milestone,
one phase):

  $ gsd-tools phase add $'Sneaky\n## [GSD.09] Hidden'   -> exit 0, written
  $ gsd-tools phase add 'Innocent follow up'            -> exit 0, written
  $ gsd-tools roadmap milestone-scope
    { "scope": "complete", "phases": ["01","1"], "phase_count": 2 }
  $ grep '^### Phase' .planning/ROADMAP.md
    ### Phase 1: Sneaky
    ### Phase 2: Innocent follow up      <- in the document, out of the window

After the fix the first add exits 1, ROADMAP.md is byte-unchanged and no
phase directory is created. The legacy twin (same text, no
`phase_id_convention`) still exits 0 — opt-in only, base behaviour preserved.

SHAPE
- `findMilestoneScopeHeadingLines(text, convention)` — REQUIRED, the same
  tripwire `scanMilestonePhaseIds` carries in the merge commit and for the
  sharper reason: a blind call here fails OPEN (the guard quietly ACCEPTS a
  window-narrowing description), so a future call site must fail to COMPILE.
  Census is one caller, which pays nothing for it. A non-bracket value takes
  the pre-existing path byte-identically. The bracket arm routes through
  `isBracketMilestoneBoundary`, the same single-owner phase-vs-milestone
  discriminator `computeBracketSectionEnd` consults; no second bracket-heading
  grammar is spelled here.
- `selectedBracketId` is deliberately `null`, so the same-milestone
  CONTINUATION exemption never fires and a value naming the ACTIVE milestone
  is flagged too. That is this predicate's third stated conservatism and
  rests on its own existing argument: which milestone is active is a property
  of the document at write time, not of the text being validated, and
  over-rejecting is one-directional.
- `assertDescriptionPreservesMilestoneScope` takes `cwd` and resolves through
  the branch's tolerant try/catch — this guard runs BEFORE `loadConfig` and
  before the ROADMAP existence check, so an unresolvable convention must
  degrade to the legacy vocabulary, never turn a rejection into a crash.
- The error's marker list gains the bracket form only when the project is on
  the bracket convention.

RED/GREEN
- Revert the bracket arm -> 2 reds (phase add; insert + add-batch), the
  non-bracket control and both no-false-positive cases stay green.
- Revert the probe threading in the merge commit -> 1 red (the probe case).
- 6 new cases in the #3262 file, each with a non-bracket or
  no-false-positive control: fenced bracket milestone heading and bracket
  PHASE heading are both non-violations.

Droppable: revert this commit and the merge stands on its own; the branch
then ships the gap as a disclosure instead of a fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changeset): narrow the enumerator claim to the directory set — per-entry rendering in progress/stats/init-manager is display-slice scope

Round-7 review confirmed 3 of the 4 surfaces named by the closing claim
parse each directory or heading with legacy-only patterns that live in
files this PR does not touch (commands.cts:1766, commands.cts:2116,
init.cts:2241). The claim now states exactly what this slice delivers:
the scoped directory set. Their per-entry conversion is display work,
deferred to the epic's display PR with the statusline/progress-card
gates it belongs beside.

* fix(#2761): de-accident the unmatched-milestone bracket fixture, re-pin to #3480's withhold contract

tests/adr-612-bracket-phase-counting.test.cjs:515 ("a milestone that
does NOT match STATE is not scoped in") asserted total_phases === 0.
Since today's merge brought in 70b5c1a1 (#3354/#3480, already in
`next`), buildStateFrontmatter withholds total_phases (omits the key,
read back as null) for a milestone that is genuinely sectioned but
matches no ROADMAP heading, instead of substituting a computed number
— the fixture's `state json` read now returns null, failing the
strictEqual(0) assertion.

The original single-section fixture's 0 only ever survived by
accident: hasMilestoneSectioning requires >=2 milestone-vocabulary
headings to call a ROADMAP sectioned, and its isPhaseHeading helper
recognizes only the legacy `Phase N:` text form — so the bracket
phase heading `### [GSD.03] 09: Not this milestone` (title containing
the word "milestone") miscounted as a second milestone heading,
tipping hasMilestoneSectioning true and routing to the disk-count
branch. With that miscount removed, the same one-section fixture
reads 1 (the non-matching milestone's phase count leaking into v2.0's
total) — proving the pinned 0 was never validating the scoping this
test claims to exercise.

Replaced the fixture with two genuine milestone sections (GSD.03,
GSD.04), neither matching STATE's v2.0, with phase titles carrying no
incidental vocabulary — the real #3354 shape. Re-pinned the assertion
to null, matching upstream's own tested contract (tests/state-
document.test.cjs, "#3354 with nothing stored, the key is omitted
rather than written from the dir count").

Verified: file green 3/3 runs (104/104), full state-suite regression
771/771 green. The isPhaseHeading gap that made the old fixture
accidental is a real, pre-existing, upstream-owned limitation
(hasMilestoneSectioning only guards >=2-section conflation, not a
single non-matching section leaking through) — not introduced by

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2761): exempt two self-authored-fixture regexes from #3441's unbounded-quantifier rule

Both sites parse the STATE.md the test itself just wrote to its own
tmpdir — fixed-size fixture output, not adversarial or document-scale
input. Exempted with the justification-comment pattern the tree's own
tests use for this exact case (settings-integrations, copilot-install,
research-agent-profiles). The sites predate the rule; #3441 landed on
next this morning and this branch picked it up in the catch-up merge.

* fix(#2761): fix forward two next-movement test regressions from the round-11 rebase

origin/next's #3573 newly threads STATE's stored `milestone:` value into
`state json`'s buildStateFrontmatter call (previously always `undefined` on
that read surface). Two pre-existing PR-2 fixtures reach code paths that
call never exercised before that merge:

- RED (repro2 case C): a version-less, all-bracket-id ROADMAP now hits
  getMilestonePhaseFilter's pre-existing (unaffected by this branch) row-5
  `versionResolved && !headingFound => SCOPE.UNSCOPED` rule, which withholds
  `progress.percent` even though the scoped total_phases/completed_phases
  are still correct. That row-5 rule is load-bearing for six other
  version-less-document pins in this same file; narrowing it broke seven of
  them in testing, so production code is untouched here. Reassert the
  test's real claim (scoping, via total_phases/completed_phases) and
  disclose the now-withheld percent instead of silently dropping it.

- PIN (repro12 LEGACY control): a fenced-example-vs-real version heading
  fixture. #3573 routes this read through sliceMilestoneWindow (fence-aware)
  instead of the legacy anyMilestonePattern raw scan (fence-blind) this pin
  was disclosing as out of scope, so the fixture no longer exercises the
  blind path — total_phases moves from the whole-doc 4 to the correctly
  scoped 2. Re-pinned with an upstream-attributed comment, matching this
  file's existing house style for prior origin/next movements (see the
  sibling "PIN (G3 LEGACY control)" comment).

Both are test-only fix-forwards: no production code changed. The remaining
12 failures in this file at 27363e9d0 (dropped seam commits' VERIFICATION.md
naming + fixture updates) are pre-existing and deferred to the stacked
follow-up per the task's own scope.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): thread phaseIdConvention explicitly at the two write-adjacent enumerator call sites (round-11 BLOCKER)

listMilestonePhaseDirs' phaseIdConvention param lost its `= null` default
(phase-locator.cts), so an omitted convention now means "resolve from
config" instead of "explicitly not bracket" — a deliberate flip, but
milestone.cts and state.cts were not in the PR-2 diff and both call the
enumerator without threading it:

- milestone.cts cmdMilestoneComplete (the single #3597 shared derivation
  feeding the stats loop, --dry-run preview, and the real archive/rename
  pass) now resolves phase_id_convention once and threads it explicitly,
  so a bracket project's `milestone complete` archives its real
  bracket-declared phase directories instead of silently inheriting
  whatever the enumerator's lazy default resolves to.
- state.cts cmdStateUpdateProgress's own enumerator call threads the same
  resolved convention. Empirically this does not change the #3217 withhold
  gate (scope is assigned before headingConvention resolves in
  getMilestonePhaseFilter, so it's convention-independent either way) or
  the reported percent (already correctly threaded via
  computeUpdateProgressPreview -> buildStateFrontmatter); it closes a
  second, silently-resolved answer to the same question this file's own
  ONCE-and-THREAD rule (~:2300) already states as policy.
- state.cts's other listMilestonePhaseDirs call site (the state-sync
  scope-only read, ~:4791) is left unthreaded on purpose, with an inline
  note explaining why: only `.scope` is consumed, and `.scope` is set
  before convention resolution in getMilestonePhaseFilter, so it cannot
  disagree with a threaded convention.

Both enumerated sets are pinned by new tests (round-11 BLOCKER block in
tests/adr-612-bracket-phase-counting.test.cjs), including a mutation-style
regression check on the milestone.cts fix (forcing the un-threaded default
back on makes the pinned test fail, catching a regression to the
pass-all-degrade legacy reading).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): correct the changeset's false archival claim and amend ADR-612 for the round-11 M2(2) split scope

The changeset (.changeset/2761-bracket-read-tolerance.md) asserted "The
archival and milestone-completion paths are unchanged" — false: both
paths reach the widened enumerator, and this PR now threads their
convention explicitly (previous commit). Replaced the closing paragraph
with an accurate description of what changes for a bracket project at
those two call sites, and notes both enumerated sets are pinned by tests.

ADR-612 (docs/adr/612-bracket-phase-id-convention.md) amended per M2(2):
the 2026-08-03 PR-2/PR-4 boundary proposal is added in-body (PROPOSED,
not stamped — mechanics per docs/contributor-standards.md:143 reserve ADR
ratification to maintainers), rescoped to what actually ships in PR-2
(#2761) now that the round-11 M1 split moved the completion-seam
threading (isPhaseArtifact/scopeToPhase, phase-id.cts:964-1090) out into
a separate, stacked follow-up PR referenced generically via epic #612:

- state.cts read-tolerance (both total_phases derivations, the #1514
  retirement filter) moves into PR-2's row, alongside milestone.cts,
  reflecting the round-11 fix above.
- the write-observability point is restated in terms of what actually
  ships: explicit threading at two named call sites, pinned by tests,
  rather than silent inheritance.
- the completion-seam threading is named and explicitly excluded from
  PR-2's scope, with a new ratify item (5) and a Negative consequence
  bullet covering the sequencing cost.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): repair round-11 response gaps — disk-side sentinel bug, test-fixture bug, scanMilestonePhaseIds caller drift

Four independently-verified gaps in the round-11 repair response:

1. isSentinelPhaseId (src/phase-id.cts) treated a bare, untagged phase
   directory under phase_id_convention: "bracket" (`0-bootstrap`, no
   `{CODE}.{MM}-` prefix) as sentinel milestone 0 by falling through to the
   legacy leading-int rule when the bracket-tag match failed. This silently
   dropped a real, on-disk, milestone-declared phase directory from
   `listMilestonePhaseDirs` (phase-locator.cts:424, the only unguarded call
   site) and undercounted completed_phases/percent. Mirrors the carve-out
   already present on both heading-side counters (state.cts's
   countRoadmapPhaseHeadings guards its bare-0 exclusion with `bracketId &&`;
   roadmap-parser.cts's scanMilestonePhaseIds composes the bare-token rule as
   999-only) — under bracket convention, milestone 0 is expressed only via an
   explicit bracket tag, so an untagged leading 0 is a real phase token.

2. tests/adr-612-bracket-phase-counting.test.cjs's writeProject fixture
   hardcoded `01-VERIFICATION.md` for every "complete" phase dir regardless
   of the dir's real phase token. Under legacy convention this file already
   correctly failed #3511's strict isPhaseArtifact match for any dir other
   than phase 01 (the fixture never modeled what it claimed to); under
   bracket convention the pre-existing ambiguity fail-safe admitted it
   regardless. The two readings' disagreement was mistaken for a missing
   completion-seam threading (isPhaseArtifact/scopeToPhase convention
   awareness, correctly split out to the stacked #3644 per round-11's M1).
   Replaced the hardcoded name with verificationNameFor(dir), deriving the
   real per-phase filename from the production normalizePhaseName the same
   way cmdScaffold does — the 7 failing "flat-legacy-twin" / CHARACTERIZATION
   assertions pass on #2867's own code with no seam threading required, and
   the genuine seam-dependent cross-phase-stray-exclusion test the fixture
   fix would otherwise have hidden lives on the stacked branch instead.

3. tests/roadmap-parser.test.cjs's two #3577 tests still called
   scanMilestonePhaseIds with the old single-Set return shape; this PR
   changed it to `{ ids, qualifiedIds }` for every other caller but missed
   these two, which don't touch #2761/bracket code at all. TypeErrors at
   runtime, not silently-wrong assertions. Updated both call sites.

4. .changeset/2761-bracket-read-tolerance.md gains a paragraph disclosing
   fix 1 above, so the changeset stays accurate to what actually ships (the
   round-11 BLOCKER was exactly this changeset going stale once).

Verified: tests/adr-612-bracket-phase-counting.test.cjs +
tests/continuation-grammar-parity.test.cjs + tests/roadmap-parser.test.cjs =
354/354. adr-612-{coherence,grammar,heading-selection,read-tolerance,
selection.property}.test.cjs + collision-characterization = 287/287
(CONFIRMED-CLEAN set, unaffected). Full unfiltered suite run separately.

PR #3643 (round-11's M1 split, opened draft per that round's explicit
requirement) was auto-closed by this repo's own draft-PR policy 11 seconds
after opening — draft PRs are unconditionally closed here. Reopened as
non-draft #3644 (same branch, same commits, DO NOT MERGE / stacked-on-#2867
marker in the body, enhancement template) since the repo's own bot confirms
non-draft contributor PRs are tolerated even off-template.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2761): narrow the state-update-progress pin claim to what mutation testing actually proved, add the test that covers the rest

Round-11 BLOCKER response gap (verified, not a guess): the changeset and
ADR-612 amendment point 2 both claimed the round-11 tests pin
`cmdStateUpdateProgress` so "a future change to the enumerator's default
cannot silently move ... what state update-progress reports without failing
a test." Mutation testing disproves this. Reverting BOTH the state.cts
explicit `phaseIdConvention` thread AND simulating the phase-locator.cts
pre-#612 hardcoded-null default (`phaseIdConvention: null` at that one call
site) leaves the existing PIN test ("writes a real percent for a
bracket-scoped milestone") green, because that percent comes from
`computeUpdateProgressPreview` -> `buildStateFrontmatter`, a separately and
already-correctly-threaded derivation the state.cts inline comment at ~:965
already candidly documents.

What the reverted thread DOES change, empirically, is `phaseDirs`/`totalPlans`
— the enumerated `.value` this call site feeds into the #3233 zero-plans
no-op gate a few lines below. That is the one place a regression at this call
site is observable in the command's output. Added a test that pins exactly
that: a bracket milestone whose declared phases carry no plans on disk,
alongside a decoy directory that plainly does not belong to the milestone
(no bracket tag, no phase token) but does have a plan. Correctly scoped, the
decoy is excluded and the #3233 no-op fires (`updated: false`). Degraded to
the pass-all legacy reading, the decoy is swept in, `totalPlans` flips
nonzero, and the no-op never fires (`updated: true`).

Mutation-tested against both scenarios:
- Reverting ONLY the state.cts explicit thread (falls back to `undefined`,
  which `getMilestonePhaseFilter` resolves via the identical
  `resolvePhaseIdConvention(cwd, undefined)` call the explicit thread also
  makes): new test stays green — confirms the single-hunk thread really is
  pure single-derivation hygiene, exactly as the existing inline comment
  claims, for this test too.
- Forcing `phaseIdConvention: null` at that call site (the combined-revert
  scenario the round-11 mutation testing actually exercised): new test FAILS
  (`updated: true, percent: 0` instead of the expected `updated: false`).
  Restoring the real code makes it pass again.

Narrowed the changeset and ADR-612 point 2 prose to match: `cmdMilestoneComplete`'s
enumerated set is pinned against a pass-all-degrade regression (genuine,
already verified in round 11); `cmdStateUpdateProgress`'s own call site is now
pinned against the #3233 zero-plans no-op specifically, not against the
reported percent, which the prose no longer claims. Added a one-line pointer
to the new test in the state.cts inline comment. No production behavior
changes — test and documentation only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(#2761): reconcile ADR PR-2 module map

* fix(#2761): restore #3639's dir-aware W007 sentinel exclusion lost in the rebase replay

The rebase replayed this file's pre-#3639 patch over next, reverting the
isSentinelPhaseId(token) -> isSentinelPhaseDir(dirName) fix: the extracted
token is milestone-stripped, so a bracket sentinel (GSD.999-07-icebox,
GSD.00-01-backlog) was invisible to the id predicate and false-fired W007.
Restores upstream's call and comment verbatim; the branch's convention-aware
token remains for the diagnostic message only.

Caught by next's own #3639 regression tests (CI shard 2). Local: the file's
18/18, health-diagnostic-rules 165/165, full npm test exit 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MeAsZNfhiUQioFPA4ygEGS

* docs(#2761): ratify scope and correct surface claims

* test(#2761): strengthen legacy selection and drift claims

* docs(#2761): correct the enumerator claim, scope list, and ratification receipt

Three documentation corrections, none touching production code.

The changeset claimed the shared phase-directory enumerator "now defaults its
convention argument to 'not yet resolved' rather than 'resolved, and not
bracket'". That is false: `phase-locator.cts:375` still destructures
`phaseIdConvention = null`, and the lazy resolve-from-config fires only on
`undefined` (`roadmap-parser.cts:941`, `:1928` — whose own comment records that
"explicit null still means 'resolved and non-bracket'"). Measured: only 4 of 17
`listMilestonePhaseDirs` call sites thread a resolved convention (`milestone
complete` and `state`'s three). The changeset went on to name `progress`,
`stats`, `phase list` and the init manager view as now receiving a correctly
scoped set — those are precisely callers that omit it. Replaced with what the
code does, and the deferral stated plainly.

The ADR's PR-2 row omitted `scripts/lint-phase-id-drift.cjs` (+141/-14) and
`scripts/lint-phase-enumeration-drift.cjs`, both changed by this PR. An
under-claim rather than an over-claim, but the row is the epic's
scope-of-record.

Ratify item 5 asserted maintainer ratification while citing only the review that
raised the question. It now cites the review that granted it (#2867 review
`5012940978`, 2026-08-24) and quotes its terms, so the claim carries its receipt.

Found by an adversarial review pass over the round-12 diff. (#2761)

* fix(#2761): drop the no-op --json that strict argv now rejects

Surfaced by the rebase onto next, not by a change in this PR's subject.

#3884 ("failure is a value — strict argv", e20744eac) made an unrecognized
flag a hard error: `state validate --json` now exits 1 with
`unknown flag "--json"; accepted: --strict` on stderr and EMPTY stdout,
where the token was previously accepted and ignored. `state validate` never
had a `--json` flag — JSON is its only output shape — so the argument was a
silent no-op from the start.

The helper parsed that empty stdout, so all five subtests in the
"state validate resolves bracket phase DIRECTORIES" suite failed identically
with `SyntaxError: Unexpected end of JSON input` at the JSON.parse, masking
what they actually assert.

Dropping the token restores the same envelope the helper already parsed. No
assertion changes. Verified against the fixture the suite builds: exit 0, and
the output carries the S005 plan-count warning the drift assertions require
with no S004 phase-directory warning — i.e. the bracket directory resolves,
which is the behaviour these tests exist to pin. 94/94 in the file.

* fix(#2761): exempt the phase-counting suite from the docs-guard lane

Surfaced by the rebase onto next. #3753 (107eb8c1d) added
lint-docs-guard-registration, which requires every test file that reads a
docs/ path to be either registered in the docs-guard lane or carry an
explicit marker. It flags adr-612-bracket-phase-counting.test.cjs, which
reads no docs/ path at all.

The file's only docs/ occurrence is the ADR-612 Decision 1 citation in a line
comment; every read call it makes targets a tmpdir .planning fixture. It trips
Detector 3, whose DOCS_TEMPLATE_LITERAL_RE sees an odd prose backtick in the
comment block above that citation as opening a template literal and reads the
span between them — citation included — as a docs/ path expression. That
detector documents this trade in its own header: it favours recall, and says a
false positive costs one docs-guard-exempt marker with a reason. The baseline
records the same class ("comment-only mentions") for 46 of its entries.

Registration was the wrong side of the trade here: it would run this suite in
the docs-guard lane on docs/ changes whose content it never reads.

Three pieces, matching what the gate requires and what its 54 existing entries
already do:
  - the marker in the file's header window, written without backticks so it
    cannot itself disturb the parity tracking findExemption does;
  - the basename in DOCS_GUARD_EXEMPT_BASELINE, since the ratchet fails a new
    marker until the baseline is deliberately updated — the reviewable diff is
    the point;
  - the FIX 3 per-file fingerprint, derived with the lint's own
    extractDocsPathReferences rather than retyped, so the exemption fails loudly
    if the set of docs/ paths this file mentions ever changes.

Gates: lint-docs-guard-registration 0 violations, ci-docs-guard-registry 51/51.

* docs(#2761): correct the bare-0 comment and state what opting into bracket costs

Round-13 review items Minor 2 and Minor 3, both still live on the previous head.

Minor 2 — src/phase-id.cts. The comment block claimed "Bare `0` is admitted
alongside it because a 0.x sentinel is a legitimate identity that predates
padding", which contradicted both the shipped constant and its own next
paragraph. BRACKET_CANONICAL_NUMERIC_SOURCE is `(?:[1-9]\d{2,}|\d{2})`;
measured against it, `0` is rejected while `00`, `05`, `99`, `100` and `999`
are admitted. The paragraph four lines below already recorded the removal
("the earlier `(?:\d{2,}|0)` ... admitted ... a bare `0` that pad2 never
produces"), so the block asserted and denied the same fact. The stale sentence
is replaced with what ships: `00` is the backlog sentinel's canonical padded
identity, `\d{2}` already covers it, and nothing needs the unpadded spelling.
No behaviour change — the constant is untouched.

Minor 3 — docs/CONFIGURATION.md. The phase_id_convention row described the
read-path widening but never said what a repo GIVES UP by opting in. Added:
on a bracket repo a heading whose bracket is followed directly by a digit is
read as a phase heading, so `### [RFC.2119] 5:`, `### [v1.0] 2024:` and
`### [ADR.612] 3:` — legal prose headings under any other convention — are
claimed as phases and move phase_count, total_phases and W006. Those three
shapes are the ones phase-id.cts's own selector comment names as the reason
the widened read is selected at construction time from this value rather than
applied everywhere; the row now carries that trade instead of only its
upside.

Gates: tsc --noEmit exit 0, npm run lint:ci exit 0, npm run test:unit
33890/33891 with the single failure being emitted-attribution's base drift
against a next that moved after the rebase (133/133 against the rebase base;
this push re-bases onto the current tip, which resolves it).

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-31 14:03:48 -04:00
Tom Boucher
6beaa66b25 enhance(#3304): gate re-verification blockers on deterministic evidence (#4085)
* test(#3304): add failing-first suite for the convergence evidence gate

Content-assertion suite for the Step 7 re-verification evidence gate
(agents/gsd-verifier.md / gsd-core/references/verifier-evidence-gate.md).
Committed before the implementation to prove RED via gsd-test.

* enhance(#3304): gate re-verification blockers on deterministic evidence

Step 7's anti-pattern scan re-runs at full, unbounded scope on every
re-verification pass, independent of the must-haves established in Step 2.
A blocker it finds — other than the self-evidencing debt-marker check —
previously reverted a completed gap-closure round and started another
--gaps cycle on nothing more than the verifier's own new judgment call,
with no bound on how many times that could repeat.

A Step 7 blocker now blocks unconditionally in re-verification mode only
if it is a carried-forward gap (present in the prior VERIFICATION.md's
gaps: list) or the flagged file was git-modified since the prior pass
(a regression; fails closed toward blocking when history is unresolvable).
Otherwise it predates the gap-closure round unflagged and needs
deterministic evidence — a named test run red, or another concrete
reproducible artifact — to stay blocking. Unevidenced, it downgrades to
a new advisory: frontmatter list and report section instead of setting
status: gaps_found, and never reverts a completed must-have.

Maintainer approval was narrowed to this evidence condition only,
explicitly rejecting the broader "advisory whenever untraceable to a
requirement/decision/prior-gap" proposal — implemented and pinned by
tests/verifier-evidence-gate.test.cjs and documented as rejected in
gsd-core/references/verifier-evidence-gate.md so it can't silently
re-expand.

Closes #3304

* fix(#3304): correct window-truncation and indentation bugs in evidence-gate tests

gsd-test's GREEN checkpoint caught 3 real bugs in the test file itself
(not the production prose): a {0,600} match window was shorter than the
724-char paragraph it was scanning (the "exclude from Step 9 Rule 1"
phrase starts at offset 662), and two regexes assumed no indentation
after a markdown list-continuation line break. All three phrases are
confirmed unique across agents/gsd-verifier.md, so the windowed
submatches are replaced with direct whole-string assertions instead of
just widening the window.

Also acknowledges the deliberate byte growth in agents/gsd-verifier.md
that the differential-attribution check (ADR-2719) correctly flagged.

Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the #3304 re-verification evidence gate (Step 7 rule, Advisory bucket, advisory: frontmatter, report section); 1488 bytes, still within the LARGE-tier 48 KiB cap (48751/49152).

* docs(#3304): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-30 16:51:52 -04:00
Tom Boucher
62b0d939b6 feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey (#4083)
* feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey

Add an optional `timeoutConfigKey` field to the reviewer lane descriptor,
resolved in `resolveLanePlan` at invocation time and falling back to the
frozen `timeoutFloorMs` when unset or invalid, in the same spirit as the
existing `promptBudgetKey`/`modelConfigKey` fields. All 12 shipped lanes
declare `review.timeouts.<slug>` on both surfaces (the descriptor and their
capability.json manifest), validated by capability-validator.cjs.

For the antigravity lane, the native `agy --print-timeout` flag — previously
a second hardcoded literal (`540s`) independent of the outer cap — is now
derived from the same resolved outer timeout in `antigravityArgv`, preserving
the existing 60-second buffer relationship (ADR-2782 D6: the outer bound is
declared data, the inner one is handler-owned).

The antigravity default timeoutFloorMs stays at 600s per the maintainer's
disposition; users raise it through the new config key instead.

* docs(#3274): document review.timeouts.* and extract resolveTimeoutMs helper

Address code-review findings on the timeoutConfigKey change: extract the
inline timeout-resolution logic into a named, exported, directly-tested
resolveTimeoutMs helper (matching the file's existing configString/
normalizeHost convention); document the new review.timeouts.* federated
config keys in docs/CONFIGURATION.md, docs/reference/capability-manifest.md,
and docs/how-to/ship-a-reviewer-lane.md; add the changeset fragment.

* fix(#3274): resolve native antigravity timeout in resolveLanePlan, not the runner

gsd-test caught two design mistakes in the prior commits:

1. SpawnPlan.argv is documented and tested as fully resolved by
   resolveLanePlan (model/effort/output/prompt already folded in) — leaving
   the antigravity '{{nativeTimeout}}' marker unresolved until the runner's
   antigravityArgv violated that contract and broke tests that read
   plan.argv directly (tests/antigravity-reviewer.test.cjs,
   tests/review-default-reviewers-workflow.test.cjs). Fix: '{{nativeTimeout}}'
   is now a fifth ARGV_PLACEHOLDER member, resolved by resolveLanePlan itself
   via the new nativeTimeoutToken() helper, exactly like the other four.
   antigravityArgv reverts to its pre-#3274 four-argument form. Also missed
   updating capabilities/antigravity/capability.json's invoke.args to match
   the descriptor, which broke the manifest/descriptor parity test.

2. tests/reviewer-config-federation.test.cjs enforces a deliberate, narrow
   invariant (#3691 narrows #2797): qwen, cursor, and coderabbit — the three
   lanes with neither a model flag nor a host — may own no config key beyond
   their own prompt-budget key. Adding review.timeouts.<slug> to all 12 lanes
   violated it. Fix: those three keep timeoutConfigKey: null and own no
   review.timeouts.* key, matching their existing modelConfigKey: null. The
   other 9 lanes are unaffected.

* chore(#3274): backfill changeset PR number (pr:0 -> 4083)

---------

Co-authored-by: sim <sim@local>
2026-08-30 14:26:02 -04:00
Tom Boucher
6eea00b707 enhance(#3301): tell reviewers the plan ids and total count, grade coverage (#4084)
* test(#3301): add failing-first plan coverage manifest tests

Failing-first regression tests for the plan-id manifest, the updated Review
Instructions, and the mechanical per-reviewer coverage check, ahead of the
review.md implementation. RED baseline before the fix lands.

* test(#3301): raise allow-test-rule-refs unverified ceiling for new marker

Adding tests/review-plan-coverage-manifest.test.cjs's source-text-is-the-product
marker grows the unverified-exemption pool by one (282 -> 283), the same
documented growth path scripts/lint-allow-test-rule-refs.cjs's own failure
output names. Confirmed clean via 'npm run lint:allow-test-rule-refs' locally.

* feat(#3301): tell reviewers the plan ids and total count, grade coverage

build_prompt now derives a plan-id manifest from each *-PLAN.md filename
(stripping the -PLAN.md suffix) and appends it, with the total plan count,
to both gsd-review-instructions.md and gsd-review-prompt.md. The Review
Instructions prose requires one heading-verbatim section per id before any
cross-plan or overall-risk content.

write_reviews grades each dispatched lane's real (non-stub, non-empty)
review against that same manifest and records an optional plan_coverage:
frontmatter block, present only when a lane is incomplete. The match
escapes regex metacharacters in the id and excludes a preceding/trailing
hyphen or word character as a boundary, closing the two traps named in the
issue (a decimal phase like 12.6 satisfied by 12X6-01; a threat id like
T-04-07 registering as coverage of plan 04-07). CodeRabbit is exempt, since
it never receives the source-grounding prompt carrying the manifest.

This closes the gap where a review that silently covers only some plans in
a multi-plan phase is indistinguishable from one that covers all of them.

* docs(#3301): add changeset fragment

* test(#3301): use t.after() instead of try/finally for cleanup

CONTRIBUTING.md bans try/finally inside test bodies. Code review caught
this in the new coverage-manifest test file; switch every fixture-cleanup
site to the approved t.after() pattern.

* test: use t.after() instead of try/finally in #3300's build_prompt tests

Pre-existing try/finally-for-cleanup pattern in this file (landed for
#3300) violates CONTRIBUTING.md's explicit ban on try/finally inside test
bodies. Surfaced incidentally while reviewing #3301's diff, which cites
this file as its extraction-pattern precedent; fixed inline per the
no-defer rule rather than deferred to a separate PR.

* test(#3301): anchor coverage-check extraction on the fence line, not prose

`.plans-manifest.md` also appears in write_reviews' own prose ahead of the
```bash fence, so indexOf found that occurrence first and the
backward-walk-to-fence-open landed on the earlier, unrelated gate-check
block instead. gsd-test caught this: coverage-check tests expecting a real
verdict got null, because the wrong block ran and never writes
.plan-coverage-<slug>.json. Anchor on the fence-only bash assignment line
instead.

Emitted-Drift-Ack-Growth: review.md — #3301 adds the plan-coverage manifest and mechanical coverage check to build_prompt/write_reviews.

* fix(#3301): route id escaping through the canonical pattern seam

ADR-3212 (epic #3212) consolidated ~44 hand-rolled regex-escape copies into
one owner, src/pattern.cts's escapeRegex, specifically to stop this exact
class of duplication. My coverage-check node -e script hand-rolled the
identical metachar-escape regex — invisible to eslint-rules/no-adhoc-regex-escape.cjs
only because it lives inside a workflow markdown file, not a .cts/.cjs
source file the shape-matching guard scans. Require the compiled seam
(gsd-core/bin/lib/pattern.cjs) instead, matching the established
node -e-requires-a-compiled-lib idiom already used elsewhere in this
workflow (code-review.md's code-review-flags.cjs/code-review-depth.cjs
calls). Verified both named traps from the issue still resolve correctly
under escapeRegex's RegExp.escape-backed implementation, which differs in
escaped-text shape (hex-escapes hyphens/leading chars) but not match
result.

* test(#3301): run coverage-check block with cwd at the repo root

The block's node -e now requires ./gsd-core/bin/lib/pattern.cjs, a path
relative to the repo root (correct for production, which always runs
from there). The test harness ran it with cwd at the fixture's own temp
dir instead, so the require failed. Add an optional cwd param to
runScript (default: root, unchanged for the plan-copy-block tests) and
pass the real repo root for every coverage-check call site. Manually
verified end-to-end before spending another remote run: the extracted
block now produces the expected {complete:true} verdict.

* docs(#3301): backfill changeset pr number (pr:0 -> pr:4084)

---------

Co-authored-by: sim <sim@local>
2026-08-30 14:25:24 -04:00
Tom Boucher
4d70b4dc43 fix(#4068): add --merge-async to c8 coverage-merge invocations (#4069)
* test(#4068): failing-first regression guard for c8 --merge-async flag

Adds a config-invariant test asserting test:coverage:unit and
test:coverage:report pass --merge-async to c8, guarding against the
coverage-merge OOM (release.yml finalize dry-run, exit 134/SIGABRT)
regressing a third time (prior stopgap: #199).

Also commits the sourced research memo backing the diagnosis.

* test(#4068): rename regression test to avoid lint-test-file-count collision

coverage-merge-async-flag.test.cjs's effective prefix (coverage-merge-
async-flag) matched the unrelated gsd-core/bin/lib/coverage.cjs module's
2-file cap under lint-test-file-count.cjs's startsWith bucketing, failing
FAIL_EXCEEDS_LIMIT (confirmed via the RED gsd-test run on 7b6e6ca9e).
Renamed to c8-merge-async-flag.test.cjs -- no colliding prefix.

* fix(#4068): add --merge-async to c8 coverage-merge invocations

The `test:coverage:unit` script OOM-crashed (exit 134, SIGABRT) in the
release.yml finalize dry-run of 1.12.0: all 1785 unit tests pass, then
c8's report/merge phase crashes ~76s later against the 6144 MB heap
ceiling. Root cause (verified against this repo's pinned c8@11.0.0
source, node_modules/c8/lib/report.js): the default sync merge path,
Report._getMergedProcessCov(), loads every raw V8 coverage file for the
whole run into memory as one array before merging. This is a recurrence
of #199 (same OOM class at ~466 tests, "fixed" by raising the heap
ceiling) -- the suite outgrew 6144 MB as it grew to 1785 tests, and
test.yml's own coverage-gate job already independently hit and
stopgap-fixed the identical class once (test.yml:614-620, 4096->8192 MB).

c8 ships the upstream fix for exactly this: --merge-async (c8 v7.14.0,
already inside the pinned c8@^11.0.0 range -- no dependency bump),
which switches to Report._getMergedProcessCovAsync(), reading and
merging one raw file at a time instead of loading them all at once.
The merge arithmetic (mergeProcessCovs()) and --all's zero-coverage-file
inclusion are unchanged between the two paths, so the coverage gate's
accuracy is unaffected.

Verified empirically (not just by source inspection): against a 223 MB
synthetic raw-coverage corpus derived from this repo's actual 235
gsd-core/bin/lib/**/*.cjs files (same order of magnitude as the 358 MB
figure documented in test.yml), c8's real Report class crashes with the
same "Ineffective mark-compacts near heap limit" signature via the sync
path at a 300 MB heap ceiling, while the async path completes at a flat
~40-50 MB peak down to a 100 MB ceiling. Full repro steps and numbers:
.gsd/bug/fix-4068-coverage-merge-oom/10-diagnosis.md.

Only `test:coverage:unit` and `test:coverage:report` are changed.
`test:coverage` (line 150) is a non-CI dev-convenience script, never
invoked as a command by any workflow. `test:coverage:unit:raw` (line
153) already runs --reporter none, which skips the merge/report phase
this flag affects, entirely -- adding it there would be a no-op.

Fixes #4068

* fix(#4068): add changeset fragment for the coverage-merge OOM fix

The changeset fragment was generated locally (npm run changeset) but
never committed -- isolated code review caught the gap (no
.changeset/*.md on the branch, confirmed via git diff --name-status).

* chore(#4068): backfill changeset PR number

pr:0 -> pr:4069

---------

Co-authored-by: sim <sim@local>
2026-08-29 22:03:28 -04:00
Dennis Kim
8487f0ed42 enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage

- pin configured, absent, and malformed branch-list behavior
- require opposite CLI and execute warning outcomes

* feat(01-01): warn on configured protected branches

- resolve the base branch union configured protected branch names
- expose exact boolean CLI comparison output for workflow callers
- keep execute-phase warning advisory and within its byte budget

* test(01-01): add failing protected branch config coverage

- cover valid list persistence and null unset
- reject hostile shapes while preserving the prior value

* feat(01-01): validate protected branch configuration

- register git.protected_branches as a canonical config key
- require a non-empty array of non-blank branch names

* test(01-02): add failing ship protected-branch controls

- Execute both workflow warning blocks with exact predicate arguments
- Require true and false results to produce opposite warning outcomes
- Preserve the none-strategy feature-branch offer contract

* feat(01-02): warn at ship on protected branches

- Reuse the typed protected-branch predicate in ship preflight
- Keep raw base resolution for PR targeting and advisory branch creation
- Prove execute and ship warning blocks with opposite-result controls

* test(01-02): add failing protected-branch docs parity

- Require the canonical schema key in both English config references
- Pin the non-empty string-array type and absent default
- Require synchronized multi-branch examples and advisory semantics

* feat(01-02): publish protected branch configuration contract

- Document the optional non-empty string-array field in both references
- Explain resolved-base union and absent-field compatibility
- Keep execute and ship warnings advisory under branching_strategy none

* fix(01): CR-01 honor active workstream branch policy

* fix(01): WR-01 assert protected config path selection

* docs: add changeset fragment for #3648

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx

* fix(#3648): resolve base_branch precedence inversion and round-1 findings

Blocker 1/2: production config resolution was flat-first, so a project
that migrated to git.base_branch but still carried a stale flat
base_branch got the old value back. Add base_branch to
normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos
pattern: canonical nested wins) and route readEffectiveGitConfig's
test seam through the same normalization so it can't silently diverge
from production again. Adds a regression test with both keys set that
fails without the fix.

Blocker 3/4/5: restore the handle_branching case-selector prose and
"none" contract sentence that #3389's tests anchor on, and revert the
unrelated prose/comment compaction in the same step — both were
drive-by edits outside #3552's scope.

Also addresses review majors/minors: delete readConfigBaseBranch and
readConfigProtectedBranches (dead in production, only self-tested);
--is-protected now fails closed (reports protected) instead of
silently answering false when the base branch can't be verified;
trim configured protected-branch names; fix HOME-without-USERPROFILE
vacuous isolation on Windows; correct the drift-ack's byte accounting.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N

* test(#3648): add failing legacy-key hoist safety coverage

Round-2 review found normalizeLegacyKeys block 5 records a normalization
carrying the DISCARDED flat value on the canonical-wins branch. Probing
that turned up a second, unreported defect in the same helper shape:
blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no
object guard, so a config whose section key holds a string is spread into
index keys —

  {"git":"main","base_branch":"release"}
    -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}}

The resolved value is accidentally still correct, so nothing fails and no
diagnostic fires. But normalizations.length > 0 sets configDirty, and
config-loader then serializes that shape back into the user's
config.json — a read that silently corrupts config.

The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"}
explicitly; this is the input it would have caught.

Covers both defects across blocks 1 and 5, with object/array/null
negative controls that must stay green in both phases, and a fast-check
property over arbitrary `git` values.

* test(#3648): pin fail-closed handling of malformed protected_branches

Replaces the test that pinned the fail-OPEN behaviour. The old
assertion — ['develop', 42] yields isProtected === false for 'develop' —
locked in the exact failure #3552 exists to close: config-set validation
is bypassable by a direct edit of .planning/config.json, so a user who
believes 'develop' is protected got a silent false and no warning.

It was also inconsistent with the fail-CLOSED direction twelve lines
away, where an unverified base reports protected and writes a
diagnostic. A protection predicate must not have two opposite failure
directions depending on which input is bad (#3648 review Blocker 3).

New coverage: a bad element drops only itself, a non-array contributes
no names, an empty list is well-formed rather than malformed, and
--is-protected surfaces the rejection. Both negative controls — a clean
list reports nothing rejected and writes no diagnostic — must stay green
in either phase, so the reject channel cannot fire unconditionally.

* fix(#3648): drop only invalid protected_branches and report them

Partition git.protected_branches instead of discarding the whole list on
one bad element, and carry the rejections out through
ProtectedBranchStatus so --is-protected can name them on stderr. Valid
names keep protecting; the user finds out the rest were ignored.

A non-array value still contributes no names — a bare string is not a
list of branch names — but is now reported rather than swallowed. An
empty array stays silent: declaring no extra protected branches is a
valid choice, not a misconfiguration.

writeDiagnostic is hoisted out of the unverified-base branch since both
arms now use it.

* test(#3648): prove the predicate diagnostic survives both call sites

The workflow bash stub now emits a stderr diagnostic the way the real
command does, which is what makes a swallowed `2>/dev/null` visible to a
test — previously the stub was silent on stderr, so discarding it changed
no observable behaviour and the call sites could drop the explanation
undetected.

Adds the Minor 2 binding check as well: ship must expose the predicate
result as IS_PROTECTED rather than only echoing a warning, asserted by
running the extracted bash and reading the bound value, not by grepping
the workflow source.

Both tests carry opposite-outcome controls — an empty diagnostic must
leave the text absent, and a false predicate must bind false.

* fix(#3648): surface the predicate diagnostic and bind ship's result

Drop `2>/dev/null` from the --is-protected call at both call sites. The
fail-closed explanation and the new rejected-entry warning both go to
stderr, so discarding it left the user with a bare "protected branch"
warning on a branch that is not protected and no way to tell a real
match from a degraded-git guess. `git branch --show-current` keeps its
own redirect — that one is genuine noise.

ship.md binds IS_PROTECTED and its prose now branches on the variable,
so the following steps have evaluable state instead of having to infer
it from warning text in tool output.

execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth
319 bytes (was 331 before the redirect came out). Baseline re-verified
against the current rebase base by blob id; the ceiling check passes
with 755 bytes of margin.

* test(#3648): restore negative space for the readFile config seam

The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm
it pinned survives verbatim in readEffectiveGitConfig's readFile branch —
the JSON.parse catch, the non-object guard, the git-section object guard,
.trim() and blank-string rejection — and the four surviving readFile
injections were positive-path only. protected_branches was never driven
through this seam at all.

Restores nine cases against the seam, including protected_branches
partitioning, plus a control proving loadConfig still wins when both
seams are supplied.

Records honestly what the suite pins. Mutating the built lib shows
.trim() is KILLED, while the non-object guard and the blank-string
rejection SURVIVE — both are unreachable through this entry point for
the same reasons the deleted suite documented against its own
equivalents: a JSON-parsed non-object carries no relevant own-property
either way, and a blank value is rejected a second time downstream by
the resolver's truthiness check. They stay as defence-in-depth and are
labelled known-unkillable rather than left looking like coverage this
suite does not provide.

* test(#3648): distinguish detached HEAD from a missing branch argument

`args[1] ?? ''` collapsed two different situations into one: a detached
HEAD, where `git branch --show-current` legitimately prints nothing, and
the flag being called with no argument at all. Both answered false, so
the right outcome arrived by an unintentional path and a caller bug was
indistinguishable from normal operation.

Asserts the detached case stays silent and the missing-argument case
reports, with a control that the two diagnostics differ.

* fix(#3648): report a missing --is-protected branch argument

Answer false either way, but say so when the flag arrives with no
argument. A detached HEAD passes an explicit empty string and stays
silent, since that is a normal state rather than a misconfiguration.

* docs(#3648): state exact-name matching and per-entry rejection

isProtected is exact string equality, so a git-flow project must
enumerate every release/* and hotfix/* by name. #3552 only asked for an
integration-branch field, so the implementation satisfies the letter of
the issue while leaving its git-flow motivation partly unserved — say so
where users will meet it rather than leaving them to discover it.

Also documents the Blocker 3 behaviour change: an invalid entry is
ignored with a warning naming it and the remaining names still apply.

Both statements land in docs/CONFIGURATION.md and
gsd-core/references/planning-config.md, and the config-field-docs parity
test asserts each in both so the two cannot drift.

* refactor(#3648): extract isValidProtectedBranches for cross-surface pinning

The `git.protected_branches` check inside `cmdConfigSet` and the resolver's
per-entry filter in `git-base-branch.cts` are deliberately different shapes —
all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot
fail the guard open. Nothing structural keeps their two definitions of "usable
branch name" in step.

Lifting the write-side check into a named, exported predicate lets a property
test ask both surfaces about the same value and assert they agree, which is the
fast-check gap the round-2 review flagged. No behaviour change: the predicate is
the same expression, called from the same place.

* fix(#3648): stop --is-protected rewriting the config it is asking about

`gsd_run query git.base-branch --is-protected` runs on every execute-phase and
every ship. It resolved config through `loadConfig`, whose normalize-then-write
path rewrites `.planning/config.json` whenever any legacy key normalizes — so a
boolean question was silently editing the user's checked-in config. This PR had
widened the trigger by adding a fifth normalization block (top-level
`base_branch` -> `git.base_branch`), making it fire for exactly the projects the
feature targets.

`loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution
is unchanged, only the two write-back side effects are suppressed. The predicate
passes `persist: false`; the ~30 other callers are untouched, so a legacy config
is still migrated by ordinary use.

Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and
reflows whitespace even when the values are equivalent. Three tests, each with
its own control: the end-to-end CLI leaves the file byte-identical while still
answering `true` from the legacy key (proving the config WAS read); an ordinary
persisting load of the same fixture DOES change the bytes (proving the fixture
is live rather than inert); and `persist:false` vs default over one directory
returns deep-equal config while differing on the write. Reverting the one-line
`persist: false` fails the first of those and only that one.

Also from the review:

- `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through
  the same precedence authority production uses". It does not, and cannot — it
  reproduces two of production's steps over a single file. The comment now names
  what the seam covers and what it does NOT (root/workstream deep merge, builtin
  and global defaults, federated merge), and the seam now applies production's
  flat-then-nested lookup so it stops disagreeing about a surviving flat key.

- The missing-argument diagnostic promised "answering false", which the
  fail-closed guard on the same call can contradict by printing `true`. It now
  states what it did with the argument and leaves the answer to stdout.

* test(#3648): re-pin block 5 on #3760's refusal contract

#3767 landed on next while this PR was in review and fixed the non-object
config-section defect properly: a present-but-non-object section now BLOCKS its
own migration — value preserved, no Normalization pushed, refusal reported via
`skipped[]` — rather than being rebuilt from a plain-object view. That supersedes
this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread
but still dropped the section value silently, and which the round-3 review
correctly called out as destruction in place of corruption. The rebase drops that
commit and routes block 5 through the upstream helper.

This file's tests asserted the superseded design, so they are rewritten to pin
block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's
suite was written — against the contract that now governs it: ordinary hoist into
an absent/null/object section, canonical-nested-wins, and refusal for each of
string/number/boolean/array sections with the exact `skipped` entry.

Two controls keep it from passing vacuously: the refusal must be scoped to block
5 (an unrelated block still normalizes in the same call), and a property over
arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually
exclusive per key, that a refusal leaves both the section and the legacy key
untouched, and that a hoist manufactures no index key the input did not carry.

* docs(#3648): correct the Git Query and Config Loader module contracts

CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct
`.planning/config.json` read. Since this PR it is the EFFECTIVE configuration
resolved by the Config Loader — a materially different authority, carrying the
root/workstream deep merge, flat-then-nested lookup and builtin/federated
defaults. The `--is-protected` predicate, `git.protected_branches`, and the two
invariants that distinguish the predicate from the plain query (fails closed on
an unverified base; must not write) were undocumented entirely.

The Config Loader entry now states that loading is not side-effect-free by
default and documents `options.persist`.

docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and
no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write`
was run and produced no diff: the manifest indexes roster NAMES, not row prose,
so a description edit cannot move it.

Also closes the global-defaults minor: `git.protected_branches` is inert in
`~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key
appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is
section-wide and predates this PR, so the fix is to state the scope where users
meet it rather than to quietly extend the resolution set for two new keys.

* fix(#3648): close four defects found by the round-4 external review

Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially
against this branch. Four findings reproduced against source; each is fixed with a
failing-first test and a control, and each fix was verified by reverting it and
watching exactly the intended test fail.

1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was
   only half closed. `loadConfigResolved` re-enters itself with a bare
   `{ workstream: null }` when a workstream has no config.json of its own, and
   that literal discarded every other option — so the recursive pass ran at the
   DEFAULT persistence and rewrote the ROOT config. Reproduced: with
   GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected`
   rewrote `.planning/config.json` despite `persist:false`. Both recursions now
   forward `options` and override only `workstream`; the explicit override still
   wins the hasOwnProperty check, so spreading cannot let `workstreamContext`
   reintroduce a workstream.

2. Both workflow call sites failed OPEN, and aborted under `set -e` (both
   reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty
   string when the query fails, so `[ "$X" = true ]` was simply false: no
   warning, no trace — a silent hole in the guard whose only job is to warn. The
   bare assignment also aborted the step under `set -e`. Both sites now degrade
   VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the
   check did not run. Deliberately not fail-closed — claiming "protected" on no
   evidence would warn on every branch whenever gsd-tools is unavailable.

3. `isValidProtectedBranches` and the resolver disagreed on a sparse array
   (antigravity). `.every()` skips holes; the resolver's `for...of` yields
   `undefined` for them, so `["main", , "develop"]` was accepted by config-set
   and rejected by the resolver. The cross-surface property passed only because
   `fc.array` cannot generate a hole. The predicate now indexes, and the
   generator punches holes so that axis is actually falsifiable. JSON cannot
   express a hole, so this is unreachable in production — but two definitions of
   one predicate must not contradict each other.

4. A top-level `protected_branches` silently outranked `git.protected_branches`
   (antigravity). Routing the key through `get(key, {section, field})` gave it
   flat-then-nested precedence, which is back-compat for keys
   `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has
   no legacy form, so that invented an undocumented alias. It now resolves
   nested-only through a new `getNested`, in production and in the test seam.
   `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's
   refusal path can leave behind — and a control pins that distinction.

Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed
only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`);
a git command that runs and exits non-zero counts as a clean negative, so a cwd
that is not a repository answers `false`, not `true`. Verified pre-existing on
next @ 738f42f4, so the documentation was over-claiming rather than the code
regressing — but an over-broad contract is exactly what the module docs must not
carry.

Both workflow byte figures re-derived after the call-site change:
execute-phase.md 92356 -> 92865 (+509), ship.md 36784 -> 37227 (+443).

* test(#3648): pin git config read parity

* docs(#3648): document git query contracts

* fix(#3648): expose protected branch default

* test(#3648): snapshot planning tree for read-only query

* test(#3648): pin planning snapshot stray-write detection

* fix(#3648): resolve merge conflict from #3078's ack-fragment sweep

next swept the fully-spent 2818/3003 ack fragments this branch had
appended to (#3078, a84f7563). Rebased onto upstream/next and took
the deletions on both, then moved the #3552 append into a new
fragment of its own.

Rebasing onto the current base also left execute-phase.md only 34
bytes under the frozen ADR-857 Phase 6 margin ceiling (93400 bytes) —
intervening next PRs consumed the rest while this PR was in review.
Extracted the "none" arm's protected-branch-warning bash block into
gsd-core/workflows/execute-phase/steps/protected-branch.md (content
unchanged, matching the existing steps/ extraction pattern used
elsewhere in this file) so the inline growth is a one-line pointer
instead of the full block. 93366 -> 93385 bytes (+19), 15 bytes
inside the ceiling.

* fix(#3648): drop stale ack entry for the new step file

The extracted execute-phase/steps/protected-branch.md needed no
acknowledgment of its own — the differential-attribution check flagged
the entry as stale once the build ran, so removed it and kept the two
growth entries (execute-phase.md, ship.md) that actually needed one.

* fix(#3648): follow the step-file reference in the bash-extraction test helper

extractProtectedBranchWarningBash() read the "none" arm's bash block
directly out of execute-phase.md. That block now lives in
execute-phase/steps/protected-branch.md (byte-ceiling extraction);
the helper follows the step-file reference and extracts from there
when no inline block is found, so the three execute-phase tests that
execute this bash for real keep exercising the actual behavior.

* fix(#3648): regenerate INVENTORY-MANIFEST.json and satisfy the CRLF-fragile lint rule

- gen-inventory-manifest.cjs --write to pick up the new
  execute-phase/steps/protected-branch.md entry (already covered by
  docs/INVENTORY.md's generic workflow_steps wildcard row, so no
  INVENTORY.md edit is needed).
- Reworked the step-file-reference lookup in
  extractProtectedBranchWarningBash() to avoid a bare-\n regex split
  on file content (local/no-crlf-fragile-split), using the same
  line-array scan the function already uses elsewhere.

* fix(#3648): regenerate golden install-tree fixtures for the new step file

npm run gen:install-tree, adding gsd-core/workflows/execute-phase/
steps/protected-branch.md to all 19 runtime install-tree fixtures.
CI's tests/golden-install-tree.test.cjs caught this on push — I'd
verified the differential-attribution and INVENTORY-MANIFEST checks
but missed this separate golden-fixture check for the new file.

* fix(#3648): add the canonical gsd_run preamble to the new step file

CI's runtime-launcher-parity suite requires exactly one canonical
resolver preamble in every workflow .md that calls gsd_run. The
inline "none"-arm block never needed one (execute-phase.md already
carried a preamble elsewhere in the same file), but the extracted
execute-phase/steps/protected-branch.md is now its own file with no
preamble of its own. Ran node scripts/sync-runtime-launcher.cjs to
insert it (execute-phase.md itself is untouched — still 93385 bytes,
inside the ADR-857 ceiling).

That preamble defines its own gsd_run(), which shadows the mock
tests/git-base-branch.test.cjs injects for the three #3648 tests that
execute this bash for real — without stripping it, those tests reached
the real gsd-tools.cjs on the machine running them instead of the
test's fixture. Preamble correctness is already covered by
tests/runtime-launcher-parity.test.cjs, so extractProtectedBranchWarningBash()
now strips the preamble line before handing the block to the harness;
it only needs to exercise the #3552 warning logic.

* fix(#3552): address PR 3648 review feedback on protected branch warnings

- Fix execute-phase handle_branching branching_strategy=none instruction
  to "Read and execute execute-phase/steps/protected-branch.md"
- Use io.error(..., ERROR_REASON.USAGE) for cmdGitBaseBranch usage errors
- Align git.protected_branches schema default to (none) without fallback []
- Relocate CONTEXT.md forward-referencing sentence into module body
- Sanitize control and ANSI characters in renderRejected diagnostics
- Clean up out-of-scope whitespace hunks in gsd-tools.cjs

Emitted-Drift-Ack-Growth: execute-phase.md — #3552: execute-phase handle_branching adds a pointer to execute-phase/steps/protected-branch.md for branching_strategy=none so the protected-branch check executes while keeping execute-phase.md within the ADR-857 Phase 6 margin ceiling (93400 bytes). 93392 bytes, 8 bytes inside the ceiling.
Emitted-Drift-Ack-Growth: ship.md — #3552: ship preflight step 3 now asks the same typed git.base-branch --is-protected predicate as execute-phase, binding IS_PROTECTED and warning without refusing execution or blocking the branching_strategy=none feature-branch offer; it degrades visibly (rather than silently reading an empty result as "not protected") when the query itself fails to run. 36841 bytes, well inside the XL cap (98304, tests/workflow-size-budget.test.cjs).

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-29 17:00:52 -04:00
0xdhx
472f585f7c fix(#3726)!: require --confirm before milestone complete mutates (#3774)
* fix(#3726): require --confirm before milestone complete mutates

`milestone complete <version>` is a one-way door — ROADMAP.md and
REQUIREMENTS.md archived, every phase directory in the milestone MOVED,
STATE.md rewritten — and ran unconditionally on first invocation through
every invocation path, including `query milestone.complete <version>`,
whose `query` meta-prefix reads as a read-only namespace but performs no
filtering (#167's invocation-compatibility shim + #3243's dotted-form
normalization).

The gate lives on the destructive command itself, not on the `query`
prefix (the prefix is an intentional invocation mechanism, not a
permission boundary — restricting it would break dozens of shipped
workflow callers). Without --confirm and without --dry-run the command
now refuses via error() before reading anything beyond its arg checks,
so an unconfirmed invocation is a guaranteed no-op on disk. --dry-run
still previews with no confirmation needed and is now documented in the
usage block (it was only documented for the sibling archive-quick).
--force keeps its narrow meaning — bypassing the TRUNCATED-scope and
unstarted-phase guards — and does not double as the mutation opt-in.
--confirm follows the existing `phases clear --confirm` idiom in the
same module.

complete-milestone.md's two invocations pass --confirm (the workflow has
gathered explicit user intent by that step). Existing tests get
--confirm appended — pre-change behavior is exactly confirmed behavior —
and a #3726 regression block covers: refusal + full-tree byte-identity
on both invocation forms, --force not satisfying the gate, --dry-run
still passing without confirmation, and --confirm proceeding. The
refusal tests fail against pre-fix code (negative control run).

Fixes #3726

* docs(#3726): document the --confirm requirement in CLI-TOOLS and COMMANDS

Cross-AI review of the fix diff (codex, pre-create) caught three shipped
doc sites still instructing the now-refused bare invocation: the
CLI-TOOLS.md milestone-complete synopsis + flag table, and COMMANDS.md's
two guard-override instructions (`--force` alone now refuses without
--confirm). Localized CLI-TOOLS copies already lag the English synopsis
(no --force/--dry-run either) and follow the translation pipeline, not
this fix.

* chore(#3726): set changeset fragment pr to 3774

* test(#3726): confirm-gate CI repairs — QA scenario caller + growth ack

Two CI reds from the --confirm gate, both this branch's own misses:

- tests/qa/scenarios/milestone-rollover.json invoked `milestone complete
  1.0 --force` as a JSON arg-array fixture — a caller shape the test
  sweep (which grepped runGsdTools/runSdkQuery in tests/*.cjs) never
  enumerated. Adds --confirm; the scenario's boundary-crossing contract
  is otherwise untouched.
- complete-milestone.md's +420-byte --confirm note trips the
  emitted-attribution growth ratchet. Acknowledged as a #3726 append to
  the existing complete-milestone.md entry in
  3409-unreachable-guard-arms.json (two ack sources may never name the
  same path, per that fragment's own precedent).

Local: lint-emitted-drift-ack ok; loop-walk.qa 115/115 green sandboxed.

* docs(#3726): CLI-TOOLS.md guard-override sentences say --force --confirm

Review Major 1: the truncated-window and unstarted-phase guard paragraphs
still told the reader to "Pass `--force` to override", which now refuses
(--force alone does not satisfy the confirmation gate), while the flag
table 470 lines later said the opposite. Mirror the docs/COMMANDS.md pair
so the file no longer contradicts itself.

* docs(#3726): synopsis renders --confirm and --dry-run as alternatives

Review Nit 1: `milestone complete <version> --confirm [--dry-run]` read as
"a dry run still needs --confirm", the opposite of AC 3. Render the pair
as `(--confirm | --dry-run)` in the CLI-TOOLS.md synopsis and the usage
docblock, and let the flag rows carry the rule.

* test(#3726): pass --confirm in base-added milestone fixtures; re-file the growth ack

Rebase onto next (26 commits) surfaced three tests the gate now refuses:
the #3685 write-flag contract pair in tests/milestone.test.cjs and the
`milestone complete` boundary fixture in tests/state-contract.test.cjs
all invoke the command bare. Each now passes --confirm (a mutating run is
exactly what they assert on).

The +420 byte complete-milestone.md growth ack rode on
3409-unreachable-guard-arms.json, which #3078 swept from next as fully
spent — hence the modify/delete conflict. Re-filed under a fresh fragment
named for this issue, never resurrecting the swept one.

* test(#3726): pin the present-but-falsy arm of the confirmation gate

Review Minor 1: the boundary triple covered absent and present but not
present-but-falsy. The gate is an exact-token match, so --confirm=false
and --confirm=0 refuse today — pinned (canonical + query forms, whole
.planning/ tree byte-identical) so a future `=`-aware or prefix-matching
parser cannot silently turn --confirm=false into a confirmed run of an
irreversible command.

* test(#3726): drop --confirm from dry-run-only invocations

Review Nit 2: --confirm was mass-appended to 14 pre-existing --dry-run
invocations that never needed it, so each stopped standing as incidental
proof that a preview needs no confirmation. Reverted to the pre-PR form;
the dedicated AC-3 test carries the explicit assertion.

* docs(#3726): sync the localized CLI-TOOLS synopsis with the confirm gate

REQ-I18N-02 (docs/features/internationalized-documentation.md) requires
translations to stay synchronized with the English source. The four
localized CLI-TOOLS.md guides still advertised a bare
`milestone complete <version>`, which now exits 1. Render the English
synopsis verbatim — `(--confirm | --dry-run)` plus the `[--force]` and
`[--archive-quick]` flags the translations had also fallen behind on.

* test(#3726): drop --confirm from the remaining preview-only invocations

Round 2 reverted the --confirm appends on --dry-run-only invocations in
tests/milestone.test.cjs, but four more sat in two files the sweep missed:
tests/milestone-archive.test.cjs (three) and
tests/milestone-window-single-owner.test.cjs (one).

Each is a preview run whose whole purpose is to document that a preview
mutates nothing, so `--dry-run ... --confirm` contradicted the semantics
the test exists to pin. Dropping the token restores each as incidental
proof that a preview needs no confirmation; the dedicated AC-3 test keeps
the explicit assertion.

No assertion added, relaxed, or removed — the change is four tokens.

* chore(#3726): migrate the emitted-drift ack from a fragment to a commit trailer

#3954 (ADR-3942) moved emitted-drift acknowledgments out of
tests/emitted-drift-acks/ and into git commit trailers, and the fragment
directory no longer exists on next. The reason this PR's fragment carried
moves verbatim into the Emitted-Drift-Ack-Growth trailer on this commit;
the fragment file is removed rather than resurrected.

Emitted-Drift-Ack-Growth: complete-milestone.md — #3726: +420 bytes (40186 -> 40606). The archive_milestone step's two `milestone complete` invocations now pass the required --confirm flag (the command refuses to mutate without it — the archive is irreversible), with a note explaining the flag and pointing at --dry-run for previews. Deliberate runtime-loaded workflow text for the new gate, not converter drift.

* fix(#3726): name --confirm in the version-required refusal

The documented arg-discovery path (gsd-tools.cjs top-level usage: invoke
the command without args and the error lists what is required) stopped at
`version required for milestone complete (e.g., v1.0)` — one required
argument short. Discovering --confirm took a second round trip through the
gate. The refusal now reads `… — and --confirm to mutate`, pinned by a test
that also asserts the version-less invocation leaves .planning/ untouched.

* test(#3726): pin the milestone complete docs against a silent regression

The changeset is `type: Fixed`, which the docs-required lint exempts, so
nothing in CI would notice a later edit that reinstated the bare-`--force`
override prose or dropped `--confirm` from the synopsis. Four tests in
tests/milestone.test.cjs now pin: the synopsis line in docs/CLI-TOOLS.md
and its four localized mirrors; the `--confirm` flag row; both
guard-override instructions in docs/CLI-TOOLS.md and docs/COMMANDS.md,
by guard name (a substring match on each instruction's `--force
--confirm` text); and — as an identity ratchet over the
milestone-complete sections — every `--force` sentence or clause that
lacks `--confirm`, so a new bare instruction in its own sentence or
clause fails whatever its wording. Named residual: a bare instruction
spliced into the same clause as a compliant one coalesces with it and
passes the ratchet; the by-name pins are what keep the four known
instructions from losing the pairing that way. The file is registered
in scripts/docs-guard-registry.cjs so the pin runs on the PR that
changes those docs, not only after merge.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-29 17:00:45 -04:00
Tom Boucher
370cfc6680 enhance(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% (#4043)
* feat(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90%

Adds two new mechanisms plus an audit-coverage extension:

- scripts/lib/ci-job-timing.cjs: shared elapsed-vs-cap arithmetic
- scripts/ci-check-job-near-cap.cjs: in-job advisory near-cap check,
  wired into test/test-full/mutate/smoke as each job's last step
- scripts/ci-timeout-report.cjs + .github/workflows/ci-timeout-report.yml:
  scheduled REST-API poll that appends new records to
  tests/ci-timeout-budget-history.jsonl and opens a small data-only PR
- tests/ci-test-job-timeout-budget.test.cjs: extended to cover mutate
  (mutation.yml) and smoke (install-smoke.yml), which previously had no
  headroom-factor gate coverage at all

Does not change any timeout-minutes value, shard composition, or shard-1
contents — those stay maintainer policy calls per the issue's own scope.

* fix(#4036): address two-orthogonal-review findings

- Parity tests guarding the two hand-duplicated literals this design
  cannot single-source through GH Actions YAML: CI_JOB_TIMEOUT_MINUTES
  vs each job's own timeout-minutes, and ci-timeout-report.cjs's
  JOB_RULES name-prefixes vs each job's actual name: template.
- Thread run.event through as runEvent on every persisted record, so
  PR-context and push-context install-smoke timings (genuinely
  different matrix shape) are distinguishable in the history rather
  than silently conflated under one job name.
- Replace the Windows near-cap start-time step's ambiguous PowerShell
  +/>> precedence with GitHub's documented string-interpolation form.
- Move github.run_id out of direct ${{ }} shell interpolation into an
  env: var in the new scheduled workflow, per this repo's own
  expression-injection-safe convention.

* test(#4036): regenerate golden install-tree fixtures for scripts/lib/ci-job-timing.cjs

npm run gen:install-tree — scripts/ ships wholesale into the installed
package (per ADR/known-defect precedent from #4012's own PR history: a
new scripts/lib/*.cjs file needs its golden entry regenerated or every
runtime's install-tree test fails). Confirmed via gsd-test: this was the
sole cause of the first real verification run's 25 failures (all in
tests/golden-install-tree.test.cjs, one per runtime). Top-level
scripts/*.cjs files (ci-check-job-near-cap.cjs, ci-timeout-report.cjs)
are not individually tracked in these fixtures — consistent with every
other existing top-level scripts/*.cjs file, so no entry was expected
or added for those two.

* fix(#4036): register new lib file with installer, fix H1 shell policy

- bin/install.js: add ci-job-timing.cjs to GSD_SCRIPTS_LIB_FILES (a
  hand-maintained registry, not generated — tests/install.test.cjs
  asserts every scripts/lib/ file is enumerated here)
- test.yml: replace the two OS-conditional "Record job start time"
  step pairs (test + test-full jobs) with a single unconditional
  `node -e` step. The prior pair's Windows variant declared an
  explicit shell: pwsh, which scripts/workflow-policy.cjs's H1 checker
  statically flags against every OS a job's matrix can realize,
  independent of the step's own if: gate. A single Node one-liner
  needs no shell override at all — it's syntactically valid and
  behaves identically under bash, zsh, and pwsh — which is both H1
  compliant and removes the last OS-specific shell syntax from this
  change entirely.

Both defects were found by a real gsd-test run, not local gates —
lint:ci and build:lib were clean throughout because neither the
scripts/lib/ install-manifest parity check nor the H1 shell-policy
baseline runs as part of lint:ci; both are gsd-test-only suites.

* docs(#4036): how-to for reading CI timeout budget signals

The phase-gate docs check correctly flagged the enablement sequence as
3 real steps (read the near-cap warning, find the accumulated trend
file, pick the right maintainer lever) — a reference table can't carry
a sequence. Adds docs/how-to/read-ci-timeout-signals.md, indexed from
docs/README.md.

* chore(#4036): backfill changeset PR number (4043)

---------

Co-authored-by: sim <sim@local>
2026-08-29 16:13:15 -04:00
Tom Boucher
519ac23ebb fix(#3839): hook tables say PreToolUse (validate-commit) and SessionStart (session-state) (#4041)
* test(#3839): docs hook tables must match surface registrations (failing first)

* docs(#3839): hook tables say PreToolUse for validate-commit, SessionStart for session-state

gsd-validate-commit.sh is registered PreToolUse (src/runtime-hooks-surface.cts;
its exit-2 block IS the contract — a post-tool hook cannot prevent a commit)
and gsd-session-state.sh is registered SessionStart (session orientation, not
post-tool tracking). Both rows said PostToolUse in ARCHITECTURE.md and the
three INVENTORY locales; the issue asked for a neighbouring-row scan, which
is how the session-state row was found. All other rows in the four tables
verify against the surface.

* fix(#3839): review fold-ins — 10 more wrong rows in ko-KR/pt-BR/zh-CN, parser authority + drift pins

Adversarial review found the same two wrong rows shipped in five more
files the issue's table missed (ko-KR ARCHITECTURE+INVENTORY, pt-BR
ARCHITECTURE+INVENTORY, zh-CN ARCHITECTURE) — all fixed; DOC_TABLES now
covers all ten shipped tables. The parity parser unioned only the Kimi
mirror list, silently exempting agent-isolation-guard (registered via
the dynamic preToolEvent push): probes are now parsed too, with bare
hook names resolved against hooks/ ground truth and dynamic event
variables resolved to their canonical (non-Gemini) events; an exact-set
pin replaces the loose size guard. allow-test-rule marker carries the
issue ref; unverified-ceiling 280→281 (audited: the new marker is
legitimate — the suite reads product docs whose text is the contract).

* fix(#3839): register the hook-table parity suite in the docs-guard lane

The new suite reads ten docs/ paths, so lint-docs-guard-registration
requires it in the docs-guard registry — the first GREEN bench run
caught the omission (the RED run's docs-guard failures were the same
signal, previously misread as marker fallout).

* chore(#3839): changeset fragment (pr number backfilled after PR creation)

* chore(#3839): backfill changeset PR number (4041)

---------

Co-authored-by: sim <sim@local>
2026-08-29 11:11:24 -04:00
Tom Boucher
80de48c319 enhance(#3914): every phase records a truthful guard ledger (#4018)
* fix(#3914): retire n/no-process-exit where its successor governs

Epic #3889 criterion 5 — no phase closes with a guard added and its
predecessor left standing — is violated in the tree by the epic that wrote it.

local/require-registered-exit was registered on gsd-core/bin/**/*.cjs and
scripts/**/*.cjs, while n/no-process-exit stayed 'error' over a nine-glob block
covering those same two. Only the hooks 'off' exemption ever came down; the
predecessor's registration never did. Both rules have been enforcing the same
property on the same surfaces since P6.

Narrowed, not deleted. Seven of those nine globs have NO successor —
eslint-rules/, bin/lib/, pi/, examples/, vscode/, .kilo/, .opencode/ — so
deleting the rule outright would silently drop enforcement on all seven. That
is the inversion this epic has already hit three times: removing a coarse guard
because a narrower one exists somewhere it does not reach. Flat config is
last-match-wins and both successor blocks come after the nine-glob block, so
'n/no-process-exit': 'off' in exactly those two retires the predecessor
precisely where the successor governs and nowhere else.

The successor is strictly more precise: it permits process.exit only inside
terminateNow in cli-exit.cts, the single sanctioned terminator (ADR-3889 §3),
where n/no-process-exit permits none and would flag terminateNow's own
generated copy.

Asserted at the consumer's altitude via ESLint.calculateConfigForFile on real
paths, with the positive control that matters: n/no-process-exit is still
'error' on six of the seven successor-less globs, so a future edit that turns
this into a blanket disable goes red. bin/lib/ has no file in this checkout and
is reported as untested rather than given an invented path. Severity is
normalized across the string/numeric/array forms the API can return, and the
normalized value asserted — not truthiness.

Verified by running calculateConfigForFile myself on both superseded globs and
four controls before trusting the test.

Found and fixed inline: the change made an eslint-disable directive at
gsd-tools.cjs:257 partially unused, which --max-warnings 0 rejects; narrowed to
the one rule still in force.

Verification runs on the remote runner.

Refs #3914

* docs(#3914): the epic added three guards, it did not remove one

The audit reconciled the epic ledger against what actually landed. The net is
+3, not -1: four lint:generated-sync --check arms (gen-scripts-cli-exit,
gen-hooks-cli-exit, gen-exit-code-registry, gen-exit-code-docs) plus one rule,
against two retirements.

An epic whose thesis was consolidation ended with a larger guard surface than
it started with. The additions are each defensible; the claim that the total
fell was never true.

Two of the three prior errors in this amendment are mine. It said "Net -1 by
count" above terms reading -1 -1 +1 +1 +1, which sums to +1 — an arithmetic
error in the paragraph directly below the sentence arguing that an ADR about
honest accounting must not pad its own ledger. And the term list omitted two of
the four --check arms, which is what turns that +1 into the real +3.

Recorded rather than quietly rewritten. This ledger has now been wrong three
times — the original -2, the -1 that replaced it, and #3914's own table, which
states -1 above terms summing to 0 — and a written claim nobody checked against
the thing it describes is the exact failure this epic exists to close.

Refs #3914

* fix(#3914): make the successor actually supersede before retiring the predecessor

An isolated security review found that the previous commit turned off a guard
that was still doing work. Reproduced by executing both rules against a
fixture, not inferred:

  const exit = 'exit';
  process[exit](1);

n/no-process-exit flags it; local/require-registered-exit did not, because it
early-returned on callee.computed. So retiring the predecessor on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs un-guarded that shape on precisely
the two globs this epic's exit contract cares most about.

This is the third time in this epic I have removed a coarse guard on the claim
that a narrower one covered it, without checking construct-level parity — after
the allowlist key-to-prefix-to-exact-membership sequence and the band
ranges-to-categories one. The rule is the same every time: a narrower guard
supersedes a coarser one only where it demonstrably reaches at least as far,
and "demonstrably" means executing both against the constructs, not reading
either.

The successor now resolves computed property access for the statically
determinable cases — a string Literal, and an Identifier bound once to a string
Literal, resolved through scope — and leaves genuinely dynamic properties
alone so the rule does not over-fire. Measured after the fix: plain
process.exit flagged, process['exit']() flagged, process[exit]() flagged,
process[globalThis.k]() not flagged. That makes it a strict superset of the
predecessor on these globs, since process['exit']() was caught by NEITHER rule
before.

The second finding is worse than the first, because it was reasoning rather
than oversight. My justification comment claimed n/no-process-exit "would flag
terminateNow's own generated copy here". It would not — that file is in the
global ignore list, so neither rule ever lints it. There was no conflict to
resolve; I wrote a rationale I had not checked, in a change whose entire
subject is written claims nobody verified. Both comment blocks now state the
real basis.

The tests that should have caught this asserted only rule SEVERITY per glob and
never construct REACH, which is exactly how a coverage hole passed. A parity
matrix now pins all five shapes, including a RED/GREEN regression pin against
an inlined reproduction of the pre-fix rule — inlined rather than loaded from
HEAD, because HEAD resolves to the fixed commit under the remote runner and
would silently stop testing anything.

Verification runs on the remote runner.

Refs #3914

* fix(#3914): the two exit rules are complementary — keep both

Reverts this branch's retirement of n/no-process-exit. The premise was wrong
twice, and the second review proved the change itself was wrong.

I claimed local/require-registered-exit was a strict superset on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs. Measured, successor vs predecessor:

  function f(exit) { process[exit](1); }         0  vs  1
  let exit='exit'; exit='exit'; process[exit]()   0  vs  1
  const { exit } = ...; process[exit](1)         0  vs  1

plus for-of bindings, let-then-assign, var redeclaration, catch params, and an
undeclared global named exit. The predecessor matches any identifier NAMED
exit however it is bound; the successor resolves only a string literal or a
single-write const. It never was a superset — I asserted the relationship after
fixing one construct and did not re-check the rest.

The justification was independently false: all three generated cli-exit copies
are in the global ignore list, so n/no-process-exit was never flagging
terminateNow. There was no conflict to resolve. I wrote a rationale I had not
verified, in the phase whose subject is written claims nobody checked.

So criterion 5 does not apply to this pair. They are not predecessor and
successor — they are complementary, each catching constructs the other misses.
The epic's criterion assumed a replacement relationship that does not exist
here, and retiring either rule loses real coverage. The ADR ledger now says so
with the measured shapes.

What survives is the genuine improvement: the computed-property strengthening.
local/require-registered-exit now catches process['exit'](1) and optional-chain
terminators like process?.[k]?.(1), which NEITHER rule caught before, while
correctly ignoring a genuinely dynamic property so it does not over-fire.

The parity tests are rewritten to assert what is true rather than what I wanted
to be true: a bidirectional matrix where each rule is shown catching shapes the
other misses. The previous matrix tested only the four shapes where the
successor wins, which is precisely why the regression shipped — a test set
selected to confirm the thesis.

Also corrected: a stale ADR sentence claiming a third wrong ledger version that
does not exist (the table it described now reads +3 over terms summing to +3),
and a changeset whose stated motivation was the false generated-copy conflict.

Verification runs on the remote runner.

Refs #3914

* fix(#3914): the exemption term was a no-op — the net is +4

Fourth correction to this ledger, and a fourth error of the same kind.

Every version counted removing the n/no-process-exit 'off' entry from the hooks
block as -1. Measured: calculateConfigForFile returns undefined for that rule on
hooks/**. It was never registered there, and no broader block sets it globally,
so the 'off' entry overrode nothing and removing it changed no enforcement at
all. A no-op removal, not a guard removal — the same category error as counting
baseline acknowledgement entries: a thing that is not a guard, in guard units.
It is misattributed too; that block came down in d98b55562 (#3910), already on
next before this branch existed.

So the epic added FOUR guards, not three.

This surfaced from a test of mine that overclaimed. I asserted n/no-process-exit
was error on "all nine CommonJS/hook globs" — but hooks is not one of the nine,
and the rule resolves to undefined there. Fixing the test to match reality is
what exposed the ledger term, which is the argument for tests that assert
identity rather than a comfortable shape.

The hooks state is now pinned explicitly rather than glossed: n/no-process-exit
unregistered, local/require-registered-exit error. It is mildly surprising and
therefore worth a test.

Also updates a pre-existing test that documented the old name-based-only
boundary as intentional. The computed-property strengthening deliberately moves
that boundary — process['exit'](0) was caught by NEITHER rule before — so the
test now asserts the new contract and cites the ADR, rather than being left to
fail or the rule weakened to satisfy it. A contract change should read as
deliberate in the test that pins it.

Verification runs on the remote runner.

Refs #3914

* chore(#3914): backfill changeset pr number to 4018

---------

Co-authored-by: sim <sim@local>
2026-08-29 00:53:33 -04:00
Tom Boucher
ac7587287b fix(#3812): document how Current Position actually resolves a duplicate field (#4017)
* docs(#3812): say that Current Position is single-valued, and pin the behavior that makes it true

#3812 shipped CLOSED with half its acceptance unmet. #3873 delivered cardinality for FRONTMATTER
keys - current_phase/current_plan render as optional at docs/reference/state-md.md:89,91, covered by
tests/gen-state-md-docs.test.cjs:374. The issue's actual ask was the ## Current Position BODY
section, and that never landed. Surfaced by an /adr-phase-coverage audit of epic #3473; the issue
was reopened rather than noted.

The section now states three things: every field is single-valued, the section is overwritten rather
than appended to, and a duplicate resolves to the FIRST occurrence with no warning - so a line
appended in good faith is silently ignored rather than winning. Progress history belongs in
## Performance Metrics, two headings down, and the text now points there.

The third claim is a behavioral promise about the reader, so it was VERIFIED BY EXECUTION before
being written rather than inferred from the issue title:

  stateExtractField(<"Phase: 1 of 5 (First)" ... "Phase: 9 of 9 (Appended later)">, "Phase")
    -> "1 of 5 (First)"

The mechanism is state-document.cjs:405 - the plain-line pattern ^<field>:[ \t]*(.+) carries flags
im with NO g, so String.match returns the first hit. Writing "first wins" without running it would
have repeated the exact error I had to retract twice in this epic already.

A test pins the reader, not the prose. Three rows in tests/state.test.cjs: T1 (load-bearing) asserts
the duplicated case resolves first; T2 asserts the ordinary single-field case still works, so a fix
that only functions when duplicated cannot pass; T3 puts a Plan: line BETWEEN the two Phase: lines
and asserts it resolves independently - negative space, because a reader returning the first line of
the SECTION rather than the first matching FIELD would satisfy T1 alone. Proven to discriminate: a
last-match variant returns "9 of 9 (Appended later)" and T1 reds.

No assertion checks that the document contains a sentence. That is what local/no-source-grep exists
to stop, and it would pin wording that is allowed to improve. The point of the test is that if that
regex ever gains g and a last-match walk, the test fails - instead of the documentation quietly
becoming a lie with nothing to notice.

Prose only, no new heading. docs-state-md-locale-parity compares heading-level sequences by LCS
rather than text, so added paragraphs cannot fail it while an added HEADING would fail all four
locales. The constraint is structural, not stylistic - confirmed by running that comparison after
the edit.

The four locale copies are translated rather than left stale. They are not gate-enforced for prose,
so "nothing fails" was available and is not the same as correct: leaving four documents asserting
something the English one now contradicts is a correctness problem. Code spans and the anchor link
stay untranslated - they name real tokens.

The whole approach rests on one fact, checked first: ## Current Position at :196-208 sits OUTSIDE
every generated marker region (:81-104, :138-151), so a hand edit survives --write. Re-confirmed
after all five edits - gen-state-md-docs --check reports all 6 targets up to date. Had that been
false the fix would have belonged in the generator, and a hand edit would have been silently
reverted.

One real gate failure fixed inline rather than reported: the new test's comments referenced
docs/reference/state-md.md, which was not in that file's registered exempt-docs paths, and
lint-docs-guard-registration failed lint:ci correctly. Registered.

Known limit, named rather than folded in: gsd-tools validate/health still do NOT warn on a
duplicated Phase:. #3812 records that as a "consider", not a requirement, and confirms none of the
nine rules in src/health-diagnostic-rules/{state-consistency,phase-structure}.cts counts
occurrences. Documenting the silent first-match is the delivered scope; making it loud is new scope
and stays unclaimed.

Closes #3812

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3812): the rule I documented was false — replace it with the measured one

An isolated review returned two blockers. Both mine, and the first is the worse
kind: I wrote a falsifiable rule into a reference page and got it wrong.

1. "A duplicate resolves to the FIRST occurrence" is FALSE.

   stateExtractField (src/state-document.cts:401-419) tries BOLD `**F:**` across
   the whole input, THEN plain `^F:`, THEN a pipe-table row. Form precedence
   beats document order. Measured against the built reader, all intra-section:

     Phase: A (plain)  /  **Phase:** B (bold, later)   -> B    LATER WINS
       Phase: A (indented) / Phase: B (plain, later)   -> B    LATER WINS
     Phase: A (plain)  /  | Phase | T (table) |        -> A    first wins

   My original verification tested plain-versus-plain, saw first-wins, and
   generalized to all forms. Measuring one case and claiming the general rule is
   the same error I have had to retract twice already in this epic.

   It is also worse than silence. The sentence told authors an appended line is
   safely ignored; a bold line appended "for emphasis" silently overrides the
   original. Someone trusting the doc would have corrupted their own state file.
   And #3812 never asked for a resolution rule - it asked for single-valued,
   overwrite-not-append, and where history goes. The rule was my unrequested
   addition.

   Replaced with the measured truth: resolution is by FORM (bold anywhere, then
   plain at line-start, then table row), and only WITHIN the winning form does
   the first occurrence win. Both consequences stated plainly - a higher-ranked
   form wins regardless of position, and an indented `Phase:` is invisible to the
   plain form. All five claims in the new paragraph verified by execution before
   being written, including the two I had wrong.

2. The tests tested the wrong case and passed for the wrong reason.

   T1/T3 put the second `Phase:` under `## Somewhere else` - the INTER-section
   case, which #2956 already fixed by scoping. #3812 says verbatim that #2956
   "fixed the inter-section case and never addressed intra-section duplication",
   so the case the new prose describes was untested, and the fixtures passed
   because of section scoping rather than field resolution. They also called bare
   stateExtractField rather than the production chain, T2 could not discriminate
   first from last at all, and no fixture mixed forms - which is precisely why the
   false claim survived to review.

   Rewritten as four rows, all intra-section, all through the real
   stateCurrentPositionSlice -> stateExtractField path: plain-then-plain (first
   wins within a form), plain-then-bold (the bold LATER value wins - the row whose
   absence let the false claim ship), indented-then-plain (indented invisible),
   and sibling-field independence. Each proven to fail against a reader that
   disagrees.

3. Two dead anchors. pt-BR and zh-CN linked `#performance-metrics` while their own
   headings are `### Métricas de Desempenho` and `### 性能指标`. Both fixed to the
   anchor their own heading generates. ja-JP/ko-KR kept the English heading, so
   theirs already resolved.

4. A ja/ko sentence inverted its own meaning. Both rendered "which is the section
   designed to grow" with a bare demonstrative whose nearest referent read as
   Current Position - saying the opposite of the point. Rewritten so the clause
   attaches unambiguously to `## Performance Metrics`.

5. Cross-locale drift, flagged by the implementing agent rather than by me: after
   fixing EN, the four locales still stated the OLD false rule. Four documents
   asserting something measured to be wrong is worse than four saying nothing.
   All four now carry a faithful translation of the corrected paragraph, with
   code spans, each file's own anchor, and the ja/ko referent fix preserved.

Verified: all five claims executed against the built reader; every rewritten test
row proven to discriminate; gen-state-md-docs --check reports all 6 targets up to
date, so the edits stay outside the generated marker regions; locale heading
parity unaffected (prose only, no headings added); build:lib, lint and lint:ci all
exit 0.

Refs #3812

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3812): second false rule on the same page — scope the ranking to the section

A second isolated review found a second false falsifiable claim, and the failure
mode is the same one twice in a row:

  attempt 1: verified plain-vs-plain, wrote a claim about ALL FORMS
  attempt 2: verified bare stateExtractField, wrote a claim about THE DOCUMENT

Both times the claim covered a wider surface than what was actually executed. The
fix each time was not a better sentence, it was executing the surface the
sentence describes.

BLOCKER — "bold `**Phase:**` anywhere in the DOCUMENT wins" is false.

  ## Current Position / Phase: 1 of 5   +   ## Archive / **Phase:** 88

    bare stateExtractField(whole doc) -> "88 (other section)"
    PRODUCTION (slice then extract)   -> "1 of 5 (in section)"

#2956's section slice means production never hands another section to the
matcher; a bold line in `## Archive`, or in the YAML frontmatter, is simply not
seen. The ranking is real but scoped: it applies WITHIN `## Current Position`.
I verified against the bare function and wrote a claim about the system.

Every existing test placed its bold line inside the section, which is exactly why
nothing contradicted the claim. T5 now puts a bold `**Phase:**` in `## Archive`
and asserts production returns the in-section plain value, with the unscoped
reader asserted to DISAGREE so the row proves the scoping rather than assuming
it.

BLOCKER — the changeset still shipped the ORIGINAL retracted claim.

I corrected the page and left the release note saying "resolves to the first
occurrence ... a second entry added in good faith is silently ignored". The note
contradicted the page it announces, and the release note is what most people
actually read. Rewritten to the corrected rule.

MEDIUM — the concession was inverted. It read "wins even if it comes FIRST in the
file", which is the vacuous direction; the surprising case, and the one the very
next clause illustrates with an APPENDED bold line, is "even if it comes LAST".
All four locales reproduced the inversion faithfully, so it was an EN-source
defect rather than translation drift.

Two sharp edges now named, both measured: a bold `**Phase:**` followed only by
trailing spaces resolves to an EMPTY STRING and does not fall through to a valid
plain line below (T6 pins it); and `| **Phase:** | 3 of 4 |` short-circuits to the
bold form and returns the literal `"| 3 of 4 |"`. A page that teaches form
ranking has to say where the ranking bites.

Also fixed: all five files labelled the link `## Performance Metrics` while the
heading is `### Performance Metrics`. Anchors resolved correctly everywhere; only
the label's level was wrong.

Every clause in the final paragraph re-verified through the PRODUCTION chain
(stateCurrentPositionSlice -> stateExtractField), clause by clause, before being
written: bold in another section does not win; bold in frontmatter does not win;
bold appended last does win; first wins within one form; trailing-space bold
yields empty. All four locales carry the same corrected rule.

gen-state-md-docs --check reports all 6 targets up to date; heading counts
unchanged at 20/20 across all five files, so locale heading-parity is untouched;
build:lib, lint and lint:ci all exit 0.

Refs #3812

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3812): backfill changeset pr number

Refs #3812

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 20:50:15 -04:00